Decisions API is in public beta
developers.openai.com380 points by chiefstorm a day ago
380 points by chiefstorm a day ago
curl https://api.openai.com/v1/decisions \
-H "Authorization: Bearer $(llm keys get openai)" \
-H "Content-Type: application/json" \
--data '
{
"model": "gpt-6-luna",
"input": [{
"role": "user",
"content": [
{"type": "input_text", "text": "I am angry about the new product feature"}
]
}],
"questions": [{
"type": "predicate",
"name": "complaint",
"instructions": "Is this a complaint?"
}, {
"type": "predicate",
"name": "compliment",
"instructions": "Is this a compliment?"
}]
}'
Returned: {
"model": "gpt-6-luna",
"answers": [
{
"type": "predicate",
"name": "complaint",
"probability": 0.91
},
{
"type": "predicate",
"name": "compliment",
"probability": 0.06
}
],
"usage": {
"input_tokens": 310,
"input_tokens_details": {
"cached_tokens": 0,
"cache_write_tokens": 0
},
"output_tokens": 0,
"output_tokens_details": {
"reasoning_tokens": 0
},
"total_tokens": 310
}
}
That https://api.openai.com/v1/decisions endpoint is notable because usually when OpenAI define an endpoint like that it ends up as a defecto standard for other providers.(I turned this all into a new llm plugin: https://github.com/simonw/llm-openai-decisions)
>defecto standard for other providers.
I know it's just a little typo but it made my morning :)
defacto: real
dejure: legal
defecto: enshittified
(adj.): The defecto way to playback music is Spotify.
How is this different from the categorisation models from ML era?
These are essentially zero-shot classifiers; they don't need to be trained for a specific classification task. You could include some natural language context on the rules for classification and it should get good enough accuracy.
At least a year ago they were not that accurate when compared to a good many-shot classifier that (1) has 10k training examples and (2) has the learning capacity to learn from 10k training examples (most fine tuned BERTs don't seem to.)
I'm skeptical of the quality of probability calibration for models if you aren't giving them training data. The issue is that there is a prior distribution that's unique to your specific data and the calibration is really sensitive to that.
How is it different than just asking an LLM with structured outputs enabled? Is the primary value-add that it gives a confidence level?
1) you can parallelize the request 2) since a structured output is still just text then every previous property of the structured output effects subsequent ones
The pitch is that it's a fully-general model, so you can skip training/tuning/selecting a particular categorisation model for each task.
That's great! So someone finally built the zero-shot model from the sales decks of 2015 =D
Specifically, it happened a few weeks ago when Typesafe released Jev; this is OpenAI's competitor to Typesafe.
This isn't really anything new it just seems like a new API but you could do the exact same thing with just a little bit of prompt engineering all the way back when GPT-3 was first released. Am I missing something?
No amount of prompt engineering will give you the true probabilities for the model producing a certain response; this is something you can only get by inspecting the internal state at inference time.
Does this give you the true probabilities for a certain response either? How does it work exactly? The probability of an overall positive answer isn't just the probability that the next token is "Yes"
For most use cases is that actually needed though? Just having it choose between predefined responses seems like enough but I'm curious about specific use cases because I do feel like I'm missing something
This is useful for classification problems; any time you need to write software that looks at some fuzzy data and needs to make a probabilistic decision. It's far more cost-efficient and performant to use this type of model instead of an LLM.
Before now you had to train a model on your specific classification problem, now these new models don't require any specific training at all to do pretty well on novel problems.