sps | 2026-10-06 23:12:05 UTC | #1 Let your app choose the right model, tool, or action in near real-time with Decisions API, now available to all developers in public beta. The Decisions API makes decisions up to 10x faster than GPT-6 Luna through the Responses API. https://youtu.be/FB6oCmrIj-Y?si=FO4rhgMOaV4DAThk Powered by GPT-6 Luna, it accepts text and image inputs and supports 3 kinds of outputs: * **Predicates:** Estimate the probability that a statement is true. * **Choices:** Select from predefined options, with confidence scores. * **Scores:** Evaluate an input against a numeric range. ### **Pricing** With `gpt-6-luna`, input costs **$0.10 per 1M tokens**. You pay only for input tokens: there are no cache-read, cache-write, or output-token charges. You can now try the Decisions API in public beta: https://developers.openai.com/api/docs/guides/decisions ------------------------- merefield | 2026-10-06 20:55:58 UTC | #2 Ah now Santa really did come early this year! Pumped to give this a try! ------------------------- platypus | 2026-10-06 21:02:06 UTC | #4 Let's see who will be the first to benchmark Decisions against Jev! ------------------------- spiroskaye | 2026-10-06 22:27:55 UTC | #5 Now this is what I was focused on from DevDay. OpenAI, really hoping you guys publish a blog explaining this and the model architecture to some extent like decider-2b. ------------------------- jefff | 2026-10-06 23:09:42 UTC | #6 I ran some quick tests vs jev for my use cases, didn't see an improvement with Luna :frowning: Looking forward to seeing Luna on some real benchmarks though - is S1MB on huggingface the best place to look? > * **Accuracy:** Jev was more accurate on nuanced judgment calls, such as whether two words are meaningfully related or what a person in a conversation is asking for. Luna made about three times as many confident wrong answers on those. > * **Simple factual questions:** on narrow yes/no checks like "is this a real English word?", the two were about equal. Luna was slightly more often right and slightly less well calibrated, and the difference wasn't large enough to matter in practice. > * **Speed:** Jev's typical response was faster. In multi-step conversational use, its median turn took about 0.3 s against Luna's 1.6 s. Luna was more consistent per call, while Jev had occasional slow outliers. > * **Cost:** Luna cost about 2–3× as much for the same questions. Both cost fractions of a cent per call. > > **Bottom line:** Jev is the better default. It's cheaper, faster and more accurate on judgment-heavy questions. Luna didn't show a clear advantage anywhere we tested. > > This is based on a few hundred test questions from word games and one conversational game, so treat it as a strong early signal rather than a general benchmark. ------------------------- sam.saffron | 2026-10-07 01:42:46 UTC | #7 Image understanding is something Jev does not support yet, I can see decisions API quite useful for those cases, eg: - Is image appropriate for sites guidelines? - Does image contain source code? ------------------------- JustAutomaidIt | 2026-10-07 08:56:01 UTC | #8 This should really give Jev it run for the money. ------------------------- Poupanka | 2026-10-07 09:26:36 UTC | #9 It seems that the decisions api has no cached input tokens. Is that expected? ------------------------- VeitB | 2026-10-07 09:32:48 UTC | #10 Yes, there’s currently no caching available for the Decisions API. The Decisions API is still in beta, so I wouldn’t be surprised if caching is introduced in the future. > With gpt-6-luna, input costs $0.10 per 1M tokens. You pay only for input tokens: there are no cache-read, cache-write, or output-token charges. https://developers.openai.com/api/docs/guides/decisions?utm_source=chatgpt.com#pricing-and-availability ------------------------- Poupanka | 2026-10-07 09:42:54 UTC | #11 Thank you, that is unfortunate as it makes classification tasks economically unattractive to switch from luna responses api to decisions api. I hope caching is enabled soon :crossed_fingers: ------------------------- JustAutomaidIt | 2026-10-07 09:46:17 UTC | #12 Has anyone been able to test this ? ------------------------- VeitB | 2026-10-07 09:51:45 UTC | #13 Yes, I’m implementing a small test project where GPT models play a real-time 3D game. On a slow internet connection of around 3 Mbps, the Decisions API responds to image inputs in about 0.8 seconds. There's no caching of input tokens. Is there anything specific you would like to know? ------------------------- platypus | 2026-10-07 10:06:01 UTC | #14 I've been playing around with it, to see if it's susceptible to the same issues I found with Jev. One big finding is that there is quite a difference between predicate and choice based formats. The predicate seems to follow more closely expected results, while choice seems to put too much probability mass on the most likely outcome. As an example, I ran a classic loaded coin test, where **the coin was biased to show heads 70% of the time**. Using predicate setup, across 1000 trials, I get heads 70% of the time - as expected. When I setup the question as a choice, I would expect heads 70% of the time, **but instead I got it 98% of the time**. ------------------------- VeitB | 2026-10-07 10:09:01 UTC | #15 So it's optimizing for the most likely outcome and then it should actually be predicting Heads 100% of the time? And it's clearly not estimating the probability of the next flip being heads at ~70%. ------------------------- platypus | 2026-10-07 10:09:48 UTC | #16 Exactly, it's not estimating the true posterior at all. ------------------------- platypus | 2026-10-07 10:18:39 UTC | #17 So in the game demo Romain showed, it doesnt matter, because you want it optimise for the most likely choice (the "opening" on the road). But let's say you make the road wider, and you give it multiple openings, which one will it choose? And let's say only one of those openings will lead to an optimal next-step outcome, and lets say Luna/Decisions can "see" that, will it choose the right one? **EDIT (can't do more than 2 consecutive replies):folded_hands:** The plot thickens - choice order changes the probabilities :scream: Should I make a dedicated post regarding the results? I mean to me this is unreliable, but Jev isn't any better. ------------------------- VeitB | 2026-10-07 10:23:19 UTC | #18 Please go ahead. I'm also looking into the best approach here to explore the probability distribution apart from predicting the most likely outcome. ------------------------- platypus | 2026-10-07 10:26:01 UTC | #19 My verdict right now is don't use choice questions. ------------------------- oskorovii | 2026-10-07 11:37:08 UTC | #20 When it will be available on Azure or AWS? ------------------------- platypus | 2026-10-07 11:51:38 UTC | #21 I think that's normally the question for Microsoft and AWS, if you have reps/account people there you can ask them. ------------------------- _j | 2026-10-07 12:24:33 UTC | #22 Issue to be aware of: The playground field descriptions for *decisions* are currently misaligned from the underlying API, and are not what is sent nor documented. This makes the experience of setting up a call non-instructive and divergent when reading documentation alongside playground experimentation. - "prompt" - actually the `questions` array [0] object's "instructions" field - "options" - actually the `questions` array [0] object's "choices" [0] value array items - chat-like input - actually "input" or filling the fields as their appear to be described, and viewing the code, for a "choice" type from ["predicate", "choice", "score], the "code" button to view the request shape delivers: ``` { "model": "gpt-6-luna", "input": "Chat box input", "questions": [ { "type": "choice", "instructions": "prompt-field input", "choices": [ { "value": "type_choice-options_1_input" }, { "value": "type_choice-options_2_input" } ] } ] } ``` ------------------------- marc_omni | 2026-10-07 16:31:23 UTC | #23 Love this. Any plans on AWS Bedrock support? ------------------------- platypus | 2026-10-07 17:19:50 UTC | #24 I forwarded that question to AWS peeps. ------------------------- platypus | 2026-10-07 17:23:36 UTC | #25 For those interested here is the post on my little toy experiment: https://community.openai.com/t/results-from-toy-experiments-with-decisions-api/1404010 The takeaway here is that how you frame the question, and how you order your choices, matters. It seems like ordering is heavily influenced by the LLM (Luna), i.e. it follows the autoregressive scheme, where some weight is given on the logprobs, in the same way logprobs for next-token change depending on the ordering. ------------------------- sabbadin12 | 2026-10-07 20:18:58 UTC | #26 What is the context window size ? ------------------------- VeitB | 2026-10-07 20:23:32 UTC | #27 The model behind the Decisions API is based on GPT-6 Luna The standard context tier is up to 272K input tokens. The 1M-token context window is also available. ------------------------- merefield | 2026-10-07 21:12:03 UTC | #28 For decision models, more context isn’t necessarily better. Beyond the relevant evidence, extra context can actually make decisions less accurate. Instead you are better to prepare a more succint context on which to make a decision. Build that up front and present that to the model instead of swamping it. That will also save you money potentially. Typesafe say: "Decision models benefit from focused, relevant context rather than maximally large context. Excessive or irrelevant information can dilute the evidence needed for accurate judgments. TypeSafe AI's recommended patterns emphasise selecting relevant evidence, independently evaluating candidate passages, and decomposing complex decisions into focused questions." Appreciate that's a different provider. ------------------------- _j | 2026-10-07 21:31:56 UTC | #29 [quote="merefield, post:28, topic:1403877"] For decision models, more context isn’t necessarily better [/quote] For some cases, the original context length may be better. *"Is this chat input going to lead up to producing biological weapons or code exploits"* A question that moderations API won't answer - but OpenAI's internal inspections of API use will certainly catch and score against your organization. ------------------------- curt.kennedy | 2026-10-09 02:49:58 UTC | #30 A post was merged into an existing topic: [Introduce yourself! 😊 ‎ ‎ ‎ ‎](/t/introduce-yourself/45/1130) -------------------------