The OpenAI Decisions API is now in public beta, and it charges $0.10 per million input tokens with nothing at all for output. OpenAI opened the endpoint to every developer on October 6, 2026, a week after previewing it at DevDay, and it runs on one model only: GPT-6 Luna.

That price puts OpenAI in an awkward spot. It is more than double what TypeSafe charges for Jev, and the open-weight Clef-flash from Cloudflare costs a cent less per million tokens. What OpenAI brings instead is the compliance paperwork most enterprises already signed.

I covered the first wave of these models in what AI decision models are and how Clef, Jev and Decider compare, when OpenAI had no public price. This post fills that gap from OpenAI’s announcement, the API docs and Vercel’s integration notes. I have not benchmarked the endpoint myself, and nobody has published an independent accuracy comparison yet.

What Is the OpenAI Decisions API?

The OpenAI Decisions API is an endpoint that answers developer-defined questions about text or images with typed probabilities instead of text.

You send a state, such as a support ticket or a screenshot, plus one or more questions. OpenAI has not documented a maximum. Luna returns a number for each one. It never writes a sentence, which is the whole point: you can threshold a probability in code, and you cannot threshold a paragraph.

OpenAI’s Decisions API guide defines three question types:

TypeReturnsExample use
predicateProbability (0 to 1) that a condition is true”Does this passage explain how to rotate an API key?”
choiceOne option, probabilities for every option, a confidence value”Which team handles this ticket: billing or other?”
scoreProbability-weighted average of ordered level indices, plus confidence”Rate documentation coverage: none, partial, complete”

(Source: OpenAI Decisions API guide, October 2026. Any answer can also be a refusal.)

If you used Jev or Clef, the shape is familiar but the names are not. Those models call the yes/no type noul and take a state field. OpenAI uses predicate and input. Moving between them is a mapping job, not a rewrite.

OpenAI’s launch video walks through the three types:

Introducing the Decisions API · OpenAI · YouTube

In short, the Decisions API is OpenAI’s version of a decision model, delivered as a special endpoint on GPT-6 Luna rather than as a separate model.

How to Call the Decisions API

Here is the docs example, trimmed to one question. It routes a support complaint to a department:

curl https://api.openai.com/v1/decisions \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-6-luna",
    "input": "I was charged twice for my order.",
    "questions": [{
      "type": "choice",
      "name": "department",
      "instructions": "Which department should handle this complaint?",
      "choices": [
        {"value": "billing", "description": "Payments, invoices, and refunds."},
        {"value": "other", "description": "Requests outside these categories."}
      ]
    }]
  }'

The response holds an answers array in question order, each echoing its name. Check each answer’s type before you read numbers from it, because a refusal has no probabilities.

A few details I would have missed without reading the docs closely:

  • SDK versions: the examples need the Python SDK 3.26.0 or the JavaScript SDK 7.30.0 or later
  • Images: only inline base64 data URLs; hosted image URLs and file_id inputs are rejected
  • Dependent questions: independent questions can share a request, but a question that depends on another’s answer needs a second call, per Unite.AI’s summary of the docs
  • Always add an escape hatch: OpenAI recommends an other choice so the model is not forced into a wrong bucket

If you are on Vercel, Vercel’s guide shows the same call through AI Gateway with the model ID openai/gpt-6-luna-decisions and the AI SDK’s experimental_decide function, from ai 7.0.128. One catch: Gateway’s OpenAI-compatible endpoint is text only and rejects image parts.

For anyone already on the OpenAI SDK, adopting the Decisions API is a version bump and a new request shape, and that low switching cost is its strongest argument.

Decisions API Pricing vs Clef, Jev and Perplexity Decider

The pricing is simple. You pay $0.10 per million input tokens and nothing for output, cache reads or cache writes. There is no prompt caching yet, which an OpenAI moderator confirmed in the forum thread. Regional processing premiums and long-context multipliers still apply.

Here is how that sits next to the other decision models, all of which bill input only:

ModelInput $/1MCost per 1M decisions (300 tokens)Weights
OpenAI Decisions API (GPT-6 Luna)$0.10$30Closed
Perplexity pplx-decider-v1-27b$0.04$12Apache 2.0
TypeSafe Jev 1.13$0.042$12.60Closed
Cloudflare Clef-flash$0.09$27Apache 2.0
Cloudflare Clef$0.24$72Apache 2.0

(Sources: OpenAI Decisions API guide; Cloudflare Clef and Clef-flash model pages; Jev and Perplexity prices as verified for this blog in October 2026. The 300-token figure is my assumption for a short ticket plus a few questions.)

Cost per million decisions at 300 input tokens (list price)
  • Cloudflare Clef 72$
  • OpenAI Decisions API 30$
  • Cloudflare Clef-flash 27$
  • TypeSafe Jev 1.13 12.6$
  • Perplexity Decider 12$

Source: OpenAI and Cloudflare docs, October 2026

Honestly, at these prices the difference between $12 and $30 per million decisions will not decide anything for most teams. A product doing ten million routing calls a month pays about $300 on OpenAI and $120 on Perplexity’s model.

The more interesting comparison is with plain LLM calls. Claude Haiku 5.5 and GPT-6 Luna both list at $0.10 input and $0.50 output per million tokens, as I noted in the Claude Haiku 5.5 breakdown. For a one-word classification, the free output saves almost nothing. What you actually gain is calibrated probabilities and, according to OpenAI, speed. All current list prices sit on the AI model pricing tracker.

On cost alone the Decisions API is mid-pack: cheaper than Clef, about the same as Clef-flash, and roughly 2.4x Jev and Perplexity Decider.

Speed, Accuracy and the Calibration Catch

OpenAI’s speed claim is relative: the docs say the endpoint is “about 10x faster than the Responses API.” There are no absolute latency numbers in the docs, and no published accuracy benchmark against Jev, Clef or anyone else.

That leaves us with Cloudflare’s own figures for the competition. In Cloudflare’s Clef launch post, Clef-flash had a 38.8 ms median latency, Clef 209.3 ms and Jev 524.1 ms. Those are vendor numbers from a vendor that sells two of the three models.

The forum thread has the first community tests, and they are worth reading as anecdotes, not data:

  • One developer reported image-input answers in about 0.8 seconds on a 3 Mbps connection
  • One tester found Luna less accurate than Jev on nuanced judgment calls and about even on simple yes/no checks
  • A biased-coin test, where the true answer was heads 70% of the time, returned about 70% as a predicate but about 98% as a choice
  • Another tester reported that changing the order of choices changed the probabilities

That coin test is the part that would worry me in production. If choice pushes probability mass onto the most likely option, its confidence numbers are not calibrated, and any threshold you set on them is wrong.

My practical rule: use predicate when you need a probability you will threshold on, and treat choice confidence as a ranking, not a likelihood. Test with your choices in a few different orders before you ship.

Until someone publishes a proper evaluation, treat the Decisions API as fast and convenient, with calibration you must measure on your own data before you trust it.

What the Decisions API Does Not Document Yet

The beta docs leave real gaps, and Unite.AI’s coverage flags the same ones:

  • Maximum questions per request is not listed. Clef documents 1 to 64.
  • Context window for this endpoint is not listed. Clef and Clef-flash document 65,536 tokens.
  • Image limits are not listed. Clef allows up to four images of up to 4 MiB each.
  • Rate limits and general availability pricing are not published.
  • Fine-tuning is not mentioned. Cloudflare offers RL fine-tuning for Clef through an engagement, and the open-weight models can be tuned yourself.

What OpenAI does document is the enterprise side. Zero Data Retention and HIPAA use are available to eligible customers, with data residency in the United States and in Europe (EEA and Switzerland). For a regulated company that already has an OpenAI agreement, that alone may settle the choice.

The Decisions API is the easiest decision model to get through procurement and the least documented one to engineer against.

Which Decision Model Should You Use in 2026?

Here is how I would choose today, from the published facts.

Pick the OpenAI Decisions API if you already run on OpenAI, need ZDR, HIPAA or EU residency under an existing contract, or want one vendor for both generation and routing. Budget time to test calibration.

Pick Cloudflare Clef-flash if latency is the constraint, such as a check that runs on every request inside a Worker. It is the fastest published option at almost the same price, and the weights are Apache 2.0.

Pick TypeSafe Jev or Perplexity Decider if volume is high and cost per call matters, or if you want the cheapest hosted option at about $12 per million short decisions.

Self-host an open model if data cannot leave your network at all.

Decision models matter most inside agent loops, where a cheap, fast call decides which expensive model or tool runs next. Model choice really does move the bill: the tsc-rs TypeScript compiler port reports about $24,000 of Claude spend against more than $400,000 of failed GPT attempts on the same task. A router that picks well is worth more than its own token price.

OpenAI says general availability is “in the coming weeks.” I will update this post when GA pricing and limits land, because the beta price is not a promise.