OpenAI's new Decisions API, now in public beta, is built for one job: getting fast, structured answers to simple questions instead of free-form text. Opened to all developers on October 6, 2026, it lets apps classify content, route requests or choose an AI agent's next action in near real time, with OpenAI claiming answers roughly 10x faster than the Responses API. Here is a practical guide to what it does, how to use it and when it makes sense.

What the Decisions API Does

Most AI applications contain dozens of small judgement calls: Is this message spam? Which department should handle this ticket? How severe is this complaint? Traditionally, developers asked a chat model and parsed its prose. The Decisions API replaces that with typed answers returned from a dedicated endpoint, POST /v1/decisions.

According to OpenAI's documentation, it supports three question types:

  • Predicate: estimates the probability, from 0 to 1, that a statement is true. Example: "Does this product photo show visible damage?"
  • Choice: selects one option from a fixed list and returns the selection with probabilities and a confidence value. Ideal for unordered categories such as departments or intents.
  • Score: rates input against ordered levels and returns a probability-weighted score. Useful for severity, priority or quality ratings.

Step-by-Step: Making Your First Decision

1. Update your SDK

The API is available through OpenAI's official SDKs. The docs list minimum versions including Python 3.26.0 and JavaScript 7.30.0, with Go, Ruby and Java also supported.

2. Experiment in the Playground

Before writing code, OpenAI recommends testing questions and inputs in the Decisions Playground on its platform. It is the quickest way to see how the model interprets your wording.

3. Build the request

Every request has three parts:

  • Model: currently only gpt-6-luna is supported. GPT-6 Luna is the reasoning model OpenAI released alongside GPT-6 Sol on September 22, 2026.
  • Input: text, images or both, placed in user messages.
  • Questions: an array of evaluation tasks, each with a unique name so you can match answers back.

4. Read the answers

The response contains an answers array that mirrors your questions. Each answer echoes the question's name and returns a probability, a chosen value or a score, plus confidence and probability distributions for choice and score questions.

Limits to Know Before You Ship

A few constraints can trip up first-time users:

  • Images must be inline base64 data URLs. Hosted image links and uploaded file IDs are refused.
  • Combine image and text in a single user message when they should be evaluated together.
  • Dependent decisions need separate requests. Ask independent questions in one call, but if one answer determines the next question, sequence them.
  • No caching yet: a user on OpenAI's developer forum reported that caching is not currently available.

Pricing and Compliance

Pricing is simple: $0.10 per million input tokens with GPT-6 Luna, and you pay only for input. There are no output-token or cache charges, though regional processing premiums and long-context multipliers can apply.

For regulated industries, the beta ships with Zero Data Retention support, HIPAA eligibility and data residency in the United States and Europe. OpenAI says it expects general availability in the coming weeks.

Best Practices From the Docs

  • Calibrate thresholds using labelled examples from your own application instead of guessing a cut-off like 0.5.
  • Weigh error costs: decide whether false positives or false negatives hurt more for each workflow, and tune thresholds accordingly.
  • Add a fallback option such as "other" to choice questions so the model is not forced into a bad fit.
  • Batch independent questions in one call to save latency.

Why It Matters

As agentic AI systems grow, a large share of their work is not writing paragraphs but making quick, repeated decisions about what to do next. A purpose-built, input-only-priced endpoint for those decisions can cut both cost and latency, and its probability outputs make it easier to set confidence thresholds and send uncertain cases to a human.

It also arrives as small-model pricing collapses, with Anthropic's new Claude Haiku 5.5 now matching GPT-6 Luna's token rates. One caveat: independent observers note OpenAI's 10x speed claim has not been published with a benchmark, so test it against your own workloads.

Use the Decisions API for the yes, no and pick-one moments in your app, and save full chat models for work that genuinely needs prose.

Sources