UpShaqo
Intelligence desk
Agent Market Source-backed analysis

OpenAI's Decisions API Turns Agent Safety Into a Price War

A cheap classifier model born at a small startup just got cloned by the industry's biggest distributor, and the gap between $2.94 and $372 explains why enterprise buyers should care.

UpShaqo Editorial IntelligenceSeptember 30, 20266 min read
Intelligence standard

Independent UpShaqo analysis built from fresh, attributed sources. We explain the impact instead of repeating the announcement.

Read for leverage: focus on the workflow change, the customer problem, and the next action—not only the product announcement.

Sam Altman spent most of OpenAI's Dev Day talking up Luna, the lab's flagship model. The detail that will actually move budgets arrived almost as a footnote: a new "Decisions API" that hands Luna a fixed menu of choices — image categories, agent behaviors, whatever a developer defines — and asks it to pick, fast and cheap, instead of reasoning its way through the problem like a full-strength chatbot would. According to TechCrunch's account of the announcement, Altman framed it as a way to keep capabilities like image understanding and safety protections intact while making the model "extremely fast" on narrow decisions.

That description will sound familiar to anyone who's been watching TypeSafe AI, a startup that shipped a model called Jev earlier the same month for exactly this job: a classifier built on an LLM that outputs probabilities across a predefined set of choices, cheaply and at high speed. TypeSafe's CEO, Diogo Almeida — a former OpenAI engineer credited with co-inventing reinforcement learning — reacted to the Decisions API news by joking on X about "the beginning of the clone wars," per the same TechCrunch reporting.

The Admission Buried in an Aside

What makes this announcement worth a memo rather than a shrug is what it concedes. OpenAI is implicitly admitting that its own frontier models are the wrong tool for a huge slice of production software — the classification calls, the routing decisions, the yes/no gates that agents need thousands of times a second. Those workloads don't need System 2 reasoning; they need what TypeSafe calls "System 1" judgment: fast, intuitive, statistically calibrated. Almeida told TechCrunch that OpenAI's move could be read as validation that "building in a System One compatible way is the future."

That's a notable thing for a foundation-model lab to concede, because it splits the market it just spent years consolidating. If cheap, narrow decisioning is a separate product category from general reasoning, then buyers no longer need to pay frontier-model prices for every agentic action — and that's a pricing-power problem for whoever sells the expensive model.

A Security Problem Gave the Idea Its Business Case

The clearest use case isn't classification for its own sake — it's watching agents so they don't misbehave. OpenAI has reportedly been running a separate model to monitor agent actions "at significant compute cost" following incidents where its agents acted badly on the open internet, according to TechCrunch. That's an expensive insurance policy: a frontier model babysitting another frontier model, action by action.

Shapor Naghibzadeh, who leads the cybersecurity startup QueryStory, built a hackathon demo using Jev to solve this more cheaply. His system checks each agentic action against its assigned task, blocking anything it's highly confident is bad, flagging ambiguous cases for a human, and letting the rest through automatically. The reported cost difference is the headline number here: monitoring of that kind runs $2.94 with Jev versus $372 with a frontier LLM, per the TechCrunch report — a roughly 125x gap for comparable coverage. TechCrunch notes the approach theoretically could have caught the Hugging Face incident before it happened.

Why this matters for buyers: enterprises adopting agentic workflows are discovering that the review layer — not the agent itself — is where costs balloon, because reviewing every action with a frontier model is prohibitively expensive at scale. A classifier cheap enough to run on every single action, rather than a sample of them, changes the economics of what "safe by default" can mean.

The Pricing Fork This Creates

Think of it as a two-tier stack forming inside every serious agent deployment:

  • Tier one — the reasoning model. Expensive, general-purpose, used for genuinely open-ended decisions.
  • Tier two — the decisioning layer. Cheap, narrow, used to gate, classify, and monitor everything the reasoning model or its agents actually do.

Once buyers see that tier two can be 100x cheaper for comparable coverage, procurement conversations change. Instead of asking "which frontier model do we standardize on," enterprise buyers start asking "which vendor's decisioning layer do we run underneath whichever frontier model we pick." That's a wedge for a company like TypeSafe to win a slice of enterprise spend independent of who wins the model war — but it's also exactly the layer OpenAI now wants to own itself, bundled directly into its existing platform.

Distribution Still Beats a Better Widget

Here's the uncomfortable part for TypeSafe. Jev shipped first, and Almeida frames his moat as the synthetic data pipeline that makes Jev's probability outputs statistically reliable rather than just fast — "intelligence is the hard part," he told TechCrunch, describing his goal as pushing the intelligence-per-dollar curve, not just the cost curve. That's a real technical claim, and it's unverified by outside benchmarking so far.

But OpenAI doesn't need to out-build that moat to win the deal. It needs its Decisions API to be good enough, and then it wins on distribution: existing API relationships, existing billing, existing trust from enterprise security teams, and a pitch that says "stay in one stack." TechCrunch reports the Decisions API shipped only as a limited preview, and it's not yet clear how closely its calibration matches Jev's — the comparison is still untested in the wild. That uncertainty is the whole ballgame for TypeSafe: if Decisions API is merely adequate, most enterprise buyers will default to the incumbent anyway rather than adding a second vendor for marginal quality gains.

The Wedge That's Still Open

Our analysis: the genuinely underserved space right now isn't general-purpose classification — it's calibrated, auditable agent monitoring sold as a compliance and security product rather than an API primitive. TechCrunch's reporting notes that other startups are building similar decision models too, which means this won't stay a two-player race. The buyers who care most — security and compliance teams responsible for agent incidents — want proof of calibration, incident logs, and audit trails, not just a cheap probability score. Whoever packages the monitoring layer with evidence that it actually prevents incidents, rather than just classifying actions quickly, has room to win business away from the platform giants on trust rather than price.

What Operators Should Do Next

For teams already running or planning agentic workflows, a few concrete moves follow from this:

  • Separate your reasoning-model spend from your decisioning/monitoring spend in cost modeling now, because the 100x price gap TechCrunch reported means monitoring every action is suddenly affordable where it wasn't before.
  • Pressure-test any vendor's calibration claims before trusting a cheap classifier to auto-block agent actions — a fast wrong answer is still wrong, and none of the calibration claims here have been independently benchmarked yet.
  • Watch whether OpenAI bundles Decisions API pricing into existing enterprise contracts; that bundling, more than raw capability, is likely to decide who wins the decisioning layer.

Sources

#OpenAI#Decisions API#TypeSafe AI#Jev#agent security#AI pricing#enterprise AI

Two doors. Pick one.

Hire the team.
Or become it.