OpenAI Just Gave Every Business a Checklist for Buying an AI Agent
OpenAI's enterprise AI agent platform, Presence, reveals exactly what to check before buying any AI agent for your business.

TL;DR
OpenAI gave every business owner shopping for an AI agent tool a free evaluation checklist this week. It isn't packaged that way, and OpenAI almost certainly wasn't thinking about your five-person clinic when it built it. On July 22, 2026, the company launched Presence, an enterprise platform for deploying AI voice and chat agents across customer support, sales development, procurement, IT, and HR (OpenAI). You cannot buy it. What's worth your time is what OpenAI decided it had to build before it trusted its own agents anywhere near a real customer.
What is OpenAI Presence, exactly?
Presence is a deployment platform, not a new model. It bundles the pieces a company needs to run an agent in production: written policies and standard operating procedures, guardrails that step in when a conversation drifts outside approved territory, a defined list of actions the agent is allowed to take on its own, and a testing process before anything reaches a real customer (VentureBeat).
Company teams connect their own systems and define, in writing, what the agent can and cannot decide on its own. New agents get tested against ordinary requests, edge cases, and the higher-risk scenarios someone had to sit down and think through in advance. After launch, OpenAI's Codex reviews live interactions and proposes improvements, but a person still has to approve a change before it goes live (Help Net Security).
Why doesn't this apply to most small businesses directly?
Because it is not for sale to most businesses. Presence is in limited general availability for enterprise customers only, deployed through OpenAI's own Forward Deployed Engineers or a short list of systems integrators, not something you sign up for on a pricing page (PYMNTS). OpenAI has not published a price, a geographic limit, or an estimate of the integration work a deployment actually takes.
The early customer list backs that up. BBVA, IAG, and SoftBank have signed on to build "trusted" voice and chat agents with OpenAI's team, not exactly companies sizing up a self-checkout SaaS tool (SiliconANGLE). If a vendor pitches you "the same tech OpenAI just launched," ask whether they mean Presence itself or something merely built on OpenAI's models, since those are very different claims.
So why should an SMB owner care at all?
Because Presence is the clearest public look yet at what a frontier AI lab thinks "responsible" actually requires, coming from the company with every incentive to claim its agents just work on their own. I have sat through plenty of AI vendor demos aimed at small businesses, and the pitch almost always skips this part. The agent answers the phone, books the appointment, and the slide moves on before anyone asks what it does when a caller wants something outside the script.
Here is the part worth noticing: even OpenAI, selling agents to companies with far bigger budgets and risk teams than yours, still ships Presence with an explicit list of what the agent may decide alone, a rule for when it has to stop and hand off to a person, a simulation phase before real customers touch it, and a human sign-off before its permissions change. That is not the behavior of a company confident its agents can run themselves. It is the behavior of a company that knows exactly how an ungoverned one goes wrong, and built the fence first. It's the same instinct behind our own approval-gate framing for AI agents at any size business.
What's actually verifiable about the 75% number?
OpenAI says it already runs its own English-language phone support line on Presence and resolves 75% of inbound calls without a human stepping in (OpenAI; Help Net Security). That is a real number, and worth noting. It is also, so far, the only number.
No published methodology explains what counts as "resolved," no comparison to before-Presence performance, no independent audit, and one company reporting on its own line. We could not find a third party that verified the figure independently while researching this piece, and none of the coverage we checked cited one either. Treat it as a marketing claim from a company with obvious reason to publish a flattering one, not a benchmark you should expect from any agent you deploy, including one running on OpenAI's own models.
The checklist: what to ask any AI agent vendor before you sign
None of this is really about OpenAI. It is a usable list for evaluating any AI agent tool a vendor puts in front of you, from an enterprise platform like Presence down to a $40-a-month AI receptionist:
- What can it decide on its own, and what does it have to hand to a person? Get the actual list, not "it's smart, it figures it out."
- What triggers a handoff, and how long does that take when it happens? Ask to see it happen, not just hear it described.
- Can you see a full log of what it did and said, not a rolled-up summary?
- Did it run against real (or realistic) transcripts before it ever touched a paying customer?
- Who has to approve a change to what it is allowed to do, and how fast can you shut it off?
- What does it do when it does not know the answer? "Guesses confidently" is a disqualifying answer.
A vendor who cannot answer most of these in plain language, without you prompting them, has not built the fence yet. That is the actual signal, not the demo.
So what does this mean for your business?
The direct mechanism is money protected, not money made. An AI agent with no defined boundary does not fail quietly: it quotes the wrong price, promises a refund policy that does not exist, or escalates a frustrated customer straight into a public one-star review before anyone at your business even knows the call happened. A $40-a-month tool that does that twice in a month has already cost more than a front-desk hour ever would have saved.
The second mechanism is time saved on evaluation itself. Walking into a vendor call with the six questions above turns a multi-week trial-and-error vendor search into one direct conversation, because you will know within minutes whether they have actually thought about the failure modes or just the happy path.
The third mechanism is growth you can actually trust. Once you know where the human handoff line sits, and you've watched a vendor test for it, you can hand off real intake, scheduling, and first-response work with confidence instead of a leap of faith. That's the same bounded-autonomy principle behind how we build Perpetua: an agent gets exactly as much room to act as someone deliberately gave it, never more.
If you are evaluating your first AI agent and want a second set of eyes on the vendor's answers to that checklist before you sign anything, that is exactly the kind of conversation a free AI Opportunity Call with Arios is for. For more on how we think about giving an AI agent the right amount of autonomy in the first place, see our AI Operations Blueprint and how to build an AI-ready tech stack without over-trusting any single tool. And if you are still mapping which tasks in your business are even candidates for an agent, start with the 7 most automatable processes in every company.
Frequently asked questions
What is OpenAI Presence?
A deployment platform, launched July 22, 2026, that packages the policies, guardrails, approved actions, testing, and update process an enterprise needs to run AI voice and chat agents for customer support, sales, procurement, IT, and HR.
Can a small business buy OpenAI Presence?
No. It is limited general availability for enterprise customers only, deployed through OpenAI's Forward Deployed Engineers or select systems integrators, with no self-serve option and no published pricing.
Is the claimed 75% call-resolution rate independently verified?
Not that we could find. It is OpenAI's own reported figure from its own English-language support line, with no published methodology and no third-party audit located during research for this piece.
What questions should I ask before buying any AI agent tool for my business?
At minimum: what it can decide alone, what triggers a human handoff and how fast, whether you can see a full action log, whether it was tested on real scenarios before going live, who approves changes to its permissions, and what it does when it does not know the answer.
Does this connect to what Arios builds with Perpetua?
Yes. Perpetua's design starts from the same principle Presence demonstrates: an AI agent should only ever have as much autonomy as someone explicitly decided to give it, with a clear, tested boundary for when it hands off to a person.


