Search

What Does CUA Stand For? — Sorting Out Computer-Using Agent, Computer Use, and Computer-Use Agents

Tadashi Shigeoka · Wed, September 30, 2026

After I started working with trycua/cua, an open-source Computer Use project, I kept running into the word “CUA.” It is in the repository name and in package names, so I vaguely took it as some Computer Use–related name and kept using it without checking what it stood for.

When I finally went looking for something that explained CUA, I landed on OpenAI’s announcement “Computer-Using Agent” and realized that CUA is short for Computer-Using Agent. I suspect I’m not the only one who has been using the acronym without knowing its origin, so here is what I confirmed, organized as a short glossary.

CUA Stands for Computer-Using Agent

On January 23, 2025, OpenAI announced a research preview of Operator, an agent that performs tasks on the web on the user’s behalf. The model announced alongside it, as the one powering Operator, was Computer-Using Agent (CUA) (Computer-Using Agent).

Summarizing the announcement, CUA is a model that:

  • Combines GPT-4o’s vision capabilities with reasoning trained through reinforcement learning
  • Is trained to interact with GUIs (buttons, menus, text fields on screen) the way a person does
  • Performs tasks through a shared interface of screen, mouse, and keyboard, without OS- or web-specific APIs

In other words, CUA originally names a specific OpenAI model. At launch, it was available to ChatGPT Pro users in the US through Operator, and the announcement said OpenAI planned to make it available to developers via the API. Operator’s announcement page has since been updated to say that, on July 17, 2025, Operator was fully integrated into ChatGPT as ChatGPT agent (Introducing Operator).

How CUA Works: A Perception, Reasoning, and Action Loop

According to the announcement, once CUA receives a user’s instruction, it works through an iterative loop of perception, reasoning, and action:

  • Perception: a screenshot of the computer is added to the model’s context, giving it a snapshot of the current screen state
  • Reasoning: it uses chain-of-thought over current and past screenshots and actions to decide the next step
  • Action: it clicks, scrolls, types, and so on, until it decides the task is done or user input is needed

For sensitive actions such as entering login details or answering a CAPTCHA, it asks the user for confirmation.

flowchart LR
  U["User instruction"] --> P["Perception<br/>take a screenshot"]
  P --> R["Reasoning<br/>decide the next step"]
  R --> A["Action<br/>click, scroll, type"]
  A -->|"screen changes"| P
  R -->|"done, or confirmation needed"| E["Return to user"]

Because it operates by looking at screen pixels rather than calling APIs, the same approach applies to apps and websites that have no dedicated API.

Benchmarks at Launch

The announcement reports the following results. The numbers are as of the January 2025 launch.

BenchmarkWhat it measuresCUAPrevious SOTA (same universal interface)Previous SOTA (web browsing agents)Human
OSWorldOperating a full OS such as Ubuntu, Windows, or macOS38.1%22.0%-72.4%
WebArenaTasks on self-hosted websites58.1%36.2%57.1%78.2%
WebVoyagerTasks on live sites such as Amazon, GitHub, and Google Maps87.0%56.0%87.0%-

On OSWorld, CUA scored 38.1% against a human baseline of 72.4%, and the announcement itself acknowledges that full-OS tasks still have a large gap to close.

Once you know what CUA means, the surrounding terms are easier to tell apart. This post uses them as follows:

TermWhat it refers toExamples
CUA (Computer-Using Agent)The name of OpenAI’s GUI-operating modelThe model behind Operator
Computer UseThe name of a feature or tool where a model looks at the screen and operates the mouse and keyboardAnthropic’s computer use, OpenAI API’s Computer use
computer-use agentA generic noun for any agent that completes tasks by operating a screenOpen-source agents and frameworks in general

Computer Use is also the name of the feature Anthropic announced in October 2024 alongside Claude 3.5 Sonnet. OpenAI’s CUA announcement lists that announcement among its references.

On the OpenAI side, the current API documentation uses “Computer use” as the feature name. At the same time, the sample app linked from those docs lives in a repository named openai/openai-cua-sample-app, whose description reads “CUA (our Computer Using Agent),” so the CUA name is still around.

trycua/cua, for its part, describes its framework as “Computer-Use Agent (Cua)” in places such as the libs/python/som README (libs/python/som/README.md). That is one word off from OpenAI’s “Computer-Using Agent,” yet both abbreviate to CUA. trycua’s blog post A Story of Computer-Use introduces Operator as “powered by their Computer-Using Agent (CUA) model,” which shows that trycua itself refers to CUA as OpenAI’s model name.

Same Name, Not Necessarily the Same Implementation

Even though both are “CUA,” OpenAI’s CUA as a model name and Cua as the name of the trycua/cua project were given by different organizations. A matching name alone does not tell you whether they share technology or implementation.

The follow-up post, “Codex Drove My Mac and “cua” Showed Up — Tracing ChatGPT.app’s Bundled cua_repl and How It Relates to the OSS cua-driver,” traces the mcp__cua_repl tool that Codex’s Computer Use called through the files bundled in ChatGPT.app, and examines how it relates to trycua’s cua-driver.

That’s all from learning that the CUA I kept seeing in trycua/cua stands for OpenAI’s Computer-Using Agent and sorting it out from Computer Use and computer-use agents, from the Gemba.

References