Spoke at Yurutech Beer Bash LT 2026.10 — Letting AI Drive a Mac with Computer Use
I gave a lightning talk at Yurutech Beer Bash LT 2026.10. My talk, “Letting AI Drive a Mac with Computer Use,” packed into five minutes the core loop behind an AI that watches the screen and operates apps, together with a live demo where Codex builds a cheers slide in Keynote.
The content reworks three earlier posts on this blog into a five-minute beer-bash LT: Notes on choosing a Computer Use agent for macOS and Windows, Notes on handing QA testing to an AI agent, and What Does CUA Stand For? — Sorting Out Computer-Using Agent, Computer Use, and Computer-Use Agents.
Event Overview
- Name: Yurutech Beer Bash LT 2026.10
- Date: Friday, October 2, 2026, 19:00–21:00 JST
- Location: Kumamoto City, Japan
- Format: Beer-bash-style lightning talks
My Talk: “Letting AI Drive a Mac with Computer Use”
The talk introduced Computer Use, which I use regularly in Codex, with a beer-bash-appropriate motif: asking the AI to prepare the toast.
What Is Computer Use
Computer Use is a mechanism where an AI watches the screen and operates an app. Clicks, text input, and scrolling, the things we usually do with a mouse and keyboard, get handed to the AI. When you describe the goal in natural language, such as “build a cheers slide in Keynote,” the AI drives the actual app to get there.
The Basic Loop: See, Act, Verify
Underneath, Computer Use is a repetition of three steps.
| Step | What happens |
|---|---|
| See | Pull a screenshot or UI-element information to understand the current screen state |
| Act | Perform a click, text input, or scroll |
| Verify | Observe the screen again and decide the next move |
The observation surface varies by implementation; some setups add UI metadata on top of screenshots. The step that matters most is verify: checking the screen again after each action. I come back to it with a concrete example later in “After You Act, Check the Result.”
flowchart LR U["User instruction"] --> P["See<br/>Capture screenshot"] P --> R["Act<br/>Click, type, scroll"] R --> V["Verify<br/>Observe the screen again"] V -->|"More to do"| P V -->|"Done or need confirmation"| E["Hand back to user"]
Pick Based on What It Can Touch
Not every Computer Use touches the same surface. Some stay inside the browser; others go all the way into OS-level apps.
| Scope | Can operate |
|---|---|
| In-browser | Web pages, web apps |
| OS-level apps | Native apps like Keynote and Finder |
The pick depends on what you actually want to do: open a web page, or drive a desktop app like Keynote. For this demo, I used a Codex setup that can operate macOS native apps. For a per-product comparison, see my earlier post Notes on choosing a Computer Use agent for macOS and Windows.
A Dev Use Case: Exploratory QA
The use case I find most interesting on the development side is exploratory QA. You hand the agent an open-ended goal like “play with this new screen and tell me where you got stuck,” and let it explore.
| Role | What it is good at |
|---|---|
| Computer Use | Take a goal and explore an unfamiliar screen |
| Existing E2E | Repeatedly verify a fixed set of conditions |
AI judgment is not deterministic: the same screen can be flagged as a bug on one run and pass on another. So suspected bugs come back with reproduction steps and screens attached, and a human verifies. Existing E2E keeps its job of verifying fixed conditions on repeat; Computer Use takes the role of touching unfamiliar screens. The design discussion is written up in Notes on handing QA testing to an AI agent.
After You Act, Check the Result
The failure that is easiest to miss with Computer Use is “the input happened, but the screen does not add up.” Even when text makes it into Keynote, the sentence can be long enough to overflow the box. Do not end at “input completed”; include “look back at the screen” as part of the completion condition.
For this demo, I made the completion condition explicit with three checks: exactly one slide, the text “みなさん、乾杯!🍻” entered as is, and no text overflowing its box. Say up front what “done” looks like.
Demo: Asking Codex to Prepare the Toast
The main-content closer was a demo where Codex builds a single Keynote slide reading “みなさん、乾杯!🍻” (“Cheers, everyone! 🍻” in Japanese). The two things I asked people to watch for were that Japanese text gets typed correctly, and that the final slide gets re-inspected after being produced.
During the demo, I let the see / act / see-again loop play out live. I had a backup screenshot of the finished slide, captured earlier on the same Mac, in case we ran out of time, but the live run finished within the LT window and the backup stayed in the deck unused.
Wrap-Up
- Computer Use is a loop: see, act, verify, then repeat
- Pick in-browser or OS-level based on what you actually need to touch
- After acting, confirm the result from the screen itself
- On the development side, start with exploratory QA layered on top of existing E2E
The closing line I ended on: “Every operation we do on a PC, now handed off to AI.” Start with the small, everyday operations you keep repeating.
I Joined the After-Party Too
After the main program, I joined the after-party organized by willing attendees. When you normally work fully remote, meeting fellow engineers in person (not through a screen) is rare, and we were able to dig into operational topics: how each company is folding Computer Use and AI agents into their work, where they draw the line and hand control back to humans. The topics I did not have time to cover on stage got their own slow discussion over drinks.
Thank you to the organizers and to everyone who attended.
That’s all from speaking at Yurutech Beer Bash LT 2026.10 and running the Codex-prepares-the-toast demo, from the Gemba.
References
- Yurutech Beer Bash LT 2026.10 (TECH PLAY)
- Computer use | OpenAI API
- Computer-Using Agent | OpenAI
- Setting up computer use with Codex | OpenAI
- Notes on choosing a Computer Use agent for macOS and Windows
- Notes on handing QA testing to an AI agent
- What Does CUA Stand For? — Sorting Out Computer-Using Agent, Computer Use, and Computer-Use Agents