Search

Spoke at Yurutech Beer Bash LT 2026.10 — Letting AI Drive a Mac with Computer Use

Tadashi Shigeoka · Fri, October 2, 2026

I gave a lightning talk at Yurutech Beer Bash LT 2026.10. My talk, “Letting AI Drive a Mac with Computer Use,” packed into five minutes the core loop behind an AI that watches the screen and operates apps, together with a live demo where Codex builds a cheers slide in Keynote.

The content reworks three earlier posts on this blog into a five-minute beer-bash LT: Notes on choosing a Computer Use agent for macOS and Windows, Notes on handing QA testing to an AI agent, and What Does CUA Stand For? — Sorting Out Computer-Using Agent, Computer Use, and Computer-Use Agents.

Event Overview

  • Name: Yurutech Beer Bash LT 2026.10
  • Date: Friday, October 2, 2026, 19:00–21:00 JST
  • Location: Kumamoto City, Japan
  • Format: Beer-bash-style lightning talks

My Talk: “Letting AI Drive a Mac with Computer Use”

The talk introduced Computer Use, which I use regularly in Codex, with a beer-bash-appropriate motif: asking the AI to prepare the toast.

What Is Computer Use

Computer Use is a mechanism where an AI watches the screen and operates an app. Clicks, text input, and scrolling, the things we usually do with a mouse and keyboard, get handed to the AI. When you describe the goal in natural language, such as “build a cheers slide in Keynote,” the AI drives the actual app to get there.

The Basic Loop: See, Act, Verify

Underneath, Computer Use is a repetition of three steps.

StepWhat happens
SeePull a screenshot or UI-element information to understand the current screen state
ActPerform a click, text input, or scroll
VerifyObserve the screen again and decide the next move

The observation surface varies by implementation; some setups add UI metadata on top of screenshots. The step that matters most is verify: checking the screen again after each action. I come back to it with a concrete example later in “After You Act, Check the Result.”

flowchart LR
  U["User instruction"] --> P["See<br/>Capture screenshot"]
  P --> R["Act<br/>Click, type, scroll"]
  R --> V["Verify<br/>Observe the screen again"]
  V -->|"More to do"| P
  V -->|"Done or need confirmation"| E["Hand back to user"]

Pick Based on What It Can Touch

Not every Computer Use touches the same surface. Some stay inside the browser; others go all the way into OS-level apps.

ScopeCan operate
In-browserWeb pages, web apps
OS-level appsNative apps like Keynote and Finder

The pick depends on what you actually want to do: open a web page, or drive a desktop app like Keynote. For this demo, I used a Codex setup that can operate macOS native apps. For a per-product comparison, see my earlier post Notes on choosing a Computer Use agent for macOS and Windows.

A Dev Use Case: Exploratory QA

The use case I find most interesting on the development side is exploratory QA. You hand the agent an open-ended goal like “play with this new screen and tell me where you got stuck,” and let it explore.

RoleWhat it is good at
Computer UseTake a goal and explore an unfamiliar screen
Existing E2ERepeatedly verify a fixed set of conditions

AI judgment is not deterministic: the same screen can be flagged as a bug on one run and pass on another. So suspected bugs come back with reproduction steps and screens attached, and a human verifies. Existing E2E keeps its job of verifying fixed conditions on repeat; Computer Use takes the role of touching unfamiliar screens. The design discussion is written up in Notes on handing QA testing to an AI agent.

After You Act, Check the Result

The failure that is easiest to miss with Computer Use is “the input happened, but the screen does not add up.” Even when text makes it into Keynote, the sentence can be long enough to overflow the box. Do not end at “input completed”; include “look back at the screen” as part of the completion condition.

For this demo, I made the completion condition explicit with three checks: exactly one slide, the text “みなさん、乾杯!🍻” entered as is, and no text overflowing its box. Say up front what “done” looks like.

Demo: Asking Codex to Prepare the Toast

The main-content closer was a demo where Codex builds a single Keynote slide reading “みなさん、乾杯!🍻” (“Cheers, everyone! 🍻” in Japanese). The two things I asked people to watch for were that Japanese text gets typed correctly, and that the final slide gets re-inspected after being produced.

During the demo, I let the see / act / see-again loop play out live. I had a backup screenshot of the finished slide, captured earlier on the same Mac, in case we ran out of time, but the live run finished within the LT window and the backup stayed in the deck unused.

Wrap-Up

  • Computer Use is a loop: see, act, verify, then repeat
  • Pick in-browser or OS-level based on what you actually need to touch
  • After acting, confirm the result from the screen itself
  • On the development side, start with exploratory QA layered on top of existing E2E

The closing line I ended on: “Every operation we do on a PC, now handed off to AI.” Start with the small, everyday operations you keep repeating.

I Joined the After-Party Too

After the main program, I joined the after-party organized by willing attendees. When you normally work fully remote, meeting fellow engineers in person (not through a screen) is rare, and we were able to dig into operational topics: how each company is folding Computer Use and AI agents into their work, where they draw the line and hand control back to humans. The topics I did not have time to cover on stage got their own slow discussion over drinks.

Thank you to the organizers and to everyone who attended.

That’s all from speaking at Yurutech Beer Bash LT 2026.10 and running the Codex-prepares-the-toast demo, from the Gemba.

References