โ† Back to blogยทTrendsยท6 min read

Google Moved Its Screen-Clicking Agent Into Production, OpenAI Locked Its Sharpest Security Model Behind a Vetting Tier, and a Chinese Lab Gave Away the Defensive Version

Gemini 3.6 Flash shipped with Computer Use as a native, production-ready tool for controlling browsers and desktops, OpenAI expanded Daybreak with a gated GPT-5.6-Cyber tier that already found a real Chrome V8 vulnerability, and Z.ai released GLM-5.3, an open-weight model built for the same defensive-security agent workloads with no waitlist at all. Three different bets on who should get to build agents that act, not just chat.

By Maya Brennan ยท Writer, Smillee AI
August 22, 2026

Three releases from the last two weeks aren't about which chatbot answers trivia best. They're about how much a model is allowed to do once it's not just replying in a text box โ€” click through a browser, patch a vulnerability, or run exploit code โ€” and who gets to decide that. Here's what shipped, and what it says about where agent access is heading.

1. Gemini's Computer Use Tool Graduates From Demo to Default

Google's Computer Use tool โ€” the one that lets a model look at a screenshot and decide where to click, type, or scroll inside a real browser or desktop session โ€” launched as a public preview back in June bolted onto Gemini 3.5 Flash. With Gemini 3.6 Flash now generally available and priced below its predecessor, Computer Use ships as a native, built-in capability alongside thinking, structured outputs, and function calling, rather than a separate preview endpoint developers had to opt into. It's also wired into the Gemini Enterprise Agent Platform's model picker, meaning teams building internal automation can point an Agent Designer workflow at a model that fills out forms, navigates internal tools, and completes multi-step UI tasks without a specialized browser-automation library sitting in between.

The shift that matters for anyone shipping a product isn't the benchmark bump โ€” it's that "an agent that operates software the way a person does" just moved from research-preview novelty to a line item in a production model card. If your roadmap has a task like "log into this vendor portal and reconcile the invoice," the honest question this month is no longer whether a model can attempt that, but whether your team has built the guardrails โ€” sandboxed sessions, action logging, a human checkpoint before anything irreversible โ€” to let it.

2. OpenAI Puts Its Best Hacking Model Behind a Waitlist

OpenAI expanded its Daybreak cybersecurity program on August 10 with GPT-5.6-Cyber, a variant of GPT-5.6 Sol trained to stop refusing dual-use security requests โ€” exploit development, vulnerability chaining, offensive tooling โ€” that the general-purpose model declines by design. On OpenAI's own Advanced Cybersecurity Completion Rate eval, GPT-5.6-Cyber completes 95% of such requests versus 1.5% for the standard model. Access isn't open: Daybreak Blue gives vetted partners the base model without its usual system-level guardrails, while Daybreak Red โ€” the tier that unlocks GPT-5.6-Cyber itself โ€” is reserved for organizations validating exploits and doing advanced vulnerability research, with Accenture, IBM, CrowdStrike, Cisco, Sophos, and Cloudflare named as initial partners. OpenAI says the model has already earned its keep, surfacing two previously unknown vulnerabilities in V8, Chrome's JavaScript engine, since patched as CVE-2026-15903.

The design is a bet that the safest way to ship a model this capable is to not ship it broadly at all โ€” gate the capability behind an approved-partner list instead of a refusal trained into the weights. It's a coherent answer to the "how do we let defenders use offensive-grade tooling without arming everyone else" problem, but it also means the sharpest security assistance is only as available as your procurement relationship with OpenAI.

3. Z.ai Skips the Waitlist Entirely

Four days later, Chinese lab Z.ai released GLM-5.3, an open-weight successor to June's GLM-5.2 aimed explicitly at coding agents, tool use, automation, and โ€” per Z.ai's own framing โ€” defensive security work. There's no partner tier, no completion-rate gate, no vetting process: the weights are downloadable under an unrestricted license, and anyone with the hardware to run a large mixture-of-experts model can point it at the same class of problem OpenAI is metering through Daybreak Red.

That's not a like-for-like substitute โ€” an open model trained for defensive triage and patching is a different tool than one OpenAI is explicitly measuring on offensive completion rates. But it's the same underlying tension the industry keeps landing on: as security-capable agents become table stakes, one lab's answer is tighter access control, and another's is to remove the gate altogether. Whichever your organization ends up depending on, the resourcing question is now less "can we afford the API calls" and more "can we afford the compliance overhead of getting approved," or the infrastructure to self-host something that skips that step.

The Common Thread

All three stories are the same question asked three different ways: once a model can act โ€” click through software, or write and validate an exploit โ€” who decides what it's allowed to touch, and how do you prove it to an auditor, a partner, or your own security team? Google answered by folding action-taking into a general-purpose model with no separate approval flow. OpenAI answered by keeping its most capable variant behind a partner list. Z.ai answered by removing the list. None of those is obviously wrong, but they lead to very different deployment checklists โ€” and if your team is scoping an agent project this quarter, "which of these three postures does our compliance story actually support" is a more useful planning question than which model wins the next leaderboard.

โ€” Maya

Frequently asked questions

What is Computer Use in Gemini 3.6 Flash and how is it different from the earlier preview?

Computer Use is a built-in Gemini tool that lets the model view screenshots of a browser or desktop and decide where to click, type, or scroll to complete a task. It launched as a public preview attached to Gemini 3.5 Flash in June 2026; with Gemini 3.6 Flash generally available as of the following month, Computer Use ships as a native capability alongside thinking, structured outputs, and function calling, and is selectable directly in the Gemini Enterprise Agent Platform's model picker rather than requiring a separate preview integration.

What is OpenAI's Daybreak Cyber program and who can access GPT-5.6-Cyber?

Daybreak is OpenAI's cybersecurity access program, expanded on August 10, 2026 with GPT-5.6-Cyber, a variant of GPT-5.6 Sol trained to reduce refusals on dual-use security work like exploit development. It has two tiers: Daybreak Blue gives vetted partners the base model without its system-level cyber guardrails, and Daybreak Red โ€” reserved for organizations doing exploit validation and advanced vulnerability research โ€” unlocks GPT-5.6-Cyber itself. Initial partners include Accenture, IBM, CrowdStrike, Cisco, Sophos, and Cloudflare; the model has already been credited with finding a real Chrome V8 vulnerability, CVE-2026-15903.

How does Z.ai's GLM-5.3 compare to OpenAI's gated Daybreak Cyber models?

GLM-5.3, released by Chinese lab Z.ai on August 14, 2026, is an open-weight model aimed at coding agents, tool use, automation, and defensive security work, available under an unrestricted license with no partner vetting or waitlist. That's a different starting point than OpenAI's Daybreak Cyber tiers, which gate a completion-rate-boosted, offensive-capable model behind an approved-partner program โ€” but both are responses to the same trend: security-capable agents becoming standard infrastructure, with labs choosing either tight access control or open availability as the safeguard.

Maya Brennan
Writer, Smillee AI

I'm Maya โ€” I write most of what you'll read here. I spent years as a copywriter before I got a little obsessed with what these AI tools can actually do, so now I spend my days poking at chatbots, breaking them, and writing up what's worth your time. Everything here is something I've actually tried. If a prompt didn't work for me, it doesn't make the cut.

Want to try any of this?

Smillee's free and there's no signup โ€” open it and paste in whatever you're working on.

Start chatting โ†’

More from the blog