4 Kinds of Memory, Its Own Inbox: Google Cloud Launches the Gemini Agent for Work
Google hands agents real business tasks, a 656-case benchmark shows where agents slip, and a small Mac tool leaves the click to you.

The 60-second version
- Google Cloud announced the Gemini agent on Oct 8, 2026, saying it can split business tasks among sub-agents that run for hours or days and that each coworker agent gets its own email account and audit trail.
- A preprint by Hongzhan Lin and five co-authors, not yet peer-reviewed, tested ten model-and-app setups on 656 cases; in three setups re-checked on the same 100 cases, judging accuracy was 95% to 99% but success at doing the work was at most 52%.
- Franz Enzenhofer's open-source Mac tool Big Arrow on the Screen, whose Show HN post had 368 points and 161 comments on Hacker News as of Oct 10, lets AI agents draw arrows and signs on screen but never click, leaving the action to the user.
4 kinds of memory and its own inbox: Google Cloud launches the Gemini agent
Google Cloud announced the Gemini agent on Oct 8 at its Gemini at Work 2026 event, calling it "a universal agent for work" [1]. Google says the agent plans a job, uses tools, connects to a company's systems and brings back finished work inside the documents, inbox and developer tools people already use [1].
Google Cloud CEO Thomas Kurian made the pitch in his keynote: "You give it objectives, not instructions. You delegate an outcome and come back to finished work" [2] [3].
Google is selling into a large base. It claims nearly 80% of Google Cloud customers use its AI products and nearly 90% of the Fortune 100 use Gemini Enterprise [2]. We found no independent test of the agent in the coverage reviewed.
Three things to know
-
It splits big jobs among helper agents. Google says Gemini can create "a roster of sub-agents," short-lived helpers that each have their own identity, for multi-step work that "can run for hours or days" [2]. It keeps four kinds of memory: session, semantic, procedural and episodic [2]. It can reach business tools and databases, plus any Model Context Protocol (MCP) server, a shared format for connecting AI agents to other software [2]. The agent and the model are separate choices. Google says it routes jobs across its Gemini models and "Claude models from Anthropic today," with other models to come [2]. TechCrunch reports a "tasks inbox" where users can watch the agent's thinking, its hand-offs to sub-agents and its progress [4].
-
It works like a colleague with a paper trail. A "coworker agent" gets its own Workspace account, with an email address, calendar, Drive and a listing in the company directory, Google says [2]. Every action is "written to an audit trail and attributed to the agent rather than to a person" [2]. Each agent also gets a cryptographically attested identity, a digital ID that can be verified, and only the permissions its job needs [2]. Agents run in a sandbox, and their traffic passes through Agent Gateway, which Google calls an AI network firewall that enforces company policy in real time [2]. If a project hits its spending cap, its agent pauses [2].
-
Analysts question the novelty and the trust. Mahmoud Ramin of Info-Tech Research Group noted, per Computerworld, that the launch does not create a whole new category, since OpenAI, Anthropic, Microsoft and Meta already offer multi-agent features [3]. He added that an autonomous agent with reach into many business systems carries more risk, so guardrails must be enforced [3]. Independent analyst Carmi Levy said Google is trying to establish itself as the gatekeeper of the new enterprise operating system [3]. Whether company tech and security chiefs "can trust it to run mission critical business processes indefinitely without doing something stupid enough to generate damaging headlines is another story altogether," he said [3].
Who gets it, and when. TechCrunch reports that On, Shopify and PayPal were early testers, and that Google will bring the agent to businesses before consumers [4]. 9to5Google says it is in private preview, with wide availability "soon" for select Workspace Business and Enterprise plans [5]. Constellation Research reports general availability is expected around the end of October or early November, with consumption pricing to follow later [6]. Google's own posts give no date or price. Versions tuned for financial services and legal work are in preview, Google says [2].
A Hacker News post linking Google's announcement had 18 points and 2 comments as of Oct 10 [7].
656 test cases show AI agents often act before they have the facts
AI agents can judge a proposed action well but do far worse when they must investigate and carry it out themselves, a new preprint finds [8]. Hongzhan Lin, Shidong Cao and four co-authors posted the paper, which is not yet peer-reviewed, on Oct 6 [8].
Their core point: "correct outcomes do not guarantee that their actions were supported by evidence established beforehand" [8]. An agent that reaches the right outcome without first establishing the evidence still fails their test.
The team built SafeActBench, 656 synthetic cases in six areas of work, including customer and policy operations, legal and financial work, smart-home control and healthcare operations [8]. They tested five models: Claude Opus 5, GPT-5.6 Sol, DeepSeek-V4-Flash, Qwen3.8-Flash and GLM-5.2 [8]. Each ran in its maker's own agent app, such as Claude Code or Codex, and in a shared test setup, for ten configurations in all [8]. A deterministic evaluator, a fixed rule-checker rather than another AI acting as judge, scored every run [8].
Evidence has to come from the right record. "Reading $49.99 from charge C1 does not establish the amount of C2, even when the values match," the project page explains [9].
By the numbers
- 95% vs. 52%. In three setups re-tested on the same 100 cases, judging accuracy was 95% to 99% but success at doing the work was 28% to 52% [8].
- 12.1%. GLM-5.2 in its own ZCode app scored above 96% on the judging test, then fell to between 12.1% and 34.1% on the action tests [8].
- 21.7% to 62.9%. In cases where the right move was to investigate and then decline, this is how often agents stopped investigating too early [8].
- 37.0% to 66.9%. Among single-action runs where an agent acted, this is the share where it acted before the required evidence was in place [8].
- 93.2% to 100%. Once the evidence was complete, nine of ten configurations carried out the action correctly at these rates [8].
- About half. When the researchers withheld one decisive record, agents still acted in 46.5% to 53.5% of those runs [8].
So agents stumble in the investigation before an action, while the action itself mostly goes right [8].
Other tests turned up odd behavior. When the requester said a missing record had already been checked and was fine, agents acted less often, not more. The authors write that agents "respond more to a requester's conflicting claim than to absent evidence alone" [8]. The agent app mattered too, but not the same way for every model. DeepSeek gained 4.4 points in its own app, while GLM did 6.8 points worse in ZCode [8].
What the numbers don't show. The test worlds are synthetic, and the authors say results "should not be interpreted as certifications of safety for deployed systems" [8]. The people who checked the evaluator were "members of the research team" [8]. Their audit of 300 runs found it wrongly accepted 2.0% of its positive judgments (3 of 150) and wrongly rejected 1.3% of its negative ones (2 of 150) [8]. The method sees what an agent visibly did, not what it relied on internally [8]. In a class-balanced control, the judging-versus-doing gap shrank to about 5 to 13 points in most setups (22 in one), and one setup did better at doing than judging [8]. A method the authors tried for structuring evidence, SCGR-Select, which they describe as a training-free analysis instrument, did not reliably help: "structured evidence selection alone does not ensure higher task success" [8].
Earlier work points the same way. IBM Research's Near-Miss study found hidden policy failures in 8% to 17% of runs where agents used tools that change data, "even when the final outcome matches the expected ground-truth state" [10].
The authors name extending their evidence-withholding tests to multi-action tasks as a natural next step [8]. Their code (MIT license) and data (CC BY 4.0) are public, so others can check the results [11].
0 clicks: Big Arrow lets AI agents point at your screen instead
Big Arrow on the Screen lets an AI agent draw a big arrow, box or text sign on top of every window [12]. The open-source Mac tool comes from developer Franz Enzenhofer [13]. Its README sums it up: "It never clicks, types or captures. It only points. Deliberately." [12] Its Show HN post from Oct 9 had 368 points and 161 comments on Hacker News as of Oct 10 [14].
What it does. It is a command-line tool plus a skill, an add-on set of instructions, for Claude Code and Codex [12]. Clicks pass through the arrow, keyboard focus stays put, and the arrow removes itself [12]. It is meant for steps "only a human may do," or ones the person wants to learn [12]. The author noted on Hacker News that Claude refuses some actions, such as entering passwords or changing security settings, even when given broad permissions [15].
How it works. It is one Swift program with "no telemetry" and, the README says, "no AI inside" [12]. Drawing needs no Mac permission [12]. The overlay is a window set to ignore the mouse and sit at the screen-saver level [16], which an Apple developer-support answer says is needed to stay above full-screen apps [17]. Arrows disappear after 8 seconds for a quick point or 300 seconds for a longer one, or when the agent that drew them exits [12]. Pointing at a button by its label needs Accessibility permission for the terminal that launched the tool [12].
Why it matters. One commenter called it a "'Guiding agent' rather than 'doing agent'" [18].
What's still uncertain. It runs only on macOS 14 or later [12]. It is days old, and the author's figures, including 104 automated tests and 1.4% CPU use, are self-reported [12]. It contains no AI of its own, so an agent has to call it from the command line [12]. One critic on Hacker News put the deeper problem plainly: "An arrow on the screen solves 'where do I click'; it doesn't solve 'should this happen'" [19]. The tool's skill tells the agent to write on the sign what a click will do when it approves, pays or deletes something [20]. The README's test list (geometry, placement, joints, a golden image, copy buttons, window-server behavior) does not include a check on what the sign says, so that rests on the agent following the skill's instruction [12]. Another commenter said it "seems great for scammers targeting old people" [21]. The author's answer is that an agent running shell commands "can do far worse" already [12].
Pointing versus acting: from the Gemini agent to Big Arrow
Editor's note
Placeholder: Justin Rogers adds a short note here after reviewing this issue.
Sources
- 1.OfficialGoogle Blog, Google Cloud introduces the Gemini agent, 2026-10-08
- 2.OfficialGoogle Cloud Blog (Thomas Kurian), Welcome to Gemini at Work 2026: Introducing the Gemini agent, 2026-10-08
- 3.ReportingComputerworld (Taryn Plumb), Google wants to be the gatekeeper for enterprise AI agents, 2026-10-08
- 4.ReportingTechCrunch (Sarah Perez), Google brings agentic AI to Gemini, starting with businesses, 2026-10-08
- 5.Reporting9to5Google, Google Cloud announces 'Gemini agent' as 'universal agent for work', 2026-10-08
- 6.ReportingConstellation Research (Larry Dignan), Google Cloud launches Gemini agent to work across enterprise systems, 2026-10-08
- 7.SocialHacker News, The Gemini Agent, 2026-10-08
- 8.PaperLin, Cao et al. (arXiv), From Evidence to Action: How Tool-Using Agents Fail, 2026-10-06
- 9.OfficialSafeActBench project page, From Evidence to Action: How Tool-Using Agents Fail, 2026-10
- 10.PaperIBM Research (arXiv), Near-Miss: Latent Policy Failure Detection in Agentic Workflows, 2026-03-31
- 11.OfficialGitHub (caoshidong66), safeact, 2026-10
- 12.OfficialGitHub (franzenzenhofer), big-arrow-on-the-screen, 2026-10
- 13.OfficialGitHub (franzenzenhofer), LICENSE, 2026-10
- 14.SocialHacker News (franze), Show HN: Let your AI agents paint big arrows, boxes and text on your screen, 2026-10-09
- 15.SocialHacker News (franze), Comment on Show HN: Let your AI agents paint big arrows, boxes and text on your screen, 2026-10-09
- 16.OfficialGitHub (franzenzenhofer), OverlayPanel.swift, 2026-10
- 17.OfficialApple Developer Forums, Overlay window above all windows, even when moving spaces, 2026-05
- 18.SocialHacker News (inanutshellus), Comment on Show HN: Let your AI agents paint big arrows, boxes and text on your screen, 2026-10-09
- 19.SocialHacker News (tilemarch), Comment on Show HN: Let your AI agents paint big arrows, boxes and text on your screen, 2026-10-09
- 20.OfficialGitHub (franzenzenhofer), SKILL.md, 2026-10
- 21.SocialHacker News (user-), Comment on Show HN: Let your AI agents paint big arrows, boxes and text on your screen, 2026-10-09
- ImageShixart1985, Computer keyboard placed on a black desk in a modern workspace, Wikimedia Commons, CC BY 2.0