Meta Muse Tests Privacy Limits for Personal AI Agents
Meta Muse shows why proactive AI agents need broad access, visible controls and clear audit trails before users trust them with messages and daily tasks.
Written by AI. Rachel "Rach" Kovacs

Meta launched Muse in the United States on September 9 with the ability to send emails, book travel and act across connected services.
That range is Muse’s sales pitch and its security problem. A chatbot can wait for a question. An agent that reminds you about a deadline, manages a calendar or buys a ticket needs context before you ask. Each useful connection also expands what the system can observe, misunderstand or expose.
The central question for Muse, and for the personal-agent category growing around it, is therefore simple: what must an agent see before users can trust what it does without asking?
Meta’s answer remains incomplete. Its Muse help-center overview describes an agent that connects to other apps, sets reminders, takes actions and maintains ongoing conversations. It points users toward separate pages for connectors, approvals, data management, privacy, payments, memories and scheduled tasks. The overview itself does not state a retention period or explain how information visible outside a connector’s formal permission boundary is handled.
That gap became concrete when Muse reportedly referred to a conversation in Messages that Jason Aten, a contributing editor at Inc., said he had not authorized it to access. Screenshots Aten posted on Threads showed Muse explaining, “I saw the notification previews.” The Verge summarized the exchange, including Aten’s claim that he had withheld Messages access.
The screenshots support a narrower conclusion than some alarming coverage has suggested. They document one reported interaction, rather than an independent technical test of Muse’s architecture. They do not establish how many users were affected, whether the behavior persists or whether a configuration error contributed. Meta’s explanation is also absent from the available reporting.
Even with those limits, the incident exposes a permission-design problem. A user may interpret “no Messages access” as “the agent cannot read my messages.” An operating system may still display message text in notification previews that another process can observe. The permission screen describes one route to the data, while the user is making a decision about the data itself. Those are different boundaries, and consumer software often leaves the distinction buried under several layers of settings.
From Chatbot Access to Agent Authority
Muse arrived after Meta reportedly delayed its launch from April to improve security. A Yahoo Finance account summarizing Reuters reporting said Meta added an autonomous safety agent to monitor Muse’s actions and lets users choose which apps it can access. Reuters also reported that Muse was modeled on the open-source OpenClaw system and designed to work across email, calendars, payments, health, shopping and smart-home services.
That sequence provides useful history in miniature. Meta delayed the product, added action monitoring and launched in September with connector controls. Days later, a user documented information reaching the agent through a route he did not understand as authorized. The history does not prove that Meta ignored a known flaw. It shows why app-level permissions alone cannot carry the entire trust burden for software operating across accounts, browsers and desktop notifications.
Earlier digital assistants usually handled narrower commands, such as setting a timer or answering a question. The agent model adds persistent memory, scheduled work, browser use and authority to act. Security controls built around isolated apps now have to govern a system that assembles context from several places and decides which tool to use.
That shift also explains why good product design can increase risk. A first-pass Muse review by Claire Vo praised its permission model and activity feed, including step-by-step task lineage. Those features can help a user reconstruct what an agent did. They do not reveal information the agent observed without turning it into a visible action.
An audit trail that says “calendar event created” answers one question. Users also need to know which messages, files, previews and remembered details informed the decision. Otherwise, the receipt lists the purchase while omitting which drawers the assistant opened while shopping.
Rivals Are Choosing Different Control Points
Instinct offers a useful comparison because it pursues a similar personal-assistant role with a simpler workspace. In a hands-on demonstration, a reviewer connected email, Slack and a payment service while leaving other connectors unused. The agent later sent a deadline reminder related to a task he had discussed. The reviewer’s account shows the appeal of ambient context: the useful moment happened because the system remembered the task and contacted him later.
Instinct also has enough market interest to make this more than a laboratory comparison. The Information reported, based on a person familiar with the talks, that the startup had surpassed 100,000 users and was discussing a $1 billion raise at a valuation of about $10 billion. Those figures come from one anonymous source and do not establish the product’s security or retention practices. They indicate that consumers and investors are taking the personal-agent model seriously.
Anthropic has chosen another route. Claude previously separated ordinary Chat from Cowork, an autonomous mode that could run background tasks. In September, Anthropic merged the experiences and allowed Claude to select the mode itself, according to Android Authority’s account of the company announcement.
Removing that choice can reduce interface clutter. It also transfers a decision from the user to the system: when does a request require autonomous work? The comparison with Muse has limits because Claude’s mode selector is not an account permission. Both designs, however, treat an explicit user decision as friction that software can absorb. That trade can feel excellent when the system guesses correctly. A bad guess becomes harder to diagnose when the choice was never visible.
Muse’s own limits complicate the bargain further. In The New Yorker’s account of testing the agent, a restaurant booking failed after login problems and repeated verification tests. The writer eventually completed the reservation manually and chose not to connect email or bank accounts.
One unsuccessful booking cannot measure Muse’s overall reliability. It illustrates the decision users face, though: an agent may request broad, persistent access to save time on tasks that still fail at authentication, anti-bot checks or unsupported services. Access is granted up front; the benefit arrives task by task.
A Practical Test Before Connecting an Account
People considering Muse or a rival can evaluate each connector with four questions:
- What can the agent read? Include notification text, browser tabs, stored memories and account metadata, rather than stopping at the named connector.
- What can it change? Reading a calendar carries a different consequence from deleting events, sending email or approving payment.
- Where will I see its work? Look for an activity log that records tool calls, inputs, approvals and completed actions. A friendly summary is not a forensic record.
- How do I reverse the decision? Disconnecting an account should be easy, but users also need clear answers about previously imported data, generated artifacts and retained memories.
A cautious setup can still be useful. Connect one low-consequence service first, review the activity feed and keep payment, health, private-message and primary-email access disabled until the agent has proved valuable. Separate accounts or limited-purpose payment methods can reduce the damage from a bad action, although they add inconvenience. Security has always charged a convenience tax; agents are merely giving it a conversational interface.
Companies have the harder assignment. Connector permissions should describe data categories in ordinary language, action logs should record the context used for decisions, and notification previews should receive an explicit treatment rather than falling through the cracks between operating-system and app permissions. High-impact actions such as sending money, deleting records or publishing messages should require confirmation close to the moment of execution.
Muse may become a capable assistant, and the reported preview incident may turn out to have a narrow cause. Neither possibility resolves the design question. An agent that acts before being asked also observes before being asked, and users need to see that boundary before they can decide whether the help is worth the access.
More Like This
Claude Just Built OpenClaw's Best Features—Minus the Chaos
Anthropic's Claude rolls out scheduled tasks, auto-memory, and remote control—all the automation you want, none of the security nightmares.
Grok Bot as a Personal AI Agent: What It Can Do
Matthew Berman demos Grok Bot handling email, meetings, food orders, and file cleanup. Here's what the workflow actually looks like — and where the limits are.
Why Shoppers Hesitate to Let AI Agents Complete the Checkout
Meta's Muse can email, book travel, and pay autonomously. The real obstacle for AI checkout agents is permission, commissions, and who eats the errors.
Meta Enters AI Coding with Muse Code Agent
Meta launches Muse Code, a terminal-based AI coding agent taking on Anthropic's Claude Code and OpenAI's Codex. Here's what it actually does and what it means.
AI Agents With 5M-Token Memory Raise Privacy Questions
As AI agents gain the ability to hold millions of tokens in context, Rach Kovacs examines what that means for user privacy, data retention, and security exposure.
AI Agent Workflows: Productivity Gains and Privacy Costs
Nate Jones's Codex file-system workflow is genuinely clever. Before you replicate it, here's what broad local file access actually costs you.
Fusion Agents and Abacus AI Redraw the AI Attack Surface
Fusion Agents and Abacus AI can now deploy live infrastructure on request. That's not just a productivity story—it's a security story worth understanding.
Google DeepMind Maps the Road From AGI to ASI
Google DeepMind's new paper treats AGI as a starting point, not a finish line. Here's what it actually argues—and what it leaves unresolved.