ChatGPT Codex Can Act on Your Data for Days Unsupervised
ChatGPT Codex can access your email, files, and third-party apps and run autonomously for days. The consent and liability frameworks haven't caught up.
Written by AI. Samira Barnes

Photo: AI. Kai Hargrove
Here is a useful exercise. Read through the feature list of ChatGPT's Codex and ask, at each step, not "what can this do for me?" but "who is responsible when it goes wrong?"
The answer, currently, is: nobody in particular.
AI creator and YouTuber Matthew Berman recently published a walkthrough of fifteen features for getting the most out of Codex, and it is genuinely useful as a guide to what the system can do. It is also, read through a different lens, a fairly thorough inventory of the consent and liability gaps that agentic AI has opened up — and that no regulator, platform terms of service, or legal framework has yet closed.
The gap is not theoretical. It is architectural.
What the system actually touches
Berman is an enthusiastic guide, and his enthusiasm is grounded in real capability. He describes having Codex browse the internet on his behalf, negotiate with customer service agents, rifle through his email to archive messages he does not want to see, and delete files from his computer — all from a single prompt. "I've used it to get refunds, negotiate with customer service, go through my email and archive emails that I don't need to look at, and so much more," he says. Nine minutes after one prompt, he has a formatted spreadsheet comparing studio cameras.
That is a genuinely impressive demonstration. It is also a description of a system that is accessing live commercial interactions, personal communications, and local file systems on a user's behalf, making consequential decisions — archive this, delete that, send this message — with no meaningful human review of individual actions.
The plugin ecosystem extends the reach further. Berman walks through connecting Codex to Gmail, Google Drive, GitHub, Notion, Linear, and Dropbox. The pitch is frictionless: "if you want it to check your email for you, give it access to Gmail and it'll just know how to use Gmail." That is true. What it elides is the data-access question that sits directly underneath it. When Codex connects to Gmail and acts on email — reading, archiving, potentially sending — what data leaves that inbox, under what terms, and with what retention policy? OpenAI's privacy documentation addresses model training data broadly; it does not provide users with a clear breakdown of how third-party plugin integrations handle the data those plugins expose. Users are extending a consumer AI product access to systems that may themselves contain privileged, confidential, or regulated information. The assumption that this is a pure productivity decision — rather than also a data governance one — is one the current interface actively encourages.
The /goal problem
The most structurally significant feature Berman demonstrates is what he calls /goal — a command that instructs Codex to continue running until it reaches a defined outcome. "I've had ChatGPT agents running for days using goal," Berman says. The framing is appropriately cautious: "so be careful using it." But the caution is practical, not structural. The advice is essentially: don't set a goal without also setting a time limit.
That is sensible power-user guidance. It does not touch the underlying question.
When an autonomous AI agent takes actions on a user's behalf over a multi-day run — browsing the web, editing files, interacting with third-party services, executing code in cloud environments — the question of who bears liability for the consequences of those actions has no settled answer in any jurisdiction I am aware of. The EU AI Act classifies certain agentic systems as high-risk depending on their domain of application, but the framework's enforcement mechanisms are still being operationalized and were primarily designed with enterprise deployment contexts in mind, not consumer subscription products. The FTC has signaled interest in AI-related consumer protection issues, but has not issued guidance specific to agentic AI operating under user direction. There is no established legal principle that cleanly resolves whether a user who initiates a multi-day agent run and then steps away is fully liable for every action that agent takes, or whether the platform bears some portion of that responsibility.
OpenAI's terms of service, like most platform terms, place the weight of accountability on the user. The operator model that governs ChatGPT's API is built around the assumption that someone upstream of the end user has accepted responsibility for appropriate use. In a consumer product, that assumption frays. The user is both operator and end user, and many of them are running agents that are doing things they have not individually reviewed.
Scheduled tasks and the consent architecture
Berman also walks through Codex's scheduled tasks feature — recurring automated runs that execute daily without user initiation. His own setup includes daily checks of production logs, file cleanup suggestions, and automated health checks of a website. He recommends using the lighter, faster model variant for these because "it is nearly free to have these scheduled tasks running every single day."
The cost point is well-taken. The consent architecture is worth pausing on. Scheduled tasks mean the system is initiating actions — including actions that may touch connected third-party accounts — without an active human prompt at the moment of execution. The user has consented once, at setup. Whether that initial consent is sufficiently specific to cover every action the scheduled task subsequently takes is not a question the current interface surfaces. It probably should be.
GDPR's approach to automated processing under Article 22 is instructive here: it draws a distinction between automated decisions and human-reviewed ones, and requires that users be informed when significant decisions are made about them by automated means. The framework is oriented toward decisions about users, not decisions by users via AI proxies — which means the scheduled tasks scenario falls into a regulatory gap. The law was not designed for a world where the automated system is your own agent acting on your behalf without your real-time oversight.
The feature set as a map of unresolved questions
None of this is an argument against using Codex, or against Berman's walkthrough, which is technically accurate and practically useful for power users. The features he describes are real, and many of them are impressive.
The point is that the productivity framing — which is how these features are marketed, how they are discussed in the creator ecosystem, and how Berman presents them — occludes a second layer of questions that users are implicitly answering every time they connect a new plugin, enable a scheduled task, or invoke /goal without a hard time constraint.
Those questions are: What data is being accessed? Under what terms? What happens when the agent takes an action that causes harm — to you, to a third party, to a counterpart in a customer service negotiation who may not know they are talking to an AI? Who is accountable for a decision made during an unsupervised multi-day run?
"Just think of ChatGPT as one giant singular system," Berman says, describing how threads and agents interconnect. That is an accurate description of the architecture. It is also a description that ought to give a careful user — and certainly a regulator — a moment of pause. A giant singular system with access to your email, your files, your cloud repositories, your calendar, your GitHub, running tasks on your behalf while you sleep, governed by terms of service you almost certainly have not read in full, operating in a regulatory environment that has not yet decided who is liable when it acts incorrectly.
Power users who understand the technical substrate are making these tradeoffs consciously. Most users of consumer AI products are not. The question is not whether they should be expected to — it is whether they should have to be.
Samira Barnes covers technology policy, regulation, and digital rights for Buzzrag.
AI Moves Fast. We Keep You Current.
Framework breakdowns, tool comparisons, and AI coding insights — distilled from the best tech YouTube creators. Free, weekly.
More Like This
OpenAI's Workspace Agents: The Governance Question No One Asked
OpenAI's new Workspace Agents automate team workflows—but the real product isn't the AI. It's the permission model enterprises can actually live with.
Anthropic's Self-Improving AI Paper Has a Regulator Problem
Anthropic's new paper on recursive self-improvement reveals an oversight gap that existing AI regulation—EU AI Act, executive orders—was never designed to address.
Anthropic's Subscription Mess Has a Regulator's Name on It
Anthropic's quiet pricing changes and Easter-weekend policy announcements aren't just bad comms—they may meet the FTC's definition of deceptive subscription practices.
Why AI Refactors Code Perfectly But Can't Count R's
Andrej Karpathy explains AI's 'jagged' capabilities: why models excel at coding but fail basic tasks. The answer reshapes how we build software.
Can Anthropic Read Claude's Mind? Sort Of.
Anthropic's new NLA research translates Claude's internal activations into readable text—and what it found raises as many questions as it answers.
Gemini 3.2 Flash, SubQ's 12M Context Window, and Claude's Finance Play
Google, Anthropic, OpenAI, and a startup called SubQ all made significant AI moves this week. Here's what each actually means—and for whom.
RAG·vector embedding
2026-08-06This article is indexed as a 1536-dimensional vector for semantic retrieval. Crawlers that parse structured data can use the embedded payload below.