Skip to main content
← /writing
  • #genai
  • #productivity
  • #side-projects

The Return of the Personal Agent

I've used five personal agents to help with the work of everyday life. A restaurant call, a changed appointment, and a playoff-ticket search show what I can hand over, what still needs my decision, and how I know the job is finished.

Vinny Carpenter14 min read2.7k words

never the stack · audio edition

The Return of the Personal Agent

20:01

High Stakes by Bartolotta opened on Water Street on September 17, and I wanted a table for three. The earliest OpenTable could offer on any night was 9:30 p.m., so I handed the job to Fo. I haven't called the restaurant once. Fo has.

Fo is the personal agent from Wajo. A couple of calls after I handed it the request, it sent me an 80-second recording. It introduces itself, says the line is recorded, and asks for Wednesday, October 7, around 6:30. The hostess has a line of people in front of her and asks it to call back after 8. Fo says it will, and asks what time works best for her.

The exchange is ordinary. That's why it caught my attention. Someone has to remember to make the next call, and I had handed that responsibility to a service I could reach in Messages.

Instinct is handling a quieter responsibility. It is watching my wife's MyChart portal. When her physical therapy schedule changed, it caught the change and tracked it without her having to do anything. There was no new request at the moment the schedule changed. The useful work was remembering to look.

In another exchange, Fo found three seats for the Brewers' NLDS opener and stopped to ask before buying them. Muse found two subscriptions I'd forgotten I had. Roci, the agent I built on OpenClaw in the spring, is still answering my Telegram messages from an EC2 instance.

These are small pieces of everyday work, but they're also the pieces that keep coming back into my head. Did I check the appointment? Did I call the restaurant? What were the tickets going for? A personal agent earns its place by taking on that follow-up reliably, with clear limits on the decisions it can make.

A year ago none of these six existed as products. OpenClaw went viral in January as an open-source project you ran on your own machine. Since then Meta shipped Muse on September 8, Wajo opened Fo to the public on September 28, and OpenAI announced dots the next day. Instinct is still invite-only at a $2.5 billion valuation. The personal agent came back as something you sign up for.

I've used five of the six agents in this post. OpenAI announced dots on September 29 and began a staged rollout; that's the one I'm assessing from its documentation. This is a report from using the others, with the uneven access, histories, and configurations that implies. It isn't a controlled test of their models.

An agent where I already look

I open Messages without deciding to. I open apps on purpose, and I forget most of them.

That makes Instinct and Fo feel different in my day. A request, a question about budget, and a recording of a call can arrive in a conversation I'm already checking. Instinct's public description says you can text or call it, and that it can follow up on dropped threads. My experience with it is through iMessage. Fo also works through messages, email, group chats, and calls.

Messaging isn't exclusive to those two. Muse works in its app and WhatsApp. OpenClaw and Hermes support messaging channels too. Dots work in ChatGPT, Slack, and Teams; although the launch post listed texting as coming soon, the current Help Center describes a limited texting beta for US Pro users.

The distinction for me is how naturally a task stays in view. An agent can continue working after I close its app. Getting its question or result into a place I check makes that work easier to use.

There is another distinction worth keeping visible. Wajo says Fo can bring in trained executive assistants when a task needs human help. I can describe what the service did for me. I can't tell you that every step was performed entirely by AI. Delegating a schedule or an inbox to an isolated virtual machine is one security posture; routing a task through an operational loop backed by human contractors is another.

Six agents, two separate questions

Where I talk to an agent and who operates it are different questions. Muse is hosted and can live in WhatsApp. Roci is self-hosted and lives in Telegram. A familiar thread tells me very little about where the data sits or what can enforce a permission.

They sort into three places to live: your thread, their app, or a box you run. This is the practical map, based on my experience and the public documentation available on September 29.

Where it livesAgentWhere I can reach itExecution & infrastructureCore tradeoffMy use
Your threadInstinctiMessage, WhatsApp, callsHosted on its own cloud computer; ongoing follow-upLives where you already look; controls still filling inHands-on
Your threadFo / WajoiMessage, group chats, email, callsHosted by Wajo; dedicated phone number, email, single-use card; human fallback availableLives where you already look; controls still filling inHands-on, daily
Their appMuse / MetaMuse app, WhatsAppDedicated secure cloud VM, browser, and Sentinel security authorityThe most careful sandbox I've used; you have to go to itHands-on
Their appdots / OpenAIChatGPT, Slack, Teams (limited texting beta)Dedicated cloud computer, browser, and read-only proactive researchMost documented controls; you have to go to itDocumentation only
Your boxOpenClawAny connected chat channel (Telegram for Roci)Self-hosted on your machine or VPS; explicit tool policies and approvalsMemory is a file you own; defaults and safety are on you to turn onHands-on since spring
Your boxHermes AgentTerminal and 20+ messaging platformsLocal machine or VPS; inspectable skills, scheduled tasks, and local filesMemory is a file you own; defaults and safety are on you to turn onHands-on, local

Six agents, three places to live: two live in your thread, two in their own app, and two on a box you run

Roci's EC2 instance costs me about $25 a month. That covers the instance, before model usage. Running it myself also means maintaining it. A hosted product takes much of that work away, while making me dependent on its access rules, service limits, and continued availability. Hosted services that spin up a browser container and place phone calls for free are being subsidized for now, by investors or by a parent company. Running your own box shows you the operational bill.

Those tradeoffs matter more to this comparison than a launch-week ranking. Fo has finished more of what I've handed it than Roci has, but their tools and operating arrangements differ. I wouldn't turn that experience into a model benchmark.

I gave them the same expectations

I keep a short charter that I paste into the agents I use. Fo, Instinct, Muse, and Roci got the same description: chief of staff for my day, inbox, calendar, and tasks; research and editorial partner; technical collaborator; project and attention manager. (Hermes ran locally with its default system prompt, so I could evaluate its skill execution and migration tools.)

The rules are the useful part. Turn ideas, information, and unfinished work into outcomes. Leave consequential decisions to me. Challenge weak assumptions, separate evidence from speculation, and be concise.

That gives me consistent expectations. It doesn't equalize their tools, memory, permissions, or available human help. It does make the gaps easier to see. Can this agent take responsibility for the task? Where do I still have to intervene? What does it bring back as evidence?

Fo asked for my ticket ceiling and paused before buying. Muse recommended subscription cancellations. Those behaviors fit the charter, but I can't attribute them to the charter alone. Product defaults, permissions, and my specific requests also shaped what happened.

The charter travels well because it is a description of my intent. It is also an incomplete control. Writing that an agent should ask before spending money is useful; I still need to know what prevents it from spending without that approval.

That is the consumer version of the argument I keep making about platforms. Clear intent improves the work. Permissions and review mechanisms determine whether the limits hold when I'm not paying attention.

Identity needs more than a phone number

Fo calls from its own number and emails from its own address. Wajo says it uses a single-use payment card for each purchase and keeps the underlying card details hidden. Those are useful pieces of separation between the service and me.

Other products solve parts of this differently. Muse uses Link from Stripe, including purchase-scoped virtual cards where needed. Dots can use saved passwords on supported websites without exposing them to the model.

A separate address, protected access to my account, and a scoped payment credential do different jobs. A phone number makes the caller reachable. A wallet can limit a purchase. A hidden password still authorizes access to an account that belongs to me. Watching an appointment portal like MyChart still means handing the agent a login or a live session, because the portal offers nothing like the read-only scope a mailbox can grant. None of those, on its own, establishes an independent agent identity with clearly bounded authority.

Four pieces an agent can hold, a phone number, an email address, a single-use card, and a cloud computer, each doing a different job, with Roci as the contrast: accounts provisioned by hand, no phone number, no card. None of them alone establishes a bounded identity.

When I set up Roci, I provisioned separate accounts by hand. It has no phone number and no card. Hosted products are taking on more of that provisioning and integration work, making delegation practical for people who will never configure a gateway.

In The Agent Is Wearing Your Badge, I argued for giving agents authority that can be identified, limited, and revoked. The restaurant and ticket threads bring that question home. I want to know who is making the request, whose account is being used, and what exactly I authorized. The agent's name in my contacts is the beginning of that conversation.

A charter cannot enforce a boundary

The MyChart change is the strongest example here of an agent noticing something without a fresh prompt. Muse's subscription audit was different: I asked it to look. With access to my inbox, it flagged Chegg and Quizlet, two subscriptions nobody in the house was using anymore. That was useful research followed by a recommendation. I had asked it to recommend cancellations, and it left the cancellations to me.

Those are different kinds of initiative. Monitoring an existing responsibility, responding to a new request, and taking an external action shouldn't get bundled into one claim about being proactive.

OpenAI distinguishes research a dot initiates itself from work I assign. The unsolicited research uses read-only tools. Assigned work can continue in the background and take actions subject to authorization and the product's checks. Meta describes Sentinel as a separate permission authority for Muse's connector actions and network access.

What interests me is where the control lives. A boundary enforced outside the agent's own conversation has a different role from an instruction inside it.

The same distinction applies to Roci. Its charter states when I expect it to ask. It cannot guarantee that it will. OpenClaw's own documentation says it plainly: the safety guardrails in the system prompt are advisory, and hard enforcement comes from tool policy, exec approvals, sandboxing, and channel allowlists.

I want an agent that follows up while I'm doing something else. That makes it more important to understand which limits keep holding while I'm doing something else.

The restaurant has a person on the other end

The Brewers exchange shows a useful handoff, even though it doesn't demonstrate a completed purchase. I asked Fo what playoff games were next and what tickets were going for. It asked how many seats I wanted and my ceiling per ticket. When I said three in the Diamond Box, it reported two sets for Game 1: section 117, row 6, at $430 a seat, and section 117, row 8, at $484.

It recommended section 117 and quoted about $1,290 for the three. It identified the listings as SeatGeek and stopped with: "Say go and I'll grab the three."

The game and marketplace details check out: the Brewers' opener is Saturday, October 3, and SeatGeek is MLB's official fan-to-fan marketplace. The prices, available seats, and quoted total are a snapshot of what Fo returned. I would still need to confirm availability, fees, and the purchase details before authorizing a charge.

Fo had taken on the search and brought me a decision. It hadn't proved that it could get the tickets into my account. That distinction is easy to lose when an agent produces a confident summary.

The restaurant is harder because the interface on the other end is a person. Fo is the service I've used to make those calls. Calling isn't unique to it: Instinct advertises calls, and OpenClaw supports a voice-call plugin.

The hostess at High Stakes didn't need an agent. She needed the line in front of her to move. I keep wondering what that host stand looks like when 40 diners have agents, each with a phone number and each willing to call again after 8.

My saved attention can become somebody else's repeated interruption. If I delegate the calls, I also need to define a reasonable retry cadence and a point at which the agent should stop and come back to me. More attempts don't necessarily mean better judgment.

Businesses are already deciding how to handle this traffic. GeekWire reported Amazon blocking Muse in September, citing Amazon's Conditions of Use. Restaurants and retailers have several choices: ordinary booking channels, limits on automated requests, approved agents, or dedicated integrations. They will need to decide whose permission counts and how an agent identifies itself.

That is the same question behind Your Platform Has a New User: The Agent. The agent can only work with the access and context the other side is prepared to provide.

What I want next

Coordination across people is the possibility I'm most curious about. Fo can already participate in group chats. I'd like a household's agents to help find a dinner time without carrying everyone's private context into the negotiation. A proposed time, the relevant constraints, and a clear decision might be enough. I want to read that transcript before I trust it to commit anyone.

Portable memory matters just as much. Hermes offers an OpenClaw migration path for supported settings, memories, and skills. With deployments I operate, I can inspect and back up the files myself. Muse also lets users inspect and download memory files; self-hosting isn't the only route to inspectability. What I haven't found is a supported way to carry my experience with Fo into Instinct, or the other way around.

The thing I want to preserve is what the agent has learned about working with me: preferences, corrections, and the decisions I expect it to bring back. That should survive a change of service.

Evidence is the feature I'd pay for first. The recording lets me hear the call. Dots' Activity View and Muse's visible browser let me inspect work as it happens. But I also want to know whether the promised callback happened, whether the reservation was confirmed, or whether the tickets arrived. A record of activity helps me assess the work; the outcome tells me whether it's finished.

I'm still less sure about how much access people will grant, what these services will cost once the early offers change, and how much supervision a busy household will tolerate. Five agents are interesting to compare. Five agents asking me to manage them would miss the point.

As I write this, three tickets for seats in section 117 are on my phone, and we have dinner at High Stakes next Thursday. I didn't make the calls or search the listings. The decisions were the part I kept, and that's the arrangement I wanted.

// found this useful? share it

Post on X Share to LinkedIn
Vinny Carpenter

Written by Vinny Carpenter

VP Engineering · 30+ years building software

I lead engineering teams building cloud-native platforms at a Fortune 100 company. I write about engineering leadership, AI-assisted development, platform strategy, and the hard lessons that come from shipping at scale.

keep reading