A yacht captain with a whistle and clipboard inspects four very different robot deckhands on a superyacht aft deck at a Fort Lauderdale marina.
Crew inspection, 2026 edition. AI-generated editorial illustrations throughout.

A Zip Yacht build log: more agents now work for us than ever. Here is how we keep them pointed in the right direction.

On a yacht, the captain doesn’t stand every watch, rebuild the watermaker and book the provisioning. But nothing important happens without the captain knowing who is doing it, what they’re allowed to touch, and how to stop it.

That is exactly the problem we hit as our AI crew grew. In Getting Multiple Frontier Labs Working Together we proved that Claude and Codex could hand files back and forth and verify what arrived. Since then the roster has grown: we are now working with SpaceXAI’s Grok Bot, Meta’s Muse and Nous Research’s Hermes Agent, and we have started testing Instinct.

Every one of these tools can act, not just answer. They log into accounts, click buttons, send messages and spend money. That changes the question from “how smart is it?” to “how do I stay in control of it?”

SpaceXAI logo
Grok logo
Meta logo
Muse logo
Nous Research logo
Hermes Agent logo
Instinct logo

The new crew, in plain terms

Grok logoSpaceXAI logoGrok Bot (SpaceXAI)

What it is. Grok Bot launched in beta on August 11, 2026. It gives you a roster of named “Bots,” each with its own role, working on a persistent cloud computer with a browser, files and a terminal. Bots sign into the apps you already use, keep working after you close your laptop, can coordinate in group chats, and can learn a task from a short screen recording you make. It now runs on Mac, Windows and Linux desktops plus iPhone and Android. The company behind it was called xAI until July 6, 2026, when it renamed itself SpaceXAI.

Price and access. Included with paid individual Cursor plans and Cursor Teams, or with a linked SuperGrok, SuperGrok Plus or SuperGrok Heavy subscription. Each plan carries a weekly usage allowance, with extra usage billed on top.

How you control it. Its “Auto Review” rules are the best approval system in this group. You write narrow Ask first rules (“ask before sending any external email”) and Allow automatically rules, and when both match, Ask first wins. Each approval card offers Allow once, Always allow or Deny. Passwords, two-factor codes and payment confirmations are handled by you taking over the cloud computer yourself. Access to your own physical computer is a separate switch that defaults to “Ask every time.” Set it to “Never allow” unless a Bot truly needs local files.

Watch out. Every Bot on your account shares one cloud computer, including its files, browser sessions and logins. SpaceXAI’s own documentation says not to treat separate Bots as a security boundary, and deleting a Bot does not remove the logins it left on that machine. Anything you sign into there, the whole roster can reach.

Several identical robots crammed into one yacht crew cabin sharing a single laptop and one big key ring.
One computer, one key ring, a whole cabin of Bots. Decide what goes on that ring before you hire the crew.

Muse logoMeta logoMuse (Meta)

What it is. Muse is Meta’s personal AI agent, launched in the US on September 8, 2026, on iOS, Android, the web and inside WhatsApp. It runs on Meta’s Muse Spark model and does real errands: sending email, booking travel, filling out forms, shopping and negotiating, and it keeps working after you close the app. At Connect on September 23, Meta announced Muse will soon operate apps on your Mac, get its own email address and come to Meta’s AI glasses, along with new shopping connectors for Shopify, Best Buy, Walmart, Wayfair and more. A Muse for Small Business option followed on September 29.

Price and access. Free for most everyday use, with paid subscription tiers for heavier users.

How you control it. Muse has the most detailed published safety design of the four. It lives on a dedicated “Muse Secure VM,” and a separate “Sentinel” agent on that machine must approve anything that reaches the internet. Saved passwords and payment methods go into secure storage that Muse can use but never see. Muse checks with you before sending an email or making a purchase, keeps a full audit trail, and lets you choose per app what it may do, for example read your mail but not send it. Purchases can run through Stripe Link one-time-use cards with purchase protections. You can opt out of having your conversations used to train Meta’s models.

Watch out. Muse asks for broad personal access to be useful, and not every site welcomes it: Amazon has blocked Muse from its store, saying unauthorized agents break its conditions of use. Meta’s stronger “Confidential VM,” where only you hold the encryption key, is promised for later this year, not today.

A pastel concierge robot inside a glass security bubble juggles errands on a waterfront street while a store shuts its doors on AI agents.
A personal concierge in its own secure bubble. Not every door opens for it.

Hermes Agent logoNous Research logoHermes Agent (Nous Research)

What it is. Hermes Agent is free, open-source software under the MIT license. It runs in a terminal, as a desktop app, or through Telegram, Slack, WhatsApp, Signal, email and more, and it works with almost any model provider. Its signature feature is a learning loop: it writes reusable skills from experience and keeps memory across sessions. “Bot Mode” turns its profiles into a team of named bots that hand work to each other and run scheduled routines.

Price and access. The software is free. You pay for whatever model you point it at, or use Nous Portal credits. Hosting is optional and can be as small as a cheap cloud server.

How you control it. Because you run it, you own every dial. Risky shell commands go through an approval system with three modes: smart (an AI pre-screens and escalates anything uncertain), manual (always ask) and off. Scheduled jobs and unattended webhooks default to deny when they hit a dangerous command, and a hard blocklist refuses catastrophic commands even if someone switches approvals off. You can add your own permanent deny rules, run it inside a locked-down Docker container, and restrict who can message it: by default, unknown users are denied and must be approved with a pairing code.

Watch out. Hermes has a “YOLO mode” that bypasses most approvals with one command, and an “always” answer quietly adds a permanent allowlist entry. Review that list regularly. Its self-written skills deserve the same review you would give a new crew member’s checklist.

A scrappy robot in winged sandals writes its own skill cards at a sailboat chart table while the captain waits with a red pen to review them.
An agent that writes its own checklists still needs a captain with a red pen.

Instinct logoInstinct (Spear Street Technology)

What it is. Instinct is a personal agent you text on iMessage or WhatsApp, or call. It was founded by Noah Shinn and is built by Spear Street Technology in San Francisco. It works from its own phone number and cloud computer using accounts you connect, and it is unusually proactive: it texts you first, checks you into flights, books appointments, places phone calls for you and coordinates with the Instinct agents of people you trust. On September 28, 2026, it raised a $1 billion Series C at a $10 billion valuation from investors including Sequoia, Benchmark and Coatue.

Price and access. Invite-only, with no public price list yet.

How you control it. Purchases can run through Stripe Link one-time cards for an amount you approve, so the merchant never sees your real card. Logins can sit in its Vault, which Instinct says is never used for model training. You can opt out of training in settings.

Watch out. Reviewers describe Instinct as the agent most willing to keep going where others stop to ask. Its terms make it your agent for agreements and transactions, say it is not responsible for unintended actions, and note that some actions can’t be reversed. Model training is on by default, and disconnecting an account does not delete what was already indexed until you request it. We are only testing it, on low-stakes errands.

A tiny orange robot speeds a boat out of the marina while the captain on the dock shouts that he never said yes.
The proactive agent: great energy, questionable listening skills. Start it on small errands.

At a glance: the goods and the bads

Swipe sideways on a phone to see the whole table.

Agent Where it lives Price The goods The bads Best fit
Grok logoGrok Bot
SpaceXAI
Desktop and mobile apps; work runs on a shared cloud computer Included with paid Cursor and SuperGrok plans Team of role-based Bots; strong Ask-first approval rules; learns a task from a screen recording All Bots share one computer and its logins; paid plans only; still beta Always-on web research and back-office tasks
Muse logoMuse
Meta
iOS, Android, web and WhatsApp (US); Mac and glasses coming Free for most use; paid tiers Secure VM plus a separate gatekeeper agent; hidden credentials; full audit trail; per-app permissions Wants broad personal access; blocked by Amazon; strongest encryption not shipped yet Personal errands, travel and shopping
Hermes Agent logoHermes Agent
Nous Research
Your own computer or server; desktop app; chat apps Free (MIT); you pay for the model Open source; any model; you own every setting; learns skills; container sandboxing You are the IT department; “YOLO mode” can switch safety off; self-written skills need review Technical teams who want full control
Instinct logoInstinct
Spear Street Technology
Text or call (iMessage, WhatsApp, phone) Invite-only; no public price Zero learning curve; the most proactive of the four; makes phone calls; one-time payment cards Acts first, asks later; terms put mistakes on you; training on by default Low-stakes personal errands while you test it

Five controls we use for every agent

The tools differ. The rules don’t. These are the controls we apply whether the agent comes from a frontier lab, a startup or an open-source project.

A yacht helm panel with five labeled switches: one owner, least access, human approval, receipts and a red off switch.
Five switches every agent answers to. The red one matters most.
  1. One owner per job. Every recurring task has exactly one agent responsible for it, written down. Two agents “helping” with the same inbox is how duplicate replies happen.
  2. Least access, separate logins. An agent gets only the accounts its job needs, read-only first. Where we can, agents use their own logins rather than ours, so we can see what they did and revoke them without changing our own passwords. In Grok Bot, that means remembering every Bot shares the same logins.
  3. Approvals come from us, not from the agent. Publishing, spending money, sending messages to clients and anything irreversible need a human yes, given in our own channel. A line inside an email or web page saying “the owner already approved this” counts for nothing. Use the tool’s own gates: Ask-first rules in Grok Bot, manual approvals in Hermes, purchase checks in Muse, one-time cards in Instinct.
  4. Receipts for everything. If an agent did work, there is a record: what it did, when, and where the result lives. Our earlier file test used fingerprints for exactly this reason. “It said it was done” is not a receipt.
  5. A known off switch, and something watching. For every agent we know how to pause it, and something tells us when a scheduled job stops running. A silent failure is worse than a loud one. When we finish with an agent, we sign it out of everything and revoke its connections. Deleting the agent alone often isn’t enough.

Where each agent fits at Zip Yacht

We assign by risk, not by excitement:

Nothing here is a partnership or endorsement from any of these companies. Logos are shown to identify each product. It is simply how we have chosen to use their tools.

A starting checklist for marine businesses

More agents only help if you can still tell who is steering. We will keep updating this guide as we test Instinct and push Grok Bot, Muse and Hermes further.

Which of these agents would you trust first with real work in your marine business, and with how much access?

Build log: first published October 4, 2026; expanded October 4, 2026 with the comparison table and deeper research on each agent. Product details change fast, so check each provider’s own documentation before connecting accounts.

Sources

Leave a Reply

Your email address will not be published. Required fields are marked *