orfloat

ORF-R-2026-005 · Preview Project

A WhatsApp assistant that runs the studio's front desk, one number serving a lead and the founders, that briefs them the moment a booking lands

Can one markdown-conductor agent staff a studio's front desk on WhatsApp, hold real Q&A and a booking conversation with a lead, write to the founders' real calendar behind code rather than prompt hope, and the instant a booking lands turn around and brief the founders, all through a single business number serving two roles?

Finding

Yes, and the proof is felt rather than asserted: the moment a lead confirms, the founders' screen fills with a booking brief, a tappable calendar link, and a ready welcome draft, while every consequential write sits behind a structural gate the model cannot talk past, not a prompt it might forget.

Orfloat’s own front desk is a WhatsApp assistant. It answers a prospect’s questions, runs the booking conversation, and writes to the founders’ real calendar, all through one business number. The recording below is one continuous take of two phones at once: a lead on the left, the studio on the right.

lead

studio

transcript of the recorded demo, word for word

[two phones, recorded at once and played in step. left, the lead's phone in
 light theme, the chat titled "Orfloat Applied-AI". right, the studio's phone
 in dark theme, the chat titled "Orfloat Assistant". the studio's screen stays
 idle through the whole conversation, until the booking lands.]

[lead screen]

lead:   Hello
agent:  Welcome to Orfloat! I'm the studio's assistant, built on Anthropic's
        Claude. What brought you here?
lead:   What does Orfloat do exactly?
agent:  Orfloat is a forward-deployed applied-AI engineering lab. We embed
        inside a business, learn the work end to end, and ship AI systems built
        on Claude. Not consulting from a distance, and no agents without evals.
        Based in Muscat, working with founder-run and family-led businesses
        across Oman and the GCC.
        Want the longer version, or shall I map it to your business?
lead:   Okay, I run a freight forwarding and distribution business. It's ours
        here in Oman. Fleet, warehouses, customs clearance etc, a lot of moving
        parts. I want to understand where AI integration genuinely fits?
agent:  A lot of that lands in what we call internal operations: demand
        forecasting, reorder triggers, scheduling, and daily operational
        briefings that pull from the systems you already run, MCP-first so the
        integration is a contract, not a brittle scrape. Customs paperwork
        intake can fit too.
        Where does it bite most right now: fleet utilisation, warehouse stock,
        or the customs clearance paperwork?
lead:   Coordination mostly is the bottleneck, everything happens on WhatsApp
        and paperwork eats hours every day
agent:  That's the pattern we see most. Coordination living in WhatsApp is
        workable: an agent can sit in those threads, hold your brand voice, and
        turn the back-and-forth into structured actions. Paperwork that eats
        hours is exactly where document intake earns its place. Which of the
        two costs you more is a discovery question, answered on-site, not
        guessed from here.
        If it helps to go deeper, I can set up a 30-minute intro with the
        founders. No obligation. One note: it's currently outside their hours,
        Sun to Thu, 9 to 5 Muscat, so I'd book it for when they're back. Want
        me to find a time?
lead:   Can you give me a ballpark figure for how much this integration will
        cost?
agent:  Honest answer: Orfloat doesn't publish pricing, and I won't invent a
        number, that wouldn't be fair to you. Every engagement is scoped from
        what an on-site discovery finds, so the figure follows the evidence,
        not a template. That's exactly the founders conversation. Shall I get
        that intro on the books?
lead:   Yeah okay book an intro
agent:  Good. So we're framing this around AI for your freight and distribution
        operation, mainly the WhatsApp-based coordination and the paperwork
        load. Anything to add before I set it up?
lead:   This is fine proceed
agent:  What's your name?
lead:   Shaikh
agent:  Thanks Shaikh. And the company?
lead:   Shaikh Enterprises
agent:  What's the best email for the calendar invite?
lead:   niotebook@gmail.com
agent:  How would you like to meet: phone, video, or in person?
lead:   Video call
agent:  Noted, video. I'll put a Google Meet link on the invite. Checking
        founders' availability now, one moment.
agent:  Tue, 16 Jun - 11:00 Muscat
        Wed, 17 Jun - 14:00 Muscat
        Thu, 18 Jun - 10:00 Muscat
        Which suits? Or tell me a window and I'll find something close.
lead:   Tuesday 11am
agent:  To confirm:
        - Meeting: Shaikh / Shaikh Enterprises with the founders
        - Purpose: AI for freight and distribution, WhatsApp coordination and
          paperwork load
        - When: Tuesday, 16 Jun 2026 - 11:00 Muscat
        - Duration: 30 min
        - Modality: Video (Google Meet)
        - Invite to: niotebook@gmail.com
        Shall I lock it in?
lead:   Yes lock it pls
agent:  Booking that in now.
agent:  All set. The invite is on its way to niotebook@gmail.com, with the Meet
        link on it.
        Talk on Tuesday.
lead:   Thank you got it

[the signature moment. the instant the lead confirms, the studio's phone, idle
 until now, fills: a google calendar invitation banner reading "[ORF] Intro:
 Shaikh (Shaikh Enterprises) - Video", then the agent's post-booking brief.]

[studio screen]

agent:  New booking confirmed

        Lead: Shaikh · Shaikh Enterprises
        Purpose: AI for freight and distribution, WhatsApp coordination and
        paperwork load
        When: Tue, 16 Jun · 11:00 Muscat · 30 min
        Modality: Video (Google Meet)
        Contact: niotebook@gmail.com

        Open in Calendar:
        https://www.google.com/calendar/event?eid=...

        Welcome-email draft is in the studio Gmail Drafts.
Figure 1. the recorded demo, two phones recorded at once and played as one: left, the lead’s phone (light); right, the studio’s phone (dark), each named by the bold label beneath it. play, pause, or seek either and both move together, so the signature moment, the instant the lead confirms and the studio’s screen fills, stays in step.

One number, two roles

The number is one WhatsApp business line, and it answers to two people at once: the lead who messages in, and the studio, the two founders, on the other side of the same number. The agent never sees a phone number. A thin channel server resolves every inbound message to a role and a number-free alias before the model reads a word, and the real numbers never leave the server’s environment. The agent routes on the role and replies to the alias, nothing more.

The two roles are not the same conversation. The studio sees its own calendar, titles and all. The lead, when a proposed time turns out to be taken, is told only that there is a conflict, never whose. The boundary is not a request in the prompt; it is where the channel sits, between the people and the model.

The moment the studio runs itself

This is the whole point, and the demo shows it in one beat. The instant the lead says yes, three things land together. The booking becomes a real calendar event, prefixed [ORF and carrying a Google Meet link. A brief arrives on the studio’s phone, the one in the dark frame: who booked, for what, when, and a calendar link the founders can tap. And a welcome email sits written and waiting in the studio’s Gmail drafts, composed but never sent.

The founders did nothing, and the studio briefed itself. That is the feeling the preview is built to produce, and it is why the demo is two screens rather than one: the left phone is the work, the right phone is the result arriving on its own. The calendar event is the source of truth, so a brief that fails to send or a draft that fails to write can never unwind a booking that already happened.

The brain is a markdown program

There is almost no application code. The brain is a local Claude Code session, and no Anthropic API is called: the session is the agent. A single CLAUDE.md conducts it, routing on the role and the intent to the files it needs and reading them only when it needs them, seven flows, five context files, and five files of voice. The program is the markdown.

The code is the channel. Two small servers run beside the session: one speaks WhatsApp’s Cloud API and resolves identity, and one holds the calendar and the inbox. Everything a visitor would call the product’s behaviour, the answers, the judgement, the booking conversation, lives in prose the founders wrote and can read, not in a codebase they would have to trust on faith.

The boring parts, enforced in code

The safety here is not asked for in the prompt, where a model can forget it. It is built into the tools, where it cannot. Every event the agent creates must start with [ORF, enforced in the server and backed by a hook; an event that is not [ORF, one of the founders’ own, the agent cannot move or delete. Before it writes a booking, it re-reads the calendar and re-checks the slot, so a time that went stale between proposal and confirmation is caught at the last moment rather than double-booked. And the Gmail tools are create-draft and delete-draft and nothing else. There is no send. Of the eight Google tools the model can reach, not one can put a message into the world.

Above all of it sits the oldest gate of all, a read-back and an explicit yes before any write. You can watch the agent hold a softer line too, in the moment the lead asks for a price: it declines to invent a number, says so plainly, and routes to the discovery conversation instead. The booking is gated by code; the restraint is gated by character. The preview leans on both.

Drawn as one picture, the whole preview is a short pipeline: a number, a channel that turns identity into a role, one shared session that does the thinking, and one server where every write is gated.

the preview’s architecture: one channel, one shared session, one gated MCP serverone WhatsApp numberchannelresolves identity to a role, never a phone numberleadstudiothe two roles on one numberone Claude Code sessionthe shared braina CLAUDE.md conductor routes on role and intent7 flows5 context5 voicegoogle MCP server8 tools: 6 calendar, 2 gmail-draft, zero send[ORF prefixre-check slotno sendto the leadanswers and the bookingto the studiobrief, link, and draft
the preview's architecture (one shared session, two roles)
  one WhatsApp number
    -> channel: resolves identity to a role + number-free alias (lead | studio);
       the model never sees a phone number
    -> one Claude Code session, the shared brain: a CLAUDE.md conductor routes on
       role and intent across 7 flows, 5 context files, 5 voice files
    -> google MCP server: 8 tools (6 calendar, 2 gmail-draft), zero send;
       every write gated in code ([ORF prefix, re-check the slot, no send)
  -> to the lead:   answers and the booking conversation
  -> to the studio: the post-booking brief, a calendar link, and a welcome draft
  one session serves both roles. this is the smallest working unit.
Figure 2. the preview’s shape. one WhatsApp number, one channel that resolves identity to a role before the model reads a word, one shared Claude Code session whose program is markdown, and one MCP server where every consequential write is gated in code. the same single session answers both the lead and the studio. it is the smallest unit that works, and the sections below are about why a single shared session is a unit, not yet a system.

Talking to the proof

This assistant did not appear from nothing. It is the successor of an earlier WhatsApp demo built for a marketing principal, the same architecture with the context and the character swapped out. And it is the working proof of the appointment-agent case study (ORF-R-2026-003): the shape that piece describes, a Claude Code session serving a principal and their leads through one number, every calendar write behind a hard gate and a read-back, is exactly this. A prospect messaging the assistant is talking to the thing the case study is about.

It is a sibling, too, of the voice agent (ORF-R-2026-004). Different channel, a phone call rather than a chat thread, but the same conviction underneath: the consequential actions belong behind code, and the model is trusted with the conversation, not the keys.

Where this honestly stands

What this is not, yet. It serves one shared founders’ number, not a fleet of them. It answers an allowlist, not open intake, so a stranger does not reach the founders by guessing the line. WhatsApp’s 24-hour session window bounds how the agent can reach back out on its own, so a message that needs to land a day later lands by other means. The deployment is private.

And the lead in the recording is a fictional persona, booked into a real calendar with a burner address: the flow is real, the prospect is not. None of this is hidden. It is the honest edge of an internal preview, shown so the working part can be believed.

The smallest unit that works

The preview is exactly that, a preview, and it is also something more precise: the smallest unit of the thing we are really building. One number, one session, two roles, every consequential write held in code. It settles the first question end to end, on real infrastructure, against a real calendar. A demo that books a real meeting outweighs a paragraph claiming one could.

But a unit is not a system, and the very shape that makes the preview legible is the shape that does not scale. One shared session is one context window and one transcript. It holds a single conversation beautifully. Ask it to hold a hundred leads at once and the seams open: the conversations crowd the same context, the transcript that is its only memory grows until it has to be compacted and the early turns are summarised away, two leads who arrive in the same second contend for one brain, and a single crash takes every conversation with it. The 24-hour messaging window sharpens the point, because memory that should outlive a session has nowhere durable to live.

None of that is a fault in the preview. It is the definition of a unit. The preview answers “does the conviction hold?” The production question is different: what carries that conviction to many businesses and many thousands of leads without the founders touching it? That is not a longer prompt. It is a different harness.

The production shape: sessions, memory, dreaming

We argued the general case already, in the harness is the half you own: an agent is a model plus a harness, and the harness is the only half you build. The preview’s harness, a markdown conductor and a gated channel, is the right harness for a unit. A production harness needs three things the unit fakes with a single session: isolation between conversations, memory that persists outside any one of them, and a way to get sharper between conversations rather than only within them. The note’s discipline is the guide here: do not hand-build and then ossify those primitives, lean on the ones the platform now provides, so each frontier gain flows through instead of breaking against scaffolding shaped for a weaker model.

The Claude platform now names all three. We are precise about status, because these are recent and several are beta or research preview: what follows is the shape we are building toward, not a system we have shipped.

  • Sessions become a managed primitive. The preview is one local Claude Code session; the Claude Agent SDK is that same agent loop offered as a library, where one session maps to one isolated process with its own resumable transcript; and Managed Agents hosts the loop so each lead gets a session of their own instead of sharing one brain. Isolation stops being something we engineer and becomes the default.
  • Memory becomes a store, not a transcript. A shared memory store sits beside the sessions: plain files the agents read and write, with read and write scopes (a read-only store of studio knowledge, a read-write store for leads and bookings), optimistic concurrency so concurrent sessions do not clobber each other, and every change versioned and attributed to the session that made it. What had to be compacted away inside one session now persists, legibly, outside all of them.
  • Dreaming closes the loop. Between sessions, an out-of-band batch process reads the day’s transcripts together with the memory store and produces a new, curated one: duplicates merged, stale entries replaced, fresh patterns surfaced. The founders sleep, the system organises what it learned, and the next day’s sessions attach a sharper memory than the day before left behind.
the production memory system: a session per lead, a shared store, and dreaming between sessionsMemoryreal-time, as sessions runDreamingbetween sessionslead-1sess_a17clead-2sess_b04elead-Nsess_f9d1memory storeshared, file-addressedread-onlyread-writestudio.mdleads/lead-1.mdbookings.mdoptimistic concurrency,versioned and attributedsession transcripts1 to 100 past sessionsdreamingverify · organise · enrichout-of-band, between sessionsstore + transcripts in, curated out
the production memory system (a session per lead, a shared store, dreaming between sessions)
  Memory  (real-time, as sessions run)
    one isolated, resumable session per lead: lead-1, lead-2, ... lead-N
    each reads and writes the shared store live, instead of sharing one session
  memory store  (shared, file-addressed)
    read-only:  studio.md                       (studio knowledge)
    read-write: leads/lead-1.md, bookings.md     (leads + bookings)
    optimistic concurrency, versioned, attributed to the writing session
  Dreaming  (out-of-band, between sessions)
    in:   session transcripts (1 to 100 past sessions) + the store
    work: verify, organise, enrich
    out:  a new, curated store, so the next sessions start sharper
Figure 3. the production memory system, in the visual language of the Claude managed-agent primitives. left, a memory region with one isolated, resumable session per lead, each reading and writing the shared store in real time rather than sharing one session. centre, the shared memory store: a read-only scope for studio knowledge and a read-write scope for leads and bookings, with optimistic concurrency and versioned, attributed changes. right, a dreaming region that runs out of band between sessions, taking the transcripts and the store and returning a curated one so the next sessions start sharper. sessions, memory, and dreaming are the primitives that carry the unit to scale.

The shape is the same conviction wearing different primitives. Identity still belongs to the channel. Consequential writes still sit behind code, and the platform agrees: its own guidance is to denylist destructive tools and keep a human confirmation step before any state change, which is the gated calendar write and the read-back restated as a platform default. What the unit proved by hand, production inherits from the harness.

This is not a single-vendor story. The same shape is forming on the other side of the frontier: OpenAI’s Agents SDK, the production successor to its experimental Swarm, gives agents, handoffs between specialists, guardrails, and sessions that manage history with compaction for long runs. Two labs, one direction. The model is the bought half, and the harness, the sessions and the memory and the orchestration and the gates, is the half a builder owns, increasingly assembled from primitives the platforms provide rather than scaffolding each shop reinvents.

We are building toward this in the open, and we are not there yet. The preview is the unit we can show today, working and uncut. The production system is that unit multiplied: a session for every lead, a memory that persists and is curated while no one is watching, the same gates in code. When it is real and shippable, it gets its own recording.

What transfers

The lesson is portable, and it is not about WhatsApp. An agent can be handed consequential writes, a real calendar, a real inbox, the moment its guardrails live in code rather than in a prompt it might forget. Identity belongs to the channel, not the model, so the agent works in roles and never holds a number. And what earns trust is not a claim of capability but a felt result: the studio briefing itself the instant a booking lands. Build that, and the demo does the arguing.

References

The systems this preview is built on, and the platform primitives its production successor leans on, each linked to its primary documentation:

Provenance · The Assay Mark

Tier
Sealed
Built
12 to 15 June 2026
Build window
4 days (12 to 15 Jun 2026) · as of 2026-06-15
Commits
18 · as of 2026-06-16
Roles on one number
2 (lead, studio) · as of 2026-06-16
Conversational flows
7 · as of 2026-06-16
Knowledge and voice files
10 (5 context, 5 voice) · as of 2026-06-16
Google tools exposed to the model
8 (6 calendar, 2 Gmail), zero send · as of 2026-06-16
Tests
149 deterministic, credless (312 assertions) · as of 2026-06-16
Application code
1 channel server, 1 google MCP server · as of 2026-06-16
Disclosure
Built as an internal research preview on a private repository, behind an allowlist and a number-free identity boundary, so it is not publicly linkable; the founders' real number and any real client identities stay withheld by design, and the demo's lead, including its name, email, and calendar invite, is a fictional persona generated for the recording. The figures here are the repository's own dated output, read from the filesystem and git history. The recorded demo below is the public evidence of the working flow: one canonical, uncut two-screen take, permanently reviewable.