012 — Agent Tools
Defines the tools the agent can call during a turn to send a tenant-built ManyChat flow, add or remove a tag, and set a custom field, and when those calls take effect. It deliberately leaves out page-level and account-global endpoints, creating or updating subscribers, and MCP. Reading the current contact and writing free-text notes were left out here too, and are added by 024-contact-read-and-notes.md.
A tool stages an action; the server performs it
The AI SDK's default tool has an execute that does the work: here, that would mean calling ManyChat mid-generation. That default is rejected (ADR-0010). At that point the model has not yet decided whether to escalate, and the guardrails have not run, so a turn that ends in a handoff would already have sent media and tagged the contact.
Every tool's execute therefore records an action and returns { staged: true }. No ManyChat request is made inside the model call. Staged actions travel with the model's result and are performed only as described below.
The model is told an action was staged, never that it succeeded. The persona must not have it claim "I've sent it" as fact. It says what it is sending, the way a person does before pressing send.
send_flow on an inbound turn is the exception: it sends the flow when called and returns { sent: true | false }, so the reply follows the flow (029, ADR-0019).
The mechanism is built here; only the choices are the tenant's
Everything in this spec is code in this repository, identical for every tenant: the tools, staging, the tools.json schema and its startup checks, the delivery order on both paths, the outbox payload, and the per-turn action record. A deployment never implements a tool. It supplies:
config/tools.json, which says which flows, tags and field values exist and when each should be used (gitignored, C1);- the flows themselves, and the media inside them, which live in the tenant's ManyChat account.
The implementing pull request commits config/tools.json.example for the fictional demo tenant, as 003-config-schema.md requires of every config file. A deployment with no tools.json gets no tools, and behaves exactly as it does today.
Six tools, parameterised by tenant configuration
| Tool | ManyChat endpoint | Model supplies |
|---|---|---|
send_flow | POST /fb/sending/sendFlow | a flow id, contactAsked |
add_tag | POST /fb/subscriber/addTagByName | a tag id |
remove_tag | POST /fb/subscriber/removeTagByName | a tag id |
set_field | POST /fb/subscriber/setCustomFieldByName | a field id and a value id |
get_contact | GET /fb/subscriber/getInfo | nothing |
write_note | POST /fb/subscriber/setCustomFieldByName | a note id and text |
get_contact and write_note are 024's. get_contact is the one tool that reads, and the one whose execute makes a request (ADR-0016); everything below about staging applies to the other five.
025 adds a fifth, schedule_nudge, which takes a delay id and is performed as a row in this service's database, not as a ManyChat request.
When the tenant marks an intent field (034), every write but set_field on that field and write_note may return reason: "not_prospect": until the contact is a prospect, it is refused with no request and takes no slot of the turn's cap. get_contact is unaffected.
Each parameter is a z.enum built from config/tools.json at load, so the model can only name something the tenant configured (C3). A tool whose list is empty is not offered at all. The file reloads on SIGHUP with the rest of the tenant config, and is added to 003-config-schema.md in the pull request that implements this spec. Shape, with values invented for the demo tenant:
{
"flows": [
{
"id": "intro_course_brochure",
"flowNs": "content00000000000000_000001",
"description": "PDF brochure for the intro course. Send when the contact asks for details, a syllabus or something to read.",
},
],
"tags": [
{
"id": "interested_intro",
"tag": "interested-intro-course",
"description": "Contact showed interest in the intro course.",
},
],
"fields": [
{
"id": "preferred_shift",
"field": "preferred_shift",
"values": ["morning", "evening"],
"description": "Which shift the contact said suits them.",
},
],
}The description fields are the "what it contains, when to use it" guidance the model reads. They are tenant copy, so they may be in the tenant's language (005-language.md). The model sees id and description and never sees flowNs, tag or field. Those name objects in the tenant's ManyChat account, and keeping them server-side means renaming one is a config edit that no prompt depends on.
The subscriber is never a parameter
No tool takes a subscriber id. Each tool closes over the current turn's subscriber. Contact text cannot change which contact an action lands on (C4), and one contact's turn cannot write to another contact's record, which would be a disclosure under C5 rather than just a bug.
Free-text field values are refused
set_field takes a value from the field's configured values, never a string the model composed. A free-text write would put unvalidated model output into a field that a ManyChat flow may later render to the contact, which bypasses C3. It would also invite the model to copy the contact's own words (names, phone numbers) into the CRM.
This holds for fields[]. Free text is permitted only in the notes[] of 024, a separate kind of field that the tenant declares no flow renders, whose text is bounded and cleaned before it is written (ADR-0017).
Guardrails run before any action is performed
Staged actions are performed only if the final reply, afterapplyGuardrails, has escalate: false. Every other outcome discards them, and the count discarded is logged. The one exception is a note marked onEscalation when the model or the confidence threshold escalated (024 § A handoff summary survives the escalation it describes):
- the model set
escalate: true; - confidence fell below the threshold;
- a prompt or fence leak was detected;
- the output failed schema validation, or the model call threw;
- the model call hit
MODEL_ABORT_MS.
Turns where the model never runs (the scripted opening, escalation keywords, budget, rate and turn caps) have no tools and so stage nothing.
A flow sent during an inbound turn (029) has already gone out when any of these escalates the turn. It keeps its outcome on the record.
The loop is bounded at four steps
Steps one to three may call tools, several in parallel. Step four offers no tools, so it must produce the AgentReply. If no reply is produced, the turn fails closed exactly as a schema failure does today.
A turn stages at most 8 actions. A call past that returns { staged: false } and is dropped. Identical staged actions are de-duplicated. A flow sent during an inbound turn (029) counts against the same 8; one past the limit is not sent and returns { sent: false } with the reason. A turn reads at most twice (024).
Both numbers were chosen, not measured, and replace this spec's original two steps and three actions (024 § The loop grows to four steps and eight actions). Four steps fit read, act, read again and reply. Eight actions bound the burst a single turn can put through the ManyChat rate limiter, where each action is one request against the existing token bucket. Change them when an eval shows a need, and record the measurement date here.
The whole loop runs inside the race
The race deadline and model abort in 002-channel-contract.md are unchanged and apply to the entire loop, not to each step. A tool turn is more likely to lose the race. That is accepted, since the outbox exists for exactly that. The reply is deferred, never dropped, and its staged actions are deferred with it.
Actions follow the text, on both delivery paths
- Race won: actions are performed after the Dynamic Block response has been sent, never before.
- Race lost: the outbox payload carries the staged actions alongside the messages. The worker performs them only after the text is delivered. If the row is dead-lettered, the actions are dropped with it, because media arriving without the reply that introduces it is worse than neither.
Actions are performed in the order the model staged them, one request each, through the existing ManyChatClient and its rate limiter.
On an inbound turn, flows are not among them: they were sent during the turn, before the text (029). A reply held in the outbox for a flow still playing carries its staged actions with it, as on the race-lost path (030).
The flow set may not include the reply flow or field
flows[].flowNs may not equal MANYCHAT_REPLY_FLOW_NS, and fields[].field may not equal MANYCHAT_REPLY_FIELD or MANYCHAT_TOKEN_FIELD. The reply flow renders whatever the reply field holds, so firing it as a tool would resend a stale reply, and writing that field as a tool would overwrite a reply in flight (002 § The two calls are one delivery). Writing the token field would replace the contact's token (019). Any collision is a startup failure, not a runtime surprise.
A failed action is logged, never retried
Each action gets one attempt. A failure is logged at warn with the tool, the configured id and the ManyChat error, with the subscriber redacted (C5), and is not retried. The text is the reply of record. Actions are enhancements to it, and a retry that lands minutes later, after the conversation has moved on, is worse than a missing tag.
On the deferred path, an outbox retry re-delivers text only. Actions are performed once, after the first successful delivery.
On the inline path, actions run in-process after the response. A crash between the response and the actions loses them. This is accepted for the same reason.
Every staged action is recorded on its turn
A log line answers "did something fail?" and nothing else. It cannot answer "what did the agent do for this contact?", and it only covers failures. So the agent's row in turns gains an actions column (jsonb): one entry per action the model staged, whatever became of it.
[
{ "tool": "send_flow", "id": "intro_course_brochure", "status": "performed" },
{
"tool": "set_field",
"id": "preferred_shift",
"value": "evening",
"status": "failed",
"error": "…",
},
]status | Meaning |
|---|---|
staged | Not yet performed; waiting for the response or outbox |
performed | ManyChat accepted the request |
failed | ManyChat rejected it; error holds the reason |
discarded | The turn escalated (see "Guardrails run before…") |
A note is recorded by its length, never its text (024). | dropped_over_cap | Staged past the per-turn limit and never sent | | dead_lettered | Its outbox row was dead-lettered, so it was never sent |
A flow sent during an inbound turn (029) is written with its outcome, performed or failed, in its place among the staged entries, and a payment link's follow-ons after it. On both paths the staged entries are written as staged with the turn and updated once the actions have run: on the inline path after the response, and on the deferred path by the outbox worker after it has sent them. An inline turn whose process dies before its actions run therefore stays staged. The column is null on turns where no tool was offered, so "no tools" and "tools offered, none chosen" ([]) are distinct.
Entries hold only configured ids and values, never flowNs, tag or field names, or contact text, so the record needs no redaction under C5.
A server-performed follow-on is recorded beside the action that carried it: the link_sent write after the payment-link flow (023), and a { "tool": "send_event", "id": … } entry after a funnel write that fired a conversion event (027). send_event is a record kind, not a tool. A flow the server sent without the model calling it carries origin, opening or stage, and a payment link sent before prepared carries contactAsked: true (032).
Performed actions reach the model on later turns
Today the history the model sees is text only. Without more, the model on the next turn cannot tell that it already sent the brochure, and will send it again when the contact says "thanks, and the price?".
Each agent turn in the history therefore carries a server-written note of the actions recorded as performed, e.g. [actions performed: send_flow intro_course_brochure]. The note:
- is built from the
actionscolumn, never from the model's own text; - sits outside the contact fence, since it is system-authored rather than untrusted (C4);
- lists
performedonly. A discarded, failed or still-staged action is not mentioned, so the model never believes something reached the contact that did not. A deferred action that has not yet been performed when the next turn starts is absent from that turn, which errs toward a repeat, not a false claim; - uses ids, not descriptions, so a long
descriptionis not repeated on every turn of history. - omits
send_evententries (027): an event reached the ad platform, not the contact.
The note's format is English and system-facing, never shown to the contact, so it is not customer copy under C9.
Verification
- A unit test asserts each tool's parameter schema rejects an id absent from
config/tools.json, and that a tool with an empty list is not offered. - A unit test asserts that
executemakes no request: afetchImplspy records zero calls for the whole model step. - A test drives each escalation path listed under "Guardrails run before any action is performed" and asserts zero ManyChat requests.
- A test asserts a fourth staged action returns
{ staged: false }and is not performed. - A unit test over
fetchImplpins each action's request body. Each carries the current turn'ssubscriber_id, andset_fieldcarries a configured value. - On both paths, a test asserts the order: text is delivered before any action request. On the deferred path, a retried row performs its actions once, and a dead-lettered row performs none.
- A config test asserts that a flow equal to
MANYCHAT_REPLY_FLOW_NS, or a field equal toMANYCHAT_REPLY_FIELD, fails at load. - The mock model gains a tool-calling case. Golden eval cases assert the tool choice for a request that should send a flow, and for one that should not.
001 § Verificationitem 5 (p95 latency) now covers tool turns. - An integration test against PGlite asserts each
actionsstatus is written on the path that produces it. It also asserts that a deferred turn's entries move fromstagedtoperformedorfailedafter the worker drains, and todead_letteredwhen the row is. - A unit test asserts the next turn's prompt carries the note for a
performedaction and omits discarded, failed and staged ones.
What this misses: as in 002, ManyChat returning success is not WhatsApp delivering. A flow that was renamed, is unpublished or points at a missing file passes every check here. Ordering is also only enforced on this side. Once the Dynamic Block response and a sendFlow are both with ManyChat, the order in which the contact sees them is ManyChat's. And no test can tell whether a description leads the model to the right flow. Evals sample that, and a person reading real conversations is the actual check.
For the same reason, performed means ManyChat accepted the request, not that the contact received it. The action record and the history note both inherit that gap. They say what this service did, not what the contact saw.