A multi-tenant AI agent platform in production · Imply · 03/2026 – Present
Context
Human-only customer service does not scale. Before the platform, the median time to first response was 2 minutes and 39 seconds, and every conversation consumed an operator from the first message to the last.
The goal was not a chatbot. It was a platform where each company configures its own AI agents, connects its own channels and knowledge, and keeps a human in the loop for the cases that need one — with the cost of every conversation visible.
Architecture
Multi-channel intake lands on a RabbitMQ queue; an asynchronous, idempotent orchestrator runs 14 stages — agent routing, RAG retrieval, LLM call, tool execution and guardrails — before delivering the answer or handing the conversation to a human. The five stages shown here summarise the 14.
Results
85% of conversations resolved end-to-end by AI, with no human intervention — 6,570 of 7,725 closed in Aug/2026, and 83.3% over the last 90 days.
Time to first response cut from 2m39s to 4.7 seconds (median, −97%), measured across 61k question-answer pairs in production.
~153k messages and ~7.8k conversations per month — 24× growth in 4 months.
Inference held at $0.12 per completed conversation across 810M tokens/month, accounted per tenant and per model against a versioned price table.
Sustained at scale: a ~107k-line monorepo across 1,019 files, 69 Prisma models, 109 migrations and ~3,250 test cases with an 80% coverage threshold and an architecture check in CI.
Engineering decisions
01
A multi-provider LLM layer behind a common interface
WhyModel quality, price and availability move every few months. Business logic must not know which provider answered.
What it costAn extra abstraction to maintain, plus per-model capability detection — a provider that lacks a feature cannot silently degrade the pipeline.
02
An in-house RAG pipeline instead of a framework
WhyChunking, embedding cache, versioning and knowledge-base access logging are product requirements here, not implementation details. A framework would have to be fought to expose them.
What it costMore code owned by us, including the query embedding cache and the retrieval tuning that a framework would have shipped for free.
03
An asynchronous, idempotent orchestrator over a queue
WhyA queue consumer will see duplicates. Any stage that is not idempotent eventually sends the same message to a customer twice.
What it costEvery one of the 14 stages must be written to be safely re-run, which is slower to build and harder to reason about than a synchronous call chain.
04
Agent routing with 7 priority levels
WhyThe right agent depends on context: an active conversation must not be hijacked, explicit routing rules must win over defaults, and an interactive menu must win over inference.
What it costRouting became the most sensitive part of the system — it needs an intent router and a message debouncer in front of it to behave under bursts.
05
Guardrails as a mandatory pipeline stage, not a prompt instruction
WhyPrompt rules degrade. Content filtering, unauthorized URL blocking, internal ID leak protection and tool argument validation have to hold even when the model misbehaves.
What it costFalse positives. A fence that blocks code blocks outright would break legitimate content — a Pix copy-and-paste key is plain text a customer genuinely needs.
What broke, and what I learned
A NUL byte silently ate messages in the dead-letter queue
What happenedMessages disappeared with no error surfaced. They had reached the dead-letter queue, but a NUL byte in the payload meant the content could not be persisted or read back — the failure looked like nothing had happened at all.
The fixSanitise the payload before persistence and make dead-letter contents readable, so a silent drop becomes a visible failure.
The agent answered outside its scope
What happenedNothing in the system constrained the subject of a conversation. Asked for help with an unrelated PHP function, the agent obliged — in a customer service channel.
The fixAn explicit scope guard. The subtlety: the fence cannot simply block fenced code, because legitimate customer content looks like that.
Conversations were returned to the queue on a false presence signal
What happenedAn expired presence heartbeat was treated as absence, so conversations were pulled from operators who were actively working and redistributed.
The fixSeparate presence from availability, and give an offline operator a 24-hour window instead of an immediate return.