← All work

Case study

Vera: a WhatsApp AI sales agent, and how I know it works

A bilingual agent that answers customers, qualifies leads and updates the CRM, measured by my own evaluation harness.

My role
Sole engineer: design, build, evaluation and production operation
When
Early 2026 – present
Status
In production
Stack
n8n, Claude API, Qdrant, Chatwoot, WhatsApp Cloud API, Python, FastAPI, pytest
The evaluation results: every prompt version scored against the same 37 cases. Recorded on demo data.

The problem

Everything comes in through one WhatsApp number: buyers asking about boats, but also job seekers, rental enquiries and everything else, in English and Arabic. The sales team was spending its time sorting that first contact instead of selling, and a high-value sale needs a person's attention at the right moment, not at the first message.

What I built

Vera handles the first contact on WhatsApp. Job seekers are sent to the application form, and their conversations show up on the portal's recruitment page. Buyers get answers from the company's product documents, brochures and videos, while Vera finds out what they need, qualifies them and writes it to the CRM, then hands them to a salesperson for the personal side of the sale. It is a multi-step workflow rather than one prompt:

  • merges a customer's burst of short messages so it answers the whole thought once
  • transcribes voice notes, and replies with a voice note when the customer sent one
  • reads photos the customer sends
  • retrieves the relevant product knowledge from a vector database
  • calls Claude with a cached system prompt, then parses structured tags from the reply
  • applies CRM labels and routes the conversation: recruitment, or a salesperson once the buyer is qualified
  • a spam gate that switches the agent off for a conversation when it trips

How I know it works

For the first versions, the test suite was me chatting to it, and customers found what I missed. So I built an evaluation harness in Python that runs 37 cases against the real production workflow in a test mode, where nothing reaches a customer or the CRM.

Fact cases check knowledge, such as a superseded specification reaching customers. Scenario cases are scripted multi-turn conversations, because the agent's real failures were never single facts: old chat history overriding what the customer asked now, or confusion about which channel it was on.

Code checks run first and override the LLM grader, because that is exactly where a model judge is generous. Retrieval and answer accuracy are scored separately: a wrong answer with a retrieval miss is a knowledge problem, one with a retrieval hit is a prompt problem. Each run records latency and splits results by tag, such as Arabic versus English.

Accuracy went from 74% to 89% over two prompt revisions. The harness caught the agent implying a classification approval the boats don't carry by default, and two prompt rules that contradicted each other.

The hard part

Of the early failures I investigated, seven were bugs in the harness and two were real: thin reference answers, an acronym matching inside a longer word, a number matching inside a longer number, a rubric that treated true but unlisted facts as made up. I found each by checking a red result against the source before acting on it, and each fix is now pinned by a test. The discipline that made the tool useful was refusing to treat its output as a verdict.

Limits, and what's next

37 cases is a starting set, not coverage: every new failure found by hand becomes a permanent case. Next is running the suite automatically before each prompt change ships, and adding cost per conversation to every run.