Testing a chatbot before you publish it: the three states
Why "we tested it internally" is not enough
Because you already know what you meant when you wrote the instruction.
Testing a chatbot before you publish it with your own team has a problem at the root: your team asks what the configuration expects. Your customers ask something else, in other words, and sometimes in a two-minute voice note.
What helps is not a longer test. It is not having to choose between "off" and "answering everything".
The three states, in order
A service passes through three states, and you can stay in any of them as long as you need.
| State | Who sees the replies | What it is for |
|---|---|---|
| Sandbox | Only you | Seeing how it answers before it talks to anybody |
| Inbox with the AI off | Your team | Getting customer service organised before automating |
| Gradual adoption | Only the conversations you tag | Opening up slowly, with the handbrake on |
It is not a compulsory staircase. Some businesses start at the second state and stay there for months.
State 1: the sandbox, with real contacts
The sandbox is a test chat where you see how the service replies before it serves anybody.
What makes it useful is not the chat: it is being able to pick a real contact from your address book. There you see how it would answer with that person's context, not with a generic case invented for the demo.
One honest limit: sandbox runs have a daily quota depending on the plan, separate from the monthly message quota. When the day runs out, it stops.
Notes on the contact to simulate the odd case
The case that breaks everything is never the usual one.
You can add notes to a contact to simulate specific situations: the customer with an open complaint, the one who buys on special terms, the one who has already asked the same thing three times.
It is the cheap way to test the scenario that worries you without waiting for it to happen.
State 2: the inbox with the AI off
This state always gets skipped and it is the one that pays off most at the start.
You connect the channel and work in the inbox with the service paused. Nobody automates anything: your team handles messages, but it handles them in an organised way, with conversations distributed and the history in one place.
Automation gets switched on once the team is comfortable. And by then you have something you did not have before: real conversations to read and find out what people actually ask.
State 3: gradual adoption, conversation by conversation
This is what separates it from "switch it on and pray".
In gradual mode, the service answers only the conversations you tag. Everything else keeps being handled by your team, as always. You choose where to start: one type of query, one time slot, a few customers.
In full mode the opposite happens: it answers everything except what you mark as off limits.
The tag that switches it off mid-case
A conversation tagged for human handling receives no automated replies, in any mode.
It is there for the uncomfortable case: the complaint that turned ugly, the customer who wants to talk to a person, the negotiation that is nobody's but yours.
And it is not a whim of ours. A Gartner survey of 3,566 customers found that 87% say a company using generative AI in customer service has to provide access to a human agent. The emergency exit is part of the product, not a concession.
Pausing one part without switching everything off
The panic button is almost always too big.
You can pause one specific skill on one service: stop it taking orders, or stop it reading the catalogue, without touching what your other services do or the rest of the replies.
That is what lets you fix a contained problem without going back to "nobody is answering".
What to look at before opening up fully
Four things, and none of them needs a dashboard.
- The handoffs. Which topics end up with a person, and whether they are the ones you expected.
- The extra answers. It answered something that was not its business, even if it answered well.
- What it did not know. Unanswered questions are the list of documents that are missing.
- The tone. Whether it sounds like your business or like a manual.
Key takeaways
- Testing a chatbot before publishing it is not a stage: it is three states the same service passes through.
- The sandbox lets you test with real contacts from your address book, and with notes to simulate odd cases.
- Using the inbox with the AI off organises customer service before anything gets automated.
- Gradual adoption means only the conversations you tag get answered.
- A human-handling tag suspends automated replies in any mode.
- A single skill can be paused without switching off the rest of the service.
Frequently asked questions
- Can I test the service without connecting my WhatsApp?
- Yes. The sandbox is a test chat inside the panel and does not need the channel connected. You can pick contacts from your address book to see how it would answer with that person's context.
- How many tests can I run per day?
- There is a daily quota of sandbox runs depending on the plan, separate from the monthly message quota. When the day's quota runs out, it stops until the next day.
- What is gradual adoption mode?
- It is the setting where the service replies only to the conversations you tag for it. The rest keep being handled by your team. Full mode does the opposite: it answers everything except what you mark.
- Can I stop the AI in one specific conversation?
- Yes, by tagging that conversation for human handling. It stops receiving automated replies regardless of which mode the service is in.
- Is it worth connecting the channel even if I do not want to automate yet?
- Yes, and it is usually the best first step. You work in the inbox with the service paused, get your team's handling organised, and switch automation on when you are ready.
Related posts
Get twice the output from your team without asking more of them
It answers WhatsApp, Instagram and web. It learns your business, takes the orders, and hands off to a person when it matters. You see every action noted in the chat.
Free Pro trial. No credit card required.