Skip to main content
AI Tests lets you check how your AI agent answers before it talks to real customers. You save customer questions with the answer or behavior you expect, run them all at once, and an AI judge scores each reply. Run the same test case again after you change your agent to catch answers that got worse.

Prerequisites

  • You already created an AI agent in AI Studio.
  • You know the questions your customers ask most, and what a good answer looks like.

How AI Tests works

  • A test case is a group of questions, for example “Returns policy answers”.
  • Each question is independent. The agent answers it as a new message, not as part of a chat.
  • For each question, you write the Expected answer or behavior: what the agent must say or do.
  • When you run a test case, your active AI agent answers every question, and the AI judge compares each answer with what you expected.
To test a full back-and-forth conversation, use the AI Playground. Use AI Tests for questions you want to re-check every time you change your agent.

Step 1: Create a test case

  1. Go to AI Studio > AI Tests.
  2. Click Add new test case.
  3. Enter a Test case name and, optionally, a Test case description.
  4. Under Add your first question, fill in:
    • Customer question: the message a customer would send, for example “Can I return a sale item?”
    • Expected answer or behavior: what the agent must say or do. Include limits, exceptions, or steps.
  5. Optional: select Add customer profile, then pick a contact in Customer profile. The agent answers as if that customer asked the question.
  6. Click Create test case.
Create test case form in AI Tests The test case opens so you can add more questions.

Step 2: Add more questions

  1. Click Add new item.
  2. Fill in Customer question and Expected answer or behavior. Add a customer profile if needed.
  3. Click Save question.
Repeat for each question you want to test.
Write expected answers the AI judge can check. “Says sale items can be returned within 14 days for store credit only” works better than “Gives a good answer about returns.”

Step 3: Run the test

  1. Open the test case and click Run test.
  2. A progress card shows how many questions are done. If you stay on the test case, a message shows the pass count when the run finishes.
You can leave the page while the test runs. The test keeps running, and the AI Tests page shows Running until the result is ready. While a test runs, you can’t add, edit, or delete questions.

Step 4: Review the results

When the run finishes, the summary at the top of the test case shows how many questions passed, for example 8 of 10 passed. Each question also shows a Last result. Test case page with the pass summary and the Last result of each question To see the details, click a question. The result page shows:
  • Customer asked: the question that was sent.
  • Expected answer or behavior: what you wrote.
  • AI answered: your agent’s actual reply.
  • AI judge: a short note on why the answer passed or failed.
Question result page showing the customer question, expected answer, AI answer, and AI judge note Use the run history on the left to compare the same question across earlier runs. If a failed answer shows a gap, update your agent’s instructions, skills, or Knowledge Base, then run the test again.

Manage test cases and questions

  • Edit a test case: on the AI Tests page, click the edit icon to rename it or change its description.
  • Edit or delete a question: open the test case, then use the edit or delete icon on the question’s row.
  • Delete a test case: on the AI Tests page, click the delete icon.
Deleting a test case or a question also deletes its run history permanently.

Troubleshooting

The run didn’t finish. You see This run didn’t finish at the top of the test case. Your questions and earlier results are not changed. Click Run again. Run test is turned off. The test case has no questions yet, or a run is still in progress. Add a question, or wait for the current run to finish.