Skip to content

19. Test Center

Test Center stress-tests the real Chat Advisor. It asks typical questions, edge cases, and security attack prompts, then checks whether the answer stays grounded, safe, and in scope.

Test Center can run:

  • normal customer questions,
  • edge cases,
  • security tests,
  • custom rules added by your team,
  • an optional 3-judge panel for deeper review.

Trigger types include:

  • manual,
  • weekly,
  • after product sync,
  • negative feedback.

Each run reports:

  • total questions,
  • passed checks,
  • failed checks,
  • notes or soft warnings,
  • last run timestamp,
  • trigger type,
  • answer source per case,
  • failed assertions,
  • optional judge summary and score.

Answer sources include:

SourceMeaning
Product catalogThe answer used product data.
BIQ matchA curated BIQ answered the question.
MaterialThe answer used connected material.
World knowledgeThe answer relied on general model knowledge.
ErrorThe answer could not be generated or evaluated.

The built-in security corpus includes attack patterns across these classes:

  • direct prompt injection,
  • indirect injection / RAG poisoning,
  • system-prompt extraction,
  • jailbreak or roleplay,
  • data leak / GDPR,
  • hallucination provocation,
  • lead or conversion abuse,
  • scope overreach.

Expected reactions can include:

  • refuse,
  • escalate,
  • answer only from real shop data,
  • clarify,
  • answer normally.

Use custom rules for organization-specific expectations.

Examples:

  • “Never promise same-day delivery unless material says so.”
  • “Do not offer discounts unless the discount-code system provides one.”
  • “Escalate legal warranty questions to support.”
  • “Answer certification questions only from source material.”

Custom rules run with the standard tests in the next run.

Run Test Center:

  • before activating Chat Advisor,
  • after changing Brand Voice rules,
  • after major product syncs,
  • after new material imports,
  • after SafeGuard corrections,
  • after negative feedback,
  • before important launches.

A failure means the chat behavior should be reviewed before relying on it. Common causes:

  • missing or weak source material,
  • BIQs not published,
  • unsupported claims,
  • product catalog gaps,
  • prompt-injection weakness,
  • overly broad Brand Voice instructions,
  • discount or price promises without source support.

Fix the underlying content or settings, then rerun the tests.

Trust Center is the activation gate. Test Center is the stress-test report. Use both:

  • Trust Center confirms notices and self-test.
  • Test Center checks real answer behavior.

For a production Chat Advisor, run Test Center before final activation and after every major content or policy change.