19. Test Center
Test Center stress-tests the real Chat Advisor. It asks typical questions, edge cases, and security attack prompts, then checks whether the answer stays grounded, safe, and in scope.
What Test Center Runs
Section titled “What Test Center Runs”Test Center can run:
- normal customer questions,
- edge cases,
- security tests,
- custom rules added by your team,
- an optional 3-judge panel for deeper review.
Trigger types include:
- manual,
- weekly,
- after product sync,
- negative feedback.
Results
Section titled “Results”Each run reports:
- total questions,
- passed checks,
- failed checks,
- notes or soft warnings,
- last run timestamp,
- trigger type,
- answer source per case,
- failed assertions,
- optional judge summary and score.
Answer sources include:
| Source | Meaning |
|---|---|
| Product catalog | The answer used product data. |
| BIQ match | A curated BIQ answered the question. |
| Material | The answer used connected material. |
| World knowledge | The answer relied on general model knowledge. |
| Error | The answer could not be generated or evaluated. |
Security Test Corpus
Section titled “Security Test Corpus”The built-in security corpus includes attack patterns across these classes:
- direct prompt injection,
- indirect injection / RAG poisoning,
- system-prompt extraction,
- jailbreak or roleplay,
- data leak / GDPR,
- hallucination provocation,
- lead or conversion abuse,
- scope overreach.
Expected reactions can include:
- refuse,
- escalate,
- answer only from real shop data,
- clarify,
- answer normally.
Custom Test Rules
Section titled “Custom Test Rules”Use custom rules for organization-specific expectations.
Examples:
- “Never promise same-day delivery unless material says so.”
- “Do not offer discounts unless the discount-code system provides one.”
- “Escalate legal warranty questions to support.”
- “Answer certification questions only from source material.”
Custom rules run with the standard tests in the next run.
When To Run Tests
Section titled “When To Run Tests”Run Test Center:
- before activating Chat Advisor,
- after changing Brand Voice rules,
- after major product syncs,
- after new material imports,
- after SafeGuard corrections,
- after negative feedback,
- before important launches.
Interpreting Failures
Section titled “Interpreting Failures”A failure means the chat behavior should be reviewed before relying on it. Common causes:
- missing or weak source material,
- BIQs not published,
- unsupported claims,
- product catalog gaps,
- prompt-injection weakness,
- overly broad Brand Voice instructions,
- discount or price promises without source support.
Fix the underlying content or settings, then rerun the tests.
Relation To Trust Center
Section titled “Relation To Trust Center”Trust Center is the activation gate. Test Center is the stress-test report. Use both:
- Trust Center confirms notices and self-test.
- Test Center checks real answer behavior.
For a production Chat Advisor, run Test Center before final activation and after every major content or policy change.