Content Moderation
Check text against Mistral's moderation model and see the flagged categories.
LLM Quirks
Small, reproducible cases where an LLM's behavior isn't what you'd expect.
Token Efficiency
Classify-then-answer vs. one big prompt: compare real token usage side by side.
Tool Calling
Define your tools to steer what the LLM says next — then watch it decide when to use them.
Model Evaluation
A router step that sorts customer messages into departments — run the test suite and see every result and the overall accuracy.