Suppose a support assistant gives a customer a refund deadline that conflicts with the current policy. A production trace holds the input and output, and may show which intermediate step led to the answer. MLflow tracing captures those details. The trace gives you a concrete failure to investigate.
Choose the failure
Open the trace and identify the input that exposed the problem. Check the answer and any retrieved material before deciding what the case should test. Keep the relevant input and context. Remove details that are not needed for the test, such as a customer’s name or account number. Write down why this example matters so another reviewer can understand the choice.
MLflow evaluation datasets can be built from existing traces or from cases written by hand. The documentation recommends choosing traces that represent important problems, including low-quality outputs and edge cases. Export the selected trace to a dataset through the UI, or find and add it with the SDK.
State the expected behavior
The failed answer alone cannot tell a future test what success looks like. Have a person who knows the policy define the expectation. For this example, it might require the assistant to use the current policy text and avoid giving a deadline when that text is unavailable. Make the criterion specific enough that a reviewer can say whether an answer meets it.
MLflow expectations attach the desired answer or behavior to a trace. The docs describe factual, structured, and behavioral expectations. They also recommend specific, measurable criteria and metadata explaining the reason for an expectation. This turns an observed failure into a case with an explicit answer key.
Check the next change
Run the curated case against the proposed prompt, model, or application change. Inspect the new output and the test result. MLflow’s regression testing guide describes pinning the input, scoring the desired behavior, and asserting a pass or fail in a pytest test. Run that check when changes are proposed so the known failure stays visible.
What to do
Pick one real failed trace. Review its input and output. Save a safe, representative version in an evaluation dataset. Ask a domain expert to write the expected behavior. Add a pass or fail check, run it against the next change, and keep the case when the fix lands.

The Campfire
No commentsNobody has pulled up a log by this one yet. Be the first to say what you make of it.
Held for the desk. It appears after a look.