Conversational metrics
The conversational metrics can be used to score the quality of conversational threads collected by Opik through multiple traces. They also apply to conversations sourced outside of Opik when you want to analyse the performance of an assistant across turns.
Opik provides two families of conversation metrics:
- Conversation-level heuristic metrics – lightweight analytics that inspect the transcript itself (for example, knowledge retention or degeneration). Use these when you only have the production conversation and no gold reference.
- LLM-as-a-judge conversation metrics – call an LLM to reason about conversation quality, user goal completion, or risk in the latest assistant responses.
Conversation-level heuristic metrics
Section titled “Conversation-level heuristic metrics”| Metric | Description |
|---|---|
| KnowledgeRetentionMetric | Checks whether the final assistant replies retain earlier user-provided facts. |
| ConversationDegenerationMetric | Detects repetition and degeneration patterns across the conversation. |
Knowledge Retention Metric
Section titled “Knowledge Retention Metric”KnowledgeRetentionMetric operates on a conversation and compares how well the last assistant message preserves facts the user injected earlier. This is useful for guardrailing agents that should respect instructions or keep important constraints.
from opik.evaluation.metrics import KnowledgeRetentionMetric
metric = KnowledgeRetentionMetric(turns_to_consider=5)
score = metric.score(conversation=my_thread)
print(score.value, score.reason)Conversation Degeneration Metric
Section titled “Conversation Degeneration Metric”ConversationDegenerationMetric detects repetitive phrases, lack of variance, or low-entropy responses across a conversation. It is a lightweight guard against models that fall into loops or short-circuit the dialogue.
from opik.evaluation.metrics import ConversationDegenerationMetric
metric = ConversationDegenerationMetric()
score = metric.score(conversation=my_thread)LLM-as-a-judge conversation metrics
Section titled “LLM-as-a-judge conversation metrics”| Metric | Description |
|---|---|
| ConversationalCoherenceMetric | Evaluates coherence and relevance across sliding windows of the dialogue. |
| SessionCompletenessQuality | Checks whether the user’s high-level goals were satisfied. |
| UserFrustrationMetric | Estimates how frustrated the user was across the interaction. |
| ConversationComplianceRiskMetric | Applies the Compliance Risk judge to the last assistant response. |
| ConversationDialogueHelpfulnessMetric | Rates how helpful the final assistant reply is. |
| ConversationQARelevanceMetric | Checks whether the final answer addresses the user’s request. |
| ConversationSummarizationConsistencyMetric | Scores how faithful a conversation summary is to the transcript. |
| ConversationSummarizationCoherenceMetric | Scores the structure and flow of a conversation summary. |
| ConversationPromptPerplexityMetric | Estimates prompt difficulty at the conversation level. |
| ConversationPromptUncertaintyMetric | Flags ambiguous prompts in threaded evaluations. |
These metrics are based on the idea of using an LLM to evaluate the turns of the conversation between user and assistant. Opik ships a prompt template that wraps the transcript, criteria, and evaluation steps for you. By default, the gpt-5-nano model is used to evaluate responses, but you can switch to any LiteLLM-supported backend by setting the model parameter. You can learn more in the Customize models for LLM as a Judge metrics guide.
ConversationalCoherenceMetric
Section titled “ConversationalCoherenceMetric”ConversationalCoherenceMetric evaluates the logical flow of a dialogue. It builds a sliding window of turns and asks an LLM to rate whether the final assistant message is coherent and relevant. It returns a score between 0.0 and 1.0 and can optionally return detailed reasons.
from opik.evaluation.metrics import ConversationalCoherenceMetric
conversation = [
{
"role": "user",
"content": "I need to book a flight to New York and find a hotel.",
},
{
"role": "assistant",
"content": "I can help you with that. For flights to New York, what dates are you looking to travel?",
},
{
"role": "user",
"content": "Next weekend, from Friday to Sunday.",
},
{
"role": "assistant",
"content": "Great! I recommend checking airlines like Delta, United, or JetBlue for flights to New York next weekend. For hotels, what's your budget range and preferred location in New York?",
},
{
"role": "user",
"content": "Around $200 per night, preferably in Manhattan.",
},
{
"role": "assistant",
"content": "For Manhattan hotels around $200/night, you might want to look at options like Hotel Beacon, Pod 51, or CitizenM Times Square. These are well-rated options in that price range. Would you like more specific recommendations for any of these?",
},
]
metric = ConversationalCoherenceMetric(model="gpt-5-nano", window_size=8, include_reason=True)
result = metric.score(conversation)
print(result.value)
print(result.reason)SessionCompletenessQuality
Section titled “SessionCompletenessQuality”SessionCompletenessQuality captures whether a conversation fulfilled the user’s top-level goals. The metric asks an LLM to extract intentions from the thread, judge completion, and aggregate the results.
from opik.evaluation.metrics import SessionCompletenessQuality
conversation = [
{
"role": "user",
"content": "I need to book a flight to New York and find a hotel.",
},
{
"role": "assistant",
"content": "I can help you with that. For flights to New York, what dates are you looking to travel?",
},
{
"role": "user",
"content": "Next weekend, from Friday to Sunday.",
},
{
"role": "assistant",
"content": "Great! I recommend checking airlines like Delta, United, or JetBlue for flights to New York next weekend. For hotels, what's your budget range and preferred location in New York?",
},
{
"role": "user",
"content": "Around $200 per night, preferably in Manhattan.",
},
{
"role": "assistant",
"content": "For Manhattan hotels around $200/night, you might want to look at options like Hotel Beacon, Pod 51, or CitizenM Times Square. These are well-rated options in that price range. Would you like more specific recommendations for any of these?",
},
]
metric = SessionCompletenessQuality(model="gpt-5-nano")
result = metric.score(conversation)
print(result.value)
print(result.reason)UserFrustrationMetric
Section titled “UserFrustrationMetric”UserFrustrationMetric estimates how likely it is that the user became frustrated (e.g. because of repetition or ignored requests). It scans windows of the conversation with an LLM and reports a value between 0.0 (not frustrated) and 1.0 (very frustrated).
from opik.evaluation.metrics import UserFrustrationMetric
conversation = [
{
"role": "user",
"content": "How do I center a div using CSS?",
},
{
"role": "assistant",
"content": "There are many ways to center elements in CSS.",
},
{
"role": "user",
"content": "Okay... can you show me one?",
},
{
"role": "assistant",
"content": "Sure. It depends on the context — are you centering horizontally, vertically, or both?",
},
{
"role": "user",
"content": "Both. Just give me a basic example.",
},
{
"role": "assistant",
"content": "Alright. You can use flexbox, grid, or margin auto. All of them work well.",
},
{
"role": "user",
"content": "Could you please just write the code?",
},
{
"role": "assistant",
"content": "Here’s one way: set the container to display:flex and then add justify-content and align-items so the child centers both ways.",
},
{
"role": "user",
"content": "But this doesn’t even center anything! This is incomplete.",
},
{
"role": "assistant",
"content": "You're right. You also need `justify-content` and `align-items`.",
},
]
metric = UserFrustrationMetric(include_reason=True, model="gpt-5-nano")
result = metric.score(conversation)
print(result.value)
print(result.reason)Using retrieved context
Section titled “Using retrieved context”ConversationalCoherenceMetric is context aware. When an agent message carries the documents it was generated
from, they are shown to the judge right next to that message, so it can see each answer alongside what the agent had
available when it wrote it. What the metric measures does not change - every turn is still judged on whether its
answer is relevant to the turns preceding it. Populate the documents with the trace_context_transform argument of
evaluate_threads, or attach them directly to the agent messages under a context key:
from opik.evaluation.metrics import ConversationalCoherenceMetric
conversation = [
{
"role": "user",
"content": "What is the fee for going overdrawn?",
},
{
"role": "assistant",
"content": "It is 5% of the overdrawn amount, charged monthly.",
"context": ["The overdraft fee is 5% of the overdrawn amount, charged monthly."],
},
]
metric = ConversationalCoherenceMetric(model="gpt-5-nano")
result = metric.score(conversation)
print(result.value)A window in which no message carries documents is evaluated with the original prompt, so a conversation carrying no context scores exactly as it did before.
Next steps
Section titled “Next steps”- Read more about conversational threads evaluation
- Learn how to create custom conversation metrics