Task 2 — Evaluate answer quality
Part of the Observe, evaluate, and secure your agents lab. New here? Start with Getting started.
Set up (start here): This task needs a Foundry project, the starter code, and a grounded agent to measure. If you haven’t already, complete Getting started to create your project, clone the code, set
PROJECT_ENDPOINTandMODEL_DEPLOYMENT_NAMEinPython/.env, and either pointAGENT_NAMEat your Lab B agent or create one withpython ../setup/bootstrap_agent.py. Then, from thePythonfolder you opened in VS Code, verify you’re ready:
python ../setup/check_env.py --task 2
Continuing from a previous task? If you just finished another task in the same
Pythonfolder, your project, virtual environment, and.envare already set — go straight to Look at the dataset first below.
The Caldova knowledge agent answers questions about capacity, contract manufacturers, and suppliers. It sounds confident every time. That’s the problem: you can read ten answers, feel good about them, and still have no idea whether the eleventh invents a returns window that doesn’t exist.
Evaluation replaces that feeling with a number. You take a set of questions you already know the right answers to, run them through the agent, and have a second model grade what comes back.
What are groundedness, relevance and similarity?
They’re three different ways an answer can be wrong, so they’re measured separately:
- Groundedness — is the answer supported by the context it was given? A low score means the agent made something up, even if what it made up happens to be true.
- Relevance — does the answer actually address the question? An answer can be perfectly grounded and still not what was asked.
- Similarity — how close is the answer to the ground truth you wrote? This is the one that needs a human-authored correct answer.
Each is scored 1–5 by a model acting as a judge. Groundedness and relevance also return a
reason, which is usually more useful than the score.
Open the Python folder and activate the virtual environment from Getting started (.\labenv\Scripts\Activate.ps1), then continue below.
Look at the dataset first
Open data/caldova_eval.jsonl. Each line is one test case, and evaluation is only ever as good as this file:
{"query": "A planner wants to raise a capacity request with a complete program brief. How long does review take?", "context": "Standard Request Window: Requests with a complete program brief: reviewed within 5 business days ...", "ground_truth": "Five business days for a request with a complete program brief ..."}
query— what you ask the agent.context— the source material the answer should be based on. Groundedness grades against this.ground_truth— the answer a knowledgeable human would give. Similarity grades against this.
The agent’s own response isn’t in the file: you generate it at evaluation time by pointing
the evaluation at a target.
Open agent_target.py and read it — you don’t edit it. It’s a callable class that takes one
query and returns {"response": ...} from your agent. That’s the whole contract a target has
to meet.
Write the evaluation
Open evaluate_agent.py and add code at each commented placeholder.
-
Add references:
# Add references from azure.ai.evaluation import ( AzureOpenAIModelConfiguration, GroundednessEvaluator, RelevanceEvaluator, SimilarityEvaluator, evaluate, ) from agent_target import CaldovaAgentTarget -
Configure the model that grades the answers — the evaluators are themselves model calls, so they need a deployment to run on. You’ll reuse the same one your agent uses:
# Configure the model that grades the answers model_config = AzureOpenAIModelConfiguration( azure_endpoint=evaluator_endpoint(), azure_deployment=model_deployment, api_version=os.getenv("AZURE_OPENAI_API_VERSION", "2024-10-21"), )No API key: with
az logindone, the evaluators authenticate as you.evaluator_endpoint()is provided at the top of the file — it derives the resource endpoint from yourPROJECT_ENDPOINT, or usesAZURE_OPENAI_ENDPOINTif you set one. -
Create the evaluators:
# Create the evaluators groundedness = GroundednessEvaluator(model_config) relevance = RelevanceEvaluator(model_config) similarity = SimilarityEvaluator(model_config) -
Run the evaluation —
evaluate()reads the dataset, calls the target once per row, and passes each evaluator exactly the columns it needs.column_mappingis how you say which column is which:${data.x}comes from the file,${target.x}comes back from the target:# Run the evaluation result = evaluate( data=str(DATASET), target=CaldovaAgentTarget(), evaluators={ "groundedness": groundedness, "relevance": relevance, "similarity": similarity, }, evaluator_config={ "groundedness": { "column_mapping": { "query": "${data.query}", "context": "${data.context}", "response": "${target.response}", } }, "relevance": { "column_mapping": { "query": "${data.query}", "response": "${target.response}", } }, "similarity": { "column_mapping": { "query": "${data.query}", "ground_truth": "${data.ground_truth}", "response": "${target.response}", } }, }, azure_ai_project=project_endpoint, output_path=str(OUTPUT), )azure_ai_projectis optional. Passing it uploads the run to your project so you can see the results in the portal alongside everything else. -
Print the aggregate scores:
# Print the aggregate scores print("\nAggregate scores (1-5, higher is better):") print(json.dumps(result["metrics"], indent=2)) print(f"\nRow-level detail: {OUTPUT.resolve()}") if result.get("studio_url"): print(f"View in the Foundry portal: {result['studio_url']}") -
Save the file (Ctrl+S).
Run and test
-
In the terminal, sign in and run the evaluation:
az loginpython evaluate_agent.py -
It asks the agent all ten questions and then grades each answer three ways, so give it a couple of minutes. You should see something like:
Aggregate scores (1-5, higher is better): { "groundedness.groundedness": 4.6, "relevance.relevance": 4.4, "similarity.similarity": 4.1 } Row-level detail: ...\eval_results.jsonYour numbers will differ. Exact values matter far less than being able to reproduce them after a change.
-
Open eval_results.json and find the lowest-scoring row. Read its
groundedness_reasonorrelevance_reason— the judge explains itself, and that explanation is what you’d act on. -
Decide whether the score is the agent’s fault. Sometimes a low similarity score means the agent gave a better answer than your ground truth, or that your
contextwas too thin for the question. Evaluation grades your dataset as much as your agent.
Make a change and prove it
This is what evaluation is actually for.
-
Add a bad test case to the end of data/caldova_eval.jsonl — a question whose
contextdoesn’t support theground_truth, so the agent has nothing to answer from:{"query": "What is the fast-track window for Halden Biologics?", "context": "Planning Desk Hours: Monday-Friday 8:00 AM - 6:00 PM.", "ground_truth": "Four months to first commercial batch."} -
Run
python evaluate_agent.pyagain and look at that row. Groundedness should drop sharply even if the agent’s answer is correct — because the answer isn’t supported by the context it was given. That distinction is exactly what a hallucination looks like in a metric. -
Remove the row again when you’re done.
✅ Checkpoint: You have a repeatable score for your agent’s answers, row-level reasons for every score, and a way to tell whether tomorrow’s prompt change made things better or worse. That’s the Core of this lab — the remaining task is optional.
When you’re finished, enter deactivate to exit the virtual environment.
Next (optional): Task 3 — Red team your agent