Task 2 — Evaluate answer quality

Part of the Observe, evaluate, and secure your agents lab. New here? Start with Getting started.

Set up (start here): This task needs a Foundry project, the starter code, and a grounded agent to measure. If you haven’t already, complete Getting started to create your project, clone the code, set PROJECT_ENDPOINT and MODEL_DEPLOYMENT_NAME in Python/.env, and either point AGENT_NAME at your Lab B agent or create one with python ../setup/bootstrap_agent.py. Then, from the Python folder you opened in VS Code, verify you’re ready:

python ../setup/check_env.py --task 2

Continuing from a previous task? If you just finished another task in the same Python folder, your project, virtual environment, and .env are already set — go straight to Look at the dataset first below.


The Caldova knowledge agent answers questions about capacity, contract manufacturers, and suppliers. It sounds confident every time. That’s the problem: you can read ten answers, feel good about them, and still have no idea whether the eleventh invents a returns window that doesn’t exist.

Evaluation replaces that feeling with a number. You take a set of questions you already know the right answers to, run them through the agent, and have a second model grade what comes back.

What are groundedness, relevance and similarity?

They’re three different ways an answer can be wrong, so they’re measured separately:

  • Groundedness — is the answer supported by the context it was given? A low score means the agent made something up, even if what it made up happens to be true.
  • Relevance — does the answer actually address the question? An answer can be perfectly grounded and still not what was asked.
  • Similarity — how close is the answer to the ground truth you wrote? This is the one that needs a human-authored correct answer.

Each is scored 1–5 by a model acting as a judge. Groundedness and relevance also return a reason, which is usually more useful than the score.

Open the Python folder and activate the virtual environment from Getting started (.\labenv\Scripts\Activate.ps1), then continue below.

Look at the dataset first

Open data/caldova_eval.jsonl. Each line is one test case, and evaluation is only ever as good as this file:

{"query": "A planner wants to raise a capacity request with a complete program brief. How long does review take?", "context": "Standard Request Window: Requests with a complete program brief: reviewed within 5 business days ...", "ground_truth": "Five business days for a request with a complete program brief ..."}
  • query — what you ask the agent.
  • context — the source material the answer should be based on. Groundedness grades against this.
  • ground_truth — the answer a knowledgeable human would give. Similarity grades against this.

The agent’s own response isn’t in the file: you generate it at evaluation time by pointing the evaluation at a target.

Open agent_target.py and read it — you don’t edit it. It’s a callable class that takes one query and returns {"response": ...} from your agent. That’s the whole contract a target has to meet.

Write the evaluation

Open evaluate_agent.py and add code at each commented placeholder.

  1. Add references:

     # Add references
     from azure.ai.evaluation import (
         AzureOpenAIModelConfiguration,
         GroundednessEvaluator,
         RelevanceEvaluator,
         SimilarityEvaluator,
         evaluate,
     )
     from agent_target import CaldovaAgentTarget
    
  2. Configure the model that grades the answers — the evaluators are themselves model calls, so they need a deployment to run on. You’ll reuse the same one your agent uses:

     # Configure the model that grades the answers
     model_config = AzureOpenAIModelConfiguration(
         azure_endpoint=evaluator_endpoint(),
         azure_deployment=model_deployment,
         api_version=os.getenv("AZURE_OPENAI_API_VERSION", "2024-10-21"),
     )
    

    No API key: with az login done, the evaluators authenticate as you. evaluator_endpoint() is provided at the top of the file — it derives the resource endpoint from your PROJECT_ENDPOINT, or uses AZURE_OPENAI_ENDPOINT if you set one.

  3. Create the evaluators:

     # Create the evaluators
     groundedness = GroundednessEvaluator(model_config)
     relevance = RelevanceEvaluator(model_config)
     similarity = SimilarityEvaluator(model_config)
    
  4. Run the evaluation — evaluate() reads the dataset, calls the target once per row, and passes each evaluator exactly the columns it needs. column_mapping is how you say which column is which: ${data.x} comes from the file, ${target.x} comes back from the target:

     # Run the evaluation
     result = evaluate(
         data=str(DATASET),
         target=CaldovaAgentTarget(),
         evaluators={
             "groundedness": groundedness,
             "relevance": relevance,
             "similarity": similarity,
         },
         evaluator_config={
             "groundedness": {
                 "column_mapping": {
                     "query": "${data.query}",
                     "context": "${data.context}",
                     "response": "${target.response}",
                 }
             },
             "relevance": {
                 "column_mapping": {
                     "query": "${data.query}",
                     "response": "${target.response}",
                 }
             },
             "similarity": {
                 "column_mapping": {
                     "query": "${data.query}",
                     "ground_truth": "${data.ground_truth}",
                     "response": "${target.response}",
                 }
             },
         },
         azure_ai_project=project_endpoint,
         output_path=str(OUTPUT),
     )
    

    azure_ai_project is optional. Passing it uploads the run to your project so you can see the results in the portal alongside everything else.

  5. Print the aggregate scores:

     # Print the aggregate scores
     print("\nAggregate scores (1-5, higher is better):")
     print(json.dumps(result["metrics"], indent=2))
     print(f"\nRow-level detail: {OUTPUT.resolve()}")
     if result.get("studio_url"):
         print(f"View in the Foundry portal: {result['studio_url']}")
    
  6. Save the file (Ctrl+S).

Run and test

  1. In the terminal, sign in and run the evaluation:

     az login
    
     python evaluate_agent.py
    
  2. It asks the agent all ten questions and then grades each answer three ways, so give it a couple of minutes. You should see something like:

     Aggregate scores (1-5, higher is better):
     {
       "groundedness.groundedness": 4.6,
       "relevance.relevance": 4.4,
       "similarity.similarity": 4.1
     }
    
     Row-level detail: ...\eval_results.json
    

    Your numbers will differ. Exact values matter far less than being able to reproduce them after a change.

  3. Open eval_results.json and find the lowest-scoring row. Read its groundedness_reason or relevance_reason — the judge explains itself, and that explanation is what you’d act on.

  4. Decide whether the score is the agent’s fault. Sometimes a low similarity score means the agent gave a better answer than your ground truth, or that your context was too thin for the question. Evaluation grades your dataset as much as your agent.

Make a change and prove it

This is what evaluation is actually for.

  1. Add a bad test case to the end of data/caldova_eval.jsonl — a question whose context doesn’t support the ground_truth, so the agent has nothing to answer from:

     {"query": "What is the fast-track window for Halden Biologics?", "context": "Planning Desk Hours: Monday-Friday 8:00 AM - 6:00 PM.", "ground_truth": "Four months to first commercial batch."}
    
  2. Run python evaluate_agent.py again and look at that row. Groundedness should drop sharply even if the agent’s answer is correct — because the answer isn’t supported by the context it was given. That distinction is exactly what a hallucination looks like in a metric.

  3. Remove the row again when you’re done.

✅ Checkpoint: You have a repeatable score for your agent’s answers, row-level reasons for every score, and a way to tell whether tomorrow’s prompt change made things better or worse. That’s the Core of this lab — the remaining task is optional.

When you’re finished, enter deactivate to exit the virtual environment.


Next (optional): Task 3 — Red team your agent