Observe, evaluate, and secure your agents

Level ▰▰▰▱▱ L300 (L100 beginner → L500 expert)

You can build an agent in an afternoon. Knowing whether it’s any good — and whether it behaves when someone tries to make it misbehave — is a different job. This lab is about that job: seeing inside a running agent, measuring the quality of its answers, and attacking it before someone else does.

Anton
Meet Anton, your AI guide.
You’ll spot Ask Anton tips throughout this lab. Want more interactive, hands-on help? Chat with Anton in the Ask Anton app.

About the Ask Anton app Ask Anton is a generative AI agent that can answer questions about AI concepts and Microsoft Foundry technologies. It's available in two versions at https://aka.ms/choose-anton:
  • Azure-based: Best experience (requires an Azure subscription and deployment of a model in a Foundry project).
  • Browser-based: Use a small language model in your browser (reduced functionality - may be slow or work only in "basic" mode in older/lower-spec devices).
Ask Anton is not a supported Microsoft product or a component of Microsoft Learn or AI Skills Navigator.
Why can't I just read the output?

Because the output is the one part of an agent that looks fine when everything else isn’t. An answer can be fluent and confidently wrong, or correct but produced by three retries and a tool call that timed out. Tracing shows you what happened on the way to the answer. Evaluation scores the answer against something you already know to be true. Red teaming tells you what the agent does when the question is hostile.

Your scenario: you work at Caldova, a pharmaceutical manufacturer preparing an accelerated product launch. The supply chain assistant you built in earlier labs is now answering real questions from planning teams — and the IT compliance lead is asking harder questions about it. Why was that answer slow? Is it making things up about the capacity policy? What happens if someone tries to talk it into something it shouldn’t say? In this lab you answer all three with evidence rather than opinion.

You’ll start with the Core tasks, which get you from “it runs” to “I can prove how well it runs”. The Optional task then goes after safety.

Note: Some of the technologies used in this exercise are in preview or in active development. You may experience some unexpected behavior, warnings, or errors.

What you’ll learn

By completing the Core tasks of this exercise, you’ll be able to:

  • Trace an agent with OpenTelemetry, export the traces to Azure Monitor, and read them — Foundry’s own server-side trace and your custom spans are both there, in different views.
  • Evaluate answer quality against ground truth using built-in evaluators (groundedness, relevance, similarity) and a JSONL dataset.

The Optional task lets you additionally:

  • Red team your agent with the AI Red Teaming Agent: run adversarial attack strategies and your own seed prompts against a deployed agent, and read the attack success rate.

How this lab is organized

This lab is modular. Each task is written to be completed on its own, starting fresh — so you can pick a single task and do just that one. Every task also shares one starter folder, one virtual environment, and one .env, so if you’d rather work straight through, you can.

  1. Start with Getting started — create your Microsoft Foundry project, connect Application Insights, get the starter code, and set up your .env. Every task begins from here; if you’re doing the whole lab in one sitting, you only need to do this once.
  2. Do any task. Each task lists the setup it needs so you can start it independently. If you’re moving straight from the previous task, a short “Continuing from a previous task?” note at the top lets you skip the repeated setup and keep going.

Lab at a glance

Complete the Core tasks first — they end with an agent you can see inside and a scorecard for its answers. Then add the Optional task if you want to test how it stands up to attack.

Section Task Level Time
Core Task 1 – Trace your agent ▰▰▰▱▱ L300 ~25 min
Core Task 2 – Evaluate answer quality ▰▰▰▱▱ L300 ~35 min
Optional Task 3 – Red team your agent ▰▰▰▰▱ L400 ~35 min

Core tasks: about 60 minutes. Full lab, including every optional task: about 1 hour 35 minutes.

Choosing your path — pick the tasks that fit the time you have:

  • Core only (~1h): do Tasks 1–2.
  • Everything (~1h 35m): add Task 3, the red team scan.

One agent, three questions: all three tasks point at the same grounded knowledge agent — Task 1 traces it, Tasks 2 and 3 measure it. If you haven’t done Lab B, one command creates an equivalent agent so this lab stands alone — see Getting started.

Measure, don’t guess

The three techniques in this lab answer different questions, and it’s worth being clear about which is which:

  • Tracing answers “what happened?” It’s a record of one run: which spans took how long, which tools were called, what the model was sent. Use it when something is slow or broke.
  • Evaluation answers “how good is it, on average?” It’s a score over a dataset, so it’s the only one of the three that tells you whether a change made things better or worse.
  • Red teaming answers “what can I make it do?” It’s an adversarial probe, and a clean result is a floor, not a guarantee.

None of them replaces the others, and all three are cheap compared to finding out in production.

Summary

Across this lab you:

  • Instrumented an agent with OpenTelemetry and exported traces to Application Insights — reading Foundry’s automatic server-side trace and your own custom spans, side by side.
  • Evaluated a grounded agent against a ground-truth dataset with built-in groundedness, relevance and similarity evaluators, and got a score you can compare across changes.
  • (Optionally) Red teamed the agent with adversarial attack strategies and your own seed prompts, and read the resulting attack success rate.

Together these turn “the demo worked” into evidence you can show someone.

Clean up

If you’re finished, delete the resources you created to avoid unnecessary Azure costs.

  1. In the Azure portal, navigate to the resource group that contains your Foundry resource.
  2. On the toolbar, select Delete resource group, enter the resource group name, and confirm.

All three tasks measure the same caldova-knowledge-agent, so deleting the resource group removes it along with everything else. If you provisioned with azd, run azd down instead — but note that Application Insights, if you created it from the Foundry portal, is a separate resource and is deleted with the resource group rather than by azd.