Skip to content

How you're judged 🚧 ​

Still being finalized

How projects are judged isn't locked yet. Check back closer to the event - this page will lay out exactly what we're looking for.

What we care about ​

This isn't about polish or points. The goal is a working agent you understand well enough to keep using. Broadly, a strong result:

  • Works - it runs on a fresh prompt and produces something useful for the scenario.
  • You can explain it - you know which rule, sample, or boundary drives its behavior.
  • Shows the change - a before, a fix, and an after.

Exact criteria and any scoring will land here before hack day.