For grading at real scale — a full cohort's worth of answers on one milestone — the review screen has two tools beyond grading one card at a time: Validate All with A.I., and a CSV round-trip for grading in a spreadsheet. The basics of the review screen are in Grading & AI-Assisted Feedback.
Validate All with A.I.
On a milestone with Validate on, in an Adventure with a Claude key, the review screen has an A.I. Review panel showing how many submissions the A.I. has decided so far. Validate All with A.I. shows you a plan before anything is spent:
- The counts — submissions, already reviewed, to judge, already judged (free), and nothing written.
- What it will judge against — the milestone's content, then each question's prompt, strictness and reviewer context. A question with no reviewer context is flagged, because without one the A.I. only judges whether an answer is a genuine attempt, not whether it's right. This is the moment to catch that, not afterwards.
- Validate at or above — the A.I. grade (70 by default) at which a submission is validated. Below it, the submission is invalidated.
- An option to leave alone any submissions that already have a comment from a human. Untick it and the A.I.'s comment replaces theirs.
- The cost — an estimate for the submissions it hasn't judged before, either measured on this milestone or estimated from the prompt that will be sent. Ones it has judged before are reused and cost nothing.
Press Review N submission(s) and a console reports each one as it goes: validated, invalidated, skipped, or left for a human. It ends with a tally and the approximate cost.
Stopping, Resuming, and One Run at a Time
- Cancel stops the work, not just the window, and everything decided so far is saved.
- Resuming is automatic. Run it again and anything already reviewed is skipped. A verdict that was bought but not applied is reused rather than paid for again. Each verdict is tied to the exact answer it judged, so an answer the player has since edited gets judged afresh.
- One run per milestone. If another GM is already running one, you're told who. Starting again in a second tab of your own takes the run over.
Nothing Applied on Doubt
An API error, an unreadable reply, or a submission with nothing written is recorded and left for a human, never quietly validated. A temporary error, such as a rate limit, is retried up to three times. Three failures in a row stop the run. An error that would repeat for every submission, like a rejected key, stops it straight away.
How the A.I.'s Grade Is Written
- No grade scale — only the validate or invalidate decision is recorded. The A.I.'s number is never stored as a score, so it can't shrink anyone's BLOO.
- Percentage — the A.I.'s grade is stored as given.
- GPA (letters) — the grade drops to the letter breakpoint at or below it, so an 82 records as a B.
The Cohort Summary
Once the A.I. has judged anything on a milestone, the A.I. Review panel shows how it graded the cohort. It refreshes when a run finishes:
- Decided (out of the total), validated, invalidated, the average grade, and overruled by a human.
- What it gave them — a spread across Weak (0–49), Borderline (50–69), Solid (70–84) and Strong (85–100).
- How many couldn't be judged, and how many were judged but not acted on yet.
- The strictness it ran at, the cost so far in dollars and tokens, and when it last ran.
Overruled is the number worth watching. It counts the times a person changed the decision after the A.I. made it. Everything else is the A.I. reporting its own tally back to you. A run that graded sensibly and a run that graded nonsense look the same from its side. If the overruled count keeps climbing, the step's strictness or reviewer context is set wrong, not the submissions.
The CSV Round-Trip
Download CSV, in the Milestone Data panel, exports every submission with the columns player_id, quest_id, adventure_id, player_name, email, grade, validation_status, comment and entry. That file is the import template. Grade and comment in your spreadsheet tool, then choose the file and press Upload Reviewed CSV. Only grade, validation_status and comment are read; everything else in the file is ignored.
- A blank cell leaves that value alone. A blank comment never wipes an existing one.
- Untouched rows don't count. A row where nothing actually changes isn't applied, and the result tells you how many entries were updated and how many rows were skipped.
- grade takes a number from 0 to 100, or a letter (A, A-, B+, B, B-, C+, C, D, F). It's ignored if the Adventure has no grade scale. On a Validate milestone, a grade above zero validates and a zero invalidates.
- validation_status validates only on the exact word
valid. Anything else in the cell —invalid, a typo — invalidates. It overrides whatever the grade implied on the same row, and it keeps an existing grade rather than resetting it. - A row whose
quest_idoradventure_idpoints at a different milestone is skipped rather than applied here.
Other Tools on the Review Screen
- Create Images Zip, then Download Zip — every uploaded image from the milestone's submissions in one archive.
- Pending Validation Reminders (CSV) — a plain list of the email addresses of players who've submitted but aren't validated yet, one per line, for a nudge from your own mail tool.
Grading Scale
The scale comes from Resource Mechanics in Adventure Settings: None, Percentage, or GPA (letters). Whichever you choose, it applies across the review screen, the A.I. and the CSV.