The End of the Red Pen: Inside NishithP2004/AEGIS
How a multi-agent pipeline uses Google ADK to bring peer review, quality assurance, and empathy to automated grading.
- AEGIS replaces fragile one-shot LLM grading with a multi-agent academic department built on Google's Agent Development Kit.
- The system uses a LoopAgent architecture, allowing a Validator agent to programmatically reject and re-route subpar grading before the student sees it.
- A dedicated Mentor agent translates cold, clinical JSON evaluation metrics into empathetic, growth-oriented student feedback.
The One-Shot Grading Fallacy
Most AI grading tools are reckless. They operate on a "one-shot" fallacy. A developer passes a student's essay into an LLM prompt, asks for a score out of ten, and surfaces the immediate result. This approach lacks a chain of thought, internal pushback, and the nuanced deliberation that actual human grading requires.
AEGIS treats grading not as a strict calculation, but as a human process. It recognizes that subjective evaluation requires checks and balances. Instead of a single API call, the system orchestrates a multi-stage, stateful pipeline where different AI personas critique and refine the assessment before a final grade is issued.
Hiring a Virtual Academic Department
The architecture of AEGIS relies heavily on Google's Agent Development Kit (ADK). This is a significant technical choice. ADK provides the necessary scaffolding to build structured agentic workflows rather than chaotic autonomous agents.
The codebase enforces a strict separation of concerns. A FastAPI backend drives the agents, while a Streamlit frontend handles OCR and multimodal artifact ingestion. The backend defines four distinct personas: the Arbiter, the Scrutinizer, the Validator, and the Mentor. Together, they form a virtual academic department.
The Validator's Veto
The technical climax of the project lies in agents/aegis_agent/agent.py. This is where the Scrutinizer and Validator form a continuous loop using the ADK's LoopAgent pattern. The Scrutinizer refines the Arbiter's initial pass, but the Validator has the final say.
The Validator doesn't just check the work. It has a programmatic exit_loop tool. If the grading isn't up to standard, it rejects the initial assessment and forces the Scrutinizer to try again. This creates a self-correcting QA loop. The cycle only breaks when the Validator deems the evaluation perfect, saving compute and protecting the student from hallucinated penalties.
The Mentor in the Machine
A perfectly accurate grade is useless if it discourages the student. The final stage of the pipeline is the Mentor agent. It takes the cold, clinical JSON output of the Validator and translates it into encouraging, growth-mindset feedback.
| Standard AI Grader Output | AEGIS Mentor Agent Output |
|---|---|
| You scored 6/10. You missed the historical context in paragraph two and your conclusion is weak. | Great start on your thesis! You have a strong grasp of the primary sources. To push this to the next level, let's look at how adding historical context to your second paragraph could strengthen your final argument. |
This final translation step proves that AEGIS is built for education, not just data processing. By delegating empathy to a specific node in the pipeline, the system ensures that automated grading finally treats the student like a human being.