The End of the Red Pen: Inside NishithP2004/AEGIS

How a multi-agent pipeline uses Google ADK to bring peer review, quality assurance, and empathy to automated grading.

7 min read • View on GitHub • More from NishithP2004

A traditional red grading pen shattering into geometric nodes, representing the shift from analog to algorithmic grading.
AEGIS replaces the subjective red pen with a structured, multi-agent evaluation pipeline.
Key Takeaways

The One-Shot Grading Fallacy

Most AI grading tools are reckless. They operate on a "one-shot" fallacy. A developer passes a student's essay into an LLM prompt, asks for a score out of ten, and surfaces the immediate result. This approach lacks a chain of thought, internal pushback, and the nuanced deliberation that actual human grading requires.

AEGIS treats grading not as a strict calculation, but as a human process. It recognizes that subjective evaluation requires checks and balances. Instead of a single API call, the system orchestrates a multi-stage, stateful pipeline where different AI personas critique and refine the assessment before a final grade is issued.

Hiring a Virtual Academic Department

The architecture of AEGIS relies heavily on Google's Agent Development Kit (ADK). This is a significant technical choice. ADK provides the necessary scaffolding to build structured agentic workflows rather than chaotic autonomous agents.

The codebase enforces a strict separation of concerns. A FastAPI backend drives the agents, while a Streamlit frontend handles OCR and multimodal artifact ingestion. The backend defines four distinct personas: the Arbiter, the Scrutinizer, the Validator, and the Mentor. Together, they form a virtual academic department.

The multi-agent pipeline uses a LoopAgent for quality assurance, mimicking a head TA reviewing a junior TA's work.

The Validator's Veto

The technical climax of the project lies in agents/aegis_agent/agent.py. This is where the Scrutinizer and Validator form a continuous loop using the ADK's LoopAgent pattern. The Scrutinizer refines the Arbiter's initial pass, but the Validator has the final say.

A close-up of a mechanical eye scrutinizing a paper, while a larger mechanical eye hovers above, inspecting the first eye's notes.
The Validator acts as a quality assurance guardian, ensuring the grading meets strict standards before proceeding.

The Validator doesn't just check the work. It has a programmatic exit_loop tool. If the grading isn't up to standard, it rejects the initial assessment and forces the Scrutinizer to try again. This creates a self-correcting QA loop. The cycle only breaks when the Validator deems the evaluation perfect, saving compute and protecting the student from hallucinated penalties.

The Mentor in the Machine

A perfectly accurate grade is useless if it discourages the student. The final stage of the pipeline is the Mentor agent. It takes the cold, clinical JSON output of the Validator and translates it into encouraging, growth-mindset feedback.

A cold metal printing press churning out data ticker tape on the left, while warm human hands weave that tape into a soft tapestry on the right.
The Mentor agent translates rigid evaluation data into empathetic student feedback.
Standard AI Grader OutputAEGIS Mentor Agent Output
You scored 6/10. You missed the historical context in paragraph two and your conclusion is weak.Great start on your thesis! You have a strong grasp of the primary sources. To push this to the next level, let's look at how adding historical context to your second paragraph could strengthen your final argument.

This final translation step proves that AEGIS is built for education, not just data processing. By delegating empathy to a specific node in the pipeline, the system ensures that automated grading finally treats the student like a human being.