Skip to main content
AgentAddaAgentAdda
All Articles
AI in EducationShunyaSaarthiGuruSaarthiAssessmentResponsible AI

When AI Gives the Answer, What Are We Assessing?

An Agent Adda point of view on governed AI in education: teacher-authored pedagogy, learning provenance, and the independent capability students retain after the conversation ends.

22 September 202610 min read·AgentAdda Collective

Open the designed HTML edition.

A student submits a correct solution. The reasoning is fluent, the equations are tidy, and the conclusion is convincing.

What has the institution learned about the student?

If an AI system produced most of the work, the answer may be: much less than the submission suggests. Yet prohibiting AI in every learning activity would also discard opportunities to explain difficult ideas, challenge misconceptions and practise with feedback.

The educational task is to decide where assistance belongs, what students must contribute, and how their understanding will be established. That demands a more precise institutional response than a single rule about whether AI is allowed.

Agent Adda's position is that governed AI should help institutions develop and assess two capabilities: what students can do independently, and how responsibly they can work with AI.

This is the direction we propose for ShunyaSaarthi and its GuruSaarthi pedagogy layer. It is a product thesis to test, not a claim that the complete system or its educational effectiveness has already been demonstrated.

An unresolved tragedy requires restraint

The death of IIT Bombay student Sahil Wakode has brought this debate into painful public focus. The institute's initial account referred to alleged ChatGPT use during an examination. On September 21, it apologised for its earlier communication and withdrew its characterisation of events before the investigation had established them. Reporting on September 22 described a continuing Mumbai Crime Branch investigation. The New Indian Express, September 21; The Times of India, September 22.

That record does not establish that AI use, a ban, or any single reported incident caused his death. The circumstances and contested accounts deserve investigation, and his family deserves privacy and respect.

This tragedy cannot serve as evidence that a particular education product would have prevented it. Our argument for governed AI stands separately. Institutions owe students clear expectations, fair procedures and accessible human support, whatever technology is involved.

Different failures demand different responses

Research offers firmer ground for examining educational design.

In a randomised mathematics study involving nearly 1,000 high-school students in Turkey, access to a general-purpose GPT-based interface improved performance during assisted practice. But when assistance was removed, that group performed 17% worse than the control group on the subsequent exam—a relative difference, not 17 percentage points. A teacher-informed version designed to guide students largely removed that deficit, but did not establish a positive effect on unaided exam performance. These were short-term results in a particular setting, not proof about every AI tutor. Bastani and colleagues, Generative AI without guardrails can harm learning.

The distinction matters: successful task completion and learning are different outcomes. A classroom can appear more productive while students become less able to perform without help.

Assessment presents another problem. In a 2024 University of Reading study, researchers inserted AI-written submissions into examinations across five undergraduate psychology modules; 94% went undetected. That was a controlled test of a particular examination system, not an estimate of how many students cheat. It shows why a finished answer alone can be insufficient evidence. Scarfe and colleagues, PLOS ONE.

Detection introduces its own risks. In August 2023, Vanderbilt University disabled Turnitin's AI detector, explaining concerns about false positives, transparency and fairness. This historical decision does not establish the accuracy of every detector available today. It does support an institutional caution: a software signal should never, by itself, become a verdict against a student. Vanderbilt's published guidance.

Beyond education, the 2023 Mata v. Avianca sanctions order illustrates the consequences of trusting fluent output without verification. Lawyers submitted fictitious authorities generated by ChatGPT and continued defending them after their validity was questioned. The failure involved professional verification and candour, not simply the presence of AI. The court's sanctions order.

These situations have different mechanisms: dependence on assistance, weak assessment evidence, unreliable suspicion and unverified claims. An education system needs a response to each.

Start with a clear learning agreement

Our proposed role for ShunyaSaarthi is the institutional learning environment: the place where curriculum scope, permitted assistance, evidence, privacy and human responsibility are made explicit.

GuruSaarthi is the teacher-authored pedagogy layer within that environment. It translates teaching judgment into instructions and activities that guide the student–AI interaction.

For every activity, the teacher should be able to answer four questions: What is the student learning? What help is permitted? What evidence is required? Who makes the final judgment?

The same learner may move through three clearly labelled modes:

ActivityPermitted assistanceEvidence that matters
ExplorationExplanations, examples, references and open questionsThe learner's questions, comparisons and source checks
Guided practiceHints and explanations within the teacher's pathwayInitial attempts, corrections and the assistance needed
Independent assessmentNo solution-generating help; declared accessibility accommodations remain availableA fresh solution, explanation or demonstration by the learner

AI-free assessment remains legitimate when independent performance is the objective. AI-assisted assessment is also legitimate when the objective includes evaluating suggestions, checking sources or using tools responsibly. Students should know the distinction before they begin, and have equitable access to whatever tools an activity requires.

Make teaching judgment executable

Consider an illustrative mechanics lesson about friction.

A student receives an incline problem and three proposed solutions. One reaches the correct numerical answer through faulty reasoning. The task is to identify the sound solution, explain the others' errors and defend the choice. GuruSaarthi and the textbook are permitted during practice.

The teacher's pathway might specify:

Ask for a free-body diagram before introducing equations. If the student assumes friction always equals μN, ask whether the surfaces are sliding and whether static friction has reached its limiting value. Offer a worked explanation when repeated hints are no longer helping. Finish with a changed situation the student must solve independently.

If the student immediately asks which solution is correct, the tutor begins by eliciting their model of the forces. When the student makes an error, it offers a targeted question or counterexample. When the student understands, assistance recedes.

Socratic teaching must include judgment about when to explain. Endless questioning can frustrate a learner who lacks a prerequisite. A useful pathway specifies when to probe, when to model a step, when to provide an explanation and when to involve the teacher.

This is what teacher-authored pedagogy could capture: prerequisites, likely misconceptions, permissible hints, branching questions, escalation conditions and evidence of mastery. Prompts alone cannot guarantee compliance; the system would also need tested controls, reliable reference material and teacher review of failures.

Learning provenance should make assessment more useful

The resulting record should distinguish the quality of the answer, the assistance used and the capability subsequently demonstrated.

An illustrative teacher summary might read:

The student corrected the friction model after two conceptual hints. On a new incline problem, they independently selected the appropriate model and explained why the limiting-friction equation did not apply.

That gives a teacher something concrete to examine. It also makes room for a student who needed considerable support during practice but can now work independently. Asking for help should not automatically reduce a learning judgment.

We call this learning provenance: a bounded record of relevant attempts, feedback, revisions and demonstrations. It is evidence available for review, not direct access to a student's thinking, proof of authorship or a guarantee that outside assistance was absent.

The limits are important. Students can rehearse explanations or use other tools. A log can be incomplete, and an AI summary can misread it. Higher-stakes decisions need independent tasks, proportionate human review and an opportunity for the student to challenge the record. Oral follow-ups should be accessible and assessed against a clear rubric, rather than becoming surprise interrogations.

Collecting every conversation would create unnecessary risk. Students should know what is recorded, why, who can see it and when it will be deleted. Personal disclosures and unrelated exploration should not become a permanent disciplinary dossier. Teachers need a concise, inspectable summary—not hundreds of transcripts to read.

Human responsibility cannot be delegated

Institutional governance includes the way people are treated when something goes wrong. Suspected misconduct calls for evidence, an opportunity to respond and proportionate decisions. The teacher retains final grading authority; discipline and appeals belong to accountable human processes.

A learning assistant also needs boundaries around its relationship with students. In September 2025, the US Federal Trade Commission launched an inquiry into companion-style chatbots, including how companies evaluate their effects on children and teenagers. An inquiry is not a finding that the companies investigated caused harm. It nevertheless identifies questions that educational deployments should take seriously. FTC announcement.

Our design position is that an educational assistant should identify itself honestly, encourage appropriate human relationships and help the learner become more independent. It should not cultivate exclusivity or treat prolonged emotional dependence as success. Human support routes must be clear; a tutor is not a substitute for qualified care, and no system should promise to detect every crisis.

A pedagogy marketplace must earn its credibility

GuruSaarthi's longer-term opportunity is to let teachers share this instructional judgment.

A teacher could publish a reviewed friction-diagnostic pathway with its intended level, prerequisites, misconceptions and mastery checks. Another could adapt it for a different class, language or learning need. Each version would preserve authorship and make its assumptions visible.

This would be a pedagogy marketplace, where discovery is organised around teaching problems and evidence of usefulness. Popularity alone would be a weak quality signal. A pathway needs review, version history, failure reporting and evidence about the learners and contexts for which it works. Teacher expertise requires attribution and a credible incentive to contribute.

The difficult question is whether these pathways actually improve learning at an acceptable workload. That must be measured.

The first test should be a classroom, not a claim

Begin with one bounded activity: an initial attempt, teacher-guided AI practice, a fresh independent task and a short explanation of the student's reasoning. Compare outcomes with a suitable existing teaching approach, and include a delayed check rather than measuring only immediate completion.

Track independent transfer, persistent misconceptions, teacher review time, accessibility, student experience and the frequency of incorrect or inappropriate tutor guidance. More messages and longer sessions are not sufficient success measures. Neither is an impressive demonstration.

If independent performance fails to improve, the learning claim must narrow. If the review burden is unreasonable, the workflow must change. If students experience the record as surveillance, its collection and use need redesign.

The promise of ShunyaSaarthi and GuruSaarthi is a learning environment in which assistance has a purpose, teaching expertise shapes the interaction and educational judgments remain accountable.

The central question is simple:

After working with AI, what can the student understand, verify and do for themselves?


Agent Adda Point of View · Education & Responsible AI · Evidence cutoff: September 22, 2026.

Disclosure: This article presents Agent Adda's proposed direction for ShunyaSaarthi and GuruSaarthi. Classroom dialogue and teacher summaries are illustrative. The complete workflows, marketplace and educational outcomes described here are proposals to evaluate, not claims of fully deployed or independently validated capability. The IIT Bombay investigation remains unresolved at the stated cutoff.

AgentAdda is a collective of data practitioners sharing honest insights on AI, data engineering, and enterprise transformation.

Back to All Articles