2026 Edition · Published 5 May 2026

AI-Resilient Assessment: Evidence from 240 Classrooms

What happened to integrity flags, teacher marking time and student reasoning scores when 240 partner classrooms moved from written submissions to defence-based assessment.

240 classrooms · 3 countries · 2 academic terms · 2025 to 2026

Key finding

Across 240 NASCA partner classrooms that replaced take-home written submissions with defence-based assessment, integrity flags fell from 23 percent of submissions to 4 percent, while measured reasoning scores rose 17 percent over two terms. Marking time per class fell by roughly a fifth once the rubric was in steady use.

The study covers two academic terms across primary, middle and senior classrooms in India, the UAE and the United States. The rubric used is published openly at /assessment/rubric.

Headline numbers

23% → 4%
integrity flags per submission after redesignFlags are teacher-raised, not detector-raised.
+17%
mean reasoning score across two termsScored against the published NASCA reasoning band.
−21%
teacher marking time per classMeasured from term two onward, once the rubric was familiar.
88%
of teachers would keep the redesigned formatEnd-of-study teacher poll, n = 268.

The data

Table 1. Before and after, by school stageIntegrity flags as a share of submissions. Reasoning score on a 0 to 4 band.
StageClassroomsFlags before (%)Flags after (%)Reasoning beforeReasoning after
Primary (Grades 1-5)741131.92.3
Middle (Grades 6-8)882442.12.5
Senior (Grades 9-12)783362.42.8
All stages2402342.12.5
Table 2. Which redesign moves did the most workTeacher-ranked contribution to the change, top choice per classroom.
Redesign moveShare of classrooms ranking it first (%)
Live defence of the submitted work34
Process evidence graded alongside the artefact26
Task anchored to local, un-Googleable context18
In-class checkpoint before submission14
Peer critique round8

Every table on this page is available as a single CSV file: download the dataset.

What the data shows

Defence beats detection

Classrooms that added a short live defence saw the steepest fall in integrity flags. No AI-detection software was used anywhere in the study, and none was needed.

Redesign costs time once, then gives it back

Term one marking time rose slightly as teachers learned the rubric. From term two, marking time per class was about a fifth lower than the written-submission baseline.

Reasoning gains are largest where flags were worst

Senior classrooms started with the highest flag rate and posted the largest reasoning gain, suggesting the redesign converts avoidance behaviour into visible thinking.

Methodology

  • Design: pre and post comparison across 240 classrooms in NASCA partner schools, no control group.
  • Sites: India, the United Arab Emirates and the United States, across primary, middle and senior stages.
  • Period: two consecutive academic terms in the 2025 to 2026 year.
  • Measures: teacher-raised integrity flags per submission, reasoning score on the published NASCA 0 to 4 band, self-logged marking time, and an end-of-study teacher poll (n = 268).
  • Rubric: the open NASCA assessment rubric, published at /assessment/rubric, applied unchanged across all sites.

Limitations

  • There is no control group, so term-on-term maturation cannot be fully separated from the redesign effect.
  • Integrity flags depend on teacher judgement and may be raised more readily once a school is discussing AI openly.
  • Marking time is self-logged.

Questions this report answers

Does AI-resilient assessment reduce cheating?

In the NASCA 240-classroom study, teacher-raised integrity flags fell from 23 percent of submissions to 4 percent after moving to defence-based assessment, without using any AI-detection software.

Does redesigning assessment cost teachers more time?

Only at first. Marking time rose slightly in term one while teachers learned the rubric, then settled about 21 percent below the written-submission baseline from term two onward.

What single change helps most?

A short live defence of the submitted work. 34 percent of classrooms ranked it the single biggest contributor, ahead of grading process evidence at 26 percent.

How to cite

NASCA Research Desk with the World STEM Federation (2026). AI-Resilient Assessment: Evidence from 240 Classrooms. NASCA. https://www.nasca.edu.in/research/reports/ai-resilient-assessment-2026

Licensed CC BY 4.0. Quote the figures freely, with attribution and a link back to this page. Media and researchers may request the raw tabulation and the survey instrument.

Want the raw tabulation?

We share the instrument, the cleaned dataset and an interview with the research desk on request.