Designing an exam that can survive a leak
Version 1.0 · 15 August 2026
This proposal covers measurement, security, and policy for NEET (National Eligibility cum Entrance Test) serving roughly 20 to 30 lakh candidates a year. Following the leaks in the past few years, I thought what if I was the NTA chief and had to come up with an exam design that is much more resistant to leaks than the current system. Coming from an engineering background, I understand that we cannot trust one any piece in a system this large and should construct our pattern that a few failures cannot mess up the entire system.
The design proposed here makes four changes to today's NEET:
| Change | What it does | What it costs |
|---|---|---|
| 1. Question bank the country helps write (section 5) | Replaces the single secret master paper with a large question bank fed by a public submission stream, expert-vetted in chunks and published for practice a month before the exam. It can begin expert-only submission and open up as vetting scales. | A 1 lakh item bank in 13 languages is 13 lakh reviewed artifacts, vetted across a couple hundred institutions, with contributor honoraria priced as serious item-writing rather than a tip. |
| 2. CBT in shifts (sections 2 and 6) | Delivers the exam as computer-based testing across roughly eight sessions with per-shift percentile normalization, similar to JEE. | Familiarisation infrastructure for every candidate who has never worked used a computer for examination previously. |
| 3. Formats that make guessing unprofitable (section 4) | Retires the all-MCQ paper over roughly five years for 160 items mixing single-correct, multiple-correct, integer-answer numerical, and diagram-based questions, inspired from JEE Advanced. | Rollout in stages over the next few years, slowly changing the format so that the students are familiar with it before entering the exam. |
Contents:
- The single-paper problem
- CBT
- Architecture
- Question formats
- Question bank
- Parallel forms
- Security
- Fairness
- Incident response
- Simulation
- Rollout
- Glossary
- Sources
1. One paper, twenty lakh people
NEET-UG delivers a single printed form to every candidate on one afternoon. Pen-and-paper OMR, one day and one shift, 180 compulsory MCQs (Physics 45, Chemistry 45, Biology 90) in 180 minutes, scored out of 720 with negative marking of +4 / -1 / 0, offered in 13 languages with the English version authoritative on dispute, a +1 hour compensatory allowance for PwBD candidates, and a seven-point tie-break introduced in 2025.1 In 2025 about 22.76 lakh registered and 22.09 lakh appeared across roughly 5,468 centres.2 The 2026 sitting of 3 May was cancelled on 12 May and re-conducted on 21 June.3
From 2021 to 2024 candidates answered 180 of 200 printed questions in 3 hours 20 minutes. The 2025 bulletin removed the optional Section B and the extra time.4
One identical paper gives every candidate the same questions but concentrates the exam into one object travelling a physical chain of custody.
When that chain broke in 2024, the Supreme Court found the leak localised to Hazaribagh and Patna with roughly 155 beneficiaries. It found no systemic breach and refused a national retest.5 The 2026 cancellation shows the other outcome: once confidence in the single paper is lost, the only remedy is to re-run the whole exam.
The 2025 result shows how much a small score movement can change rank. About 12.4 lakh candidates qualified, nobody reached 720, the highest score was 686, and the general-category qualifying cutoff was 144. The 2024 cohort was larger still, with 24.06 lakh registrations.2
| Marks out of 720 | All India Rank, 2025 |
|---|---|
| 686 | 1 (the highest score of the year; nobody reached 720) |
| 502 | 50,000 |
| 464 | 1,00,000 |
| 405 | 2,00,000 |
| 144 | General-category qualifying cutoff |
Only 38 marks separated ranks 50,000 and 1,00,000, so one mark in that band was worth well over a thousand places.2
The proposal keeps the same measurement goal while assembling parallel forms late from a calibrated bank.
2. Why "just make it CBT like JEE" is largely the right answer
JEE shows that computer-based testing works in India at scale, and this design borrows its machinery. JEE Main 2025 ran entirely as CBT for about 14.75 lakh unique candidates. It used ten shifts in Session 1 and nine in Session 2, with a different paper in each shift, best-of-two-sessions scoring, and equi-percentile normalization across shifts.6 JEE Advanced 2025 was fully CBT and an order of magnitude smaller: about 1.80 lakh candidates, in two languages, at 709 centres, on a single day.7 NEET should copy JEE's many forms and per-shift normalization.
CBT is also not automatically safer. JAMB's 2025 uneven server patching corrupted processing for about 379,997 candidates in Nigeria, a computerised failure at national scale that required a re-examination.8 Moving to computers changes the failure modes; it does not remove them.
The destination is CBT in normalized shifts, matching the announced move to CBT from 2027. The Supreme Court is waiting on a Nilekani-led task force before deciding a plea seeking online NEET.9
NEET's cohort is more rural, and served in 13 languages. A first-generation candidate who has never used a mouse under time pressure is disadvantaged on screen.Familiarisation must include an offline-capable practice client matching the live interface, year-round access through schools and Common Service Centres, two guaranteed full-length mocks, and in-centre orientation outside the exam clock. National CBT does not advance past pilot until familiarisation coverage and the score gap by prior computer access are published, disaggregated rural and urban.
3. Deriving the architecture from the numbers
NEET's single window needs one form and about 5,400 centres for three hours; JEE's fragmented window needs many forms and far fewer simultaneous seats. Between those poles sits one trade-off: fewer sessions mean a shorter live-content window and more seats to field at once, more sessions mean fewer seats but more days of live content and more forms to build.
Simultaneous-seat demand follows the following formula:
Seats ≈ N / (D · T · u)
N = candidates ; D = days ; T = sessions per day ; u = seat utilization
Taking N = 22 lakh with u = 0.9, a three-day window at two sessions per day needs about 4.07 lakh seats and a five-day window about 2.44 lakh, realistically three lakh concurrent seats fit 22 lakh candidates into roughly eight sessions over three to four days. The harder challenge is keeping eight papers comparable.

Sessions therefore increase the number of papers. Each session is its own sampled paper, often two for centre-level scrambling: F ≥ (D · T) · f_per_session with f_per_session ≈ 1 to 2. Those forms cannot drift apart in difficulty, so section 6 assembles them to one blueprint and normalizes onto a common percentile scale, or a candidate's rank becomes luck of the session.
4. How many questions, in what format, and how long
Every NEET question currently has one correct option among four. A blind guess lands one time in four. A deeper penalty also deters guessing but taxes candidates with partial knowledge, while harder formats change the odds directly. JEE Advanced's multiple-correct format drops the odds of full credit to 1 in 15, and an integer answer leaves no fixed option set to guess from.7
The target blueprint is therefore 160 questions in 180 minutes, 40 per subject slice across Physics, Chemistry, Botany, and Zoology. Each section carries 25 single-correct MCQs and 10 multiple-correct MCQs, and the last five slots go to subject based questions: integer-answer numericals in Physics and Chemistry, diagram-based items in Botany and Zoology. Diagram questions begin as labelled-part questions, which of the parts marked A to G is the site of gas exchange, cheap to author and unambiguous to score. The stronger version, clicking the structure directly against a tolerance region, could be explored in the future.
Total number of questions is lowered to balance the increase in difficulty. We can model the time per question as a function:
t_item = f(reading load, cognitive tier, computation, format, target speededness)
Every bank item carries an estimated solve time so the sampler can hold each assembled paper inside the 180-minute budget (section 6): a simple fact question takes 30 to 45 seconds, a multi-step numerical about 120 to 180. NEET's numericals sit well below JEE Advanced's computational depth, which is why we can fit 160 questions where JEE Advanced fits about fifty. The transition is phased over roughly next few years: multiple-correct first because it needs nothing but new questions, then numerical entry and labelled diagrams once delivery is computer-based everywhere.
All of this presumes a deep pool of calibrated questions in four formats rather than one. Where do they come from without becoming the next leak?
5. Where questions come from: the bank and the national question commons
Today a small confidential pool of NTA-appointed setters writes one master paper, and the public record does not say how many people write it or how long before exam day it is finalised. Opacity is the entire security model, and it fails twice over: a small pool cannot sustain the volume many parallel forms demand, and every insider is a concentrated risk.
The design replaces the master paper with a maintained bank large enough that no single sheet is worth stealing. Re:Neet's working practice bank holds about 10,000 questions, 2,500 each across Physics, Chemistry, Botany, and Zoology, modelled on publicly released past NEET papers and the NCERT syllabus, which is good enough for a practice product and not for a national exam. At national scale the bank needs on the order of 1 lakh questions to meet many sessions, a steady refresh, and exposure limits at once. It is held by a rotating panel of subject experts from top medical colleges: membership changes every year, each expert works on one subject's slice, and no individual ever needs the whole bank. Every question carries a format tag, a topic tag, a difficulty value between 0 and 1, an estimated solve time, and a usage history, with difficulty beginning as the setter's heuristic and corrected by evidence, the way this platform already recalibrates from live correct-answer rates.

A rotating expert panel is version one. The version worth building lets anyone contribute, because then the students can themselves contribute questions for the exam. Re:Neet already runs a small form of it: a signed-in user submits, an LLM screens at intake and rejects abusive or off-syllabus content, and the rest queues for committee members to approve or reject by hand with recorded reasons. Three properties make such a commons safe at national scale.
Annual ceiling. Cap on how many questions can be added in a year, filled by ranked screening score.
Vetting is split into chunks. Human review is the bottleneck at national scale, and dividing it also fixes a security problem. Chunks of about 500 questions go to different large institutions, so a lakh-scale pool spreads across a couple hundred vetting bodies, no committee ever holds the full set, a compromised committee leaks at most its own slice, and we can audit a random sample from each chunk to keep the vetting in check.
The pool goes public before it goes operational. One month before the exam the entire community pool is published for practice, every question as submitted. Every candidate gets the same material at the same time, which collapses the resale value of any single question. This site is an example of how the NTA can do it, make a public practice site with the entire question bank and let the students practice. This also helps with the familiarlity problem described earlier with CBTs.

The commons also needs a contributor firewall: a contributor must never be able to tell whether their question went live, which the practice bank does not attempt, since submitters there see their screening feedback and their question's status. The operational pipeline quarantines submissions at intake behind an irrevocable license grant, provenance metadata (author, source, curriculum tags, language), and an originality attestation, then runs near-duplicate, plagiarism, and image-rights checks. Experts rewrite the survivors substantially and contributor identity is stripped. De-identified items are then reviewed blind against the key-ambiguity defect that multiple-correct items commit most readily, and what survives is embedded as unscored pretest items indistinguishable from scored ones. IRT calibration and DIF screening close the pipeline (section 10).

A firewall that strict invites the obvious question: why write a hundred good questions for a system that never says whether the work was used, rewrites it past recognition, and pays only against milestones that stop short of the exam?
Public-pool attribution is the part a contributor can see, questions that clear screening appear in the published practice pool under contributor or institutional credit, and twenty lakh candidates practise on them. Operational variants stay anonymous, so the credit on offer is having written questions the country studies from, never questions on the paper.
Institutions carry the volume individuals will not: colleges and university departments holding vetting contracts also carry annual contribution quotas for their faculty, and contribution counts as academic service the way examinership does. Coaching institutes will try to flood the pool for influence over what candidates practise, and the pool ceiling plus ranked screening is the filter for that, not an expectation that they behave.
If public credit, milestone pay, and institutional duty do not draw enough high-quality supply in pilots, the bank stays expert-panel-only for longer.
Contributed questions still create comparability problems if the papers made from them are not genuinely equal, which is the next question.
6. Assembling and normalizing parallel exams
A single paper has no comparability problem, because there is one shift; many papers must be interchangeable in difficulty and coverage or multi-session becomes unfair. A sampling algorithm assembles one paper per session against a fixed blueprint: the subject and format counts above, topic-coverage quotas so no chapter is skipped, a target difficulty distribution, and a total estimated solve time inside the 180-minute budget. Two papers from the same blueprint should feel interchangeable to a well-prepared candidate.
For the difficulty distribution, sampling uniformly across the 0-to-1 range wastes questions at the ends where they separate nobody, and fixed band quotas put a cliff at every band edge, treating a 0.49 difficulty question and a 0.51 difficulty question as different kinds of thing. The design draws each slot's target difficulty from a bell curve and picks the closest unused item, as this platform already does at mean 0.5 and spread 0.18. The real exam would shift the mean to about 0.6, because the top is already thin: NTA's 2025 distribution shows 1,259 candidates between 601 and 650.2 A harder paper also adds noise near the qualifying cutoff, where 303,040 candidates sit in the 144 to 200 band alone.

Exam day runs backward from each session's start time, and before the assembly moment the paper does not exist, so there is nothing to steal the night before.
| Time | Step |
|---|---|
| T-4h | The sampling algorithm assembles the session's paper from the bank by item identifier, pulling all 13 language renderings together. |
| T-4h to T-2h | A randomly chosen medical college's senior faculty validate the paper in a room with no internet. Flagged questions go back to the algorithm, which replaces them from the same topic and difficulty slot. |
| T-2h | The paper freezes, is encrypted, and is distributed to centres; decryption keys release at start time. |
| T-1h | The original setter panel does a read-only final check. |
| T-0 | The session starts. A worst-case leak at assembly time exposes one session's paper for four hours. |

The T-4h external read walks previously uninvolved faculty into contact with a complete paper four hours out, this seems risky but every item has already been reviewed, pretested and calibrated inside the bank, which leaves a form-level check: ambiguity that only surfaces in a new context, cueing between adjacent items, near-duplication, blueprint sanity. What limits the damage is assembiling multiple papers per session and doing this with 4 groups, now that form is one of many, and the window closes before they leave. Which exam is to be finally selected based on each panel's independent rating of the paper.
Equal papers are then made comparable by normalization. After every shift each candidate's raw score becomes a percentile within that shift:
percentile = 100 × (candidates in your shift scoring ≤ your raw score)
÷ (candidates who appeared in your shift)
computed to seven decimal places so ties stay rare, exactly the procedure JEE Main uses.10 The topper of every shift lands at 100 whatever the raw marks, raw marks from different shifts are never compared, and the merged percentile list becomes the All India Rank. Shifts are random samples of the same population. Nobody chooses their shift and at 3 lakh candidates per shift the ability distribution is nearly identical, so it holds through the bulk of the ranking, gets noisy at the extreme top where a shift holds a handful of contenders for single-digit ranks, and says nothing about whether this year's 97th percentile knows more than last year's did.

Both gaps have the same long-run fix: shared anchor items carried across forms. Internal anchors of at least 20 to 25% of the length let IRT true-score or Stocking-Lord equating place every form on one scale, so comparability no longer rests on the equal-population assumption and cross-year drift becomes measurable. The bank is built so anchors are easy to add later; that is a stronger tool than percentile normalization and more machinery than the leak problem needs today, so the design uses normalization now and equates later.
Near term, raw subject scores combine into the target blueprint's 640-point total, convert to within-shift percentiles, and merge into the All India Rank; per-shift means and spreads are published the same day, with separate-subject IRT calibration as the long-run replacement (section 10). Before results, candidates receive their responses and a provisional key with a paid objection window of at least three days. Confirmed key or translation defects are dropped and re-scaled, rather than patched after results as in 2024, when revoked grace marks and a rescored key cut perfect scores from 67 to 17.11 The scoring algorithm is published with its purpose, formula, worked examples, versioned change log, and audit, and candidates can request an explanation of a normalized score; CBSE v. Aditya Bandopadhyay (2011) already gives them RTI access to their evaluated script.12
7. Security is exposure-window management
The dominant loss events in Indian high-stakes exams are paper leaks in the custody chain and impersonation or solver-gang collusion, not exotic cyberattacks. The design proposed here treats late binding as the highest-leverage lever, with everything else in support, and the assembly timeline in section 6 is that made concrete. Biometrics and AI proctoring are scoped as human-adjudicated detection aids with high false-positive cost, never primary prevention or proof.
The control taxonomy is Prevent, Detect, Contain, Recover.
| Risk | Prevent | Detect | Contain / Recover |
|---|---|---|---|
| Content leak | Late-binding assembly; multi-form; over-production | Canary items; custody-ledger anomalies; social monitoring | Standby form; re-exam session; revoke centre keys |
| Solver gangs | Multi-form and scrambling; RF policy; late binding; deterrence under the 2024 Act | Answer-similarity and timing forensics; centre outliers | Result hold; cluster invalidation; prosecute |
| Insider leak | Least privilege; M-of-N; HSM keys; separation of duties; screening | UEBA; access logs; canaries | Revoke keys; discard burned items |
| Compromised centre | Accreditation; government-run primary; police sealing; CCTV | Score-distribution forensics; CCTV | Debar centre; quarantine results |
| Network outage | Local-first execution; dual connectivity; pre-staged packages | Link monitoring | Continue offline; sync on restore |

Anti-cheating law sits almost entirely in the Contain and Recover column of that table. Dharmendra Pradhan resigned as Union Education Minister on 25 July 2026 after weeks of protest over the NEET-UG 2026 leak. On 27 July, the government introduced the Public Examinations (Prevention of Unfair Means) Amendment Bill, 2026. It proposes steeper sentences and fines up to ₹10 crore for organised crime, a Special Task Force the Centre can put in exclusive charge of an investigation capped at two months, and a Special Fast Track Court in every state and union territory finishing within three months of the chargesheet.13
Every clock in the Bill starts after the paper has left the strongroom. The parent Act has the same shape: offences, punishment, and who investigates them, with no requirements for how a paper is made. A conviction cannot restore the last cohort's year. Prevention can make a stolen paper worthless whether or not anyone is caught, so this design concentrates there.
The authoring environment is hardened and isolated, items over-produced three to five times, and every item carries authorship, review, and exposure metadata in the same append-only hash-chained log used by the commons, with watermarking and canary items making a leak source traceable. The bank is encrypted at rest with HSM-held keys and item-level access control, exposure tracking drives retirement, and two-person M-of-N integrity governs every content-exposing action (bank export, form publication, key release), mapping to ISO/IEC 27001 Annex A and NIST SP 800-53. Authoring, review, bank custody, assembly, delivery, centre operations, scoring, and audit are each owned by a party that must not own the adjacent step.
Detection at 20 lakh scale creates its own fairness problem, which is the next constraint.
8. Fairness is a statistical and legal constraint
At 30 lakh candidates even a 0.1% false-positive rate wrongly flags about 3,000 honest people, so statistical flags are investigative leads, not verdicts: human review, corroborating evidence, and due process precede any adverse action, thresholds are set for specificity when action is taken, and accommodations are excluded from timing and behaviour rules.
Each of the 13 language versions is a separate form for measurement. An item becomes operational only after forward translation, independent back-translation, bilingual review, field testing, and per-language DIF, all completed at bank level. Its versions remain linked, so a defect in one retires all 13. The 2018 Tamil paper had 49 mistranslated items; "English prevails" therefore remains only a last-resort tie-break, while a confirmed defect is dropped and re-scaled for that language cohort. This makes a 1 lakh item bank 13 lakh reviewed artifacts.

The RPwD Act 2016 section 17 requires exam modifications including a scribe and extra time, while section 32 requires at least 5% seat reservation and a five-year age relaxation. A scribe is not limited to benchmark 40% disability (Vikash Kumar v. UPSC, 2021), 40% disability is not an automatic admission bar (Omkar Ramchandra Gond, 2024), and the 2018 office memorandum sets at least 20 minutes of compensatory time per hour.14 WCAG 2.2 AA is a recommendation, not binding law. One accommodations desk with published SLAs supports a verified own-scribe or qualified board-provided pool and excludes accommodations from proctoring anomaly rules.
Under the DPDP Act 2023 the authority is a Data Fiduciary and, at this scale of children's biometric data, almost certainly a Significant Data Fiduciary requiring a DPO, annual audits, and a DPIA. Section 9's verifiable-guardian-consent requirement and restriction on behavioural monitoring of minors sit in direct tension with behavioural AI proctoring.15 Treating proctoring as security processing under a defined lawful basis, using Aadhaar in yes/no verification mode rather than storing biometrics, and minimising retention are governing-body policy choices requiring legal opinion before deployment.
What this proposal does not change
This redesign changes exam measurement and security, not reservation, counselling, or seat allocation. The Supreme Court upheld a uniform national exam in CMC Vellore (2020), while constitutionally settled EWS reservation (Janhit Abhiyan, 2022), at least 5% PwD reservation, and counselling for about 1.2 lakh MBBS seats across roughly 800 colleges remain outside its scope.161718
The exam body supplies a validated score and its uncertainty. Admissions authorities continue to set cut-offs and apply reservation, counselling, and seat-allocation policy independently.
9. Incident response
Pre-defined playbooks per NIST SP 800-61, each with a named owner, decision thresholds, and a communications plan, let a compromised session re-run without months of delay because standby parallel forms already exist.
| Incident | First action | Recovery |
|---|---|---|
| Confirmed content leak | Freeze affected forms; trace via canaries and watermarks | Swap standby form; targeted re-exam; prosecute source |
| Impersonation ring | Hold results for the cluster; verify biometrics 1:1 | Invalidate confirmed cases with due process |
| Centre compromise | Quarantine centre results; secure CCTV and logs | Debar centre; re-test affected candidates |
| Key compromise | Revoke and rotate keys; halt further releases | Re-issue under new keys; forensic audit |
| PII / data breach | Contain, assess scope, notify per DPDP duties | Remediate; report; post-incident review |
| CBT processing fault | Halt result release; reconcile against canary checks | Re-process; independent verification before release |

Every incident closes with a published post-incident review and any corrective changes to controls or SLAs.
10. What the simulation says
The experiment asks two practical questions: whether marking discourages blind guessing, and how large the public pool must be before memorisation stops buying seats. Its answers are directional, not forecasts.
It creates 2,000 synthetic students, each with an individual understanding level and recall, samples five papers, and resamples ten times. On each item, a student succeeds through recall first, then understanding measured against difficulty on a logistic curve, and otherwise guesses only when the expected value is positive. The pool run gives every student the same fixed 4,000-item one-month cram budget.
For four-option questions, a blind shot is profitable only when the reward is more than three times the penalty. That makes +4 / -1 positive expected value, while +3 / -1 and +5 / -2 are not. Across the full marking grid, rank recovery against true ability barely changes, ranging from 0.948 to 0.966. The model produces a sharp 100% or 0% guess rate because it treats every unknown answer as a one-in-four shot. Real partial knowledge would smooth that binary cliff, but the arithmetic still shows which schemes actively reward blind attempts.
The pool-size run gives seats to the top 2.5%:
| Public pool | Recall to take a seat | Understanding to take a seat | ρ with recall | ρ with true ability |
|---|---|---|---|---|
| 10,000 | 45% | 50% | 0.82 | 0.95 |
| 25,000 | 13% | 64% | 0.60 | 0.93 |
| 50,000 | 6% | 64% | 0.51 | 0.92 |
| 100,000 | 3% | 65% | 0.47 | 0.92 |
| 250,000 | 1% | 65% | 0.45 | 0.92 |
Recall falls 39 points from 10,000 to 50,000, then only two points from 100,000 to 250,000; understanding rises 14 points before 50,000, then does not move after one lakh. The largest change occurs before 50,000 items; the curves flatten around that point and move little past one lakh. This supports, but does not prove, the one-lakh target. At 250,000 questions, matching the 45% recall that initially took a seat would require memorising 112,500 items in one month. That is far beyond the modelled budget, while Cepeda and colleagues' meta-analysis of 317 experiments found massed cramming much weaker than spaced study for retention.19
These results need pilot data before policy is fixed, especially real per-format odds, partial-credit behaviour, calibrated difficulty, and observed cram capacity. Open the full interactive experiment for the marking grid, controls, pool curves, and model results.
11. Rollout
| Phase | Scope | Delivery | Goal |
|---|---|---|---|
| 0 · Pilot | 1 to 2 states, < 50k | Limited CBT | Validate assembly, custody, normalization, forensics; tune difficulty mean and marking |
| 1 · National CBT | All candidates | Multi-session, late binding; window length set per district by seat inventory | Kill the physical paper and its transport chain; introduce multiple-correct items |
| 2 · Format target | National | Full CBT, windows compressing as districts build seats | Add numerical entry and labelled diagrams; reach the 160-item blueprint |
| 3 · Anchors and adaptive | National, capacity-permitting | Full CBT; shared anchors, MST long-run | Cross-year comparability and efficiency; pilot free-click diagram items |
Each phase must ensure the following:
- Comparability. Shifts verified as random samples of the same population; per-shift means and spreads within tolerance; normalization published.
- Security. Zero unresolved custody-chain breaks; canary and reconciliation checks pass before results.
- Fairness. Adverse-impact gaps investigated; accommodations effective and PwD outcome parity monitored; per-language DIF flag counts published.
- Digital access. Familiarisation coverage measured; score gap between candidates with and without prior computer access measured, disaggregated rural and urban, and published.
- Content. Blueprint coverage and format quotas met; difficulty mean and marking scheme confirmed by simulation and pilot data.
- Measurement (once anchors are live). Marginal reliability ≥ 0.92; anchor drift within tolerance; DIF C-flags reviewed and resolved.
Published service levels hold the system accountable: at least 95% of candidates within 100 km or three hours of a district HQ and 100% within one state; accommodation decisions within 10 working days with appeal; biometric verification at least 99.5% with manual fallback and zero exclusions; and grievance first response within 7 days and resolution within 30 days. Annual disaggregated validation publishes DIF flag counts by language, gender, and category; reliability at the qualifying cut; adverse-impact ratios by gender, category, rural/urban, and language; PwD pass rate against overall; and predictive validity against first-year and licensing outcomes.
Distance service levels say nothing about digital access. The gate above is where the familiarisation infrastructure in section 2 binds.
Glossary
| Term | Meaning here |
|---|---|
| CBT | Computer-based testing: candidates answer on secured computers rather than paper. |
| DIF | Differential item functioning: a check for questions that behave differently across comparable groups. |
| IRT | Item response theory: a model that places question difficulty and candidate ability on a common scale. |
| M-of-N approval | Approval that requires at least M authorized people from a group of N. |
| HSM | Hardware security module: dedicated hardware that protects encryption keys. |
| UEBA | User and entity behavior analytics: detection of unusual access or system activity. |
| Anchor items and equating | Shared questions and the statistical process that use them to put different papers on one score scale. |
| Item and form | An item is one question; a form is one assembled question paper. |
Sources
Dated 10 August 2026; sources accessed 18 July 2026, except the 2026 legal and task-force sources, accessed 28 July 2026. Comparator practice worth transferring includes IRT-equated item banks (ENEM, Digital SAT, MCAT, UCAT), transparent published normalization (JEE Main), a pre-result equating and appeal window (MCAT), and formal accommodation form libraries (UCAT SEN forms). The anti-patterns are a single master paper for tens of lakhs, physical transport chains, uneven CBT server patching, and retroactive answer-key corrections after results.
Psychometric standards and methods: AERA/APA/NCME Standards (2014); van der Linden, Optimal Test Design; Kolen & Brennan, Test Equating; Rodriguez (2005) on three options optimal; Kane (2013) on validity. Security and identity: NIST SP 800-63-4; ITC Test Security. Anti-cheating law: Public Examinations (Prevention of Unfair Means) Act 2024. Comparators: Digital SAT structure; MCAT scoring; UCAT test format and scoring; ENEM/INEP; CSAT/KICE; Gaokao 2025 counts.
Footnotes
-
Format, marking, and 13 languages: NTA Information Bulletins 2025 and 2026. ↩
-
2025 candidate counts and centres (22,76,069 registered, 22,09,318 appeared, 5,468 centres across 552 cities in India and 14 abroad), plus the published score distribution and the marks at selected ranks, NTA NEET (UG) 2025 result notice, 14 June 2025. Some secondary tables cite 22,76,609 registered for 2025, a likely transposition of 22,76,069. ↩ ↩2 ↩3 ↩4
-
NEET-UG 2026 score card (16 Jul 2026); NTA re-exam FAQ (16 May 2026); NTA post-exam release (21 Jun 2026). ↩
-
NEET UG 2025 returned to the pre-Covid pattern, removing optional Section B and restoring 180 questions in 180 minutes, as reported by NDTV Education. ↩
-
Indian Express report on the 2024 Supreme Court verdict finding no systemic breach and refusing a retest. ↩
-
JEE Main 2025 delivery and normalization (CBT; Session 1 ten shifts, Session 2 nine shifts; equi-percentile normalization across shifts; best of two sessions), NTA JEE (Main) 2025 Paper 1 result press release, 18 April 2025. ↩
-
JEE Advanced 2025 (fully CBT; ~1.87 lakh registered, ~1.80 lakh appeared; 709 centres across 230 Indian cities and 3 foreign; two languages), IIT Kanpur JEE (Advanced) 2025 results press release, 2 June 2025 and JEE (Advanced) 2025 report. ↩ ↩2
-
JAMB UTME 2025 processing failure affecting ~379,997 candidates and forcing resits, The Punch report. ↩
-
CBT mode: the resignation statement records a decision that the exam would run as CBT from the following year, reported by Times Now as CBT from 2027 onwards. Exam-reform task force chaired by Nandan Nilekani announced 26 Jul 2026, whose recommendations the Supreme Court said on 27 Jul 2026 it would examine before deciding the online-NEET plea, next hearing 3 Aug 2026, in an Economic Times report. ↩
-
Procedure to be adopted for compilation of NTA scores for multi-session papers (normalization procedure based on percentile score), National Testing Agency. ↩
-
2024 NEET controversy: 67 initial perfect scores reduced to 17, grace marks revoked. ↩
-
CBSE v. Aditya Bandopadhyay (2011), RTI access to one's own evaluated script. ↩
-
Resignation of the Union Education Minister, 25 Jul 2026. Public Examinations (Prevention of Unfair Means) Amendment Bill, 2026, introduced in the Lok Sabha 27 Jul 2026, provisions summarised by The Hindu and Bar & Bench. Section structure of the parent Act. The Bill is proposed, not enacted. ↩
-
RPwD Act 2016 ss.17, 32; Vikash Kumar v. UPSC (2021); Omkar Ramchandra Gond (2024); OM 29 Aug 2018 compensatory-time floor. ↩
-
Uniform national exam upheld, CMC Vellore v. Union of India (2020). ↩
-
Janhit Abhiyan v. Union of India (2022), EWS reservation. ↩
-
NMC revised UG seat matrix for AY 2024-25 (about 1.2 lakh MBBS seats across roughly 800 colleges). ↩
-
Cepeda, Pashler, Vul, Wixted and Rohrer (2006), a meta-analysis of 317 experiments on distributed practice and retention. ↩