A national medical entrance exam, derived from its constraints
Version 1.0 · 8 August 2026
The full measurement, security, and policy blueprint behind the exam specification.
This is a measurement, security, and policy blueprint for a national undergraduate medical entrance test at roughly 20 to 30 lakh candidates a year, written with NEET-UG as the live referent. It is meant to be read as one long line of reasoning rather than a reference binder: start from how the exam works today, let each constraint bite in turn, and see what system those constraints push you toward. Where a claim is fact, the sentence says so and cites a source; where it is a design choice, it says "the design proposed here"; where a decision really belongs to a governing body, it says "this is a policy choice"; and where the public record is silent, it says so rather than inventing anything.
Everything quantitative here is a planning estimate to test through pilots and simulation, not a settled number, and nothing in it is legal advice or a conclusion about the 2026 events. Please read it as a careful proposal to argue with rather than a finished answer.
Stated plainly, the design proposed here makes four changes to today's NEET. First, it replaces the single secret master paper with a large question bank the country helps write: a capped stream of public submissions, each vetted by experts and published for practice a month before the exam, able to begin expert-only and open up as vetting scales. Second, it delivers the exam as CBT in shifts with per-shift normalization, which is JEE's machinery borrowed wholesale rather than reinvented, and it commits to computers outright, with no paper fallback. Third, it retires the all-MCQ paper: over roughly five years the blueprint moves to 160 questions mixing single-correct, multiple-correct, integer-answer numerical, and diagram-based items, formats borrowed from JEE Advanced because they make guessing structurally unprofitable rather than merely penalised. Fourth, it fixes the genuinely open numbers, the bank size, the marking scheme, and the format quotas, by simulating the exam on the live bank rather than arguing them on a slide. The first, third, and fourth are the real contribution; the second is borrowed openly because it already works at scale. The rest of this document is the reasoning that gets to each.
1. One paper, twenty lakh people
NEET-UG currently delivers a single printed form to every candidate on one afternoon. The 2025 and 2026 sittings share an identical design: pen-and-paper OMR, a single day and single shift, 180 compulsory MCQs (Physics 45, Chemistry 45, Biology 90) answered from 2:00 to 5:00 PM IST in 180 minutes, scored out of 720 with real negative marking of +4 / -1 / 0, offered in 13 languages with the English version authoritative on dispute, with a +1 hour compensatory allowance for PwBD candidates and a seven-point tie-break introduced in 2025.1 In 2025 about 22.76 lakh registered and 22.09 lakh appeared across roughly 5,468 centres.2 The 2026 original sitting of 3 May was cancelled on 12 May and re-conducted on 21 June, with score cards published 16 July.3
One identical paper for everyone is the fairest-sounding sentence in Indian education, and it is also a single point of failure. The measurement is fine; the delivery concentrates the entire exam into one artifact travelling a physical chain of custody. When that chain broke in 2024 the damage was contained: the Supreme Court found the leak localised to Hazaribagh and Patna with roughly 155 beneficiaries, found no systemic breach, and refused a national retest.4 The 2026 cancellation and re-run show the same structural fact from the other side: when confidence in the single form is lost, the only remedy is to re-run the whole exam.
The design proposed here keeps the measurement goal unchanged (rank readiness to begin UG medical study) and removes the single point of failure. Instead of one printed form, it assembles many parallel forms from a large, calibrated, continuously refreshed bank, so that no single sheet is worth stealing. A leak is only valuable if it predicts your form; when your form is one of many, drawn late and normalized against the rest, a stolen paper buys almost nothing.
The obvious objection is that another national exam already runs on computers at similar scale. Why not simply make NEET behave like JEE?
2. Why "just make it CBT like JEE" is largely the right answer
JEE shows computer-based testing works in India at scale, and the design proposed here borrows its machinery rather than treating it as a compromise. JEE Main 2025 ran entirely as CBT: roughly 15.4 lakh registered and about 14.75 lakh unique candidates appeared, spread over ten shifts in Session 1 and nine in Session 2 across many days, each shift sitting a different paper, with best-of-two-sessions scoring and equi-percentile normalization across shifts to make the scores comparable.5 JEE Advanced 2025 was also fully CBT but an order of magnitude smaller: about 1.87 lakh registered and 1.80 lakh appeared, in two languages, at 709 centres across 230 Indian cities plus three foreign locations, on a single day.6
The parts NEET should copy are exactly the parts sometimes called compromises: many forms and per-shift normalization. Different candidates answering different papers, ranked by where they fall within their own shift, is already routine and its statistics are well understood (section 6). What JEE does not demonstrate, and what NEET actually adds, is twofold. First, scale and language multiplicity: NEET is about 22 lakh candidates in 13 languages, larger than JEE Main and vastly larger than JEE Advanced, and neither JEE exam demonstrates a same-window sitting at that size in that many languages. Second, a CBT-seat inventory nobody has published: how many secure, uniform, invigilated CBT seats India can field in a compressed window is not established in any primary source. This document therefore does not assert a national CBT-seat count. It treats seat capacity as something to build and measure, never a number to quote.
CBT is also not automatically safer. JAMB's 2025 uneven server patching corrupted processing for about 379,997 candidates in Nigeria, a computerised failure at national scale that forced resits.7 Moving to computers changes the failure modes; it does not remove them.
So the destination is CBT in shifts with per-shift normalization, and CBT alone, which is now also the announced direction: after the 2026 cancellation the government said NEET would run in CBT mode from 2027, and the Supreme Court is waiting on a Nilekani-led task force's recommendations before deciding a plea seeking online NEET.8 The tempting hedge is to keep paper OMR as a guaranteed fallback wherever a district cannot yet field the seats, and the design proposed here deliberately refuses it: a printed form cannot be assembled four hours before a shift and still travel to centres, cannot shuffle questions per candidate, and cannot carry the diagram format section 4 most wants, so a fallback paper would be a second, different exam wearing the same name, and two delivery modes that measure different things is a worse unfairness than a longer window. Where seats are short, the release valve is time, more sessions across more days, never paper. The binding constraint is not "can we run CBT" but "can each district field the seats, and will scores stay comparable across shifts". The rest of this document derives the exam from those two questions.
3. Deriving the architecture from the numbers
NEET's single window needs one form and about 5,400 centres for three hours; JEE's fragmented window needs many forms but far fewer simultaneous seats. Between those poles sits one trade-off: fewer sessions mean a shorter live-content window, which is safer, but more simultaneous seats, which is harder to field; more sessions mean fewer seats but a longer window during which content is live across days, and more forms to build.
The simultaneous-seat demand for a compressed window follows from an assumption-driven identity, useful for intuition rather than as a capacity claim:
Seats ≈ N / (D · T · u)
N = candidates ; D = days ; T = sessions per day ; u = seat utilization
Every input below is an assumption to validate, not a published fact. Taking N = 22 lakh with u = 0.9: a three-day window at two sessions per day needs about 22 lakh / (6 · 0.9) ≈ 4.07 lakh seats; a five-day window at two sessions per day needs about 22 lakh / (10 · 0.9) ≈ 2.44 lakh seats. Halving the window roughly doubles simultaneous seat demand. The concrete planning figure used throughout is deliberately round: at an illustrative three lakh concurrent seats, about 22 lakh candidates fit in roughly eight sessions over three to four days. These are numbers to build toward and measure, not current capacity, and nothing here claims any of them exists today. It is tempting to dwell on this arithmetic, but the harder and more interesting problem is not the seats at all, it is keeping eight different papers comparable, which the next sections take up.

Sessions therefore multiply forms. Each distinct session is its own sampled paper, and often two per session for centre-level scrambling: F ≥ (D · T) · f_per_session with f_per_session ≈ 1 to 2. Those forms cannot drift apart in difficulty, so section 6 assembles them all to one blueprint and normalizes their scores onto a common percentile scale, or a candidate's rank becomes luck of the session.
The transition follows the same seat identity applied per district:
Seats_needed(district) ≈ N_district / (D · T · u)
A district's window in a given phase follows from its verified secure-seat inventory: where seats are plentiful the window compresses, and where they are scarce the same candidates sit across more days. This turns "phased rollout" from a calendar into a measurable per-district gate with seats as the only currency: building them is the one way to shorten a district's window, and a long window's cost, more days of live content, is exactly the exposure problem section 7's late binding exists to manage.
Comparable forms assume each form is the same length and time budget. So how many questions, in what format, and how long?
4. How many questions, in what format, and how long
NEET is 180 questions in 180 minutes, 45 each in Physics and Chemistry and 90 in Biology, and every one of them is the same kind of thing: one stem, four options, exactly one correct. The leak problem does not force that shape to change, and an earlier instinct was to hold it fixed for legibility. The measurement goal argues otherwise. A blind guess on a single-correct MCQ lands one time in four, and section 11 shows that even an aggressive penalty deters only part of that guessing, because the format keeps the gamble cheap; the sharper instrument is the question itself. JEE Advanced has run the harder formats for years: a multiple-correct item drops the odds of full credit by blind guessing from 1 in 4 to 1 in 15, and an integer answer leaves nothing to guess between at all.6
The target blueprint is therefore 160 questions in 180 minutes, 40 per subject slice across Physics, Chemistry, Botany, and Zoology. Each slice carries 25 single-correct MCQs and 10 multiple-correct MCQs, and the last five slots go to the format the subject earns: numerical items with integer answers in Physics and Chemistry, and diagram-based items in Botany and Zoology, where a labelled structure is the natural unit of the subject. Diagram items begin as labelled-part questions, which of the parts marked A to G is the site of gas exchange, cheap to author and unambiguous to score. The stronger version, click the structure directly and score the click against a tolerance region, measures something no option list can, and it enters as a pilot (section 12) because a disputed region boundary at 22 lakh scale is an answer-key objection problem to rehearse before trusting.
Five fewer questions per subject is not generosity; the harder formats buy their time. There is no universal minutes-per-question, since time per item is emergent:
t_item = f(reading load, cognitive tier, computation, format, target speededness)
Every bank item carries an estimated solve time precisely so the sampler can hold each assembled paper inside the 180-minute budget (section 6): a recall one-liner runs about 30 to 45 seconds, a multi-step numerical item about 120 to 180 seconds. The format quotas make that constraint do real work, with single-correct slots skewing quicker to make room for items that think slower, and NEET's numericals sit well below JEE Advanced's computational depth, which is what lets 160 items fit where JEE Advanced fits about fifty. The transition is gradual, roughly five years end to end, phased so no cohort meets a format it never practised against (section 12): multiple-correct arrives first because it needs nothing but new questions, numerical entry and labelled diagrams follow once delivery is computer-based everywhere, and the free-click diagram waits for its pilot. Whether the exam should eventually shorten itself adaptively the way the Digital SAT does remains a later measurement upgrade (section 11).
All of this presumes a deep pool of calibrated questions in four formats rather than one. Where do they come from without becoming the next leak?
5. Where questions come from: the bank and the national question commons
Today a small confidential pool of NTA-appointed setters writes one master paper, and the public record does not say how many people write it or how long before exam day it is finalised. Opacity is the entire security model, and it fails in two ways: a small pool cannot sustain the volume that many parallel forms demand, and every insider is a concentrated risk.
The design replaces the master paper with a maintained bank large enough that no single sheet is worth stealing. Re:Neet's working practice bank holds about 10,000 questions, 2,500 each across Physics, Chemistry, Botany, and Zoology, and a national exam at scale would need far more, on the order of 1 lakh questions or more (roughly 25,000 per subject), so that many sessions, a steady refresh, and exposure limits can all be met. The bank is held by a panel of subject experts drawn from a rotating group of top medical colleges; panel membership changes every year, each expert works on one subject's slice, and no individual ever needs the whole bank to do their job. Every question carries metadata: a format tag (single-correct, multiple-correct, numerical, diagram), a topic tag, a difficulty value between 0 and 1, an estimated solve time in seconds, and a usage history. Difficulty begins as the setter's heuristic and gets corrected by evidence; this platform already recalibrates its practice questions from live correct-answer rates, and the real exam would do the same from unscored pretest slots.

A rotating expert panel is version one. The version worth building lets anyone in the country contribute, because the supply of people who can write one good MCQ dwarfs any panel. Re:Neet already runs a small form of this: a signed-in user submits a question, an LLM screens it at intake and rejects abusive or off-syllabus content immediately, and everything else queues for committee members to approve or reject by hand with recorded reasons. Three properties make such a commons safe at national scale.
A submission cap bounds the pool. Contributions are limited to 100 questions per mobile number per year, while reading the pool is unlimited. The cap sits on input, not on reading, because without it a coaching operation could dump 10 lakh questions in and force every student to grind through all of them just to be safe. A per-person yearly cap keeps the public pool at a size a student can actually practise against.
Vetting is split into chunks. Human review is the bottleneck at national scale, and dividing it also fixes a security problem. The pool is split into chunks of about 500 questions handed to different large institutions, so a lakh-scale pool spreads across a couple hundred vetting bodies and no committee ever holds the full set. A compromised committee leaks at most its own slice, and auditing a random sample from each chunk keeps the vetting honest.
The pool goes public before it goes operational. One month before the exam the entire community pool is published for practice, every question shown as submitted rather than in its rewritten operational form. Every candidate gets the same material at the same time, which collapses the resale value of any single question, and operational papers never lift questions verbatim because the expert-rewrite step means live items are unpublished variants.

It is worth being precise about what the full firewall would still take, since the product does not implement it yet. Today submitters see their screening feedback and their question's status, which is fine for a practice bank but impossible for a real exam, where a contributor must never be able to tell whether their question went live. The operational pipeline therefore quarantines submissions at intake behind an irrevocable license grant and provenance metadata (author, source, curriculum tags, language) with an originality attestation, runs near-duplicate, plagiarism, and image-rights checks, has experts rewrite items substantially so the operational item differs from the submitted one, strips contributor identity and places items into enemy sets, reviews de-identified items blind against the recurring key-ambiguity defect, which multiple-correct items commit most readily, and embeds survivors as unscored pretest items indistinguishable from scored ones. Only then, keyed off pre-air-gap milestones and never off operational use, can any recognition or honorarium be paid, because a reward tied to going live would itself leak status. IRT calibration and DIF screening by language, gender, and category sit at the end of that pipeline as the measurement upgrade (section 11); the near-term bank runs on heuristic-then-recalibrated difficulty.

Contributed items still create fairness and comparability problems if the forms they populate are not genuinely equal. So the next question is how different papers are made to measure the same thing.
6. Assembling and normalizing parallel forms
A single paper has no assembly problem and no comparability problem, because there is one shift. Many forms must be interchangeable in difficulty and coverage, or multi-session delivery is unfair. A sampling algorithm assembles one paper per session against a fixed blueprint: the subject and format counts above (40 Physics, 40 Chemistry, 80 Biology, drawn as 40 Botany and 40 Zoology from the bank's two biology slices, each slice at 25 single-correct, 10 multiple-correct, and 5 numerical or diagram items), topic-coverage quotas so no chapter is skipped, a target difficulty distribution, and a total estimated solve time that fits the 180-minute budget. Two papers from the same blueprint should feel interchangeable to a well-prepared candidate.
For the difficulty distribution, sampling uniformly across the whole 0-to-1 range wastes items at the ends where they separate nobody, and fixed band quotas (say 30% easy, 50% medium, 20% hard) put a cliff at every band edge, treating a 0.49 item and a 0.51 item as different kinds of thing. The design draws each slot's target difficulty from a bell curve and picks the closest unused bank item. This platform already does this for its practice papers at mean 0.5 and spread 0.18; the real exam would shift the mean to about 0.6, because a paper with more genuinely hard questions stretches out the top of the score distribution and leans less on tie-breaking (in 2024, 67 candidates sat at a perfect 720 before revision).9 The exact mean is a choice to simulate, and section 11 is the plan for choosing it.

Exam day runs backward from each session's start time. Before the assembly moment the paper does not exist, so there is nothing to steal the night before.
| Time | Step |
|---|---|
| T-4h | The sampling algorithm assembles the session's paper from the bank. |
| T-4h to T-2h | A randomly chosen medical college's senior faculty validate the paper in a room with no internet. Flagged questions go back to the algorithm, which replaces them from the same topic and difficulty slot. |
| T-2h | The paper freezes, is encrypted, and is distributed to centres; decryption keys release at start time. |
| T-1h | The original setter panel does a read-only final check. A genuine defect switches the session to a standby paper assembled the same way. |
| T-0 | The session starts. A worst-case leak at assembly time exposes one session's paper for four hours. |

Equal forms are then made comparable by normalization, not by comparing raw marks. After every shift each candidate's raw score becomes a percentile within that shift:
percentile = 100 × (candidates in your shift scoring ≤ your raw score)
÷ (candidates who appeared in your shift)
computed to seven decimal places so ties stay rare, exactly the procedure JEE Main uses.10 The topper of every shift lands at 100 whatever the raw marks, raw marks from different shifts are never compared, and the merged percentile list becomes the All India Rank. One assumption carries all of it: shifts must be random samples of the same population. Nobody chooses their shift, and at a lakh candidates per shift the ability distribution is nearly identical across shifts, so the assumption holds well through the bulk of the ranking. It gets noisy at the extreme top, where a shift holds only a handful of contenders for single-digit ranks, and it says nothing about whether this year's 97th percentile knows more than last year's.

Both gaps have the same fix, and it is the long-run upgrade rather than the near-term default: shared anchor items carried across forms. A NEAT design with internal anchors of at least 20 to 25% of the length, representative of the blueprint, lets IRT true-score or Stocking-Lord equating place every form on one scale, so comparability no longer rests on the equal-population assumption and cross-year drift becomes measurable and monitored. The bank is built so anchors are easy to add later. That is a strictly stronger tool than percentile normalization, and it is also more machinery than the leak problem needs today; the design ships normalization now and equates later.
Marking starts from today's +4 correct, -1 wrong, 0 blank on single-correct and diagram items, and borrows JEE Advanced's partial-credit scheme for multiple-correct ones: full marks for the complete set of correct options and nothing else, partial credit for a correct subset, -2 the moment a wrong option is marked. Numerical items score +4 or 0 with no penalty, because an integer field leaves nothing to guess against, which puts the target blueprint at a 640 maximum. The design also studies a harsher +5 / -2 on the single-correct slice: a bigger penalty widens the gap between knowing and guessing, though it also punishes partial knowledge, so the net effect on rank quality has to be measured rather than assumed, and the whole grid has to be re-swept with partial credit in it (section 11). Remaining ties resolve by subject percentiles in the order Biology, Chemistry, Physics, then by fewer wrong answers, close to the current seven-step rule.
Comparable forms still leak if the window stays open too long. Security, then, is really about time.
7. Security is exposure-window management
The dominant loss events in Indian high-stakes exams are paper leaks in the custody chain and impersonation or solver-gang collusion, not exotic cyberattacks. The design proposed here treats the highest-leverage lever as late binding, shrinking how long content is exposed, with everything else in support. The assembly timeline in section 6 is that late binding made concrete: the paper does not exist until four hours before the session and cannot change in the final two, so a leak is worthless once the window closes. Cryptography for its own sake does not help. Biometrics and AI proctoring are scoped as human-adjudicated detection aids with high false-positive cost, never primary prevention or proof.
The control taxonomy is Prevent, Detect, Contain, Recover.
| Risk | Prevent | Detect | Contain / Recover |
|---|---|---|---|
| Content leak | Late-binding assembly; multi-form; over-production | Canary items; custody-ledger anomalies; social monitoring | Standby form; re-exam session; revoke centre keys |
| Solver gangs | Multi-form and scrambling; RF policy; late binding; deterrence under the 2024 Act | Answer-similarity and timing forensics; centre outliers | Result hold; cluster invalidation; prosecute |
| Insider leak | Least privilege; M-of-N; HSM keys; separation of duties; screening | UEBA; access logs; canaries | Revoke keys; discard burned items |
| Compromised centre | Accreditation; government-run primary; police sealing; CCTV | Score-distribution forensics; CCTV | Debar centre; quarantine results |
| Network outage | Local-first execution; dual connectivity; pre-staged packages | Link monitoring | Continue offline; sync on restore |
Anti-cheating law sits almost entirely in the Contain and Recover column of that table, and the one place it appears in the Prevent column, deterrence under the 2024 Act, is the weakest entry there. Dharmendra Pradhan resigned as Union Education Minister on 25 July 2026 after weeks of protest over the NEET-UG 2026 leak, and on 27 July the government introduced the Public Examinations (Prevention of Unfair Means) Amendment Bill, 2026 in the Lok Sabha: steeper sentences and fines up to ₹10 crore for organised crime, a Special Task Force the Centre can put in exclusive charge of an investigation, a two-month investigation cap, a Court of Session designated in every state and union territory (in consultation with the High Court) as a Special Fast Track Court trying cases day to day and finishing within three months of the chargesheet, Special Public Prosecutors, and appeals to a High Court division bench to be disposed within three months as far as possible.11 None of it engages until a paper has already left the strongroom, and the parent Act has the same shape: sections 3 to 8 define unfair means and the offences around them, 9 to 11 make those offences cognizable and set punishment, 12 fixes the rank of the investigating officer, and on the making of a paper the two statutes are silent, reaching delivery only to bar premises other than the examination centre.11 Faster prosecution is worth having and it is a cure rather than prevention: a conviction is a fact about the last cohort's year, not a way of giving it back, and deterrence has to be re-earned against every new gang. The controls in the Prevent column are the ones that make a stolen paper worthless whether or not anyone is caught, and the design proposed here spends its effort there.
The authoring environment is hardened and isolated, items over-produced three to five times, and every item carries authorship, review, and exposure metadata in the same append-only hash-chained log used by the commons. Watermarking and canary items make a leak source traceable. The bank is encrypted at rest with HSM-held keys and item-level access control, and exposure tracking drives retirement. Two-person M-of-N integrity governs every content-exposing action (bank export, form publication, key release), mapping to ISO/IEC 27001 Annex A and NIST SP 800-53 access controls, with authoring, review, bank custody, assembly, delivery, centre operations, scoring, and audit each owned by a party that must not own the adjacent step.
Detection at 20 lakh scale creates its own fairness problem, which is the next constraint.
8. Fairness is a statistical and legal constraint
Fairness at this scale is not a slogan; it is arithmetic and law. At 30 lakh candidates even a 0.1% false-positive rate wrongly flags about 3,000 honest people, so statistical flags are investigative leads, not verdicts. Mandatory human review, corroborating evidence, and due process must precede any adverse action; thresholds are set for specificity when action is taken and lowered only to open investigations; and accommodations are excluded from timing and behaviour rules.
Language is a measurement problem, not a translation problem. The design proposed here treats each of the 13 language versions as a separate form needing its own equating and DIF review, built through forward-translation, independent back-translation, bilingual expert review, field test, and DIF by language before scoring. The 2018 Tamil paper carried 49 mistranslated items, and an "English prevails" clause silently penalises candidates weak in English, so it is kept only as a last-resort tie-break and any confirmed translation defect is handled as a scored-item problem, dropped and re-scaled for that language cohort, with per-language DIF flag counts published annually.
Accessibility is the most legally loaded area, and here binding law must be distinguished from recommendation. The RPwD Act 2016 section 17 imposes a duty to modify the exam system with a scribe and extra time, section 32 requires at least 5% seat reservation with a five-year age relaxation in higher education, and the Supreme Court has held that a scribe is not limited to benchmark 40% disability (Vikash Kumar v. UPSC, 2021) and that at least 40% disability is not an automatic admission bar but requires an individualized reasoned assessment (Omkar Ramchandra Gond, 2024); the 2018 office memorandum sets a compensatory time floor of at least 20 minutes per hour.12 These are law. WCAG 2.2 AA for accessible digital delivery is a recommendation. The design runs one accommodations desk with published SLAs, allows both a verified own-scribe and a qualified board-provided pool, and excludes accommodations from proctoring anomaly rules.
Under the DPDP Act 2023 the authority is a Data Fiduciary and, at this scale of children's biometric data, almost certainly a Significant Data Fiduciary with a DPO, annual audits, and a DPIA; section 9's verifiable-guardian-consent and restriction on behavioural monitoring of minors sits in direct tension with behavioural AI proctoring.13 Treating proctoring as security processing under a defined lawful basis, using Aadhaar in yes/no verification mode rather than storing biometrics, minimising retention, and obtaining a legal opinion on the child-monitoring interaction before deployment is a policy choice for a governing body, informed by that legal opinion.
Even a perfect score is only one signal, which raises the last measurement question: what is the score for?
9. What the score is for
The single most important structural choice is a high wall between the measurement instrument (item writing, psychometrics, security, scoring) and the use of the score (cut-offs, reservation, counselling, seat allocation). This lets a measurement defect be fixed without reopening the entire admissions settlement, and the reverse.
A well-built MCQ science exam is a reliable, scalable, coaching-imperfect but comparable screen for foundational readiness, and at 20 to 30 lakh candidates comparability and logistics strongly favour a common instrument; the Supreme Court upheld a uniform national exam in CMC Vellore (2020).14 What such an exam does not measure is non-cognitive competence such as communication, ethics, teamwork, and resilience, which is why international practice (AAMC holistic review, MMI, SJT) supplements the cognitive score rather than replacing it, and why the MCAT is used as one input and never the sole signal.
This remains a policy choice for a governing body, not a measurement question. The position the design supports is to use the exam as a competency gate plus a selection rank within a constrained admissions framework, with constitutionally settled reservation and EWS (Janhit Abhiyan, 2022) and at least 5% PwD reservation layered on top, and any non-cognitive component piloted only at the counselling-shortlist stage and scaled only if it adds predictive validity without adverse impact.15 Counselling is otherwise unchanged: roughly 1.24 lakh MBBS seats across about 800 colleges were on offer for 2025-26, filled in rounds where the better rank picks first.16 The exam body supplies only the validated score and its uncertainty; it does not set admissions policy.
The machinery that produces "the validated score" is worth stating explicitly. In the near term the raw subject scores are combined into the paper's point total (640 at the target blueprint), converted to within-shift percentiles as in section 6, and merged into the All India Rank, with per-shift means and spreads published the same day; separate-subject IRT calibration (3PL, with 2PL fallback where the guessing parameter is unstable) onto an equated scale is the long-run replacement (section 11). A provisional key with candidate access to their own responses opens a paid, time-boxed objection window of at least three days before the final key and result, and confirmed key or translation defects are dropped and re-scaled rather than silently absorbed. Every scoring and allocation algorithm is published with purpose, exact formula, worked examples, and a versioned change log, with an independent pre-deployment audit and a candidate right to an explanation of a normalized score; CBSE v. Aditya Bandopadhyay (2011) already gives candidates access to their own evaluated script under RTI, without a right to re-evaluation.17
10. Incident response
Pre-defined playbooks per NIST SP 800-61, each with a named owner, decision thresholds, and a communications plan, let a compromised session re-run without months of delay because standby parallel forms already exist.
| Incident | First action | Recovery |
|---|---|---|
| Confirmed content leak | Freeze affected forms; trace via canaries and watermarks | Swap standby form; targeted re-exam; prosecute source |
| Impersonation ring | Hold results for the cluster; verify biometrics 1:1 | Invalidate confirmed cases with due process |
| Centre compromise | Quarantine centre results; secure CCTV and logs | Debar centre; re-test affected candidates |
| Key compromise | Revoke and rotate keys; halt further releases | Re-issue under new keys; forensic audit |
| PII / data breach | Contain, assess scope, notify per DPDP duties | Remediate; report; post-incident review |
| CBT processing fault | Halt result release; reconcile against canary checks | Re-process; independent verification before release |
Every incident closes with a published post-incident review and any corrective changes to controls or SLAs.
11. Quantitative baseline, simulation, and long-run targets
Every parameter here is an engineering target with a range, to be validated by pilots and simulation, not a universal law or a legal minimum.
The near-term design parameters, close to what the platform already runs:
| Parameter | Value | Note |
|---|---|---|
| Scored items | Target 160 (40 Physics, 40 Chemistry, 80 Biology) | Phased from today's 180 over ~5 years |
| Format mix | Per 40-item slice: 25 single-correct, 10 multiple-correct, 5 numerical or diagram | Numerical in Physics and Chemistry; diagram in Botany and Zoology |
| Time | 180 minutes | Harder formats paid for by fewer items |
| Bank size | Practice ~10,000 (2,500 each: Physics, Chemistry, Botany, Zoology) | Rotating panels, then a capped commons; ~1 lakh+ at national scale |
| Difficulty scale | 0 (easy) to 1 (hard), per item | Heuristic, then recalibrated from response data |
| Difficulty sampling | Bell curve, spread ~0.18; mean ~0.6 for the real exam (0.5 in practice) | Mean is a simulation output |
| Estimated solve time | ~30s (easy) to ~180s (hard) per item | Sampler holds the total ≤ 180 min |
| Marking | +4 / -1 single-correct and diagram; partial credit multiple-correct; +4 / 0 numerical; +5 / -2 under study | Penalty controls ties; grid re-swept with partial credit |
| Submission cap | 100 per mobile number per year | Bounds the public pool |
| Vetting chunks | ~500 items per institution | No committee sees the whole set; ~200 vetting bodies at lakh scale |
| Public-pool release | 1 month before the exam | As submitted, not operational form |
| Delivery | CBT in shifts, ~3 lakh seats → ~8 sessions / 3-4 days | No paper fallback; scarce-seat districts stretch the window |
The judgment calls, the difficulty mean, the marking scheme, and the format quotas, are settled by simulation rather than argument, and Re:Neet runs that simulation on the live bank. A synthetic student carries a recall percentage that varies a little by topic; for each question a coin weighted by that recall decides whether he knows it, and if not, whether he risks a guess. Run 1,000 such students, recall drawn from a bell curve centred on 40%, across sampled papers, and the score distributions fall out. Sweeping every correct mark from +1 to +5 against every penalty from 0 to -4 shows the penalty does the work: with no negative marking every setting leaves the top bunched; a single -1 roughly halves the ties; by -2 the lower correct marks reach a unique topper per paper. The same grid measures deterrence, and it is the sweep's most sobering output: under +4 / -1 a blind guess keeps positive expected value and nobody leaves a blank, and even +4 / -2 leaves about half the blind guessing in place, which is the observation that sends section 4 after the question format instead of the penalty. The positive mark barely moves the tie count, and raising it from +1 to +5 mostly rescales the spread (standard deviation grows from about 6 to about 29) without sharpening the ranking. At a 40% mean nothing reaches the ceiling, so a large positive mark looks free; at a 55% mean the +4 ceiling starts collecting students at the maximum while +5 still has room. Together the data favour +5 / -2, deep enough to keep a unique topper and high enough to survive a strong cohort, with +4 / -1 almost as good until the cohort gets strong. A follow-up already run on this platform replaces the synthetic students with nine language models of clearly different, stable strength attempting the same exported papers, so the ranking has a ground truth. The format mix adds an axis the sweep has not covered yet: synthetic students need per-format guess models, a 1-in-4 shot on single-correct, 1 in 15 for full credit on multiple-correct, effectively zero on an integer field, and the whole marking grid has to be re-run with partial credit in it before the 25/10/5 quotas are frozen.
The long-run measurement upgrade layers IRT on top of the same bank. These are the targets it would hold, independent of the near-term normalization design:
| Parameter | Target | Note |
|---|---|---|
| Target reliability | ≥ 0.92 (stretch 0.95) | Marginal reliability |
| Discrimination a | 0.8 to 2.0 (reject < 0.5) | Distribution healthier than uniformly max a |
| Difficulty b | -0.5 to +2.0 logits | Denser near merit cutoffs |
| Point-biserial | ≥ 0.20 (prefer ≥ 0.30) | Screening threshold |
| Guessing c (3PL) | ~0.10 to 0.25 | Beta prior; 4 options, chance 0.25 |
| Anchors | ≥ 20 to 25% of length | For equating and drift monitoring |
| Annual refresh | ~0.30 | Exposure-driven retirement |
| Pretest N per item | 200-500 (1PL), 500-1,000 (2PL), 1,000-2,000+ (3PL) | Calibration sample |
Information and reliability tie together through the standard relations:
Iᵢ(θ) = D²·aᵢ²·Pᵢ(1−Pᵢ) ; I(θ) = Σ Iᵢ(θ) ; SE(θ) = 1/√I(θ) ; ρ ≈ σ²θ / (σ²θ + SDz)
At 20 to 30 lakh candidates the calibration Ns are trivially met; pretest breadth, not per-item N, is the real limit. Whether the exam later moves from fixed linear forms to multistage adaptive, which buys target precision from fewer items, is an efficiency question gated on CBT capacity, not a leak-security one.
12. Phased rollout and pilot gates
Rollout is gated on evidence, not a calendar. Each phase advances only when the prior phase's pilot metrics pass pre-set gates, and the difficulty mean, marking scheme, and format quotas are confirmed by simulation and pilot data before national use. The format transition rides the same phases, roughly five years end to end, so every cohort sits a paper whose formats were public and practisable years in advance.
| Phase | Scope | Delivery | Goal |
|---|---|---|---|
| 0 · Pilot | 1 to 2 states, < 50k | Limited CBT | Validate assembly, custody, normalization, forensics; tune difficulty mean and marking |
| 1 · National CBT | All candidates | Multi-session, late binding; window length set per district by seat inventory | Kill the physical paper and its transport chain; introduce multiple-correct items |
| 2 · Format target | National | Full CBT, windows compressing as districts build seats | Add numerical entry and labelled diagrams; reach the 160-item blueprint |
| 3 · Anchors and adaptive | National, capacity-permitting | Full CBT; shared anchors, MST long-run | Cross-year comparability and efficiency; pilot free-click diagram items |
The gates each phase must clear:
- Comparability. Shifts verified as random samples of the same population; per-shift means and spreads within tolerance; normalization published.
- Security. Zero unresolved custody-chain breaks; canary and reconciliation checks pass before results.
- Fairness. Adverse-impact gaps investigated; accommodations effective and PwD outcome parity monitored; per-language DIF flag counts published.
- Content. Blueprint coverage and format quotas met; difficulty mean and marking scheme confirmed by simulation and pilot data.
- Measurement (once anchors are live). Marginal reliability ≥ 0.92; anchor drift within tolerance; DIF C-flags reviewed and resolved.
Published, reported service levels hold the system accountable: at least 95% of candidates within 100 km or three hours of a district HQ and 100% within one state; accommodation decisions within 10 working days with appeal; biometric verification at least 99.5% with manual fallback and zero exclusions; provisional key and own-response access within a pre-committed window and an objection window of at least three days; and grievance first response within 7 days and resolution within 30 days, with outcomes published in aggregate. Annual disaggregated validation publishes DIF flag counts by language, gender, and category; reliability and decision consistency at the qualifying cut; adverse-impact ratios by gender, category, rural/urban, and language over time; PwD pass rate and score distribution against overall; and predictive validity against first-year and licensing outcomes, corrected for range restriction.
13. What the public record does not disclose
Some things are deliberately withheld and are not invented here. The exact number and names of paper setters, the number of terminals or copies, how many days before the exam a paper is finalised, the printing-press identity, and the seat-allocation algorithm are not public. The full Radhakrishnan committee report is not accessible; the hybrid encrypted-delivery and local-print (CPPT) proposal is reputable secondary reporting attributed to that committee, constituted 22 June 2024 and reporting 21 October 2024, not an officially verified design.18 The named 2026 arrest narratives and the claim that a 2025 paper was compromised by a 2026 network lack primary confirmation and are excluded from the factual layer; earlier circulated figures of 200 questions in 3h20m (true only for 2021 to 2024), of questionable negative marking, and of AI proctoring in NEET are corrected against the current OMR, single-shift, 180-item design.
The load-bearing unknown behind the entire CBT-first argument is that no primary source establishes a national CBT-seat inventory for a compressed 20-lakh-plus window; it is treated throughout as capacity to build and measure, and with the paper fallback removed it gates the calendar directly: a district short of seats gets a longer window, not a different exam. Every quantitative target in this document, including the bank size, the difficulty mean, the marking scheme, and the capacity scenarios, awaits validation by pilot and simulation.
Sources
Dated 28 July 2026; sources accessed 18 July 2026, except the 2026 legal and task-force sources, accessed 28 July 2026. Comparator practice worth transferring, drawn from the sources below, includes IRT-equated item banks (ENEM, Digital SAT, MCAT, UCAT), transparent published normalization (JEE Main), a pre-result equating and appeal window (MCAT), formal accommodation form libraries (UCAT SEN forms), and an independent annual technical report. The anti-patterns are a single master paper for tens of lakhs, physical transport chains, uneven CBT server patching, and retroactive answer-key corrections after results.
Psychometric standards and methods: AERA/APA/NCME Standards (2014) (https://www.testingstandards.net/uploads/7/6/6/4/76643089/standards_2014edition.pdf); van der Linden, Optimal Test Design (https://link.springer.com/book/10.1007/0-387-29054-0); Kolen & Brennan, Test Equating (https://link.springer.com/book/10.1007/978-1-4939-0317-7); Rodriguez (2005) on three options optimal (https://doi.org/10.1111/j.1745-3992.2005.00006.x); Kane (2013) on validity (https://doi.org/10.1111/jedm.12000). Security and identity: NIST SP 800-63-4 (https://pages.nist.gov/800-63-4/); ITC Test Security (https://www.intestcom.org/files/guideline_test_security.pdf). Anti-cheating law: Public Examinations (Prevention of Unfair Means) Act 2024 (https://www.indiacode.nic.in/handle/123456789/20100). Comparators: Digital SAT (https://satsuite.collegeboard.org/sat/whats-on-the-test/structure); MCAT scoring (https://students-residents.aamc.org/mcat-scores/how-mcat-exam-scored); UCAT (https://www.ucat.ac.uk/about-ucat/test-format-and-scoring/); ENEM/INEP (https://www.gov.br/inep/pt-br); CSAT/KICE (https://www.kice.re.kr/sub/info.do?m=0205&s=english); Gaokao 2025 counts (https://english.www.gov.cn/news/202505/28/content_WS6836ac07c6d0868f4e8f2edf.html).
Footnotes
-
Format, marking, and 13 languages: NTA Information Bulletins 2025 (https://cdnbbsr.s3waas.gov.in/s37bc1ec1d9c3426357e69acd5bf320061/uploads/2025/02/2025020754.pdf) and 2026 (https://cdnbbsr.s3waas.gov.in/s37bc1ec1d9c3426357e69acd5bf320061/uploads/2026/02/202602081576322299.pdf). ↩
-
2025 candidate counts and centres (22,76,069 registered, 22,09,318 appeared, 5,468 centres across 552 cities in India and 14 abroad), NTA NEET (UG) 2025 result notice, 14 June 2025: https://cdnbbsr.s3waas.gov.in/s37bc1ec1d9c3426357e69acd5bf320061/uploads/2025/06/2025061472.pdf. Some secondary tables cite 22,76,609 registered for 2025, a likely transposition of 22,76,069. ↩
-
NEET-UG 2026 score card (16 Jul 2026): https://neet.nta.nic.in/score-card-for-neet-ug-2026/; NTA re-exam FAQ (16 May 2026): https://nta.ac.in/Download/Notice/Notice_20260516152301.pdf; post-exam release (21 Jun 2026): https://nta.ac.in/Download/Notice/Notice_20260621193322.pdf. ↩
-
2024 Supreme Court verdict finding no systemic breach and refusing a retest: https://indianexpress.com/article/education/neet-ug-2024-sc-rules-out-cancellation-and-retest-paper-leak-final-verdict-9403650/. ↩
-
JEE Main 2025 delivery and normalization (CBT; Session 1 ten shifts, Session 2 nine shifts; equi-percentile normalization across shifts; best of two sessions), NTA JEE (Main) 2025 Paper 1 result press release, 18 April 2025: https://nta.ac.in/Download/Notice/Notice_20250419125449.pdf. ↩
-
JEE Advanced 2025 (fully CBT; ~1.87 lakh registered, ~1.80 lakh appeared; 709 centres across 230 Indian cities and 3 foreign; two languages), IIT Kanpur JEE (Advanced) 2025 results press release, 2 June 2025 (https://jeeadv.ac.in/documents/Result2025PressRelease.pdf) and the JEE (Advanced) 2025 report (https://jeeadv.ac.in/reports/2025.pdf). ↩ ↩2
-
JAMB UTME 2025 processing failure affecting ~379,997 candidates and forcing resits, The Punch: https://punchng.com/mass-failure-jamb-boss-weeps-as-human-error-forces-lagos-seast-resits/. ↩
-
CBT mode: the resignation statement records a decision that the exam would run as CBT from the following year (https://frontline.thehindu.com/politics/dharmendra-pradhan-resigns-cockroach-janta-party-neet/article71265726.ece), reported as CBT from 2027 onwards (https://www.timesnownews.com/education/sc-to-examine-nilekani-led-panel-recommendations-on-neet-cbt-calls-for-out-of-the-box-suggestions-article-155185873). Exam-reform task force chaired by Nandan Nilekani announced 26 Jul 2026, whose recommendations the Supreme Court said on 27 Jul 2026 it would examine before deciding the online-NEET plea, next hearing 3 Aug 2026: https://economictimes.indiatimes.com/industry/services/education/sc-says-nilekani-led-task-force-recommendations-will-be-key-to-online-neet-decision/articleshow/132654207.cms. ↩
-
2024 NEET controversy: 67 initial perfect scores reduced to 17, grace marks revoked: https://en.wikipedia.org/wiki/2024_NEET_controversy. ↩
-
The NTA percentile formula and its seven-decimal computation, explained with the official procedure: https://engineering.careers360.com/articles/jee-main-normalisation-process-how-scores-are-calculated. ↩
-
Resignation of the Union Education Minister, 25 Jul 2026: https://frontline.thehindu.com/politics/dharmendra-pradhan-resigns-cockroach-janta-party-neet/article71265726.ece. Public Examinations (Prevention of Unfair Means) Amendment Bill, 2026, introduced in the Lok Sabha 27 Jul 2026, provisions summarised: https://www.thehindu.com/news/national/centre-proposes-bill-to-strengthen-anti-cheating-law-in-bid-to-curb-exam-malpractices/article71265665.ece and https://www.barandbench.com/news/stf-probe-harsher-punishment-and-fast-track-courts-governments-proposed-changes-to-anti-paper-leak-law. Section structure of the parent Act: https://www.indiacode.nic.in/handle/123456789/20100. The Bill is proposed, not enacted. ↩ ↩2
-
RPwD Act 2016 ss.17, 32; Vikash Kumar v. UPSC (2021): https://api.sci.gov.in/supremecourt/2019/19177/19177_2019_36_1503_26115_Judgement_11-Feb-2021.pdf; Omkar Ramchandra Gond (2024): https://api.sci.gov.in/supremecourt/2024/39448/39448_2024_3_1501_56394_Judgement_15-Oct-2024.pdf; OM 29 Aug 2018 compensatory-time floor. ↩
-
Digital Personal Data Protection Act 2023: https://prsindia.org/files/bills_acts/acts_parliament/2023/Digital_Personal_Data_Protection_Act,_2023.pdf. ↩
-
Uniform national exam upheld, CMC Vellore v. Union of India (2020): https://www.scconline.com/blog/post/2020/04/29/uniform-neet-for-admission-to-medical-dental-courses-does-not-violate-rights-of-the-unaided-aided-minority-institutions/. ↩
-
Janhit Abhiyan v. Union of India (2022), EWS reservation: https://main.sci.gov.in/supremecourt/2019/1827/1827_2019_1_1501_39619_Judgement_07-Nov-2022.pdf. ↩
-
NMC final MBBS seat matrix for 2025-26: 1,23,700 seats across 808 medical colleges: https://news.careers360.com/nmc-seat-matrix-2025-1-23-lakh-mbbs-seats-in-808-medical-colleges-6850-added-over-1000-dropped-over-renewal-approval-court-matters. ↩
-
CBSE v. Aditya Bandopadhyay (2011), RTI access to one's own evaluated script: http://hsamb.org.in/sites/default/files/documents/CBSE-Vs-Aditya-Bandopadhyay.pdf. ↩
-
Radhakrishnan committee dates via parliamentary answer (https://sansad.in/getFile/annex/266/AU1008_Qt83wr.pdf) and SC compliance order 7 Apr 2025 (https://api.sci.gov.in/supremecourt/2024/51798/51798_2024_11_41_60784_Order_07-Apr-2025.pdf); CPPT secondary reporting: https://indianexpress.com/article/education/panel-after-leak-send-test-paper-digitally-answers-on-omr-sheet-9645210/. ↩