
CASE STUDY
Designing AI evaluation for high-stakes exam writing
An AI layer that reads and assesses handwritten exam answers — where the design problem isn't the feature, it's what happens when the model gets it wrong.
Role
Product designer & lead
Timeline
~1 month, design & planning
Status
Designed · Under Productions
Tools
Figma
01 — Framing
The hard part was never the feature
NPSC Prep is a mobile-first platform for aspirants preparing for the Nagaland Public Service Commission exams.
This study is about a single layer of it: the AI. Specifically, the feature that reads an aspirant's handwritten descriptive answer and tells them how they did.
Most AI design writing is about the wrapper — the chat panel, the sparkle icon, the prompt box. That's not the hard part. The hard part is that AI is a fundamentally different design material from anything we're trained on:
A deterministic interface has one correct output. A probabilistic one is confidently wrong sometimes.
And the stakes here make that concrete. An NPSC aspirant has years and a career riding on this exam. If an AI evaluates their Mains answer and gets it wrong — misreads their handwriting, misjudges their structure, scores them higher or lower than a real examiner would — that isn't an inconvenience. It misdirects months of preparation.
The design problem
How do you build something genuinely useful on top of a system you know will sometimes be wrong — without either hiding the uncertainty or making it useless? Hide the fallibility and you build false confidence into someone's exam preparation. Foreground it too heavily and the feature becomes so hedged it tells them nothing. The design lives between those two failures.
Every decision in this layer answers one question: what happens when the model is wrong? Not as a footnote or an error state — as the organising principle. Each screen exists because of a different answer to it: one prevents wrongness before it happens, one lets the aspirant check the model's work, one turns a failure into something useful, one lets them contest the result.
02 — Why the checklist wasn't enough
The AI didn't come first
A self-evaluation checklist did — and it still ships, free, to every user.
After writing an answer, the aspirant sees the model answer structure and the points a strong response would cover, then walks through four questions about their own work. Did you directly address the question asked? Did you state your position in the first paragraph? Did you support each point with evidence? Did you proofread for clarity? Yes, partially, or no. It produces a self-score and a plain-language summary.
That's genuinely useful. It gives an aspirant a rubric where they had none, it costs nothing, and it makes them read their own answer critically — a skill the exam rewards.
But it has a ceiling built into its premise:
It asks you to evaluate work using judgment you don't yet have.
An aspirant who doesn't know their conclusion is weak will confidently tick "yes, I stated my position." Someone who thinks a point is well-evidenced because they believe it will mark it yes. Self-assessment is bounded by self-awareness — and the people who most need feedback are precisely the ones least equipped to give it to themselves. The checklist can tell you what to look for. It can't tell you whether you found it.
What the AI actually adds
Not a better rubric — the dimensions are identical in both tiers. Not more questions. The same four questions, answered by something other than the person who wrote the answer. An outside read.
A note on how this was reached: this framing came from a product insight, not from user research. The survey behind NPSC Prep validated the access problem — fragmented sources, cost barriers — not this one. Testing whether the checklist's ceiling shows up in real usage, and whether AI feedback closes it, is named in the limitations below.
03 — The decisions
Eight decisions, one question
Each answers what happens when the model is wrong from a different angle.
DECISION 01
A range, not a score
The evaluation returns "Likely 62–70%" — never a single number. A point score implies a precision the system doesn't have; in a competitive exam, a false-precise "68" invites an aspirant to treat a machine estimate as a verdict. A range does the job people actually need — am I in the passing zone? — while encoding the uncertainty honestly in its shape. And it's harder to fixate on: you can't screenshot a range and treat it as your mark.
Paired with it is a plain-language line: "A solid answer with clear points — the structure is where most of the gap sits." Calibration in numbers, direction in words.
Evaluation Summary
Mains Assessment
ESTIMATED BAND
GS-III COMPLIANT
Likely 62–70%
A solid answer with clear points — the structure is where most of the gap sits. Your coverage of the economic dimension was thin.
DIMENSION-WISE SCORE BREAKDOWN
Relevance
85/100
Directly addresses urbanisation dynamics well.
Structure
45/100
No clear conclusion paragraph detected.
Point Development
70/100
Economic impacts could be more elaborate.
Grammar & Sentence
90/100
Excellent flow and highly coherent style.
AI estimates structure & coverage — it can't predict precise human examiner mood.
"A solid answer with clear points — the structure is where most of the gap sits."
DECISION 02
The score is a decomposition, never a verdict
The range isn't a badge at the top of the page — it's constructed in front of the aspirant from the four dimensions the system can genuinely assess:
Each carries its own score and its own specific observation — "No clear conclusion paragraph detected." "Economic impacts could be more elaborated." Every number is traceable to the thing that produced it.
If you could screenshot the range without the dimensions, the design had failed. You can't — the score and its reasoning occupy the same visual unit.
Structure
Point development
Grammar & sentence
Relevance
Evaluation Reasoning
Relevance
Structure
Content
Language
WHAT WE LOOKED FOR
A balanced answer structure: Introduction, well-segregated body points (social, economic, cultural), and a constructive conclusion.
WHAT WE FOUND IN YOUR ANSWER
Introduction was clear, and body points were separated. However, we found "No concluding paragraph identified after your third point."
HOW TO IMPROVE
A brief two-line conclusion restating your main argument (e.g. summarizing the double-edged sword of modernization) typically lifts structure marks.
Review model answer & notes for urbanisation
AI estimates structure & coverage — it can't predict precise human examiner mood.
DECISION 03
Confirm the transcription before evaluating
Between upload and evaluation sits a screen most products would skip: "Check what we read." The transcribed text is shown, editable, with uncertain words flagged inline — "3 Uncertain Words" — so attention goes where the risk is rather than asking someone to proofread everything. A toggle switches to the original image for comparison.
The reasoning: evaluating a misread answer is worse than not evaluating. A wrong transcription produces a confidently wrong assessment of something the aspirant never wrote. This is where most potential wrongness gets caught before it can do damage — and it costs one tap.
Check What We Read
OCR Step
Transcription
Original Image
PAGE 1 OF 1
• 3 Uncertain Words
Urbanisation in Northeast India has accelerated rapidly. While it brings infrastructural growth, it profoundly affects the tribal social fabric. Traditional community-owned land is increasingly converted into private holdings, leading to displacement.
Furthermore, tribal customs and dialects are diluted by cosmopolitian influences, causing marginalisation of older customs. Yet, it offers access to better higher education and jobs for the youth.
Looks right — evaluate
Retake photo
3 Uncertain Words
[Toggle to Original Image]
While it brings infrastructural growth, it profoundly affects the
Tribal Social fabric
Furthermore, tribal customs and dialects are diluted by cosmopolitian influences, causing marginalisation
DECISION 04
the one I'd most want to be judged on
When the AI can't read it, that is the feedback
Reading handwriting fails sometimes — that's a given, whatever the underlying approach. The obvious design is an error state: "We couldn't read your answer." Dead end, wasted attempt, nothing gained.
Instead: "Some of this we couldn't read." The illegible segments are shown inline, marked in the aspirant's own text. And underneath, the reframe:
Handwriting impact on marks
A human examiner may struggle with these highlighted segments too. Legibility directly affects your evaluation marks — improving these letters could raise your overall score.
The system's limitation becomes genuine exam coaching, because legibility is actually graded in the real exam. What the AI can't read, a human examiner may also struggle with. The failure isn't hidden or apologised for — it's converted into the most actionable feedback on the screen.
And the aspirant is never blocked: retake the photo, or evaluate what's readable anyway with a clear partial-assessment caveat.
Illegible Segments Detected
Legibility Coach
Some of this we couldn't read
Line 1: Urbanisation in NE has accelerated...
[Unreadable segment: "trb. soc. fbrc"]
Line 3: lead to displacement and cultural...
Handwriting impact on marks
A human examiner may struggle with these highlighted segments too. Legibility directly affects your evaluation marks — improving these letters could raise your overall score.
Partial evaluation generated based on readable parts.
Retake photo with clearer focus
Evaluate what's readable anyway
[Line 4 Segment Unreadable]
"A human examiner may struggle with these highlighted segments too. Improving legibility directly affects marks..."
DECISION 05
Disagreement needs somewhere to go
An AI that can't be questioned is an AI that has to be right. Every evaluation carries "This is wrong" alongside "See reasoning." Tapping it opens a structured disagreement — misread my answer · misunderstood the question · assessment seems unfair · something else — with an optional detail field.
See the reasoning per dimension
What the model looked for, what it found in their answer — quoted directly — and one concrete improvement. Most disagreement dissolves once the basis is visible.
Flag it — and get a response
"Feedback logged. Your flag helps make the engine smarter." Silent collection feels like shouting into a void.
Fix transcription & resubmit
Offered inside the disagreement flow, because misreading is the most common cause and the one with an immediate fix.
Illegible Segments Detected
Legibility Coach
File Disagreement
What seems wrong?
Select the reason why you feel the evaluation needs another look.
Misread my answer
Misunderstood the question
Assessment seems unfair
Something else
ADDITIONAL DETAILS (OPTIONAL)
Write any specific points the AI missed...
Feedback logged. Your flag helps make the engine smarter.
Thanks — this helps us improve the model. If it misread your answer, you can fix the transcription and try again.
Fix transcription & resubmit
A proper feedback by humans to understand what went wrong when the AI evaluated the answers
DECISION 06
State the limits on the screen, permanently
Every evaluation result carries a persistent line: "AI estimates structure & coverage — it can't predict precise human examiner mood."
Not a first-run disclaimer that disappears. Not buried in settings. On the screen, every time, in the same place — naming precisely what the system does and doesn't claim to do.
AI estimates structure & coverage — it can't predict precise human examiner mood.
DECISION 07
The free tier keeps the framework — permanently
The self-checklist doesn't sunset when AI evaluation arrives. It stays free, for everyone, forever.
This isn't generosity — it follows from the product's founding premise. NPSC Prep exists because accurate preparation was gated behind ₹50–75k coaching most aspirants can't afford. A product built to end that exclusion can't turn around and lock its core value behind a paywall; it would recreate the barrier it was made to remove.
Free — forever
The framework
The four dimensions, the model answer structure, and the discipline of evaluating your own work — genuinely the thing that improves writing over time.

Premium
The outside read
An independent assessor: someone other than you, reading what you wrote. A fast track, not a gate — and an honest reason for the price, since AI evaluation carries real cost per use.

The two tiers deliberately share visual language — same four dimensions, same structure, same layout logic — so the upgrade feels like continuity rather than a different product.
DECISION 08
Map the assessment to the actual exam paper
Each evaluation is tagged to the paper it's assessed against — "GS-III Compliant" — sitting next to the estimated band.
Small element, load-bearing purpose. NPSC's papers each reward different things: General Studies III has its own expected structure, weighting, and conventions. An evaluation that assesses "good writing" in the abstract is far less useful than one assessing good GS-III writing. It's also the narrow-but-deep thesis of the whole platform surfaced in one line: national prep apps assess generically; this assesses against this exam's actual structure.
GS-III Compliant
04 — The flow
As an aspirant moves through it
Each screen carrying the decision behind it.
Where it starts
Writing on paper
The aspirant writes their answer on paper, the way they'll write it in the exam. The app gives them the question, and — on request — the expected structure: what the intro should frame, what the body should cover, what a balanced conclusion looks like, plus the specific points a strong answer would include. This is the framework, and it's free.
Analyze the role of decentralized governance in regional economic development under NPSC GS-III criteria.
EXPECTED STRUCTURE
• Intro: Define decentralized governance with Nagaland local context (Article 371A).
• Body: Outline 3 clear economic pillars supported by evidence/statistics.
QUESTION DESK
Paper I (Mains)

Write on paper and scan once finished
The free path
Structured self-evaluation
Four questions, one per dimension, answered by the aspirant about their own work. It produces a self-score and a plain summary: "Good attempt — you self-scored highly on Relevance, but noted structural gaps and missing evidence." Genuinely useful, and bounded by the ceiling above: it can tell you what to look for, not whether you found it.
Did you directly address the question asked?
Yes
Did you state your position in the first paragraph?
Partially
Did you support each point with local evidence?
No
SELF-EVALUATION
FREE PLATFORM

Upload
Quality control at the source
Photo or PDF. The guidance line before capture is doing preventive work — every legibility problem avoided here is one that never becomes a failed evaluation. And the honesty notice appears before the aspirant has committed anything: "AI-assessed — it can make mistakes, so check the transcription on the next screen."
UPLOAD HANDWRITTEN ANSWER
PDF / JPEG

AI-assessed — it can make mistakes. You will review and check the transcription on the next screen.
Decision 03, on screen
Check what we read
EDIT TRANSCRIPTION BEFORE SYSTEM EVALUATION
The agricultural reform bill was intended to boost farmers' livings [tap to edit] by ensuring direct market access and fair pricing structures across states. However, local administrative layers faltered in executing the initial framework.
Compare with original image
Looks right — evaluate
DECISION 03: TRANSCRIPTION
3 UNCERTAIN WORDS

Decisions 01, 02, 06 & 08 at once
The evaluation result
An estimated band rather than a mark. A plain-language read beneath it. Then the four dimensions, each with its own score and its own specific observation — the decomposition that makes the number inseparable from its reasoning. The GS-III tag anchors it to the actual paper. The persistent limit line sits at the bottom, every time. And three actions: See reasoning · This is wrong · Try another.
ESTIMATED SCORE RANGE
Likely 62–70%
"A solid answer with clear points — the structure is where most of the gap sits."
Relevance: High
Structure: Fair
Grammar: Great
EVALUATION PROFILE
GS-III Compliant

AI estimates structure & coverage — it can't predict precise human examiner mood.
Decision 05, first route
Show your work
Per dimension: the criterion in plain words, what the model actually found — quoted from their answer — and one concrete improvement. Plus a link out to the model answer and notes for further reading.
WHAT WE LOOKED FOR
A dedicated concluding paragraph summarizing the core economic thesis.
WHAT WE FOUND IN YOUR ANSWER
The essay ends abruptly on economic statistics without a structured final synthesis.
CONCRETE IMPROVEMENT
DIMENSION: STRUCTURE REASONING
Mains Rubric 3.2

Decision 04 — the one to look at first
When it can't read the page
The unreadable segment is shown in place, in the aspirant's own text. Underneath, the reframe that turns the system's failure into exam coaching. And the aspirant is never blocked: retake with clearer focus, or evaluate what's readable anyway with the partial-assessment caveat visible.
[Line 4 Segment Unreadable] — handwriting too condensed
A human examiner may struggle with this segment too.
Legibility directly affects your evaluation marks in the real Mains exam. Improving the alignment of these letters will raise your overall performance.
Retake Photo
Evaluate readable sections anyway
LEGIBILITY COACH
OCR FAILURE

Decision 05, routes two and three
Filing a disagreement
Structured reasons rather than a free-text void, an optional detail field, and — the part most products skip — a response: the flag is acknowledged, and where transcription is the likely cause, the fix is offered immediately.
Why is this assessment wrong?
Misread my handwriting (Transcription error)
Misunderstood the regional question context
Point assessment was too severe / rigid
Your input helps calibrate the grading engine
Submit Feedback
CONTEST ASSESSMENT
Feedback Loop

05 — Outcome
Designed, not yet built
NPSC Prep itself is in production and launching soon; this layer is specced and designed, ready for implementation. So the honest measure isn't usage — it's whether the thinking holds up.
What the work produced: a complete evaluation flow built on a single organising principle rather than assembled feature by feature. Eight decisions that each answer the same question through a different route — preventing wrongness before it happens, making the reasoning inspectable, converting failure into usefulness, and giving disagreement somewhere to go. And a tier structure where the free path isn't a demo of the paid one; it's the same framework, assessed by a different party.
What I'd measure once it ships
1
Whether aspirants actually use the transcription confirm step, or tap straight through it — because if they skip it, the main defence against confident wrongness is theoretical.
2
Whether the range genuinely resists anchoring, or people mentally collapse "62–70" into "66."
3
How often the disagreement path gets used, and what the reasons cluster around — the fastest read on where the model is actually failing.
4
Whether the legibility feedback changes behaviour: do aspirants whose handwriting is flagged actually improve it?
06 — The honest limits
Three, stated plainly
It hasn't met a single real user
Every decision here is reasoned, not validated. The self-checklist's ceiling — that self-assessment is bounded by self-awareness — is a product insight, not a research finding. It's a well-grounded one, but testing whether it shows up in real usage, and whether AI feedback actually closes it, is the first thing to do after launch.
The evaluation's accuracy is unvalidated against real examiner scoring
This is the biggest one, and it sits underneath everything else. The system estimates structure and coverage — but nobody has yet compared its bands against what NPSC examiners actually award for the same answers. Until that comparison exists, "Likely 62–70%" is an informed estimate, not a calibrated one. The design is honest about this on every result screen; the product still needs to earn the number. Running a calibration study against marked answers is the essential next step before this can be trusted at scale.
The disagreement path is unproven at volume
Flagging works as a design; whether the flags produce signal useful enough to actually improve the model, and whether aspirants trust a system that acknowledges its own errors more or less than one that projects confidence, are open questions.
"The design is built for a system that will sometimes be wrong. Whether it's wrong in the ways I anticipated — that part I don't know yet."
NPSC Prep — AI-UX Case Study
Close