CASE STUDY

Designing AI evaluation for high-stakes exam writing

An AI layer that reads and assesses handwritten exam answers — where the design problem isn't the feature, it's what happens when the model gets it wrong.

Role

Product designer & lead

Timeline

~1 month, design & planning

Status

Designed · Under Productions

Tools

Figma

01 — Framing

The hard part was never the feature

NPSC Prep is a mobile-first platform for aspirants preparing for the Nagaland Public Service Commission exams.


This study is about a single layer of it: the AI. Specifically, the feature that reads an aspirant's handwritten descriptive answer and tells them how they did.


Most AI design writing is about the wrapper — the chat panel, the sparkle icon, the prompt box. That's not the hard part. The hard part is that AI is a fundamentally different design material from anything we're trained on:

A deterministic interface has one correct output. A probabilistic one is confidently wrong sometimes.

And the stakes here make that concrete. An NPSC aspirant has years and a career riding on this exam. If an AI evaluates their Mains answer and gets it wrong — misreads their handwriting, misjudges their structure, scores them higher or lower than a real examiner would — that isn't an inconvenience. It misdirects months of preparation.

The design problem

How do you build something genuinely useful on top of a system you know will sometimes be wrong — without either hiding the uncertainty or making it useless? Hide the fallibility and you build false confidence into someone's exam preparation. Foreground it too heavily and the feature becomes so hedged it tells them nothing. The design lives between those two failures.


Every decision in this layer answers one question: what happens when the model is wrong? Not as a footnote or an error state — as the organising principle. Each screen exists because of a different answer to it: one prevents wrongness before it happens, one lets the aspirant check the model's work, one turns a failure into something useful, one lets them contest the result.

02 — Why the checklist wasn't enough

The AI didn't come first

A self-evaluation checklist did — and it still ships, free, to every user.


After writing an answer, the aspirant sees the model answer structure and the points a strong response would cover, then walks through four questions about their own work. Did you directly address the question asked? Did you state your position in the first paragraph? Did you support each point with evidence? Did you proofread for clarity? Yes, partially, or no. It produces a self-score and a plain-language summary.


That's genuinely useful. It gives an aspirant a rubric where they had none, it costs nothing, and it makes them read their own answer critically — a skill the exam rewards.


But it has a ceiling built into its premise:

It asks you to evaluate work using judgment you don't yet have.

An aspirant who doesn't know their conclusion is weak will confidently tick "yes, I stated my position." Someone who thinks a point is well-evidenced because they believe it will mark it yes. Self-assessment is bounded by self-awareness — and the people who most need feedback are precisely the ones least equipped to give it to themselves. The checklist can tell you what to look for. It can't tell you whether you found it.

What the AI actually adds

Not a better rubric — the dimensions are identical in both tiers. Not more questions. The same four questions, answered by something other than the person who wrote the answer. An outside read.


A note on how this was reached: this framing came from a product insight, not from user research. The survey behind NPSC Prep validated the access problem — fragmented sources, cost barriers — not this one. Testing whether the checklist's ceiling shows up in real usage, and whether AI feedback closes it, is named in the limitations below.

03 — The decisions

Eight decisions, one question

Each answers what happens when the model is wrong from a different angle.

DECISION 01

A range, not a score

The evaluation returns "Likely 62–70%" — never a single number. A point score implies a precision the system doesn't have; in a competitive exam, a false-precise "68" invites an aspirant to treat a machine estimate as a verdict. A range does the job people actually need — am I in the passing zone? — while encoding the uncertainty honestly in its shape. And it's harder to fixate on: you can't screenshot a range and treat it as your mark.


Paired with it is a plain-language line: "A solid answer with clear points — the structure is where most of the gap sits." Calibration in numbers, direction in words.

Evaluation Summary

Mains Assessment

ESTIMATED BAND

GS-III COMPLIANT

Likely 62–70%

A solid answer with clear points — the structure is where most of the gap sits. Your coverage of the economic dimension was thin.

DIMENSION-WISE SCORE BREAKDOWN

Relevance

85/100

Directly addresses urbanisation dynamics well.

Structure

45/100

No clear conclusion paragraph detected.

Point Development

70/100

Economic impacts could be more elaborate.

Grammar & Sentence

90/100

Excellent flow and highly coherent style.

AI estimates structure & coverage — it can't predict precise human examiner mood.

"A solid answer with clear points — the structure is where most of the gap sits."

DECISION 02

The score is a decomposition, never a verdict

The range isn't a badge at the top of the page — it's constructed in front of the aspirant from the four dimensions the system can genuinely assess:

Each carries its own score and its own specific observation — "No clear conclusion paragraph detected." "Economic impacts could be more elaborated." Every number is traceable to the thing that produced it.

If you could screenshot the range without the dimensions, the design had failed. You can't — the score and its reasoning occupy the same visual unit.

Structure

— Intro, body, conclusion

Intro, body, conclusion

Point development

— How ideas were built and supported

How ideas were built and supported

Grammar & sentence

— Construction and clarity

Construction and clarity

Relevance

— Did it address the question actually asked

Did it address the question actually asked

Evaluation Reasoning

Relevance

Structure

Content

Language

WHAT WE LOOKED FOR

A balanced answer structure: Introduction, well-segregated body points (social, economic, cultural), and a constructive conclusion.

WHAT WE FOUND IN YOUR ANSWER

Introduction was clear, and body points were separated. However, we found "No concluding paragraph identified after your third point."

HOW TO IMPROVE

A brief two-line conclusion restating your main argument (e.g. summarizing the double-edged sword of modernization) typically lifts structure marks.

Review model answer & notes for urbanisation

AI estimates structure & coverage — it can't predict precise human examiner mood.

DECISION 03

Confirm the transcription before evaluating

Between upload and evaluation sits a screen most products would skip: "Check what we read." The transcribed text is shown, editable, with uncertain words flagged inline — "3 Uncertain Words" — so attention goes where the risk is rather than asking someone to proofread everything. A toggle switches to the original image for comparison.


The reasoning: evaluating a misread answer is worse than not evaluating. A wrong transcription produces a confidently wrong assessment of something the aspirant never wrote. This is where most potential wrongness gets caught before it can do damage — and it costs one tap.

Check What We Read

OCR Step

Transcription

Original Image

PAGE 1 OF 1

• 3 Uncertain Words

Urbanisation in Northeast India has accelerated rapidly. While it brings infrastructural growth, it profoundly affects the tribal social fabric. Traditional community-owned land is increasingly converted into private holdings, leading to displacement.


Furthermore, tribal customs and dialects are diluted by cosmopolitian influences, causing marginalisation of older customs. Yet, it offers access to better higher education and jobs for the youth.

Looks right — evaluate

Retake photo

3 Uncertain Words

[Toggle to Original Image]

While it brings infrastructural growth, it profoundly affects the

Tribal Social fabric

Furthermore, tribal customs and dialects are diluted by cosmopolitian influences, causing marginalisation

DECISION 04

the one I'd most want to be judged on

When the AI can't read it, that is the feedback

Reading handwriting fails sometimes — that's a given, whatever the underlying approach. The obvious design is an error state: "We couldn't read your answer." Dead end, wasted attempt, nothing gained.


Instead: "Some of this we couldn't read." The illegible segments are shown inline, marked in the aspirant's own text. And underneath, the reframe:

Handwriting impact on marks

A human examiner may struggle with these highlighted segments too. Legibility directly affects your evaluation marks — improving these letters could raise your overall score.

The system's limitation becomes genuine exam coaching, because legibility is actually graded in the real exam. What the AI can't read, a human examiner may also struggle with. The failure isn't hidden or apologised for — it's converted into the most actionable feedback on the screen.


And the aspirant is never blocked: retake the photo, or evaluate what's readable anyway with a clear partial-assessment caveat.

Illegible Segments Detected

Legibility Coach

Some of this we couldn't read

Line 1: Urbanisation in NE has accelerated...

[Unreadable segment: "trb. soc. fbrc"]

Line 3: lead to displacement and cultural...

Handwriting impact on marks

A human examiner may struggle with these highlighted segments too. Legibility directly affects your evaluation marks — improving these letters could raise your overall score.

Partial evaluation generated based on readable parts.

Retake photo with clearer focus

Evaluate what's readable anyway

[Line 4 Segment Unreadable]

"A human examiner may struggle with these highlighted segments too. Improving legibility directly affects marks..."

DECISION 05

Disagreement needs somewhere to go

An AI that can't be questioned is an AI that has to be right. Every evaluation carries "This is wrong" alongside "See reasoning." Tapping it opens a structured disagreement — misread my answer · misunderstood the question · assessment seems unfair · something else — with an optional detail field.

See the reasoning per dimension

What the model looked for, what it found in their answer — quoted directly — and one concrete improvement. Most disagreement dissolves once the basis is visible.

Flag it — and get a response

"Feedback logged. Your flag helps make the engine smarter." Silent collection feels like shouting into a void.

Fix transcription & resubmit

Offered inside the disagreement flow, because misreading is the most common cause and the one with an immediate fix.

Illegible Segments Detected

Legibility Coach

File Disagreement

What seems wrong?

Select the reason why you feel the evaluation needs another look.

Misread my answer

Misunderstood the question

Assessment seems unfair

Something else

ADDITIONAL DETAILS (OPTIONAL)

Write any specific points the AI missed...

Feedback logged. Your flag helps make the engine smarter.

Thanks — this helps us improve the model. If it misread your answer, you can fix the transcription and try again.

Fix transcription & resubmit

A proper feedback by humans to understand what went wrong when the AI evaluated the answers

DECISION 06

State the limits on the screen, permanently

Every evaluation result carries a persistent line: "AI estimates structure & coverage — it can't predict precise human examiner mood."


Not a first-run disclaimer that disappears. Not buried in settings. On the screen, every time, in the same place — naming precisely what the system does and doesn't claim to do.

AI estimates structure & coverage — it can't predict precise human examiner mood.

DECISION 07

The free tier keeps the framework — permanently

The self-checklist doesn't sunset when AI evaluation arrives. It stays free, for everyone, forever.


This isn't generosity — it follows from the product's founding premise. NPSC Prep exists because accurate preparation was gated behind ₹50–75k coaching most aspirants can't afford. A product built to end that exclusion can't turn around and lock its core value behind a paywall; it would recreate the barrier it was made to remove.

Free — forever

The framework

The four dimensions, the model answer structure, and the discipline of evaluating your own work — genuinely the thing that improves writing over time.

Premium

The outside read

An independent assessor: someone other than you, reading what you wrote. A fast track, not a gate — and an honest reason for the price, since AI evaluation carries real cost per use.

The two tiers deliberately share visual language — same four dimensions, same structure, same layout logic — so the upgrade feels like continuity rather than a different product.

DECISION 08

Map the assessment to the actual exam paper

Each evaluation is tagged to the paper it's assessed against — "GS-III Compliant" — sitting next to the estimated band.


Small element, load-bearing purpose. NPSC's papers each reward different things: General Studies III has its own expected structure, weighting, and conventions. An evaluation that assesses "good writing" in the abstract is far less useful than one assessing good GS-III writing. It's also the narrow-but-deep thesis of the whole platform surfaced in one line: national prep apps assess generically; this assesses against this exam's actual structure.

GS-III Compliant

04 — The flow

As an aspirant moves through it

Each screen carrying the decision behind it.

Where it starts

Writing on paper

The aspirant writes their answer on paper, the way they'll write it in the exam. The app gives them the question, and — on request — the expected structure: what the intro should frame, what the body should cover, what a balanced conclusion looks like, plus the specific points a strong answer would include. This is the framework, and it's free.

Analyze the role of decentralized governance in regional economic development under NPSC GS-III criteria.

EXPECTED STRUCTURE

• Intro: Define decentralized governance with Nagaland local context (Article 371A).

• Body: Outline 3 clear economic pillars supported by evidence/statistics.

QUESTION DESK

Paper I (Mains)

Write on paper and scan once finished

The free path

Structured self-evaluation

Four questions, one per dimension, answered by the aspirant about their own work. It produces a self-score and a plain summary: "Good attempt — you self-scored highly on Relevance, but noted structural gaps and missing evidence." Genuinely useful, and bounded by the ceiling above: it can tell you what to look for, not whether you found it.

Did you directly address the question asked?

Yes

Did you state your position in the first paragraph?

Partially

Did you support each point with local evidence?

No

SELF-EVALUATION

FREE PLATFORM

Upload

Quality control at the source

Photo or PDF. The guidance line before capture is doing preventive work — every legibility problem avoided here is one that never becomes a failed evaluation. And the honesty notice appears before the aspirant has committed anything: "AI-assessed — it can make mistakes, so check the transcription on the next screen."

UPLOAD HANDWRITTEN ANSWER

PDF / JPEG

AI-assessed — it can make mistakes. You will review and check the transcription on the next screen.

Decision 03, on screen

Check what we read

The transcription is editable, uncertain words are marked inline so attention goes where the risk is, and a toggle compares against the original image. "Fix anything we got wrong before we evaluate. Evaluating a misread answer helps nobody." Then: Looks right — evaluate.

The transcription is editable, uncertain words are marked inline so attention goes where the risk is, and a toggle compares against the original image. "Fix anything we got wrong before we evaluate. Evaluating a misread answer helps nobody." Then: Looks right — evaluate.

EDIT TRANSCRIPTION BEFORE SYSTEM EVALUATION

The agricultural reform bill was intended to boost farmers' livings [tap to edit] by ensuring direct market access and fair pricing structures across states. However, local administrative layers faltered in executing the initial framework.

Compare with original image

Looks right — evaluate

DECISION 03: TRANSCRIPTION

3 UNCERTAIN WORDS

Decisions 01, 02, 06 & 08 at once

The evaluation result

An estimated band rather than a mark. A plain-language read beneath it. Then the four dimensions, each with its own score and its own specific observation — the decomposition that makes the number inseparable from its reasoning. The GS-III tag anchors it to the actual paper. The persistent limit line sits at the bottom, every time. And three actions: See reasoning · This is wrong · Try another.

ESTIMATED SCORE RANGE

Likely 62–70%

"A solid answer with clear points — the structure is where most of the gap sits."

Relevance: High

Structure: Fair

Grammar: Great

EVALUATION PROFILE

GS-III Compliant

AI estimates structure & coverage — it can't predict precise human examiner mood.

Decision 05, first route

Show your work

Per dimension: the criterion in plain words, what the model actually found — quoted from their answer — and one concrete improvement. Plus a link out to the model answer and notes for further reading.

WHAT WE LOOKED FOR

A dedicated concluding paragraph summarizing the core economic thesis.

WHAT WE FOUND IN YOUR ANSWER

The essay ends abruptly on economic statistics without a structured final synthesis.

CONCRETE IMPROVEMENT

Add a 3-sentence closing mapping local results directly to national economic indexes.

Add a 3-sentence closing mapping local results directly to national economic indexes.

DIMENSION: STRUCTURE REASONING

Mains Rubric 3.2

Decision 04 — the one to look at first

When it can't read the page

The unreadable segment is shown in place, in the aspirant's own text. Underneath, the reframe that turns the system's failure into exam coaching. And the aspirant is never blocked: retake with clearer focus, or evaluate what's readable anyway with the partial-assessment caveat visible.

[Line 4 Segment Unreadable] — handwriting too condensed

A human examiner may struggle with this segment too.

Legibility directly affects your evaluation marks in the real Mains exam. Improving the alignment of these letters will raise your overall performance.

Retake Photo

Evaluate readable sections anyway

LEGIBILITY COACH

OCR FAILURE

Decision 05, routes two and three

Filing a disagreement

Structured reasons rather than a free-text void, an optional detail field, and — the part most products skip — a response: the flag is acknowledged, and where transcription is the likely cause, the fix is offered immediately.

Why is this assessment wrong?

Misread my handwriting (Transcription error)

Misunderstood the regional question context

Point assessment was too severe / rigid

Your input helps calibrate the grading engine

Submit Feedback

CONTEST ASSESSMENT

Feedback Loop

05 — Outcome

Designed, not yet built

NPSC Prep itself is in production and launching soon; this layer is specced and designed, ready for implementation. So the honest measure isn't usage — it's whether the thinking holds up.

What the work produced: a complete evaluation flow built on a single organising principle rather than assembled feature by feature. Eight decisions that each answer the same question through a different route — preventing wrongness before it happens, making the reasoning inspectable, converting failure into usefulness, and giving disagreement somewhere to go. And a tier structure where the free path isn't a demo of the paid one; it's the same framework, assessed by a different party.

What I'd measure once it ships

1

Whether aspirants actually use the transcription confirm step, or tap straight through it — because if they skip it, the main defence against confident wrongness is theoretical.

2

Whether the range genuinely resists anchoring, or people mentally collapse "62–70" into "66."

3

How often the disagreement path gets used, and what the reasons cluster around — the fastest read on where the model is actually failing.

4

Whether the legibility feedback changes behaviour: do aspirants whose handwriting is flagged actually improve it?

06 — The honest limits

Three, stated plainly

It hasn't met a single real user

Every decision here is reasoned, not validated. The self-checklist's ceiling — that self-assessment is bounded by self-awareness — is a product insight, not a research finding. It's a well-grounded one, but testing whether it shows up in real usage, and whether AI feedback actually closes it, is the first thing to do after launch.

The evaluation's accuracy is unvalidated against real examiner scoring

This is the biggest one, and it sits underneath everything else. The system estimates structure and coverage — but nobody has yet compared its bands against what NPSC examiners actually award for the same answers. Until that comparison exists, "Likely 62–70%" is an informed estimate, not a calibrated one. The design is honest about this on every result screen; the product still needs to earn the number. Running a calibration study against marked answers is the essential next step before this can be trusted at scale.

The disagreement path is unproven at volume

Flagging works as a design; whether the flags produce signal useful enough to actually improve the model, and whether aspirants trust a system that acknowledges its own errors more or less than one that projects confidence, are open questions.

"The design is built for a system that will sometimes be wrong. Whether it's wrong in the ways I anticipated — that part I don't know yet."

NPSC Prep — AI-UX Case Study

Close