Designing AI evaluation layer

Designing AI evaluation layer

INTRODUCTION

01

this case study covers

Designing a probabilistic AI evaluation layer for handwritten NPSC Mains answers.

The Business Impact

Accessible to premium exam feedback, disrupting the ₹50,000+ traditional coaching market by offering AI-driven grading and study materials for just ₹150–200/month.

The Challenge

AI is confidently wrong sometimes. Misjudging a handwritten answer doesn’t just give wrong feedbacks; but it actively disrupt months of high-stakes exam preparation.

AI is confidently wrong sometimes. Misjudging a handwritten answer doesn’t just give wrong feedbacks; but it actively disrupt months of high-stakes exam preparation.

NOTE: The Project is Vast and there are multiple features that can be a complete case study but i choose this because it was one of the interesting feature.

THE PROBLEM SPACE

02

The Problem Space

The Problem Space

NPSC aspirants write long answers by hand. They have no way to know if the answer is good until an examiner marks it months later.

But AI is different from normal software

Normal software gives one right answer

AI gives a confident answer that is sometimes wrong

And the stakes are high

A misread page gets judged on words the person never wrote, and it leads to A wrong score misdirects months of study

A misread page gets judged on words the person never wrote, and it leads to A wrong score misdirects months of study

asking myself

03

THE QUESTIONS

Before I even start designing I need to ask some question on what I should do?

  • Do I Hide that AI can make mistakes but the Score it give should be trusted ?

  • Give them warning about the score and other suggestion it provides, to cover the mistake.

Every decision here answers one thing: what happens when the AI is wrong?

WHAT I DISCARDED

03

WHY NOT JUST A CHECKLIST


A free self-check came first. Four questions about your own answer.
It works, but it has a limit:

It asks you to judge your work using knowledge you don't have yet.

  • Someone who doesn't know their ending is weak will still tick "yes"

  • The other limitation is you can only spot what you already understand

EXPLORATION

03

Explorations & Trade-offs

Explorations & Trade-offs

There was many ways through which this could have been implemented this so here are some of the explorations.

• Trade-off 1: Instant AI Evaluation on Photos Uploads

• Trade-off 1: Instant AI Evaluation on Photos Uploads

Explored: Running the AI evaluation instantly upon photo upload for a frictionless experience.

Why discarded: If the AI misreads the handwriting, the subsequent evaluation is cascaded garbage.

The Pivot: Require users to confirm, edit, or retake the transcription before grading.

• Trade-off 2: Handling Misreads of notes

• Trade-off 2: Handling Misreads of notes

Explored: A standard “Error: Cannot Read” is a dead end when transcription completely fails.

Explored: A standard “Error: Cannot Read” is a dead end when transcription completely fails.

Why discarded: It blames the user and traps them.

The Pivot:

A dual fail-safe. If illegible, the system warns: “not legible enough to transcribe.” Users can retake the picture or proceed with a caveat.

The Pivot:

A dual fail-safe. If illegible, the system warns: “not legible enough to transcribe.” Users can retake the picture or proceed with a caveat.

• Trade-off 3: The Scoring system

• Trade-off 3: The Scoring system

Explored: Giving a precise definitive score (e.g., “68%”).

Explored: Giving a precise definitive score (e.g., “68%”).

Why discarded: It implies mathematical precision the uncalibrated AI doesn’t possess.

Why discarded: It implies mathematical precision the uncalibrated AI doesn’t possess.

The Pivot: Delivering a decomposed probabilistic range (“Likely 62–70%”), resisting numerical anchoring.

The Pivot: Delivering a decomposed probabilistic range (“Likely 62–70%”), resisting numerical anchoring.

the final solution

04

The Final Solution

The Final Solution

There was many ways through which this could have been implemented this so here are some of the explorations.

Where it all Started

The free path

The free path

Four questions: relevance, structure, evidence, clarity. Self-score and summary. Useful but limited by your own observations.

Upload

Prevent problems before they happen

Prevent problems before they happen

Submit a photo or PDF. Ensure the page is flat, well-lit, and fully framed. Unreadable pages lead to failed evaluations. A notice states AI can make errors.

Submit a photo or PDF. Ensure the page is flat, well-lit, and fully framed. Unreadable pages lead to failed evaluations. A notice states AI can make errors.

Transcription check

See what was read, fix what’s wrong.

See what was read, fix what’s wrong.

The AI reads the text, fully editable. Words it flagged are highlighted for you to check and edit if needed with A toggle compares to your original photo.

Result

The score, with its reasons attached.

The score, with its reasons attached.

Estimated band “62–70%” with a summary of the gap.

Four dimensions with scores and a note:
“No clear conclusion“
A “GS-III Compliant” tag links to the paper. Limit line at the bottom Actions:
see reasoning, this is wrong, try another.

Estimated band “62–70%” with a summary of the gap.

Four dimensions with scores and a note:
No clear conclusion
A “GS-III Compliant” tag links to the paper. Limit line at the bottom Actions:
see reasoning, this is wrong, try another.

Reasoning per dimension

Why each score landed where it did.

Why each score landed where it did.

Choose a dimension to view three aspects: AI’s search criteria, findings from your answer, and a suggested improvement, along with a link to the model answer.

Choose a dimension to view three aspects: AI’s search criteria, findings from your answer, and a suggested improvement, along with a link to the model answer.

Unreadable handwriting

The failure that teaches something.

The failure that teaches something.

Unreadable parts in your text may affect marks. A human examiner might struggle too. Retake the photo or check what’s readable with a warning.

Unreadable parts in your text may affect marks. A human examiner might struggle too. Retake the photo or check what’s readable with a warning.

Disagreement

A way to push back

A way to push back

Four reasons misread my answer and misunderstood the question. A reply confirms it was logged, and if transcription caused it, the fix is offered.

OUTCOME

05

Designed, not built yet.

Designed, not built yet.

NPSC Prep is in production; this layer is ready to build.

NPSC Prep is in production; this layer is ready to build.

What I’d measure at launch:

What I’d measure at launch:

Do people use the transcription check, or skip it?

Does the range work, or do people round it to one number?

How often do people disagree — and why?

Does the handwriting feedback change anything?

THE LIMIT

06

Honest Limits & Metrics

Honest Limits & Metrics

No real user has used it. Every decision is reasoned, not tested.

The scores aren’t calibrated. Nobody has compared them to what real examiners give.

No real user has used it. Every decision is reasoned, not tested.

The scores aren’t calibrated. Nobody has compared them to what real examiners give.