INTRODUCTION
01
this case study covers
Designing a probabilistic AI evaluation layer for handwritten NPSC Mains answers.
The Business Impact
Accessible to premium exam feedback, disrupting the ₹50,000+ traditional coaching market by offering AI-driven grading and study materials for just ₹150–200/month.
The Challenge
NOTE: The Project is Vast and there are multiple features that can be a complete case study but i choose this because it was one of the interesting feature.
THE PROBLEM SPACE
02
NPSC aspirants write long answers by hand. They have no way to know if the answer is good until an examiner marks it months later.
But AI is different from normal software
Normal software gives one right answer
AI gives a confident answer that is sometimes wrong
And the stakes are high
asking myself
03
THE QUESTIONS
Before I even start designing I need to ask some question on what I should do?
Do I Hide that AI can make mistakes but the Score it give should be trusted ?
Give them warning about the score and other suggestion it provides, to cover the mistake.
Every decision here answers one thing: what happens when the AI is wrong?
WHAT I DISCARDED
03
WHY NOT JUST A CHECKLIST
A free self-check came first. Four questions about your own answer.
It works, but it has a limit:
It asks you to judge your work using knowledge you don't have yet.
Someone who doesn't know their ending is weak will still tick "yes"
The other limitation is you can only spot what you already understand
EXPLORATION
03
There was many ways through which this could have been implemented this so here are some of the explorations.
Explored: Running the AI evaluation instantly upon photo upload for a frictionless experience.
Why discarded: If the AI misreads the handwriting, the subsequent evaluation is cascaded garbage.
The Pivot: Require users to confirm, edit, or retake the transcription before grading.
Why discarded: It blames the user and traps them.
the final solution
04
There was many ways through which this could have been implemented this so here are some of the explorations.
Where it all Started
Four questions: relevance, structure, evidence, clarity. Self-score and summary. Useful but limited by your own observations.

Upload

Transcription check
The AI reads the text, fully editable. Words it flagged are highlighted for you to check and edit if needed with A toggle compares to your original photo.

Result

Reasoning per dimension

Unreadable handwriting

Disagreement
Four reasons misread my answer and misunderstood the question. A reply confirms it was logged, and if transcription caused it, the fix is offered.
OUTCOME
05
Do people use the transcription check, or skip it?
Does the range work, or do people round it to one number?
How often do people disagree — and why?
Does the handwriting feedback change anything?
THE LIMIT
06
