OCR & Speech Workspace

A 400-page scanned PDF is where most OCR demos quietly die: one giant request, one timeout, nothing to show.

Posted onLinkedIn
OCR & Speech Workspace, slide 1 of 9
Slides
01 / 09
OCR & Speech Workspace, slide 1 of 9
01
OCR & Speech Workspace, slide 2 of 9
02
OCR & Speech Workspace, slide 3 of 9
03
OCR & Speech Workspace, slide 4 of 9
04
OCR & Speech Workspace, slide 5 of 9
05
OCR & Speech Workspace, slide 6 of 9
06
OCR & Speech Workspace, slide 7 of 9
07
OCR & Speech Workspace, slide 8 of 9
08
OCR & Speech Workspace, slide 9 of 9
09
The caption

A 400-page scanned PDF is where most OCR demos quietly die: one giant request, one timeout, nothing to show. The OCR & Speech Workspace splits it into 20-page batches, runs up to four at once, and shows progress as pages land. Speech works the same way — a realtime model for latency while you talk, a more accurate batch model that re-transcribes the moment you stop. Then every document becomes a scoped knowledge base you can chat with, citing the exact page. All on Mistral's models. Here's the architecture and the numbers. Built at NextUpgrad Web Solutions. I also build tools.scult.in — 15 free tools, 1,200+ prompts and a 50,000-skill library, free for anyone, no signup.

Posted onLinkedIn
Keep flipping

More carousels

All 59 decks →
Campus talks · hackathon judging · mentoring

Want this as a live session at your college?

Open to GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges.

Pranjul Rathour, GenAI engineer in Kanpur