
Surya : OCR and line detection in 90+ languages
Date : 2024-01-10
Description
This summary was drafted with mixtral-8x7b-instruct-v0.1.Q5_K_M.gguf
Surya is an open-source document OCR toolkit developed by Vik Paruchuri that offers accurate OCR in 90+ languages and line-level text detection in any language. It supports a range of documents, including images, PDFs, and folders of images/PDFs, and is capable of detecting tables and charts (coming soon). The toolkit includes a streamlit app for interactive use, making it accessible to users who want to try Surya on their images or PDF files. Surya's name comes from the Hindu sun god, who has universal vision.
GitHub repo here
Recently on :
Artificial Intelligence
Information Processing | Computing
WEB - 2025-11-13
Measuring political bias in Claude
Anthropic gives insights into their evaluation methods to measure political bias in models.
WEB - 2025-10-09
Defining and evaluating political bias in LLMs
OpenAI created a political bias evaluation that mirrors real-world usage to stress-test their models’ ability to remain objecti...
WEB - 2025-07-23
Preventing Woke AI In Federal Government
Citing concerns that ideological agendas like Diversity, Equity, and Inclusion (DEI) are compromising accuracy, this executive ...
WEB - 2025-07-10
America’s AI Action Plan
To win the global race for technological dominance, the US outlined a bold national strategy for unleashing innovation, buildin...
WEB - 2024-12-30
Fine-tune ModernBERT for text classification using synthetic data
David Berenstein explains how to finetune a ModernBERT model for text classification on a synthetic dataset generated from argi...