indicTranslate v1 - Machine Translation for 11 Indic languages. For latest v2, check: https://github.com/AI4Bharat/IndicTrans2
-
Updated
Jan 2, 2024 - Jupyter Notebook
indicTranslate v1 - Machine Translation for 11 Indic languages. For latest v2, check: https://github.com/AI4Bharat/IndicTrans2
A pipeline for transliteration, spell correction, POS tagging and word sense disambiguation of Hinglish code mixed data to Hindi Devanagari script.
Vyākarana: A Colorless Green Benchmark for Syntactic Evaluation in Indic Languages
Non-contextual : Word2Vec, FastText Contextual : BERT, RoBERTa, ELECTRA, CamemBERT, Distil-BERT, XLM-RoBERTa Analyzed embedding models, used the best one to build a Flask web app for Hindi NER and data collection from user feedback, deployed on AWS.
Bharat Multimodal EO AI - ISRO Satellite Vision + Indic NLP for Disaster Management, Agriculture & Climate Monitoring
Lightweight on-device Hindi TTS for Android & iOS — fine-tuned on AI4Bharat IndicVoices, ONNX export, runs offline on CPU in real-time.
Ultra-Fast Indic Text-to-Speech Engine with Zero-Shot Voice Cloning
Contextualized Topic Modeling using Zero-Shot Learning on Indic Languages (IndicCTM)
A high-performance, community-driven tokenizer for Indian languages. Built for NLP, LLMs, and multilingual AI.
HumanCTO's Indic Voice Pipeline — Download, transcribe, and translate audio/video in 12 Indian languages. 100% local, no API keys. Claude Code skills powered by OpenAI Whisper + vasista22 + AI4Bharat IndicWhisper fine-tuned models.
Python toolkit to decode legacy Hindi font-encoded PDFs (KrutiDev, Chanakya, DevLys) into Unicode Devanagari. Built for Hindi PDF & govt document ingestion pipelines.
Deterministic, explainable, CPU-only Roman-Hinglish text normalization engine (MIT).
A full-stack ML system for Hindi news classification powered by fine-tuned IndicBERTv2. Features real-time multi-modal inputs: manual text entry, live scraping from major news portals, and OCR-based heading extraction. Deployed via Docker on Hugging Face Spaces.
Multi-agent RAG system for high-quality Telugu story generation — Planner → Drafter → Critic loop
SageGPT: A ~7.2M parameter Sanskrit-only decoder-only Transformer SLM trained from scratch on a 139 MB tokenized Sanskrit corpus containing 72.8M SentencePiece model-token IDs, derived from a 105.2M-character purified Sanskrit text corpus using NVIDIA DGX Spark. Architecture: 6 layers, 8 attention heads, 256 embedding dim, 1024 context, 8K vocab.
WhatsApp-native hyperlocal labour marketplace for India - workers register free, employers pay ₹20 to get matched contacts. No app download, no login. Built on FastAPI, Supabase, pgvector, and Razorpay.
Voice-first health guidance and surveillance dashboard
Classical Telugu prosody (chandassu) identifier — detects and classifies traditional poetic meters using NLP
A production-ready, frugal, sovereign AI system that orchestrates India's open-source language models to achieve state-of-the-art reasoning on consumer hardware through Test-Time Compute (TTC) and Cognitive Serialization.
Add a description, image, and links to the indic-nlp topic page so that developers can more easily learn about it.
To associate your repository with the indic-nlp topic, visit your repo's landing page and select "manage topics."