Back to Engineering Projects
NLP Research📚 AI Research
GMSAC — Gujarati Sentiment Analysis Corpus & Benchmark
A manually curated Gujarati sentiment analysis corpus with benchmark evaluations across multiple transformer models to support reproducible NLP research for low-resource languages.
Benchmark Results & Metrics
Corpus Size
15k+ Sentences
Fleiss Kappa
0.86 Agreement
Baseline Models
5 Benchmarked
Best F1-Score
0.82 (MuRIL)
01 // Overview
GMSAC provides a benchmark and resource corpus for Gujarati NLP, containing thousands of manually labeled entries. The study evaluates transformer backbones like MuRIL and IndicBERT.
02 // The Problem
Low-resource languages like Gujarati suffer from a lack of labeled datasets, forcing developers to rely on suboptimal machine translation models that miss cultural nuances.
03 // System Architecture
NLP Dataset Benchmarking Pipeline
Data was scraped and annotated by native speakers. Standardized pre-processing scripts clean Gujarati script, feeding into fine-tuned Indic transformer models.
ARCHITECTURAL FLOW DIAGRAM:
graph TD
RawText["Raw Gujarati Text"] --> Preprocessor["Data Preprocessor"]
Preprocessor --> Encoder["MuRIL / IndicBERT Encoder"]
Encoder --> Classifier["Sentiment Classifier"]
Classifier --> Metrics["F1 Evaluation & Benchmarks"]
style RawText fill:#0f172a,stroke:#3b82f6,color:#f8fafc
style Preprocessor fill:#0f172a,stroke:#38bdf8,color:#f8fafc
style Encoder fill:#1e1b4b,stroke:#6366f1,color:#f8fafc
style Classifier fill:#0f172a,stroke:#a855f7,color:#f8fafc
style Metrics fill:#064e3b,stroke:#10b981,color:#f8fafc
Architecture Components:
Scraping & Cleaning Scripts
Human Annotation Guidelines
IndicBERT/MuRIL Fine-tuning Hub
Evaluation Suite
04 // Implementation
Built annotation pipelines using Python and custom tooling. Fine-tuned multilingual BERT variants (mBERT, XLM-RoBERTa, MuRIL) using PyTorch and Hugging Face.
05 // Execution Workflow
11. Scrape text from news portals and social media.
22. Clean noise and format text.
33. Conduct double-blind human annotation.
44. Resolve label conflicts via curator review.
55. Train and evaluate transformer models on splits.
06 // Technology Stack
NLP & Transformers
Hugging FacePyTorchMuRILIndicBERT
Dataset & Scraping
BeautifulSoupPandasScikit-Learn
07 // Technical Challenges
⚠️Standardizing spellings and regional dialects in informal Gujarati social comments.
⚠️Ensuring high inter-annotator agreement among human labelers.
Lessons Learned
- •Pre-training language models on local scripts performs much better than using English-translated inputs.
- •Rigorous cleaning of non-Gujarati characters is vital for training stability.
Future Improvements
- •Expanding dataset size to include speech audio transcripts.
- •Releasing a fine-tuned, lightweight Gujarati model on Hugging Face hub.

