SinAI — SinhalaJournalLLM
A final-year research project developing a Sinhala-focused AI writing assistant for journalism using a fine-tuned language model, with my work centered on grammar correction, data engineering and research leadership.
The problem
Sinhala journalism lacks mature language tools that can reliably correct spelling, grammar, punctuation and sentence structure while respecting real editorial language.
The approach
I led the research team and developed the grammar-correction component, using Llama-based models, LoRA fine-tuning and Sinhala-specific datasets created from large-scale news collection and error analysis.
Scraped and processed more than 700,000 Sinhala newspaper articles for language-model research.
Created a 36,000+ row correction dataset spanning 18 Sinhala grammar categories.
Used LoRA, Unsloth, TRL and PEFT to adapt Llama-based models efficiently.
Evaluated controlled sentence examples separately from real-news paragraphs.
Results
87.7% sentence-level grammar correction accuracy.
75.0% accuracy on real-news paragraphs.
A 700K+ article Sinhala news corpus and a 36,000+ row dataset across 18 grammar categories.
Explore more AI systems and product engineering case studies.