Engineering Bangladesh's Sovereign AI Pipeline
Hackules builds the research, infrastructure, and talent pipeline for a self-reliant AI future — from foundation model research to the engineers who will run it.
Sangbad-SLM0.8B params
Not a fine-tune of a general foundation model — Sangbad-SLM is trained from the ground up on our own Bangla news corpus. It's small enough to run on modest hardware, cheap enough to serve at scale, and it beats larger general-purpose models on Bangla news understanding.
Why we're building this
Most AI infrastructure Bangladesh depends on today is rented, not owned — foreign models, foreign clouds, foreign decisions about what gets built and who it serves.
Hackules exists to change that ratio. We run open research on low-resource language modeling, train the next generation of ML engineers, and build the data and compute pipelines a sovereign AI stack needs to stand on its own.
Every course we teach and every paper we publish feeds the same pipeline: local talent → local research → local infrastructure.
Vertical integration of the AI stack
Own the layer between raw Bangla data and a served model — tokenizer, corpus pipeline, fine-tuning harness, eval suite — so quality doesn't depend on a foreign provider's roadmap.
Compute that doesn't leave the country
Push training and inference workloads onto in-region or on-shore compute, with cost, latency, and data-residency numbers institutions can actually audit.
A national AI stack, not a national AI vendor
Sovereignty isn't one model — it's a reproducible pipeline other Bangladeshi teams can fork, retrain, and deploy on their own terms.
Pipelines in motion
Bangla Corpus Engine
Large-scale collection and cleaning pipeline for Bangla text, speech, and dialect data across the region.
Low-Resource LLM Training
Fine-tuning and pretraining methods tuned for limited compute — efficient tokenizers, quantization, distillation.
Sovereign Benchmarks
Evaluation suites built for local context: government, healthcare, and education use cases, not translated Western benchmarks.
National Inference Pipeline
Infrastructure blueprint for on-shore inference — latency, cost, and data-residency work for public institutions.
Automation-Assisted Dev Pipeline
Research into AI-assisted software delivery — using automation to cut build time and cost on client engineering work.
AI Security & Red-Teaming
Cybersecurity research applied to AI systems themselves — model hardening, adversarial testing, and secure deployment practices.
Data Layer
- Bangla web crawl + OCR
- Speech corpus (dialects)
- Dedup & toxicity filter
- Consent & licensing checks
Model Layer
- Custom Bangla tokenizer
- LoRA / QLoRA fine-tuning
- Distillation for edge use
- Mixed-precision on limited GPU
Evaluation Layer
- Local-context benchmarks
- Red-team & adversarial tests
- Bias & safety review
- Human-in-the-loop scoring
Serving Layer
- On-shore inference nodes
- Quantized model serving
- Latency & cost monitoring
- Data-residency guarantees
Governance Layer
- Institutional handoff
- Open reproducible pipeline
- Local team retraining rights
- Audit & compliance trail
Projects & milestones
AWS Activate Grant
Selected as one of 35 Bangladeshi startups awarded a $10,000 AWS Activate grant to build on Amazon's cloud infrastructure.
AWS Startups Day Bangladesh
Represented Hackules at AWS Startup Day Bangladesh 2023, engaging with cloud-native development and ML practitioners nationally.
AI/ML Instructional Programme
Launched a professional AI/ML bootcamp covering foundational to advanced concepts, built in collaboration with industry partners.
50% AI Development Focus
Half of all client engagements are AI development work, alongside web development, custom software, cybersecurity, and ERP consulting.
Build the pipeline's engineers
AI & ML Foundations
BeginnerPython, linear algebra, and core ML concepts — built for people starting from zero.
Data Engineering Pipelines
IntermediateDesign real ingest → transform → serve pipelines using the same tools our research team runs.
LLM Fine-Tuning Lab
AdvancedHands-on fine-tuning, evaluation, and deployment of open-weight models on limited GPU budgets.
MLOps & Sovereign Infra
AdvancedDeploying and monitoring models on-shore — the infrastructure layer of a sovereign AI stack.
How we overcome the barriers
Building sovereign AI in a low-resource setting means fighting the same four constraints at every stage of the pipeline.
Where the pipeline is headed
Corpus & Academy Launch
Bangla data pipeline goes live; first academy cohorts trained in ML foundations and data engineering.
First Bangla-Native Model
Release of an open, low-resource-optimized Bangla language model trained on the corpus pipeline.
Institutional Partnerships
Deploy evaluation and inference pipelines with government and education partners; scale the academy nationally.
Sovereign Inference Pipeline
On-shore compute and inference infrastructure fully operational — a self-reliant national AI stack.
Who's building it
A T M Ragib Raihan
Physics graduate from SUST with seven years of experience as a software engineer specializing in data science and AI. IBM-certified data scientist, and a subject-matter expert to UNICEF and the Bangladesh Army. Founded Hackules to build Bangladesh's AI talent pipeline and infrastructure from the ground up.
Research Team
Data Engineering
Academy Team
We're Hiring
Join the sovereign AI pipeline
Get research updates, course openings, and roadmap news as we build. No noise — just what's shipping.