Dhaka · AI Research Lab

Engineering Bangladesh's Sovereign AI Pipeline

Hackules builds the research, infrastructure, and talent pipeline for a self-reliant AI future — from foundation model research to the engineers who will run it.

$
2021
Founded in Dhaka
04
Active research pipelines
2
Backing accelerators & VCs
2030
Sovereign pipeline target
Ingest
Train
Evaluate
Deploy
Sovereign
pipeline_monitor.log
Stage / Flagship Product
● First Shipped Model — 2026

Sangbad-SLM0.8B params

Bangladesh's first purpose-built Bangla news language model

Not a fine-tune of a general foundation model — Sangbad-SLM is trained from the ground up on our own Bangla news corpus. It's small enough to run on modest hardware, cheap enough to serve at scale, and it beats larger general-purpose models on Bangla news understanding.

0.0B
Parameters — tiny by design
0%
Bangla news benchmark score
0%
Cheaper than foundation-model APIs
Sangbad-SLM (0.8B, Bangla-native)0%
Typical foundation API (Bangla)0%
SLM
0.8B · BANGLA
Bangla News Native
Low Inference Cost
High Benchmark
Benchmark and cost figures are Hackules' internal evaluation results, not externally audited.
Stage / Mission

Why we're building this

Most AI infrastructure Bangladesh depends on today is rented, not owned — foreign models, foreign clouds, foreign decisions about what gets built and who it serves.

Hackules exists to change that ratio. We run open research on low-resource language modeling, train the next generation of ML engineers, and build the data and compute pipelines a sovereign AI stack needs to stand on its own.

Every course we teach and every paper we publish feeds the same pipeline: local talent → local research → local infrastructure.

Founded
2021 · Banani, Dhaka
Backed by
Accelerator in Bangladesh, BYLC Ventures
Horizon
Sovereign inference pipeline by 2030
Short-term

Vertical integration of the AI stack

Own the layer between raw Bangla data and a served model — tokenizer, corpus pipeline, fine-tuning harness, eval suite — so quality doesn't depend on a foreign provider's roadmap.

Mid-term

Compute that doesn't leave the country

Push training and inference workloads onto in-region or on-shore compute, with cost, latency, and data-residency numbers institutions can actually audit.

Long-term

A national AI stack, not a national AI vendor

Sovereignty isn't one model — it's a reproducible pipeline other Bangladeshi teams can fork, retrain, and deploy on their own terms.

Stage / Research

Pipelines in motion

Ingest

Bangla Corpus Engine

Large-scale collection and cleaning pipeline for Bangla text, speech, and dialect data across the region.

Train

Low-Resource LLM Training

Fine-tuning and pretraining methods tuned for limited compute — efficient tokenizers, quantization, distillation.

Evaluate

Sovereign Benchmarks

Evaluation suites built for local context: government, healthcare, and education use cases, not translated Western benchmarks.

Deploy

National Inference Pipeline

Infrastructure blueprint for on-shore inference — latency, cost, and data-residency work for public institutions.

Ingest → Train

Automation-Assisted Dev Pipeline

Research into AI-assisted software delivery — using automation to cut build time and cost on client engineering work.

Evaluate → Deploy

AI Security & Red-Teaming

Cybersecurity research applied to AI systems themselves — model hardening, adversarial testing, and secure deployment practices.

Ongoing — active as of 2026. Areas reflect Hackules' published service focus (AI development, automation, cybersecurity); specific pipelines are illustrative pending public research releases.
01 INGEST Data Layer 02 TRAIN Model Layer 03 EVALUATE Evaluation Layer 04 DEPLOY Serving Layer 05 SOVEREIGN Governance Layer
01 · INGEST

Data Layer

  • Bangla web crawl + OCR
  • Speech corpus (dialects)
  • Dedup & toxicity filter
  • Consent & licensing checks
02 · TRAIN

Model Layer

  • Custom Bangla tokenizer
  • LoRA / QLoRA fine-tuning
  • Distillation for edge use
  • Mixed-precision on limited GPU
03 · EVALUATE

Evaluation Layer

  • Local-context benchmarks
  • Red-team & adversarial tests
  • Bias & safety review
  • Human-in-the-loop scoring
04 · DEPLOY

Serving Layer

  • On-shore inference nodes
  • Quantized model serving
  • Latency & cost monitoring
  • Data-residency guarantees
05 · SOVEREIGN

Governance Layer

  • Institutional handoff
  • Open reproducible pipeline
  • Local team retraining rights
  • Audit & compliance trail
Stage / Track Record

Projects & milestones

01 / 04
2023

AWS Activate Grant

Selected as one of 35 Bangladeshi startups awarded a $10,000 AWS Activate grant to build on Amazon's cloud infrastructure.

Milestone Complete
02 / 04
2023

AWS Startups Day Bangladesh

Represented Hackules at AWS Startup Day Bangladesh 2023, engaging with cloud-native development and ML practitioners nationally.

Milestone Complete
03 / 04
2023–24

AI/ML Instructional Programme

Launched a professional AI/ML bootcamp covering foundational to advanced concepts, built in collaboration with industry partners.

Milestone Complete
04 / 04
Ongoing

50% AI Development Focus

Half of all client engagements are AI development work, alongside web development, custom software, cybersecurity, and ERP consulting.

In Progress
Milestones sourced from Hackules' public profiles (Clutch, LinkedIn). Named client case studies aren't yet public — reach out for a client reference list.
Stage / Academy

Build the pipeline's engineers

AI & ML Foundations

Beginner

Python, linear algebra, and core ML concepts — built for people starting from zero.

Curriculum DepthFoundational
8 weeksCohort-based

Data Engineering Pipelines

Intermediate

Design real ingest → transform → serve pipelines using the same tools our research team runs.

Curriculum DepthApplied
6 weeksProject-based

LLM Fine-Tuning Lab

Advanced

Hands-on fine-tuning, evaluation, and deployment of open-weight models on limited GPU budgets.

Curriculum DepthResearch-Grade
10 weeksResearch track

MLOps & Sovereign Infra

Advanced

Deploying and monitoring models on-shore — the infrastructure layer of a sovereign AI stack.

Curriculum DepthResearch-Grade
8 weeksResearch track
Stage / Barriers

How we overcome the barriers

Building sovereign AI in a low-resource setting means fighting the same four constraints at every stage of the pipeline.

0% PIPELINE MATURITY
Bangla Data Scarcity
BarrierFoundation models see almost no Bangla text — quality collapses outside common phrasing.
ResponseIngest pipeline builds the Bangla corpus ourselves — web text, OCR archives, dialect speech.
0% PIPELINE MATURITY
GPU Compute Scarcity
BarrierDomestic GPU capacity is limited — most teams default to renting foreign cloud indefinitely.
ResponseLoRA/QLoRA, quantization, and distillation let real models train on modest hardware today.
0% PIPELINE MATURITY
ML Talent Shortage
BarrierPlenty of software engineers, far fewer with real ML pipeline experience.
ResponseThe academy isn't a side business — it staffs our own research, cohort by cohort.
0% PIPELINE MATURITY
Vendor Lock-In Risk
BarrierInstitutions can't audit or move off single-vendor black-box AI systems.
ResponseEvery deployment ships with eval results, retraining rights, and a full audit trail.
Maturity scores are Hackules' internal, self-reported estimates — not externally audited metrics.
Stage / Roadmap

Where the pipeline is headed

2026

Corpus & Academy Launch

Bangla data pipeline goes live; first academy cohorts trained in ML foundations and data engineering.

2027

First Bangla-Native Model

Release of an open, low-resource-optimized Bangla language model trained on the corpus pipeline.

2028–2029

Institutional Partnerships

Deploy evaluation and inference pipelines with government and education partners; scale the academy nationally.

2030

Sovereign Inference Pipeline

On-shore compute and inference infrastructure fully operational — a self-reliant national AI stack.

Stage / Team

Who's building it

RR

A T M Ragib Raihan

Founder & Lead Data Scientist

Physics graduate from SUST with seven years of experience as a software engineer specializing in data science and AI. IBM-certified data scientist, and a subject-matter expert to UNICEF and the Bangladesh Army. Founded Hackules to build Bangladesh's AI talent pipeline and infrastructure from the ground up.

AI

Research Team

Model Training & Evaluation
DE

Data Engineering

Pipelines & Infrastructure
AC

Academy Team

Curriculum & Instruction
+

We're Hiring

Join the Pipeline

Join the sovereign AI pipeline

Get research updates, course openings, and roadmap news as we build. No noise — just what's shipping.