Griot Edge
Ghanaian speech recognition in 60.9 million parameters.
A streaming Conformer CTC model for Akan, Dagbani, Dagaare, Ewe, Fante, Ghanaian English, and Ga. Run a recording locally, batch long audio, or keep a transcription server warm. The 151-character vocabulary preserves case, punctuation, and orthographic vowels.
Quick start · Results · Batch & server · Deployment guide · Full benchmarks
| At a glance | |
|---|---|
| Model | 60.9M parameters · 17-layer Conformer · CTC |
| Input | 16 kHz mono; helpers resample and mix down other audio |
| Context | Selectable lookahead; default 13, causal 0 |
| Decoding | Bundled multilingual KenLM; optional Flashlight or greedy |
| Execution | CPU FP32 · Apple MPS / supported NVIDIA CUDA BF16 |
| License | CC BY-NC-SA 4.0 |
Results
26.93% WER / 14.59% CER on the internal test set with 13-frame lookahead and KenLM. Lower is better. These are acoustic-model evaluation results, separate from the synthetic speed tests below.
Compared with Griot Nano 1
| Model / decoding | Parameters | Test utterances | CER ↓ | WER ↓ |
|---|---|---|---|---|
| Griot Nano 1, greedy | 153M | 99,547 | 13.66% | 29.19% |
| Griot Edge, greedy | 60.9M | 109,008 | 15.04% | 34.94% |
| Griot Edge, KenLM beam | 60.9M | 109,008 | 14.59% | 26.93% |
Edge has about 60% fewer parameters and selectable streaming lookahead. The accuracy rows are not a matched benchmark: Nano and Edge use different test populations, and the KenLM row also changes decoding. They do not establish an accuracy improvement over Nano. A shared-test-set comparison is still needed.
Within Edge's own test set, KenLM reduces WER from 34.94% to 26.93% at the default lookahead: 8.01 percentage points. Causal KenLM decoding scores 28.86% WER / 15.81% CER.
Scores normalize references and predictions with NFC, lowercase, and tone-mark removal. The evaluated beam uses width 25, LM weight 0.5, and word bonus 1.0. These accuracy scores do not validate the optional Flashlight preset. See evaluation details and source breakdown and release metrics.
Quick start
With uv installed:
uvx --from huggingface-hub hf download Qlerqly/griot-edge --local-dir griot-edge
cd griot-edge
uv venv --python 3.12 .venv
uv pip install --python .venv/bin/python -r requirements-server.txt
.venv/bin/python inference.py --model-dir . --audio recording.wav
The default uses the bundled KenLM and 13-frame lookahead. Add --greedy to skip the LM, or --lookahead-frames 0 for causal inference. Use --device mps --dtype bf16 on Apple GPU, --device cuda --dtype bf16 on supported NVIDIA GPU, or --device cpu --dtype fp32 on CPU. CUDA requires a CUDA-enabled PyTorch installation; see setup options.
Batch & server
Long recordings are split into 30-second windows with 5-second overlap. Chunks from one or many recordings share acoustic batches; overlapping output frames are trimmed before one decode per recording. Outputs keep input order.
# One-shot batch
.venv/bin/python batch_inference.py --model-dir . \
--audio recording-a.wav recording-b.flac --batch-size 8 --output results.json
# Optional persistent HTTP server
.venv/bin/python batch_server.py --model-dir . --batch-size 8 --port 8000
The server collects requests for up to 20 ms, or starts sooner at 8 requests; it then runs the available work, even if the acoustic batch is not full. New arrivals do not reset the timer. --batch-wait-ms 0 removes the collection wait; busy-worker queue time is additional.
Start at batch 8. GPU capacity was tested through 128; CPU through 8. Larger batches did not improve throughput on the tested large workloads. Python API, manifests & HTTP examples · Docker, CUDA & server limits
Faster decoding
uv pip install --python .venv/bin/python -r requirements-flashlight.txt
.venv/bin/python batch_inference.py --model-dir . --audio recording.wav \
--decoder-backend flashlight --batch-size 8
Native commands default to pyctcdecode. Docker bundles both backends and defaults to Flashlight, token beam 8, acoustic batch 8, and one decoder worker. Flashlight uses a closed word lexicon and changes transcripts; validate names, unseen words, and your languages before choosing it for accuracy-sensitive work. Decoder details
Measured throughput
128 × 30-second inputs: 64 minutes of audio, batch 128. End-to-end compute, including preprocessing and CPU decoding; BF16 GPU acoustics. Median of three runs after warmup; excludes model loading and file/HTTP I/O. Repeated synthetic English speech measures capacity, not recognition accuracy.
| Hardware | pyctcdecode | Flashlight | Speedup |
|---|---|---|---|
| Apple M5 Pro GPU + CPU, 6 Torch threads | 10.524 s | 5.056 s | 2.08× |
| NVIDIA L4 + Modal host CPU, 4 allocated cores | 26.127 s | 13.321 s | 1.96× |
Host CPUs affect decoder time, so compare backends within each row. All CPU/GPU/L4 batch-size sweeps · Benchmark methodology & raw results · Flashlight comparisons
Limitations
- Akan source quality varies. Reported WER is 19.15% on human-transcribed sources versus 49.98% on machine-transcript sources at the default tier with KenLM. These are different source subsets, not a controlled causal comparison. Counts and CER.
- Coverage is mainly Ghanaian speech. Listed languages do not imply equal accuracy; Fante has no separate row in the internal six-language evaluation.
- Lookahead needs care. Use the default 13-frame or causal tier; 6 and 26 are available. The 51-frame and full-context settings are outside the validated release tiers. Lookahead is not total service latency.
- CUDA BF16 can change transcripts across batch sizes. This occurred on L4; check representative speech at your intended batch size. Docker on macOS cannot use the Apple GPU; run MPS natively.
Training & license
Training used 44,000 audio hours of pretraining, 12,000 hours of speech/no-speech post-training, then 10,000 hours with intermediate CTC and variable lookahead. Sources follow Nano 1, adding KasaSpeech English–Twi and the University of Ghana Dagbani/Dagaare subsets of WAXAL.
Released under CC BY-NC-SA 4.0. See source acknowledgements, license audit, and license text.
- Downloads last month
- 73