Speech in. Metadata-enhanced record out.
An oral history interview is evidence — and Akoé is an AI tool for the historians and archivists who keep it. Automatic transcription, each voice labelled, every word timestamped, an editor built for careful checking: what enters the archive is what was said.
Word-timed, speaker-labelled, citable
Every word carries its own timestamp, so a quotation can point to the exact moment on the tape — and while you check, the audio plays from the word you're on.
Speakers are separated automatically and named by you. The mentions index collects the people, places, organisations and dates in the testimony — plus labels you define — and exports as CSV: raw material for a finding aid.
When a passage is uncertain, it stays flagged until a human has looked at it. The transcript that leaves Akoé is one a person has checked, not one a model has promised.
The formats archives actually ask for
TEI XML
Structured export for digital archives and text collections — corrections and timings included.
Word & plain text
An interview layout for reading copies and deposit files, or clean text for catalogues and summaries.
Word-timed data
CSV and JSON with per-word timing and speaker labels — for indexes, research datasets and custom pipelines.
Subtitles
SRT and VTT timed to the word, for access copies and listening stations.
Custodianship you can explain to a narrator
Interviewees trust you with their memories; you should be able to say plainly where the files go. Every company that touches the audio is EU-owned and EU-based, publicly listed. Recordings are encrypted at rest, per workspace.
Retention is a schedule you choose — 7, 30 or 90 days, or keep until you delete — and audio and transcript are deleted together. Projects are private to their creator by default; sharing is explicit.
For collections: annual hour blocks on invoice, purchase orders and bank transfer, a data-processing agreement, priority processing — write to billing@toolbox21.com.
Fair questions before your first upload
Can I cite a specific moment?
We have a large collection of recordings. How do we start?
What happens to the recordings after transcription?
Does it handle older or difficult recordings?
Does it mark laughter, pauses, or emotion?
Your first 5 minutes are free
A real recording or just your voice—see how it works. No card details needed.
Upload a recording Start dictating