ΕΛ Sign in Try 5 minutes free

Hear it. Read it. Cite it.

Upload audio or video: an interview, a meeting, a dictation. Akoé writes it out, labels who said what, and ties every word to its second in the recording: check anything against the sound, fix mistakes, and download the transcript in whatever form you choose.

First 5 minutes free. No card, no subscription. After the trial, we charge €5 per hour of audio pro-rated, starting from a €1 minimum charge.
EU-owned infrastructurenot just EU regions of overseas cloudsDeleted on your scheduleaudio and transcript togetherNever used to train modelsno third-party AI touches your filesExact price before you pay€5 per hour of audio, to the second
How it works

Upload. Check. Cite.

Three steps, no manual needed. Built for people who work with recordings, not for people who work with software.

01

Drop in your recording

Audio or video, minutes or hours. From video we keep only the sound. The images are dropped during processing and never kept.

02

Watch the transcript arrive

The first minutes appear while the rest is still transcribing and the voices are getting labelled. We email you when it's done.

03

Check it & export it

Fix anything in the editor, then export Word, subtitles, text, data or TEI XML, with your corrections in every format.

The editor

Automatic transcription is never perfect. The akoé platform is built for checking effectively.

Every automated transcript needs human checking. What matters is how fast you find and fix the mistakes.

Hear what you're fixing

The audio plays from the word you're checking at your preferred speed, so you hear the moment while you fix it at your own pace. Words the AI is unsure about stay flagged until you've looked; find & replace clears a repeated mistake everywhere.

Voices, untangled

Automatic speaker separation shows who said what. Wrong name, wrong voice, missed change of speaker—each is a one-click repair.

See who and what is mentioned

Akoé scans the transcript for the people, places, organisations and dates mentioned. It builds an automatic index and even lets you define your own categories for things to scan. Click a mention to jump to it.

Subtitles

Add subtitles to your videos and podcasts (SRT and VTT timed to the word) in the original language and in translation.

Translate, then verify

Machine-translated drafts open beside the original, line by line, editable, timings intact and ready to check.

Dictate straight in the browser

Tired of typing? Press record, speak, stop. Your words come back as text so you don't have to write from scratch.

Who it's for

One Product, Many Use Cases

akoe.ai does automated transcription (converts speech to text), speaker separation, dictation, translation and entity recognition—and because every transcript in your workspace is searchable and every word is timed, it works not only as a transcription service but also as a research tool: find who said what, cite it, or pull the exact soundbite for your edit.

Researchers & Universities

Interview and focus-group transcription your ethics committee can approve: EU-only processing, deletion guarantees, and a data-processing agreement (DPA) upon request.

Oral History & Archives

Speaker-labelled, timecoded transcripts. Recorded memory becomes a searchable, citable record.

Professional Transcribers

Start from a strong draft instead of a blank page. Playback, shortcuts and speaker tools built for high-volume correction. Your skill goes into perfecting, not typing.

Journalists

Sound in, quotable text out, every word checkable against the tape. Retention and processing on EU servers, gone on your schedule.

NGOs & Advocacy

Field interviews, testimonies and case documentation, deleted on your schedule.

Professionals Who Dictate

Tired of typing? Speak, then simply edit your dictated text. A dictation platform for confidential work.

Councils, Boards & Committees

Meeting recordings become draft minutes: who spoke, what was said, timestamped to the second—the secretary will check, not retype from scratch. Invoicing for public bodies and businesses available. Contact us.

Podcast & Video Makers

Word-timed SRT/VTT subtitles with free translated versions of them.

Speech Corpus Developers

Gold-standard transcripts, faster: correct a strong draft in a verbatim-first editor, then export word-timed, speaker-labelled JSON or CSV.

Languages

About 100 languages, auto-detected

Upload in whatever language the recording is in—akoé.ai detects it automatically, or you set it yourself.

Usually a lighter check is needed for

English · Spanish · French · Portuguese · German · Italian · Polish · Dutch

Usually more checking is required for

Greek · Arabic · Russian · Turkish · Vietnamese · Ukrainian · Romanian · Hungarian · Swedish · Serbian · Czech · Bulgarian · Slovak · Croatian · Danish · Finnish · Norwegian · Galician

Usually heavy checking is required for

Hindi · Estonian · Slovenian · Tamil · Latvian · Azerbaijani · Urdu · Lithuanian · Hebrew · Welsh · Persian · Icelandic · Kazakh · Afrikaans · Swahili · and dozens more, auto-detected

Grouped by how automatic transcription tends to perform on clear audio. Accuracy varies depending on the quality of the recording and the language, which is exactly why the editor exists. Use your free 5 minutes on a real file to see what to expect.

Security

We respect your material

Every company that touches your audio is EU-owned and EU-based. Not the EU regions of overseas clouds.
Deleted means gone. Audio and transcript are deleted together, automatically on the schedule you choose (7, 30 or 90 days) or by hand. You can always delete them permanently, whenever you decide. Everything derived from your file, such as indexes and translations, is deleted with it.
Your recordings never train anyone's models. Every AI model we use runs on our own EU servers; no third-party AI service ever sees your files. No trackers, no cookie banner—this site's analytics are self-hosted and cookieless.
We value transparency. The security page explains in plain language where your data goes and how it is protected. The full process concept, written to be attached to an ethics application or DPIA, comes with the data-processing agreement on request.
Pricing

Transcription without subscription.

€5 per hour of audio, pro-rated. That's it.

Billed to the second

A 90-minute interview is €7.50. A 45-minute lecture, €3.75. A 20-minute memo, €1.67. No tiers, no rounding up to the hour.

The exact price, before you pay

Your file's duration is known before checkout, so the price you see is the price you're charged. Every time. If a file can't be processed, the charge refunds itself automatically.

No subscription required

Nothing to prepay, no top-ups, nothing to cancel. Your first 5 minutes are free, without a card.

What would my recording cost?

minutes

Universities, archives & public bodies: monthly and annual hour blocks on invoice are available and so are data-processing agreements and priority processing — write to info@akoe.ai.
Frequently Asked Questions

Fair questions before your first upload

How do I convert audio to text?
Upload a recording or dictate straight into the browser and Akoé converts the audio to text automatically (automatic speech recognition or ASR, if you like the technical term): speech becomes written words, with speaker labels and a timestamp on every word. You then check the transcript in the editor, word by word, and export it as Word, subtitles, plain .txt file or TEI XML.
How accurate is the automatic transcription?
It depends on the recording: clear speech near the microphone needs lighter fixes; audio captured in noisy environments with bad use of microphones and in a low-resource language needs more. The next question lists what makes audio hard to transcribe. We do not guarantee an errorless transcript. What we promise is the ease of checking and correcting it: words for which our model returns a low confidence score get flagged for you to check, and the audio plays from the word you're checking at a speed you find comfortable, so correcting the transcript gets as easy as possible.
When will automatic transcription struggle? What makes automatic speech recognition hard?
Recordings with distant or shared microphones, people talking over each other, switching languages mid-sentence, in noisy rooms and in languages, dialects or accents that are underrepresented in public training data are all factors that impact the quality of the resulting transcript. That last one is a dataset problem, not a defect of how anyone speaks. In the 'Languages' section above, you will find which languages are more likely to lead to better transcripts. We're based in Greece, a country rich in exactly the varieties that datasets underserve, so we feel this one personally—and we encourage research on lower-resourced speech. If your work needs a custom model for a specific variety or domain, write to us: describe your use case, and tell us whether you have recordings you hold the legal rights to share for training. Your normal uploads are never used for training—a custom model is a separate, explicit, written agreement.
Does it detect laughter, pauses, or emotion?
No. Akoé transcribes speech, the words. It doesn't detect or mark non-speech sounds (laughter, coughing, crying), tone or emotion, and it doesn't transcribe silence: quiet stretches simply produce no text. If your method needs those cues, the editor is a good place to add them by hand—every word is timestamped, and the audio is right there while you work.
Can it tell who said what?
Yes, automatic speaker separation (speaker diarization) labels each voice, and telling the wizard how many people are speaking makes it sharper. Where it slips, the editor repairs it in a click: rename, reassign, split or merge speakers.
Which languages does each feature support?
Transcription and dictation: about 100 languages, auto-detected or set by you—the languages section above shows what to expect from each. Speaker separation: any language. It works on voices, not words. Translation: into 28 languages, from a transcript in any of them: Bulgarian, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, German, Greek, Hungarian, Italian, Latvian, Lithuanian, Polish, Portuguese, Romanian, Slovak, Slovenian, Spanish and Swedish (22 of the 24 official EU languages), plus Arabic, Catalan, Chinese, Norwegian, Russian and Turkish. Subtitles and exports: Word, subtitles (SRT and VTT), plain text, CSV and TEI, in the recording's language and in any translation you've made. Search by meaning: all transcription languages, and across them: ask in one language, find the moment in another. Mentions: people, places, organisations, dates and labels you define. Multilingual, indexed in the recording's language.
How long does a transcription take?
First your file waits for a processing machine—usually a couple of minutes. Then the transcription itself runs several times faster than the recording, so an hour of audio means minutes of processing, not another hour. On longer files the first minutes of text appear while the rest is still working, and we email you when the full transcript is ready—no need to keep the tab open. Sometimes a very short file doesn't feel proportionally fast, because the first wait is the same whether the recording lasts two minutes or two hours.
Can I transcribe video, and what happens to the footage?
Yes—mp4, mov and other common formats. We keep only the audio track; the video is never kept. In modern browsers the sound is extracted on your device before upload, so a 2 GB interview video travels as a few megabytes of audio; if your browser can't do this, the file is converted on our EU servers and the original is deleted the moment the audio is out.
What file types can I upload?
Audio: mp3, wav, m4a, flac, ogg/opus and most other common formats. Video: mp4, mov, webm etc.—we keep only the sound. If a file can't be read it isn't processed, and any charge refunds itself automatically.
Can I dictate instead of uploading a file?
Yes. Dictation runs straight in the browser: press record, speak, stop, and the text comes back fast — ready to play back word by word and export like any other recording. One voice, no speaker labels, the same €5-per-hour meter. It's also the fastest way to try Akoé: your first 5 minutes are free, no file and no card needed.
What does it cost?
€5 per hour of audio, pro-rated to the second, with the exact price shown before you pay. The smallest checkout is €1 (card fees), and a €1 checkout covers up to 12 minutes of audio in the next 24 hours — your file included. Failed files refund automatically. Your first 5 minutes are free, no card needed.
Why is there a €1 minimum on checkouts?
Card and payment fees make charges under €1 not worth it for anyone. When the minimum exceeds your job's price (say a 5-minute memo) the difference covers more audio over the next 24 hours, your file included, so nothing is wasted. It lapses on its own and is never stored credit.
Couldn't I just translate my transcript with an AI chatbot?
You could, the transcript is yours, export it any time and use it as you see fit. But the built-in translation is included with every transcript and keeps what a chat window loses: the draft stays aligned line by line with the original for easier review and every timestamp and speaker label survives.
Where is my audio processed? Can I use this under GDPR?
Everything runs on EU-owned, EU-based infrastructure, publicly listed. Files are encrypted at rest. Retention is a schedule you choose and immediate deletion is always available. For your DPO or ethics committee, an Art. 28 data-processing agreement and the full process concept are available on request—and the security page carries the plain-language summary.
Will my recordings be used to train AI models?
No. Your audio and transcripts are never used to train models, and no third-party AI service touches them. The models run self-hosted on EU servers.
What can I export?
Word (DOCX) in interview or dictation layout, plain text, SRT and VTT subtitles, CSV for spreadsheets, and TEI XML for archives — plus a CSV of the mentions index. Every export reflects your corrections, timings included.
Can I include the transcript in a thesis, paper or report?
Yes. The transcript is yours to keep and to publish. Researchers often add a full interview transcript as an appendix, and Akoé is made for it: export the corrected text as Word or plain text, with speaker labels and every word timed, so a reader can trace any quote back to its moment in the recording.
Can I search across my recordings?
Yes, search finds moments by meaning—not just exact words, across your recordings. Your query can even be in a different language than the one of the recording. Results are real snippets from your own transcripts, each with a timestamp: click one and it plays. Your searches are never logged or stored, a transcript's search index is deleted together with it, and search is included with every account.
Do I need a subscription?
No. Akoé is pay-as-you-go automatic transcription and dictation app. You pay per recording, the price is shown first, and there is nothing to cancel: no money parked with us, no top-ups, no monthly fee.

Your first 5 minutes are free

A real recording or just your voice—see how it works. No card details needed.

Upload a recording Start dictating