elevenLabsTranscriptToCaptions()v4.0.443
Converts an ElevenLabs Speech to Text API response or a JSON transcript exported from ElevenLabs into an array of Caption objects.
This function can be used in any JavaScript environment, but you should not use the ElevenLabs API in the browser because your API key will be exposed.
Example
When calling the ElevenLabs Speech to Text API, you must set timestamps_granularity to "word" to include word-level timing in the response.
Example usageimport fs from 'fs'; import {elevenLabsTranscriptToCaptions} from '@remotion/elevenlabs'; const form = new FormData(); form.append('file', new Blob([fs.readFileSync('audio.mp3')])); form.append('model_id', 'scribe_v2'); form.append('timestamps_granularity', 'word'); const response = await fetch('https://api.elevenlabs.io/v1/speech-to-text', { method: 'POST', headers: { 'xi-api-key': process.env.ELEVENLABS_API_KEY!, }, body: form, }); const transcript = await response.json(); const {captions} = elevenLabsTranscriptToCaptions({transcript});
API
Arguments
An object with the following property:
transcript
Pass the parsed JSON object in either of the following formats. The function converts it locally without calling the ElevenLabs API.
Speech-to-Text API response
The response must include a words array. Set timestamps_granularity to "word" when calling the API, as shown in the example above.
Speech-to-Text response (required fields){ "language_code": "en", "words": [ {"text": "Hello", "type": "word", "start": 0, "end": 0.5} ] }
Each "word" entry becomes one caption. start and end are measured in seconds from the start of the audio.
Entries with type: "audio_event" are skipped. Entries with type: "spacing" do not produce captions, but their start times are used as the start times of the following words.
JSON exportv4.0.530
You can also pass a JSON transcript exported from ElevenLabs. It must contain a language_code field (which can be null) and a segments array:
JSON export without word-level timing{ "language_code": "en", "segments": [ { "text": "Hello world", "start_time": 0, "end_time": 1 } ] }
Each segment becomes one caption unless it contains a non-empty words array. In that case, each word becomes a caption instead:
JSON export with word-level timing{ "language_code": "en", "segments": [ { "text": "Hello world", "start_time": 0, "end_time": 1, "words": [ {"text": "Hello", "start_time": 0, "end_time": 0.5}, {"text": " world", "start_time": 0.5, "end_time": 1} ] } ] }
All start_time and end_time values are measured in seconds from the start of the audio, not from the start of the segment.
Invalid input
The function throws an error if required fields are missing or invalid, timestamps are negative, or an end time is earlier than its start time.
Return value
An object with the following property:
captions
An array of Caption objects.
Compatibility
| Browsers | Servers | Environments | |||||||
|---|---|---|---|---|---|---|---|---|---|
Chrome | Firefox | Safari | Node.js | Bun | Serverless Functions | ||||