Back to Nodes
ElevenLabs Scribe V2

ElevenLabs Scribe V2

Official

Transcribe long audio or video into structured text with language detection, timestamps, optional speaker labels, and sound-event tags. Choose Scribe

Nodespell AI
AI / Audio / Elevenlabs

Transcribe long audio or video into structured text with language detection, timestamps, optional speaker labels, and sound-event tags. Choose Scribe V2 when downstream workflow steps need both the transcript and machine-readable word metadata. The response preserves text, detected language and confidence, duration, and optional timed speaker segments.

Inputs (1)

Media

String

Audio or video file to transcribe, up to the service limits.

RequiredMin: 1Max: 1
Parameters (10)

Media

String

Audio or video file to transcribe, up to the service limits.

Required
Default:

Language Code

String

ISO-639 language code, or auto for language detection.

Default: auto

Identify Speakers

Boolean

Label different speakers in the timed transcript.

Default: false

Maximum Speakers

Number

Expected maximum speaker count, or 0 for automatic detection.

Default: 0

Timestamp Granularity

String

Whether timing metadata is omitted, word-level, or character-level.

Default: word

Tag Audio Events

Boolean

Mark non-speech events such as laughter, footsteps, or applause.

Default: true

Clean Transcript

Boolean

Remove filler words, false starts, and other disfluencies.

Default: false

Key Terms

String

Comma-separated names or specialist terms that should be favored.

Default:

Temperature

Number

Sampling temperature, or -1 to use the model default.

Default: -1

Seed

Number

Optional random seed for repeatable transcription.

Outputs (1)

Response

Object

Structured transcript with language and optional timing metadata.

Nodespell Team

Creator profile

Type

Node

Status

Official

Package

Nodespell AI

Category

AI / Audio / Elevenlabs

Input

AudioVideo

Output

Text

Keywords

TranscriptionSpeech To TextDiarizationTimestampsSubtitles
Use in Workflow