Sourcelane

Whisper Speech to Text

AI & Media Processing

Operated by Sourcelane✓ verified

A speech-to-text connector that uses Whisper to transcribe audio and video into timestamped, searchable transcripts with speaker labeling (diarization), automatic chapter detection, keyword extraction, and generated subtitle files (SRT / VTT / TXT). The connector produces structured, spreadsheet-friendly output that includes minute-level transcript slices with sentence-level timing and confidence scores, speaker summaries with speaking time and quotable samples, chapter summaries with start/end timestamps and short previews, top keywords with occurrence timestamps, and file-level metadata such as language detection, duration, word counts and speaking-rate estimates. Optional features include automatic language detection across many languages, translation into English, word-level timestamps (when diarization is off), and sentence-level rows for fine-grained analysis; subtitle files include speaker tags when diarization is enabled.

Verified Sep 28, 3:55 AM

$0.04

per result

Up to 25 per call. Only pay for results returned; failed calls are refunded.

What people use it for

Trust & reliability

7d uptime trend
WindowUptimeSuccess ratep50p95Calls
24h————0
7d————0
30d————0

Calling contract

Call it through Sourcelane's gateway or MCP server. We run the connector, apply your agent's spend guardrails, and bill only the results returned.

Call it via Sourcelane

curl -X POST https://api.usesourcelane.com/v1/call \
  -H "Authorization: Bearer sl_live_your_agent_key" \
  -H "Content-Type: application/json" \
  -d '{"listing":"whisper-speech-to-text","params":{"audioUrls":["https://raw.githubusercontent.com/ggml-org/whisper.cpp/master/samples/jfk.wav"],"maxResults":5},"maxResults":10}'

Request params

FieldTypeRequiredDescription
audioUrlsarrayoptionalDirect URLs to audio or video files (mp3, wav, m4a, mp4, ogg, flac, webm, …). Up to 10 files per run, 500 MB per file.
audioUrlstringoptionalAlternative to the list above — a single direct media URL.
languagestringoptionalISO 639-1 code of the spoken language, e.g. "en", "es", "ko". Leave empty to auto-detect (99 languages supported).
translateToEnglishbooleanoptionalTranscribe speech in any language directly into English text.
wordTimestampsbooleanoptionalAdd per-word start/end times to every segment (slower, larger rows). Word times are estimated and can drift by a few hundred milliseconds.
speakerDiarizationbooleanoptionalDetects the different voices in the recording and labels every transcript segment with SPEAKER_1, SPEAKER_2, … Adds one speaker summary row per detected voice (speaking time, word count, share of speech, sample quote) and [SPEAKER_N] tags in the SRT/VTT/TXT files. Recommended for meetings, interviews and podcasts.
numSpeakersintegeroptionalIf you know how many people speak (e.g. 2 for an interview), set it here for more reliable labels. Leave 0 to detect the count automatically.
enrichmentRowsbooleanoptionalAdds per-file enrichment rows: one file summary, pause-detected chapters with YouTube-ready timestamps, and the top 10 keywords with every mention's timestamp. Each delivered row is billed like any other result row.
segmentRowsbooleanoptionalAdds one row per sentence segment (roughly 10 per minute of speech). Useful for pipelines that join on exact timestamps. Off by default because it multiplies the number of billed rows.
maxMinutesPerFileintegeroptionalTranscription stops after this many minutes of audio per file (cost guard). Longer files are trimmed and flagged with a notice row. Max 240.
maxResultsintegeroptionalUpper bound on billable result rows (minutes, chapters, keywords, summaries, speakers, segments) returned by this run. Free-plan runs return at most 25. Capped at 25 per call.

Each result contains

FieldTypeDescription
sourceUrlstringSource Url
fileIndexintegerFile Index
chargedbooleanCharged
rowTypestringRow Type
languagestringLanguage
languageProbabilitynullLanguage Probability
taskstringTask
audioDurationSecintegerAudio Duration Sec
transcribedSecintegerTranscribed Sec
transcribedMinutesintegerTranscribed Minutes
truncatedbooleanTruncated
wordCountintegerWord Count
segmentCountinteger—
wordsPerMinuteinteger—
speakerCountnull—
chapterCountinteger—
topKeywordsarray—
transcriptstring—
subtitleFilesarray—

Use Whisper Speech to Text from Claude, ChatGPT or Cursor

Pick your tool and connect in under a minute. Then just ask — for example: “Transcribe this video and give me timestamps for every product mention: <url>”

Connect Claude

Web, desktop and mobile. Paste one URL.

  1. 1Copy your personal connector URL
  2. 2In Claude open Settings → Connectors → Add custom connector
  3. 3Paste the URL and click Add — done
Manual setup (config files, REST, Python) +

Claude Code

Adds the Sourcelane MCP server with your key as a header.

terminal
claude mcp add --transport http sourcelane https://api.usesourcelane.com/mcp \
  --header "Authorization: Bearer sl_live_your_agent_key"

Claude Desktop (config file)

Alternative to the connector URL: add to claude_desktop_config.json, then restart Claude.

claude_desktop_config.json
{
  "mcpServers": {
    "sourcelane": {
      "command": "npx",
      "args": [
        "-y",
        "mcp-remote",
        "https://api.usesourcelane.com/mcp",
        "--header",
        "Authorization:${AUTH_HEADER}"
      ],
      "env": {
        "AUTH_HEADER": "Bearer sl_live_your_agent_key"
      }
    }
  }
}

Cursor & Windsurf

Add to ~/.cursor/mcp.json (or Windsurf's mcp_config.json).

mcp.json
{
  "mcpServers": {
    "sourcelane": {
      "url": "https://api.usesourcelane.com/mcp",
      "headers": {
        "Authorization": "Bearer sl_live_your_agent_key"
      }
    }
  }
}

VS Code

Save as .vscode/mcp.json. VS Code prompts for your key once.

.vscode/mcp.json
{
  "servers": {
    "sourcelane": {
      "type": "http",
      "url": "https://api.usesourcelane.com/mcp",
      "headers": {
        "Authorization": "Bearer ${input:sourcelane-key}"
      }
    }
  },
  "inputs": [
    {
      "type": "promptString",
      "id": "sourcelane-key",
      "description": "Sourcelane agent key",
      "password": true
    }
  ]
}

ChatGPT Custom GPT (Actions)

Create a GPT → Actions → Import from URL, then Authentication: API Key, Bearer.

OpenAPI schema URL
https://usesourcelane.com/openapi.json

REST

One POST. Pass maxResults to cap cost.

curl
curl -X POST https://api.usesourcelane.com/v1/call \
  -H "Authorization: Bearer sl_live_your_agent_key" \
  -H "Content-Type: application/json" \
  -d '{"listing":"whisper-speech-to-text","params":{"audioUrls":["https://raw.githubusercontent.com/ggml-org/whisper.cpp/master/samples/jfk.wav"],"maxResults":5},"maxResults":10}'

Python, LangChain, CrewAI, OpenAI Agents SDK…

Wrap the REST call as a tool, or point an MCP client at the endpoint.

python
import requests

res = requests.post(
    "https://api.usesourcelane.com/v1/call",
    headers={"Authorization": "Bearer sl_live_your_agent_key"},
    json={
        "listing": "whisper-speech-to-text",
        "params": {"audioUrls":["https://raw.githubusercontent.com/ggml-org/whisper.cpp/master/samples/jfk.wav"],"maxResults":5},
        "maxResults": 10,
    },
    timeout=300,
)
body = res.json()
print(body["receipt"]["chargedMicros"], "micro-USD for", body["receipt"]["results"], "results")
print(body["data"])

Reviews

—

0 reviews

5
0
4
0
3
0
2
0
1
0

Leave a review

Loading reviews…