A Python script to massively download video transcripts from YouTube channels. This tool fetches all videos from a given YouTube channel and downloads their transcripts in various formats (JSON, TXT, or SRT).
- Download transcripts from all videos in a YouTube channel
- Support for multiple output formats (JSON, TXT, SRT)
- Multi-language support with fallback options
- Automatic handling of generated and manual transcripts
- Progress logging and error handling
- No API key required (uses scrapetube for video discovery)
- Detailed logging to both console and file
- NEW: ScraperAPI integration to avoid rate limiting and request blocking
- Python 3.7 or higher
- pip (Python package manager)
- Clone this repository:
git clone <repository-url>
cd channel-transcript-download- Install dependencies:
pip install -r requirements.txtDownload all transcripts from a channel (default: JSON format, English language):
python youtube_transcript_downloader.py "https://www.youtube.com/@channelname"# Specify output directory
python youtube_transcript_downloader.py "https://www.youtube.com/@channelname" --output my_transcripts
# Choose output format (json, txt, or srt)
python youtube_transcript_downloader.py "https://www.youtube.com/@channelname" --format txt
# Specify preferred languages (tries in order)
python youtube_transcript_downloader.py "https://www.youtube.com/@channelname" --languages en es fr
# Limit number of videos to process
python youtube_transcript_downloader.py "https://www.youtube.com/@channelname" --limit 10
# Enable verbose logging
python youtube_transcript_downloader.py "https://www.youtube.com/@channelname" --verboseThe script supports various YouTube channel URL formats:
https://www.youtube.com/@usernamehttps://www.youtube.com/channel/UC...https://www.youtube.com/c/CustomNamehttps://www.youtube.com/user/Username
Contains complete transcript data including timing information and metadata:
{
"video_id": "...",
"title": "...",
"language": "English",
"language_code": "en",
"is_generated": false,
"transcript": [
{
"text": "...",
"start": 0.0,
"duration": 2.5
}
]
}Plain text format with video metadata and transcript text:
Title: Video Title
Video ID: abc123
Language: English (en)
URL: https://www.youtube.com/watch?v=abc123
================================================================================
Transcript text here...
Standard subtitle format compatible with video players:
1
00:00:00,000 --> 00:00:02,500
Transcript text here...
2
00:00:02,500 --> 00:00:05,000
More text...
usage: youtube_transcript_downloader.py [-h] [-o OUTPUT] [-f {json,txt,srt}]
[-l LANGUAGES [LANGUAGES ...]]
[--limit LIMIT] [-v] [--force]
[--scraperapi-key SCRAPERAPI_KEY]
channel_url
positional arguments:
channel_url YouTube channel URL
optional arguments:
-h, --help show this help message and exit
-o OUTPUT, --output OUTPUT
Output directory for transcripts (default: transcripts)
-f {json,txt,srt}, --format {json,txt,srt}
Output format (default: json)
-l LANGUAGES [LANGUAGES ...], --languages LANGUAGES [LANGUAGES ...]
Preferred language codes (default: en)
--limit LIMIT Maximum number of videos to process
-v, --verbose Enable verbose logging
--force Force refresh video list and re-download all transcripts
--scraperapi-key SCRAPERAPI_KEY
ScraperAPI key for proxy support (default: SCRAPERAPI_KEY env var)
- Channel Discovery: Uses
scrapetubeto fetch all video IDs from the channel without requiring API keys - Transcript Retrieval: Uses
youtube-transcript-apito download transcripts for each video - Language Handling: Attempts to get transcripts in preferred languages, falls back to any available transcript
- Error Handling: Gracefully handles videos without transcripts, disabled transcripts, and unavailable videos
- Output: Saves transcripts in the specified format with sanitized filenames
The script creates a download.log file in the output directory with detailed information about:
- Videos processed
- Successfully downloaded transcripts
- Failed downloads and reasons
- Overall statistics
The script handles various error cases:
- Videos with disabled transcripts
- Videos without available transcripts
- Unavailable or private videos
- Rate limiting (Too Many Requests)
- Network errors
- Some videos may not have transcripts available
- Auto-generated transcripts may have lower accuracy
- YouTube may rate-limit requests if too many are made quickly
- Private or age-restricted videos cannot be accessed
0: All transcripts downloaded successfully1: No transcripts were downloaded2: Some transcripts downloaded, but some failed130: Interrupted by user (Ctrl+C)
python youtube_transcript_downloader.py "https://www.youtube.com/@veritasium"python youtube_transcript_downloader.py "https://www.youtube.com/@3blue1brown" --format txt --limit 5python youtube_transcript_downloader.py "https://www.youtube.com/@DuolingoSpanish" --languages es enpython youtube_transcript_downloader.py "https://www.youtube.com/@crashcourse" --format srt --output course_subtitlesScraperAPI helps avoid rate limiting, IP blocks, and request failures when downloading large numbers of transcripts. This is especially useful for channels with many videos.
- Avoid Rate Limiting: YouTube may temporarily block requests if too many are made quickly
- Better Success Rate: Reduces failed requests and "Too Many Requests" errors
- Automatic Retry Logic: ScraperAPI handles retries automatically
- Geographic Distribution: Routes requests through different IPs
-
Sign up for ScraperAPI: Get a free account at https://www.scraperapi.com
- Free tier includes 5,000 API credits (enough for ~5,000 requests)
-
Get your API key: Find it in your ScraperAPI dashboard
-
Test your API key (optional but recommended):
python test_scraperapi_connection.py YOUR_API_KEY
-
Configure the script: Use one of these methods:
Method 1: Environment Variable (Recommended)
export SCRAPERAPI_KEY="your_api_key_here" python youtube_transcript_downloader.py "https://www.youtube.com/@channelname"
Method 2: Command-Line Argument
python youtube_transcript_downloader.py "https://www.youtube.com/@channelname" --scraperapi-key your_api_key_here
# Download with ScraperAPI to avoid rate limiting
export SCRAPERAPI_KEY="your_api_key_here"
python youtube_transcript_downloader.py "https://www.youtube.com/@veritasium"
# Combine with other options
python youtube_transcript_downloader.py "https://www.youtube.com/@3blue1brown" \
--scraperapi-key your_api_key_here \
--format txt \
--limit 50
# Process large channels without worrying about blocks
python youtube_transcript_downloader.py "https://www.youtube.com/@crashcourse" \
--scraperapi-key your_api_key_here \
--output course_contentWhen ScraperAPI is enabled, all HTTP requests (both video discovery and transcript fetching) are routed through ScraperAPI's proxy servers. This provides:
- Automatic proxy rotation
- Request retries on failure
- IP rotation to avoid blocks
- Geographic distribution of requests
The integration is transparent - all existing features work the same way, just more reliably.
Issue: "Too Many Requests" error
- Solution: Use ScraperAPI to avoid rate limiting:
--scraperapi-key YOUR_KEY - Alternative: Wait a few minutes before retrying. YouTube may temporarily rate-limit requests.
Issue: No transcripts found for videos
- Solution: Not all videos have transcripts. Try with
--languages ento include auto-generated English transcripts.
Issue: Import errors
- Solution: Make sure all dependencies are installed:
pip install -r requirements.txt
Issue: Many failed requests or blocked videos
- Solution: Use ScraperAPI for better reliability and to bypass IP blocks
Contributions are welcome! Please feel free to submit issues or pull requests.
This project is provided as-is for educational and personal use.
This tool is for personal and educational use only. Please respect YouTube's Terms of Service and content creators' rights. Always ensure you have the right to download and use the transcripts.