Dataset · Speech recognition
dhravani-iitpatna-test
by Shrikant Nayak shrikanth-19/dhravani-iitpatna-test
Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model.
Dataset Card
By Shrikant Nayak, published under cc-by-4.0, revision 841c8ce6a15f.
Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
Dataset Preparation Interface for Fine-tuning Whisper
A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication.
Features
- User authentication via Pocketbase
- Cloud storage support (Hugging Face Datasets)
- Multi-language support with native names
- Modern Material Design recording interface
- CSV transcript file support
- Session-based recording workflow
- Advanced recording controls
- ⌨ Keyboard shortcuts for efficiency
- Progress tracking and navigation
- Local and cloud metadata management
- Responsive, mobile-friendly UI
Getting Started
- Create a transcript CSV file with your content:
transcript
"First sentence to record"
"Second sentence to record"
# For multi-language support:
transcript_en,transcript_es
"English sentence","Spanish sentence"
- Start the Flask application:
python app.py
- Access the interface:
http://localhost:5000
Usage
- Authentication
-
Sign in using your Google account 2. Session Setup
-
Upload your transcript CSV
- Select language and recording location
- Enter speaker details
-
Click "Start Session" 3. Recording
-
Use on-screen controls or keyboard shortcuts:
R: Start recording / Stop recordingSpace: Play recordingEnter: Save recordingBackspace: Re-record←: Previous transcript→: Skip current
- Navigate using row numbers
- Adjust transcript font size as needed
Data Storage
Recordings are stored in language-specific directories:
- Storage:
datasets/
├── en/
│ ├── audio/
│ │ ├── {user_prefix}_{YYYYMMDD_HHMMSS}.wav
│ │ └── ...
│ └── en.parquet # English recordings metadata
├── es/
│ ├── audio/
│ │ ├── {user_prefix}_{YYYYMMDD_HHMMSS}.wav
│ │ └── ...
│ └── es.parquet # Spanish recordings metadata
│
Technical Details
Audio Recording
- Browser Recording Format: 48kHz mono WebM
- Storage Format: 16bit mono WAV
- Maximum Duration: 30 seconds
- Audio Processing: WebM -> WAV conversion with sample rate adjustment
- Channels: 1 (mono)
Data Management
- Metadata Organization:
- stats.json: Global recording statistics
- {language_code}.parquet: Language-specific metadata files
- File Naming:
{user_id_prefix}_{YYYYMMDD_HHMMSS}.{format} - Unicode Handling: NFC normalization for text
Authentication
- Provider: Pocketbase with Google OAuth
- Session Management: Server-side Flask sessions
Languages
- Support: 74 languages with native names
- Codes: ISO 639-1 standard
- CSV Format:
- Single language:
transcriptcolumn - Multi-language:
transcript_${lang_code}columns
Upload Management
- Queue System: Background worker thread
- Status Tracking: Real-time upload status polling
- Error Handling: Automatic retries with timeout
- Progress Updates: Toast notifications
- Temporary Storage: ./temp folder for conversions
Frontend Features
- Keyboard Shortcuts: Recording and navigation
- Real-time Status: Progress tracking and notifications
Security
- Authentication Required: All routes except static/login
- File Validation: MIME type and extension checking
- Secure Context: HTTPS recommended
Performance
- Upload Queue: Asynchronous processing
- Audio Conversion: Server-side processing
- Session Caching: Browser storage optimization
- Progress Tracking: Real-time websocket updates
Browser Support
- Chrome (recommended)
- Brave
- Edge
- Safari
Known Limitations
- Requires microphone permissions
- Internet connection needed
- Maximum recording duration: 30 seconds
- File size limits based on storage backend
Structure
bn 1 rows
| Split | Rows | Size |
|---|---|---|
| train | 1 | 103.5 KB |
mai 9 rows
| Split | Rows | Size |
|---|---|---|
| train | 9 | 1.5 MB |
ol 11 rows
| Split | Rows | Size |
|---|---|---|
| train | 11 | 1.2 MB |
ol_odi 3 rows
| Split | Rows | Size |
|---|---|---|
| train | 3 | 1.4 MB |
Details
- Repository
- shrikanth-19/dhravani-iitpatna-test
- Publisher
- Shrikant Nayak
- Task category
- Speech recognition
- Tags
- Not stated by the source
- Size category
- Not stated by the source
- Languages
- bn, mai, ol
- Revision
- 841c8ce6a15f8b628c2a68e4e7a7b2368e081b1b
- Last updated
- 2026-09-18
Files
17 files, 5.0 MB in total.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| bn/bn.parquet | Data | 49.7 KB | 719e679f8377 |
| mai/mai.parquet | Data | 466.3 KB | 15445a9faa7e |
| ol/ol.parquet | Data | 1.0 MB | 8e5b0b91fae3 |
| ol_odi/ol_odi.parquet | Data | 1.3 MB | e5973c202ade |
| README.md | Documentation | 4.6 KB | — |
| ol/audio/san_test1_nayakshrikant56_gmail_com/nayakshrikant56_20260912_095316.wav | Other | 79.2 KB | e761d222e271 |
| ol/audio/san_test1_nayakshrikant56_gmail_com/nayakshrikant56_20260912_095330.wav | Other | 112.0 KB | 0f344248790f |
| ol/audio/san_test1_nayakshrikant56_gmail_com/nayakshrikant56_20260912_095336.wav | Other | 95.6 KB | 56e1d3617562 |
| ol/audio/san_test1_nayakshrikant56_gmail_com/nayakshrikant56_20260912_095347.wav | Other | 57.4 KB | d77e5eb7cba8 |
| ol/audio/san_test1_nayakshrikant56_gmail_com/nayakshrikant56_20260912_095354.wav | Other | 79.2 KB | d2ab63cce955 |
| ol/audio/san_test1_shrikanth_22cs150_sode_edu_in/unknown_20260912_102335.wav | Other | 251.3 KB | ff57c092c15a |
| ol/audio/san_test1_shrikanth_22cs150_sode_edu_in/unknown_20260912_102339.wav | Other | 65.6 KB | ef3b8b525657 |
| ol/audio/san_test1_shrikanth_22cs150_sode_edu_in/unknown_20260912_102344.wav | Other | 38.3 KB | 5fa97900797c |
| ol_odi/audio/odiatest_nayakshrikant56_gmail_com/nayakshrikant56_20260912_105847.wav | Other | 991.3 KB | c5124b32181b |
| ol_odi/audio/odiatest_nayakshrikant56_gmail_com/nayakshrikant56_20260912_110047.wav | Other | 174.8 KB | 2c2c9da459cc |
| ol_odi/audio/odiatest_nayakshrikant56_gmail_com/nayakshrikant56_20260912_110055.wav | Other | 221.2 KB | e8808dadec42 |
| .gitattributes | Repository | 2.5 KB | — |
License and Download
- License
- cc-by-4.0
- Access
- No access gate
Released by Shrikant Nayak through its official repository on Hugging Face. Read the license.