Background Buffering
ML-Dash includes a comprehensive background buffering system that eliminates I/O blocking during training. All write operations (logs, metrics, tracks, files) are automatically batched and executed in background threads.
Overview
The buffering system provides:
- Non-blocking writes: All operations return immediately
- Automatic batching: Groups writes for efficiency
- Parallel uploads: Uses ThreadPoolExecutor for files
- Graceful error handling: Training continues even if uploads fail
- Progress feedback: Shows flush status after training
How It Works
Flush Triggers
Data is automatically flushed when:
- Time-based: Every 5 seconds (default)
- Size-based: When a queue reaches its batch size (100 logs or track entries, 1000 metric points by default)
- Manual: When you call
experiment.flush() - Context exit: When the experiment context manager exits
Configuration
Environment Variables
The settings are read from the environment when the Experiment is
constructed, so set them before you create it. There is no per-experiment
argument: Experiment(..., buffer_config=...) is not a parameter, and like any
unknown keyword it is silently ignored.
Manual Flushing
Force an immediate flush at any time:
Disabling Buffering
For debugging or special cases, you can disable buffering:
Performance Benefits
Without Buffering (Blocking)
With Buffering (Non-blocking)
Result: 10-100x speedup for high-frequency logging!
Progress Messages
When the experiment completes, you'll see:
Error Handling
The buffering system handles errors gracefully:
Thread Safety
The buffering system is fully thread-safe:
Best Practices
- Let it auto-flush: Don't call
flush()unless you need guarantees before checkpoints - Use context managers: Ensures all data is flushed on exit
- Monitor progress: Watch the flush messages to understand batching behavior
- Adjust batch sizes: Increase for very high-frequency logging
- Keep buffering enabled: It's designed for production use
Technical Details
Architecture
- Single background thread: One daemon thread per experiment
- Resource-specific queues: Separate queues for logs, metrics, tracks, files
- No timeout on close: Waits indefinitely for all data to flush
- Temp file cleanup: Automatically cleans up temporary files after upload
File Handling
When you save files via methods like save_image(), save_json(), etc.:
- Content is written to a temporary file
- File upload is queued in the buffer
- Temporary file is kept until upload completes
- Buffer manager cleans up temp file after successful upload
- All cleanup happens automatically in the background
This ensures:
- No file handle leaks
- No disk space waste
- Proper cleanup even if uploads are delayed
Troubleshooting
Data not appearing immediately
This is expected! Data is buffered and flushed periodically. If you need immediate visibility:
Memory usage growing
If you're generating data faster than it can be uploaded:
Uploads failing
Check the warnings in console output. Common issues:
- Network connectivity
- Authentication expired (run
ml-dash loginagain with the separately installed CLI; see Authentication) - Server unavailable
The buffering system will retry and warn you, but training continues.