ML-Dash

Experiments

Experiments are the foundation of ML-Dash. Each experiment represents a single experiment run, containing all your logs, parameters, metrics, and files.

Prefix Format

The prefix is a universal key that identifies your experiment:

{owner}/{project}/path.../[name]
  • owner: First segment (e.g., your username)
  • project: Second segment (e.g., project name)
  • path: Remaining segments form the folder structure
  • name: Derived from the last segment

Three Usage Styles

Context Manager (recommended for most cases):

python

from ml_dash import Experiment

# Prefix format: owner/project/experiment-name
with Experiment(prefix="alice/project/my-experiment").run as exp:
    exp.log("Training started")
    exp.params.set(learning_rate=0.001)
    # Experiment automatically closed on exit

Decorator (clean for training functions):

python

from ml_dash import ml_dash_experiment

@ml_dash_experiment(prefix="alice/project/my-experiment")
def train_model(experiment):
    experiment.log("Training started")
    experiment.params.set(learning_rate=0.001)

    for epoch in range(10):
        loss = train_epoch()
        experiment.metrics("train").log(loss=loss, epoch=epoch)

    return "Training complete!"

result = train_model()

Direct (manual control):

python

from ml_dash import Experiment

exp = Experiment(prefix="alice/project/my-experiment")
exp.run.start()

try:
    exp.log("Training started")
    exp.params.set(learning_rate=0.001)
finally:
    exp.run.complete()

Automatic Path Detection with RUN.entry

Use RUN.entry = __file__ to automatically detect the script path and generate a meaningful experiment prefix:

python

from ml_dash.run import RUN
from ml_dash.auto_start import dxp

# Set entry point to current file - enables automatic path detection
RUN.entry = __file__

# Now dxp will use the script's path to generate the prefix
# e.g., if running /home/user/project/experiments/train.py
# prefix becomes: user/project/experiments/train
with dxp.run:
    dxp.log("Training with automatic prefix")
    dxp.params.set(learning_rate=0.001)

This is useful when:

  • Running multiple scripts in the same project
  • You want the prefix to automatically reflect the file structure
  • Organizing experiments by script location

Local, Hybrid, and Remote Mode

Where an experiment writes depends on dash_url and dash_root:

ModeHow to get itWrites to
Local (default)no dash_url.dash/ on disk. No server, no login
Hybriddash_url=...the server and .dash/
Remotedash_url=..., dash_root=Nonethe server only
python

# Local: zero setup, filesystem storage
with Experiment(prefix="alice/project/my-experiment").run as exp:
    exp.log("Using local storage")

# Hybrid: the server, plus a local copy
with Experiment(
    prefix="alice/project/my-experiment",
    dash_url="https://api.dash.ml"
).run as exp:
    exp.log("Using the server and .dash/")

# Remote only
with Experiment(
    prefix="alice/project/my-experiment",
    dash_url="https://api.dash.ml",
    dash_root=None
).run as exp:
    exp.log("Using the server only")

Any mode that uses the server needs a login first. See Authentication. Runs recorded in local mode can be uploaded later with ml-dash upload.

Experiment Metadata

Add a description (readme), tags, and bindrs for organization:

python

with Experiment(
    prefix="alice/computer-vision/resnet50-imagenet",
    readme="ResNet-50 training with new augmentation",
    tags=["resnet", "imagenet", "baseline"],
    bindrs=["gpu-cluster", "team-a"]
).run as exp:
    exp.log("Training started")

Metadata fields:

  • readme: Human-readable experiment description (shown as the description on the server). Note that the argument is readme, not description. Unknown keyword arguments are silently ignored.
  • tags: List of tags for categorization (e.g., ["baseline", "production"])
  • bindrs: List of bindrs for resource/team association (e.g., ["gpu-1", "team-ml"])

Experiment Status Lifecycle

Experiments automatically track their status through the lifecycle:

  • RUNNING: Automatically set when experiment opens
  • COMPLETED: Set when experiment closes normally
  • FAILED: Set when exception occurs during experiment
  • CANCELLED: Can be set manually

Automatic status management (recommended):

python

# Normal completion - status becomes COMPLETED
with Experiment(
    prefix="alice/ml/training",
    dash_url="https://api.dash.ml"
).run as exp:
    exp.log("Training...")
    # Status automatically set to COMPLETED on exit

# Exception handling - status becomes FAILED
with Experiment(
    prefix="alice/ml/training",
    dash_url="https://api.dash.ml"
).run as exp:
    exp.log("Training...")
    raise ValueError("Training failed!")
    # Status automatically set to FAILED on exception

Manual status control:

python

from ml_dash import Experiment

exp = Experiment(
    prefix="alice/ml/training",
    dash_url="https://api.dash.ml"
)
exp.run.start()

try:
    exp.log("Training...")
    # ... training code ...
    exp.run.complete()
except KeyboardInterrupt:
    exp.run.cancel()
except Exception as e:
    exp.log(f"Error: {e}")
    exp.run.fail()

Note: Status updates only work in remote mode. Local mode doesn't track status.

Resuming Experiments

Experiments use upsert behavior - reopen by using the same prefix:

python

# First run
with Experiment(prefix="alice/ml/long-training").run as exp:
    exp.log("Starting epoch 1")
    exp.metrics("train").log(loss=0.5, epoch=1)

# Later - continues same experiment
with Experiment(prefix="alice/ml/long-training").run as exp:
    exp.log("Resuming from checkpoint")
    exp.metrics("train").log(loss=0.3, epoch=2)

Available Operations

Once an experiment is open, you can use all ML-Dash features:

python

with Experiment(prefix="alice/test/demo").run as exp:
    # Logging
    exp.log("Training started", level="info")

    # Parameters
    exp.params.set(lr=0.001, batch_size=32)

    # Metrics tracking
    exp.metrics("train").log(loss=0.5, epoch=1)

    # File uploads
    exp.files("models").save("model.pth")

Storage Structure

Local mode creates a directory structure:

.dash/
└── alice/                          # owner
    └── project/                    # project
        └── my-experiment/          # experiment name
            ├── logs/
            │   └── logs.jsonl
            ├── parameters.json
            ├── metrics/
            │   └── train/
            │       └── data.jsonl
            └── files/
                └── models/
                    └── 7218065541365719/
                        └── model.pth

Files are stored under files/{prefix}/{snowflake_id}/{filename} where:

  • prefix: Logical path (e.g., "models", "configs")
  • snowflake_id: Unique identifier for the file
  • filename: Original filename

Remote mode stores data in MongoDB + S3 on your server.


Next: Learn about Logging to track events and progress.