DataLab: analyse the experiments¶
Train with scripts, explore with a notebook. All figures below use validation results.
Run from the mini-project: jupyter lab notebooks/results.ipynb. First pull the logs and small prediction files with bash scripts/sync.sh pull. Checkpoints stay on the cluster. Prepared real runs in demo_results/ provide a queue-independent fallback.
from pathlib import Path
import sys
ROOT = Path.cwd()
if ROOT.name == 'notebooks':
ROOT = ROOT.parent
sys.path.insert(0, str(ROOT))
from src.analysis import find_runs, learning_curves, tradeoff, confusion, mistakes
import plotly.io as pio
pio.renderers.default = "plotly_mimetype+notebook"
import ipywidgets as widgets
from IPython.display import display
runs = find_runs(ROOT / 'runs') or find_runs(ROOT / 'demo_results')
assert runs, 'Pull runs or use the supplied demo_results first.'
preferred = [p for p in runs if p.name in ['cnn-full', 'resnet-full']] or [p for p in runs if p.name in ['cnn-6', 'resnet-6']] or runs
print('Available runs:', [p.name for p in runs])
Available runs: ['cnn-6', 'cnn-full', 'config-smoke', 'trial-000', 'trial-001', 'trial-002', 'resnet-6', 'resnet-full', 'resume-check']
1. Compare runs¶
Choose architectures, zoom both axes, change the metric. What explains the gap between train and validation?
selector = widgets.SelectMultiple(options=[(p.name, str(p)) for p in runs], value=tuple(str(p) for p in preferred), description='Runs')
metric = widgets.Dropdown(options=['accuracy', 'loss'], description='Metric')
@widgets.interact(selected=selector, metric=metric)
def compare(selected, metric):
display(learning_curves([Path(p) for p in selected], metric))
2. Which run would you choose?¶
These are short demonstration runs, not converged benchmarks. Check data size, epochs, seed and device before comparing. Selecting a model uses validation, never the held-out test.
display(tradeoff(preferred))
3. Inspect predictions¶
Filter by true class and inspect mistakes. The confidence shown by a softmax is not a calibration guarantee.
run_picker = widgets.Dropdown(options=[(p.name, str(p)) for p in runs], value=str(preferred[0]), description='Run')
class_picker = widgets.Dropdown(options=[('All', -1), ('airplane', 0), ('automobile', 1), ('bird', 2), ('cat', 3), ('deer', 4), ('dog', 5), ('frog', 6), ('horse', 7), ('ship', 8), ('truck', 9)], description='Class')
@widgets.interact(run=run_picker, only_errors=True, class_id=class_picker)
def inspect(run, only_errors, class_id):
fig = mistakes(Path(run), only_errors, class_id)
if fig is not None:
display(fig)
display(confusion(preferred[0]))
4. Trace one result¶
Find its command/configuration, software versions, Slurm job ID, seed and checkpoint-selection rule. Propose one controlled follow-up experiment.
import json
config = json.loads((preferred[0] / 'config.json').read_text())
display({key: config.get(key) for key in ['model', 'lr', 'seed', 'train_size', 'val_size', 'epochs', 'device_used', 'job_id', 'torch_version', 'git_commit', 'git_dirty']})
{'model': 'small_cnn',
'lr': 0.001,
'seed': 42,
'train_size': 45000,
'val_size': 5000,
'epochs': 12,
'device_used': 'cuda',
'job_id': '121733',
'torch_version': '2.0.1+cu117',
'git_commit': None,
'git_dirty': None}