tensorflow — bootstrapping an environment
Sources: examples/workflow_tensorflow — settings in workflows/tensorflow/.
workflows:
- tensorflow + epochs/*
The last example trains a small classifier for a swept number of epochs. It is the only one with an external dependency, and it handles it inside the workflow rather than requiring a pre-installed environment:
tasks:
- name: main
dependencies:
- /tmp/odatix_tensorflow_example_env/bin/python3
commands:
- /tmp/odatix_tensorflow_example_env/bin/python3 train.py
- name: /tmp/odatix_tensorflow_example_env/bin/python3
commands:
- python3.10 -m venv /tmp/odatix_tensorflow_example_env
- /tmp/odatix_tensorflow_example_env/bin/python3 -m pip install --upgrade pip
- /tmp/odatix_tensorflow_example_env/bin/python3 -m pip install tensorflow
on_failure_commands:
- python3.10 -m venv /tmp/odatix_tensorflow_example_env
- ...
Three things are going on:
- The dependency task is named after the file it produces, so the virtualenv is built once and reused by every subsequent job instead of being rebuilt ten times.
on_failure_commandsgives a fallback path when the primary one does not work — a different interpreter version, here.- The training script reports progress from inside a Keras callback, so the job monitor tracks epochs:
class ProgressCallback(tf.keras.callbacks.Callback):
def on_epoch_end(self, epoch, logs=None):
logs = logs or {}
progress = int(((epoch + 1) / EPOCHS) * 90 + 5)
report_progress(progress)
The epochs domain generates its ten configurations rather than listing them, and writes the value straight into the script:
start_delimiter: 'EPOCHS = '
stop_delimiter: "\n"
param_target_file: "train.py"
generate_configurations: Yes
generate_configurations_settings:
template: "${epochs}"
name: "${epochs}"
variables:
epochs:
type: range
settings:
from: 5
to: 50
step: 5
The metrics end on an operation that answers the actual question of the sweep — is the model overfitting?
metrics:
final_loss:
type: json
settings: { file: workflow_results.json, key: final_loss }
format: "%.6f"
final_val_loss:
type: json
settings: { file: workflow_results.json, key: final_val_loss }
format: "%.6f"
generalization_gap:
type: operation
settings:
op: "final_val_loss - final_loss"
format: "%.6f"
Plotting generalization_gap against epochs in the Explorer shows the training loss and the validation loss parting company — the same kind of curve, read the same way, as an area-versus-frequency trade-off.
This example needs a Python interpreter TensorFlow supports, available as python3.10 on the machine. It is the one workflow that will not run out of the box everywhere.