Hands-on computational biology

From peptide data
to model interpretation.

Install, configure, fit, validate, and explain. A complete command-line workflow with real synthetic-data outputs and a field-by-field reference.

6 scenarios × 2 estimatorsNested CV + external validationROC · PR · SHAP · feature tablesWorks offline

phipml: from peptide data to model interpretation

A runnable tutorial for Random Forest, XGBoost, nested cross-validation, external validation, and publication-ready plots.

Built against phipml 4.2.0, commit 932f14f, inspected on 8 September 2026. The commands below use the implemented configuration schema and were run on the supplied synthetic data.

You will start with peptide and sample metadata CSVs, fit either estimator, compare fixed/default and tuned models, evaluate an independent synthetic cohort, and produce ROC, precision–recall, classification, SHAP, and annotated feature plots. Finished results are included so you can inspect them before rerunning anything.

Start here: follow sections 1–6 for your first model and plots. Then run the six scenarios, compare the results, and use the references when adapting your own configuration. For a workshop, start with Random Forest and add XGBoost after the workflow is familiar.

What is included

Item Location
This walkthrough README.md
Every model-config argument docs/MODEL_CONFIG.md
Shared selector, RF and XGBoost hyperparameters docs/HYPERPARAMETERS.md
Every plot-config argument and heatmap command docs/PLOTTING_CONFIG.md
Observed results and verification notes reports/REFERENCE_RUN.md
Deterministic mock-data/config generator scripts/generate_examples.py
Twelve ready-to-run model recipes configs/models/{random-forest,xgboost}/01...06*.yaml
Corresponding plot recipes configs/plots/{random-forest,xgboost}/
CSV data, metadata, annotations, data hashes data/
Fitted models, predictions, and metrics results/, reports/
Actual plots, in PDF/SVG/PNG plots/
One-command workflow and repeated-run example scripts/run_all.sh, scripts/run_repeated.sh

The tutorial is a standalone folder. It can also be added under tutorials/from_data_to_plots/ in the phipml repository: relative paths inside the folder continue to work. Run its commands from that folder. It does not require changes to phipml's source.

1. Install phipml

Use Linux/macOS, or an equivalent Bash environment. Install Git and Micromamba first if needed; platform setup and Docker/Apptainer alternatives are in the repository's installation guide.

git clone https://github.com/csReynaB/phipml.git
cd phipml

# Pin the source used by this tutorial. This is a local checkout only.
git checkout 932f14ffca41ba4f24f3cc570afea27265a5893c

micromamba create --yes --name phipml --file ML_env.yml
micromamba activate phipml
python -m pip install --no-build-isolation --no-deps -e .

phipml --version
phipml -h
phipml-plot -h
phipml-heatmap -h

ML_env.yml supplies the tested scientific package pins with Python 3.10.16. --no-deps prevents pip from replacing those packages. -e links the installed package to the cloned source, so retain the checkout.

Alternative: pip in a virtual environment

In the same pinned repository checkout, use an installed Python 3.10–3.12:

python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install .
phipml --version

Although the package declares Python >=3.10, its dependency upper bounds make 3.10–3.12 the practical choices for this tutorial; do not assume every newer Python can resolve the pinned scientific stack. Pip resolves compatible ranges and can choose versions different from ML_env.yml. The included run used Python 3.12.13, sklearn 1.5.2, XGBoost 2.1.4, and SHAP 0.46.0; exact package versions are recorded in reports/requirements-reference.txt.

The virtual environment must remain active while running the tutorial. The small examples use CPU execution; no GPU is required.

Open the tutorial folder

Extract phipml-tutorial-932f14f.zip and enter the extracted folder. For example, if your downloaded archive is in the current directory:

unzip phipml-tutorial-932f14f.zip
cd phipml-tutorial-932f14f

All commands below assume this directory. Data/config files are already included. To recreate them deterministically:

python scripts/generate_examples.py

The generator rewrites its example data/configs with the documented defaults; copy a config before customizing it. It does not delete old results. If you change --data-seed or --n-iter, keep a separate tutorial copy/output directory so old artifacts cannot be mistaken for the new experiment.

2. Understand the example data

There are two datasets, perfect and noisy. Each includes 160 training samples and 80 independent external samples, with balanced Control/Case labels. Every sample has 80 binary peptide measurements; four numeric clinical predictors are supplied through metadata.

File Rows represent Required identifiers/content
data/noisy/peptides.csv Samples First column SampleName; remaining columns are named peptides containing 0/1.
data/noisy/metadata.csv Samples SampleName, group_test, cohort, Sex, Smoking, Age, BMI.
data/peptide_library.csv Peptides peptide_id, Description, Role, Species, Protein, is_demo_signal. Used for annotation, not as predictor columns.

The matching files under data/perfect/ have the same schema. Inspect the actual data with:

python - <<'PY'
import pandas as pd
X = pd.read_csv("data/noisy/peptides.csv", index_col=0)
metadata = pd.read_csv("data/noisy/metadata.csv").set_index("SampleName")
print("Peptide matrix:", X.shape)
print(X.iloc[:4, :6])
print(metadata.head(4))
print(pd.crosstab(metadata["cohort"], metadata["group_test"]))
assert X.index.is_unique and metadata.index.is_unique
assert set(X.index) == set(metadata.index)
PY

Expected counts:

Cohort Control Case Total
training 80 80 160
external 40 40 80

group_tests: [Control, Case] encodes Control=0 and Case=1. cohort selects training/external rows and is not included in the predictor matrix. The model receives 80 peptides plus the four configured extras: 84 features before pipeline feature selection.

What “perfect” and “noisy” mean

In the perfect data, six engineered peptides exactly encode the label or its inverse. This is an intentional software demonstration. The remaining 74 peptides and clinical features are independent of class in the data-generating process. Correlated perfect peptides can share model importance unevenly.

In the noisy data, six peptides have class-dependent presence probabilities: 0.70 versus 0.30 in training, and 0.65 versus 0.35 externally, with alternating directions. The external relationship is deliberately slightly weaker. The external age distribution is also shifted by four years, independently of class. There is no promise that tuning improves the result.

The generator draws training and external samples independently. Default and tuned runs reuse exactly the same input files and model seed. Synthetic annotations identify the intended signal/noise roles; they are not real taxa or biomarkers. data/generation_manifest.json records the generation settings and SHA-256 hashes.

Matrix orientation

The tutorial uses sample rows × peptide columns with transposed: false. For a peptide-row matrix, use transposed: true. Equivalent peptides_transposed.csv files are included for illustration. The first column must contain the relevant row identifiers in either orientation.

3. Read the first model config

Open configs/models/random-forest/01_perfect.yaml. The main decisions are:

# Paths below are relative to configs/models/random-forest/.
data_input: ../../../data/perfect/peptides.csv
metadata_input: ../../../data/perfect/metadata.csv
data_input_mode: matrix
transposed: false

col_sample_name: SampleName
col_target: group_test
group_tests: [Control, Case]
peptide_prefixes: [agilent_, twist_, corona2_]
extra_features_to_include: [Sex, Smoking, Age, BMI]
fillna_value: null

lib_metadata_input: ../../../data/peptide_library.csv
lib_col_peptide_name: peptide_id
random_state: 420

classification:
  model_type: random-forest
  param_grid_name: random-forest
  seed: 420
  train_filters: {cohort: training}
  run_nested_cv: true
  only_train_model: true
  use_pretrained: false
  with_oligos: true
  with_additional_features: true
  prevalence_threshold_min: 0
  prevalence_threshold_max: 100
  outer_cv_splits: 3
  inner_cv_splits: 2
  n_iter: 1
  n_jobs_outer: 1
  n_jobs_inner: 1
  classification_threshold: 0.5
  validation_sets: []
  output_dir: ../../../results/random-forest/01_perfect
  output_name: 01_perfect

param_grid: {}

This is a complete minimal working configuration; the supplied file spells out additional options documented in the model reference.

param_grid: {} means no hyperparameter search. Constant-feature removal and elastic-net peptide selection still fit inside each training pipeline. run_nested_cv: true requests outer-CV evaluation; with an empty grid there is no inner tuning. only_train_model: true then saves a model fitted on the whole training cohort and skips external validation.

External rows are present in the input files but excluded from all training and tuning by train_filters: {cohort: training}.

4. Fit a Random Forest and locate the outputs

phipml -c configs/models/random-forest/01_perfect.yaml

The log should report samples=160, features=84. This run creates:

Output under results/random-forest/01_perfect/ Use
nested_random-forest_01_perfect_420.joblib Held-out outer-CV predictions, metrics, SHAP, selected features, and fold models. Use this to plot internal performance.
training_random-forest_01_perfect_420.joblib Pipeline refitted on all training samples. Use this for later prediction/external validation.

Filename pattern: {nested|training}_{model_type}_{output_name}_{seed}.joblib. Running the same configuration/seed again replaces the same filenames.

5. Generate the first plots

phipml-plot --plot-config configs/plots/random-forest/01_perfect.yaml

Open plots/random-forest/01_perfect/01_perfect_performance.png. It should show near-perfect discrimination because the data were constructed to be perfectly separable. The plotting recipe produces nine figure types in PDF, SVG, and PNG, plus a detailed feature-importance CSV.

The essential plotting fields are:

plotting:
  results:
    - ../../../results/random-forest/01_perfect/nested_random-forest_01_perfect_420.joblib
  config: ../../models/random-forest/01_perfect.yaml
  split: train
  plots: [all]
  formats: [pdf, svg, png]
  dpi: 180
  output_dir: ../../../plots/random-forest/01_perfect
  output_prefix: 01_perfect
  class_labels: [Control, Case]
  max_display: 10
  library_metadata: ../../../data/peptide_library.csv
  library_id_column: peptide_id

--plot-config reads the plotting recipe; its config: entry points to the model/data YAML. split: train refers to the training cohort's out-of-fold predictions. Use split: test for an external validation artifact. The separate training_...joblib file is not an evaluation result to plot.

See the plotting reference for all fields, colors, annotation controls, feature ranking, and standalone-panel commands.

6. Move to noisy data, then tune

# Identical schema, imperfect signal, default hyperparameters.
phipml -c configs/models/random-forest/02_noisy.yaml
phipml-plot --plot-config configs/plots/random-forest/02_noisy.yaml

# The same noisy samples and seed, with inner-CV hyperparameter search.
phipml -c configs/models/random-forest/03_noisy_tuned.yaml
phipml-plot --plot-config configs/plots/random-forest/03_noisy_tuned.yaml

The tuned YAML adds search spaces beneath param_grid.random-forest and sets n_iter: 12. It searches the peptide selector's regularization, forest tree count/depth, split/leaf sizes, feature subsampling, and class weighting. All candidate comparisons happen in training folds. The held-out outer samples are evaluated only after a candidate has been selected.

For example, increasing min_samples_leaf makes leaves depend on more samples and generally smooths predictions; increasing selector C weakens regularization. Those are different parts of the pipeline. The hyperparameter guide explains every search parameter, the direction of its effect, and the exact YAML syntax.

What a real noisy result looks like

Random Forest noisy-data internal performance

This figure is generated by the command above. ROC/PR summarize outer-fold performance. The confusion matrix pools held-out predictions. The classification bars in this version use outer-fold means and variability when present; these can differ slightly from metrics recomputed on all pooled predictions.

Interpret the learned features

Annotated noisy-data features

The highest-ranked peptides largely recover the engineered signal in this example. Noise peptides and clinical variables can still receive nonzero importance. Class-wise prevalence/means and SHAP importance answer different questions: prevalence describes the samples; SHAP explains the fitted model. The CSV also includes column types and statistic labels omitted from the compact figure. Here peptide values are percentages and Age/BMI entries are means in their original units.

7. Evaluate the external cohort

Use run_nested_cv: false, only_train_model: false, and add the validation cohort. Each model is fitted/tuned using the training rows and then evaluated on the 80 independent external samples.

classification:
  train_filters: {cohort: training}
  run_nested_cv: false
  only_train_model: false
  validation_sets:
    - name: 05_external_noisy
      filters: {cohort: external}
  classification_threshold: 0.5
  bootstrap_validation: true
  bootstrap_n_resamples: 500
  bootstrap_confidence_level: 0.95

The fragment illustrates the change; run the complete supplied files:

phipml -c configs/models/random-forest/04_external_perfect.yaml
phipml-plot --plot-config configs/plots/random-forest/04_external_perfect.yaml

phipml -c configs/models/random-forest/05_external_noisy.yaml
phipml-plot --plot-config configs/plots/random-forest/05_external_noisy.yaml

phipml -c configs/models/random-forest/06_external_noisy_tuned.yaml
phipml-plot --plot-config configs/plots/random-forest/06_external_noisy_tuned.yaml

For example, the noisy external artifact is results/random-forest/05_external_noisy/validation_random-forest_05_external_noisy_420.joblib. Its name comes from the validation set's name, not output_name.

The current CLI selects cohorts from shared input files. Separate per-validation matrix/metadata paths are not implemented in validation_sets at this commit. For new files, see external-input guidance.

8. Run the same six scenarios with XGBoost

Use the XGBoost configurations supplied alongside Random Forest:

for scenario in 01_perfect 02_noisy 03_noisy_tuned 04_external_perfect 05_external_noisy 06_external_noisy_tuned; do
  phipml -c "configs/models/xgboost/$scenario.yaml"
  phipml-plot --plot-config "configs/plots/xgboost/$scenario.yaml"
done

The input/cohort definitions stay the same. XGBoost tuning instead searches tree count, learning rate, depth, child weight, row/column subsampling, and leaf/split regularization, along with the shared peptide selector.

XGBoost external tuned SHAP beeswarm

Each dot is one external sample. Red/blue show high/low feature values; left/right show contributions away from/toward Case. XGBoost's SHAP values here are raw-margin/log-odds contributions, not probability-point differences. Mean absolute SHAP can rank features but does not give direction, statistical significance, or a causal interpretation.

9. One command for all scenarios

After installation, from the tutorial directory:

# Six models, all plots, heatmaps, and a metric/prediction export.
bash scripts/run_all.sh random-forest

# Or use XGBoost.
bash scripts/run_all.sh xgboost

# Or run all twelve estimator/scenario combinations.
bash scripts/run_all.sh both
Scenario What it evaluates Tuning? Evaluation samples
01_perfect Strong, engineered signal via outer CV No 160
02_noisy Imperfect signal via outer CV No 160
03_noisy_tuned Same noisy data via nested CV Yes 160
04_external_perfect Full training model on independent perfect data No 80
05_external_noisy Full training model on weaker external signal No 80
06_external_noisy_tuned Training-selected tuned model on the same noisy external data Yes 80

Core recipes use three outer folds, two inner folds, one worker, and 12 search candidates for tuned runs. This is a short tutorial budget, not a recommended final analysis design. Hardware, installation startup, tuning, and plot export affect runtime; the reference execution timings are recorded separately.

10. Compare the observed results

python scripts/summarize_results.py

phipml-heatmap --manifest manifests/random-forest.csv \
  --metric roc.auc --vmin 0 --vmax 1 --palette viridis --dpi 180 \
  --output plots/random-forest/roc_auc_overview

reports/observed_metrics.csv contains actual metrics, uncertainties, and artifact paths. reports/predictions/ contains sample IDs, encoded labels, Case probabilities, and predictions at 0.5. reports/fitted_parameters.json records the selected full-cohort pipeline settings.

Reference point estimates, seed 420:

Scenario RF ROC-AUC RF AP XGBoost ROC-AUC XGBoost AP
Perfect internal CV 1.000 1.000 1.000 1.000
Noisy internal CV 0.912 0.908 0.917 0.921
Noisy tuned internal CV 0.911 0.915 0.926 0.928
Perfect external 1.000 1.000 1.000 1.000
Noisy external 0.810 0.843 0.827 0.825
Noisy tuned external 0.806 0.805 0.808 0.804

Internal ROC-AUC/AP values are means across outer folds, not metrics calculated from pooled scores. The exported table reports both. External metrics evaluate one fixed full-cohort model. Tuning improved some internal estimates but did not improve external AP in this run. This is useful: tuning on training data cannot guarantee better generalization, especially after a distribution shift. Do not select the deployment model by repeatedly inspecting this external set.

The tutorial uses AP consistently: it is not interchangeable with trapezoidal area under a plotted precision–recall curve. The baseline AP of a random ranking is approximately the positive-class prevalence, here 0.5.

11. Repeat CV to examine stability

bash scripts/run_repeated.sh random-forest

This executes seeds 420, 421, and 422 on the same noisy dataset, stores them in results/random-forest/repeated_noisy/, and aggregates them using configs/plots/random-forest/07_repeated_noisy.yaml. It skips unnecessary full-cohort model fitting for these repeat-only runs.

Equivalent model commands, showing the relevant overrides:

for seed in 420 421 422; do
  phipml -c configs/models/random-forest/02_noisy.yaml --seed "$seed" \
    --output-dir ../../../results/random-forest/repeated_noisy \
    --no-only-train-model
done
phipml-plot --plot-config configs/plots/random-forest/07_repeated_noisy.yaml

The plot recipe keeps each run's top 10 features, then displays leading features that occur in at least 50% of runs. These three repeats demonstrate the mechanics; they are not enough to characterize feature stability precisely. More repeats reuse the same participants and must not be treated as independent studies.

Use one estimator/experiment per glob. To compare RF against XGBoost, use separate panels or a heatmap, not repeated-run aggregation.

12. Reuse the saved fitted model

After 03_noisy_tuned, validate its saved full-cohort model without another search/refit:

phipml -c configs/models/random-forest/08_reuse_noisy.yaml

The key settings are:

classification:
  run_nested_cv: false
  only_train_model: false
  use_pretrained: true
  input_dir: ../../../results/random-forest/03_noisy_tuned
  input_name: training_random-forest_03_noisy_tuned_420.joblib

The complete recipe includes the training input context and external filters, and saves validation_random-forest_08_reuse_noisy_420.joblib under results/random-forest/08_reuse_noisy/. Use the matching environment when loading serialized fitted estimators. The XGBoost equivalent is supplied too.

13. Adapt this workflow to your own data

  1. Copy a supplied model YAML to another file in the same model-config folder. Replace the input paths, sample/target column names, class order, matrix orientation, and peptide prefixes. Check sample IDs and class counts.
  2. Choose numeric clinical extras. Handle genuine missingness and categorical encodings deliberately; fillna_value: 0 is not a general solution.
  3. Define disjoint training/validation cohort filters and choose the split design before inspecting external performance. This CLI uses stratified sample-level CV, not participant-group CV.
  4. Start with param_grid: {}. Then copy the corresponding estimator's tuning space, revise it for your data size, and increase the search budget if useful.
  5. Set distinct output_dir, output_name, and validation names. Fit the model.
  6. Copy the matching plot YAML; point results to the new evaluation artifact, set config, select train/test, and supply annotation column names that actually exist. Plot without retraining.
  7. Preserve inputs, configs, source commit, dependency versions, predictions, and evaluation outputs together.

Troubleshooting

Symptom What to check
phipml: command not found Activate the environment containing the installed package; run which python and which phipml.
File not found despite an existing CSV Resolve model paths from the model YAML, plot paths from the plot YAML, and plot CLI overrides from the terminal directory.
No sample IDs in common, or unexpected sample count Matrix orientation, first-column IDs, metadata SampleName, filters, and target spelling.
Missing columns/unsupported clinical data Match extra_features_to_include exactly; encode nonnumeric categorical values before fitting.
No useful peptide features survive Inspect prefixes, prevalence filters, constant columns, and selector C; verify binary measurements and labels before widening a search.
A parameter-grid key fails Use random-forest/xgboost consistently and the exact pipeline path from the hyperparameter guide.
Validation did not run Set only_train_model: false and add nonempty validation_sets.
Plotting says a metric/split is unavailable Use a nested_ artifact with train, or a validation_ artifact with test, not a training_ model file.
Plotting after moving the folder fails Set the plot recipe's config to the current local model YAML. Keep input files unchanged.
Repeated SHAP alignment fails Check identical samples/features/experiment across matched files. Do not silently merge unrelated cohorts.
Tuned results are worse This can be legitimate. Check the search objective, data size, uncertainty, and external distribution; do not search for a favorable seed.
Exact metrics differ from the table Check source/dependency versions, generation seed/data hashes, model seed, and search budget.

Further implementation details are linked from the three reference documents. This tutorial's observed results are synthetic examples for learning the tool, not estimates of diagnostic performance on patient data.

Model configuration reference

This reference describes the implemented CLI at phipml 4.2.0, commit 932f14f. Start with a complete YAML in configs/models/; use this page when adapting it.

The three parts of a model YAML

Section Controls Example
Top level Files, sample IDs, target encoding, feature sources metadata_input, group_tests, transposed
classification: Estimator, cohorts, CV, output paths, decision threshold model_type, train_filters, outer_cv_splits
param_grid: Search spaces for pipeline parameters random-forest:, xgboost:

Use spaces for YAML indentation, native true/false, null for no value, and lists in brackets or with hyphens. Quoted "false" is a string, not a Boolean. Class and column names must match the input files, including capitalization.

Path rule: model YAML paths are relative to the YAML's own directory. classification.output_dir and input_dir CLI overrides follow this same rule. In a plotting YAML, paths are relative to that plotting YAML; explicit plotting CLI paths are relative to your terminal's current directory. Absolute paths work in either case.

Explicit model CLI values override the classification: section, which overrides internal defaults. classification.seed overrides random_state. Avoid adding unrecognized configuration keys: nested mappings can contain keys that are not consumed by the CLI.

Inputs and target: top-level keys

Key What to provide and what it means
data_input Required path to the peptide matrix, directory of enriched-peptide sample files, or sample-file manifest. In this tutorial: ../../../data/noisy/peptides.csv.
metadata_input Required sample metadata CSV, TSV/TXT, XLSX, or XLS. One row per sample. Must contain the configured sample and target columns.
data_input_mode matrix, sample-files, or auto. auto treats a directory as sample files and a file as a matrix; explicitly select sample-files for a manifest file.
project Optional descriptive identifier; output filenames are controlled separately by classification.output_name.
col_sample_name Sample-ID column in metadata, here SampleName. IDs must be unique and match the matrix. The combined matrix's first column is read as its index.
col_target Metadata column containing the target classes, here group_test. The classification CLI is binary.
group_tests Two target labels in encoding order: [Control, Case] gives Control=0 and Case=1. Predicted probabilities and positive SHAP direction refer to Case. Non-selected labels are excluded; check that the intended samples remain.
col_predict Compatibility column name used by some workflows; default class1_proba. This does not change the target or the model's decision threshold. Omitted in the examples.
transposed false for sample rows × peptide columns; true for peptide rows × sample columns. Sample-file mode constructs sample rows automatically.
peptide_prefixes Prefixes distinguishing peptides from clinical features. Here [agilent_, twist_, corona2_]. A missing trailing underscore is added. All peptide IDs must use your configured prefixes.
extra_features_to_include Metadata columns to append, here [Sex, Smoking, Age, BMI]; enabled by classification.with_additional_features. Do not include target, cohort, participant ID, or outcomes measured after the prediction time.
fillna_value Whole-matrix missing-value fill after loading. null preserves missingness. A number such as 0 fills both peptide and clinical missingness, so it is not a clinical imputation strategy.
random_state General seed; model/CV runs use classification.seed if supplied. This does not regenerate synthetic data; the generator has a separate --data-seed.
lib_metadata_input Peptide annotation file: CSV, TSV/TXT, Excel, or a pandas pickle. Required for library-based peptide filters; also useful for annotated plots. Not required for unfiltered numeric modeling.
lib_col_peptide_name Library column identifying peptides, here peptide_id. If null, table loading uses the first column; pickle input can already use a peptide-ID index.
filters_metadata Base sample filter, e.g. {timepoint: baseline}. Conditions across keys use AND, a list of allowed values within one key uses OR. Used as the default when no training filter is set. Run-specific training/validation filters replace this mapping when provided, so include required conditions in each cohort filter.
combined_filters_metadata List of condition mappings: OR between mappings, AND within each mapping. This stage still applies alongside run-specific cohort filters; keep it null for the tutorial.
oligo_filters Default peptide-library filter mapping, such as {Species: Homo sapiens}. Use annotation columns, not sample metadata. The run can supply its own filters.
oligo_filter_mode all means AND across library conditions; any means OR. Defaults to all.

The loader intersects sample IDs between metadata and the matrix. A typo can therefore remove samples without meaning that the cohort definition was right. Check sample counts before fitting and the Final training data log message.

Alternative input: one enrichment file per sample

The main examples use combined matrices. For per-sample enriched-peptide lists, replace only the data-loading part of your YAML:

data_input: ../../../my_data/enriched_samples
data_input_mode: sample-files
sample_file_patterns: ["*.csv"]
sample_file_peptide_column: ID
sample_name_regex: null
Key Meaning
sample_file_patterns Glob patterns within the directory; recursive patterns such as batch_*/**/*.csv are supported.
sample_file_peptide_column Column containing enriched peptide IDs in every sample file, e.g. ID. These are presence lists, not raw sequencing-count matrices.
sample_name_regex Optional regex to extract a sample ID from the filename; preferably a named group such as 'prefix_(?P<sample>.+)_enriched'. null uses the filename stem.

A tab-delimited manifest can list sample_name and path columns, or one path per line. Paths in a sample manifest are relative to that manifest. An explicit sample name takes precedence over the filename-derived name.

Execution: classification: keys

Model and run type

Key Meaning and tutorial value
model_type Exactly random-forest or xgboost; not random_forest, rf, or XGB.
param_grid_name Which top-level param_grid subsection to read. Keep it equal to model_type in these examples. Switching models in a tuned run requires switching this too.
seed Estimator/CV/split/bootstrap seed and filename suffix; 420 here. One seed creates one CV run; multiple repeats require multiple commands.
run_nested_cv Evaluate the training cohort with outer stratified folds. Each outer model tunes on inner folds if a search space exists; with param_grid: {}, this is ordinary outer CV without tuning.
only_train_model Fit/save the full-cohort model and stop before external evaluation. If nested CV is enabled, CV runs first. Internal examples set true; external examples set false.
use_pretrained Load a fitted pipeline for full-model external evaluation. It does not stop an independently enabled nested-CV run from fitting fold models. Use run_nested_cv: false when only reusing a model.
input_dir Directory containing a pretrained artifact, used with use_pretrained: true.
input_name Direct .joblib filename or base name used to find standard saved-model names. An explicit filename is clearest.
output_dir Destination for joblib artifacts; created automatically. Give each experiment a separate directory.
output_name Base name for nested_... and training_... artifacts. External artifacts instead use each validation set's name.

Features and preprocessing

Key Meaning and practical effect
with_oligos Include peptide features. true here. At least this or with_additional_features must be true.
with_additional_features Include columns in extra_features_to_include. true here. Adding Sex/Age makes them predictors; it does not by itself establish a causal adjustment.
subgroup all retains the available library; another column name is a compatibility shortcut for selecting rows where that annotation is true. Use explicit oligo_filters for new analyses.
oligo_filters Run-specific library filters. Example: {Species: Homo sapiens, is_PNP: true}. Use this with subgroup: all. Do not combine a named subgroup with non-null filters.
oligo_filter_mode all/any across run-specific annotation conditions.
prevalence_threshold_min Minimum percentage of training samples with a nonzero, nonmissing peptide value. Inclusive; tutorial 0. 1 means 1%, not a proportion of 1.
prevalence_threshold_max Maximum percentage with a nonzero, nonmissing peptide value. Inclusive; tutorial 100.
impute_extra_numeric Fit a SimpleImputer inside each pipeline for continuous numeric clinical features. Tutorial data are complete, so false. It does not impute peptides or numeric binary extras.
extra_numeric_impute_strategy mean, median, most_frequent, or constant; median here. Active only when imputation is enabled. The CLI has no dedicated constant-fill-value option.
fill_missing_peptides_with_zero Aligning external data: add absent peptide columns as zero when true, otherwise fail. We set false because all expected features are measured. Missing clinical columns always fail. This option handles absent columns, not arbitrary NaNs inside present columns.

The peptide pipeline removes constant columns, then fits an elastic-net logistic-regression selector. Clinical variables bypass that peptide selector. Feature selection and optional continuous imputation are fitted inside CV pipelines. The current CLI applies prevalence filtering to the full training cohort before outer CV, although it excludes external samples. The tutorial uses 0–100% to make this filter inactive. Do not describe a restricted prevalence filter as being fitted separately inside every CV fold in this version.

Sex/Gender values F/Female/M/Male and 0/1 are recognized. Other categorical clinical variables require numeric encoding before fitting. Define encodings using training data and apply the same schema externally. Genuine missing measurements should not automatically become biological absences.

CV, tuning, and uncertainty

Key Meaning and tutorial value
outer_cv_splits Number of outer evaluation folds; 3 for the tutorial. Each training sample receives one held-out prediction per run.
inner_cv_splits Number of folds used to rank hyperparameter candidates; 2 here. The final full-cohort model uses an additional inner search on the entire training cohort. Ignored for fitting when there is no search.
n_iter Number of candidate configurations per Bayesian search, not number of CV repeats or trees. 12 in tuned recipes. It has no tuning effect when param_grid: {}.
n_jobs_outer Concurrent outer-fold jobs; 1.
n_jobs_inner Concurrent search/CV jobs; 1. Use a bounded positive number or -1 for available CPUs. Estimators themselves are created with n_jobs=1; avoid multiplying outer and inner parallelism on small machines.
classification_threshold Preselected Case-probability cutoff for confusion matrix, accuracy, sensitivity, specificity, F1, etc.; 0.5. ROC-AUC and AP do not depend on this cutoff.
bootstrap_validation Bootstrap the external predictions for uncertainty; true. This resamples observations with a fixed fitted model, without refitting or retuning.
bootstrap_n_resamples External bootstrap draws; 500 for a short demonstration, implementation default 1000. More draws reduce Monte Carlo noise, not the uncertainty caused by limited sample size.
bootstrap_confidence_level External interval level, as a fraction; 0.95 means 95%. It does not turn CV fold SD into a confidence interval.

The search score is average precision in this commit. It is not a YAML option. XGBoost's built-in eval_metric="auc" does not change the search scorer. Class stratification preserves class proportions; it does not group repeated samples from one participant. This CLI has no participant-group CV setting. The mock data represent independent samples.

For a larger exploratory run, a reasonable starting budget is 5 outer folds, 3 inner folds, and 30 candidates, then repeated seeds if compute permits. Choose the split design and sample sizes for the scientific study; more repeats do not create more independent participants. Never select a seed or tuning recipe by its external-cohort result.

Cohorts and optional hold-out splitting

Key Meaning
train_filters Metadata conditions selecting training samples, here {cohort: training}. Omitting this from a combined training/external dataset can include external samples in fitting.
validation_sets A list of {name: ..., filters: {...}} entries, all read from the configured inputs. Names control validation filenames. Example below.
split_filters Optional cohort to partition into train and hold-out portions; null in the six core examples.
train_size Fraction of the split cohort assigned to training; default 0.7. Active only with split_filters.
split_only true trains/evaluates only the split cohort. With false, the split's training portion is appended to train_filters, and its held-out portion is appended to each validation set. Use disjoint selections.
classification:
  train_filters: {cohort: training}
  run_nested_cv: false
  only_train_model: false
  validation_sets:
    - name: independent_site
      filters: {cohort: external}

Keep participants and any repeated samples disjoint between cohorts. Labeling rows external does not make them independent if they are copied from training.

What this version supports for external files

At the inspected commit, validation_sets accepts name and filters, not a separate per-cohort data_input or metadata_input. The six examples therefore use a combined matrix and metadata with disjoint training/external IDs. The training filter excludes external rows from fitting and tuning.

If a new cohort arrives in separate files, combine compatible tables first (with unique sample IDs, a cohort column, and checked feature/metadata schemas) and then use the filters above. Or adapt the Python API explicitly. The current pretrained CLI still prepares its training input context, so these reuse recipes retain the training rows; an external-only CLI workflow is not documented here as if it were already implemented.

Useful command overrides

Run these from the tutorial directory:

# Same noisy dataset, a different CV seed, dedicated output folder.
phipml -c configs/models/random-forest/02_noisy.yaml --seed 421 \
  --output-dir ../../../results/random-forest/repeat_421

# A longer search. Output is kept separate from the reference examples.
phipml -c configs/models/xgboost/03_noisy_tuned.yaml \
  --n-iter 30 --inner-cv-splits 3 --outer-cv-splits 5 \
  --output-dir ../../../results/xgboost/noisy_longer_search

# Untuned clinical-only baseline, on the same target/cohort/CV seed.
phipml -c configs/models/random-forest/02_noisy.yaml \
  --no-with-oligos --with-additional-features \
  --output-name clinical_only --output-dir ../../../results/random-forest/clinical_only

# Untuned peptides-only comparison.
phipml -c configs/models/random-forest/02_noisy.yaml \
  --with-oligos --no-with-additional-features \
  --output-name peptides_only --output-dir ../../../results/random-forest/peptides_only

For a tuned estimator change, use the supplied estimator-specific YAML. Merely passing --model-type xgboost to a Random Forest tuning YAML leaves its selected grid pointing at Random Forest parameters unless it is also replaced.

Source: the pinned repository's data/config loader, run settings, and training CLI.

What the hyperparameters mean

phipml fits a pipeline: peptide preprocessing/selection, optional clinical imputation, then a Random Forest or XGBoost classifier. A hyperparameter is a setting chosen before fitting; training then estimates model coefficients or trees from the training samples. The tuned examples choose hyperparameters by inner cross-validation, using average precision (AP).

These are small teaching search spaces, not claimed optimal settings for PhIP-seq. The same noisy data, target, cohort filters, and seed are used for untuned/tuned pairs. Better training or inner-CV performance does not guarantee better outer-CV or external performance.

Read a parameter path

param_grid:
  random-forest:
    estimator__max_depth:
      type: integer
      low: 2
      high: 10

classification.param_grid_name: random-forest selects this section. estimator__max_depth means “the max_depth parameter of the pipeline step called estimator.” Double underscores traverse nested pipeline objects.

The longer path preprocessor__peptides__feature_selection__estimator__C changes the logistic regression used to select peptides, not the final Random Forest/XGBoost.

Search specification Meaning Valid example
type: integer Whole-number choices between inclusive bounds {type: integer, low: 2, high: 10}
type: real Continuous values between bounds {type: real, low: 0.2, high: 0.8}
type: categorical Choose one of the listed alternatives {type: categorical, categories: [sqrt, log2, 0.5]}
prior: uniform Equal weight to equal-width intervals; default for reals Useful for a fraction such as subsample.
prior: log-uniform Equal weight across multiplicative scales Useful for C, learning rate, or regularization; both bounds must be strictly positive.

Do not use equal low/high to hold a value fixed. Use a one-element categorical space, e.g. {type: categorical, categories: [100]}. This still invokes the search machinery if param_grid is nonempty. The current CLI does not expose a separate estimator_params block for fixed overrides. param_grid: {} uses the built-in pipeline and estimator defaults; it does not disable feature selection.

Shared peptide selector

Both models fit elastic-net logistic regression with solver: saga, then retain peptides whose absolute coefficient meets the selection threshold.

Full parameter key Tutorial search Meaning Effect of increasing it
preprocessor__peptides__feature_selection__estimator__l1_ratio 0.2–0.8, uniform Mix of L1 and L2 penalties: 0=pure L2, 1=pure L1. Fixed default 0.5. Generally promotes sparser selection as the L1 share increases, though correlated predictors can change which peptide survives.
preprocessor__peptides__feature_selection__estimator__C 0.1–10, log-uniform Inverse regularization strength. Fixed default 1.0. Weaker shrinkage, often more selected peptides; a small value may eliminate nearly all peptide signal.

Fixed implementation settings are VarianceThreshold(threshold=0.0), selector threshold=1e-5, and logistic-regression max_iter=10000. The selector is supervised and learns from class labels inside the fitted CV pipeline. Clinical features do not pass through this peptide selector. Inactive peptide search parameters are ignored when no peptide branch exists.

Feature selection can choose different peptides in different folds. The saved SHAP matrices retain original feature names and assign zero to unused features for the relevant fitted model/fold. A zero SHAP value is not proof of no biological association.

Random Forest

A Random Forest averages many randomized decision trees. Each tree can capture nonlinear patterns; averaging reduces the sensitivity to any one fitted tree.

Use the exact final-model prefix estimator__ for every parameter below.

Parameter key Tutorial search Meaning Practical tradeoff
estimator__n_estimators Integer 60–160 Number of trees. Untuned sklearn default: 100. More trees usually stabilize predictions, but cost time/memory; tree count alone does not control overfitting as strongly as tree size.
estimator__max_depth Integer 2–10 Maximum depth of each tree. Untuned default: None (no explicit depth limit). Deeper trees fit finer interactions and noise; shallow trees may underfit.
estimator__min_samples_split Integer 2–10 Minimum observations required in a node before attempting a split. Untuned default: 2. Larger values suppress splits on small groups. This is different from the size of each resulting leaf.
estimator__min_samples_leaf Integer 1–6 Minimum observations in each terminal leaf. Untuned default: 1. Larger leaves smooth predicted probabilities and reduce small-sample fitting.
estimator__max_features sqrt, log2, or 0.5 Number/fraction of available features considered at a split, after preprocessing. Untuned default: sqrt. Smaller subsets decorrelate trees, but can hide useful features at a split. 0.5 is numeric: half of available features, not the string "0.5".
estimator__class_weight null or balanced Weight the contribution of classes to fitting. Untuned default: None. balanced weights inversely to class frequency. It can change decision behavior; it does not rebalance the evaluation cohort.

Other supported Random Forest parameters can be added through a valid pipeline search path. Common optional choices:

Parameter How to specify Meaning
estimator__bootstrap {type: categorical, categories: [true, false]} Whether each tree uses a bootstrap sample. Default true.
estimator__criterion {type: categorical, categories: [gini, entropy]} Criterion for scoring candidate splits; default gini.
estimator__class_weight Add balanced_subsample to categories Recalculate inverse-frequency weights separately for each tree's bootstrap sample.

The balanced mock classes are included to keep the demonstration easy to read; class weighting is present to teach its meaning, not because imbalance needs correction in these data.

XGBoost

XGBoost builds trees sequentially. Each boosting round adds a correction to the current model. Smaller learning steps often need more trees. The phipml factory sets objective="binary:logistic", tree_method="hist", eval_metric="auc", n_jobs=1, and the configured seed. Other untuned settings come from the installed XGBoost version. There is no early-stopping evaluation set supplied by this CLI.

Parameter key Tutorial search Meaning Practical tradeoff
estimator__n_estimators Integer 50–160 Maximum number of boosting rounds/trees for this tree booster. More rounds add capacity; combine with learning_rate. Without early stopping, the configured rounds are fitted.
estimator__learning_rate 0.02–0.25, log-uniform Shrinkage applied to each added tree; also called eta. Smaller steps often need more rounds and can improve generalization; too few rounds then underfit.
estimator__max_depth Integer 2–5 Maximum depth of a tree. Larger values fit more complex interactions and require more data/memory.
estimator__min_child_weight Integer 1–8 Minimum sum of second-order loss weights (Hessians) required in a child node. Larger values discourage splits. This is not a literal minimum sample count for logistic classification.
estimator__subsample 0.65–1.0, uniform Fraction of training observations used in each boosting iteration. Values below 1 introduce row subsampling; too little data per round can weaken learning.
estimator__colsample_bytree 0.65–1.0, uniform Fraction of available features sampled for each tree. Can reduce correlated tree behavior and fitting to noise; may miss useful predictors.
estimator__reg_lambda 0.1–10, log-uniform L2 regularization on leaf weights. Larger values shrink leaf contributions and make the model more conservative.
estimator__reg_alpha 0.0001–1.0, log-uniform L1 regularization on leaf weights. Larger values can set some leaf contributions to zero. The positive lower bound makes the log-uniform prior valid.
estimator__gamma 0–2, uniform Minimum reduction in objective required to accept a split. Larger values discourage low-benefit splits.

Optional estimator__scale_pos_weight increases the training weight assigned to the positive class. If exploring it, use training-fold information and validate its effect. It is omitted here because the synthetic classes are balanced. Do not add class_weight to an XGBoost grid or learning_rate to a Random Forest grid.

How tuning is evaluated

For 03_noisy_tuned:

  1. Split the training cohort into three outer folds.
  2. For each outer fold, run 12 candidate configurations using two inner folds on the outer-training subset; refit the best candidate on that subset.
  3. Predict and explain the held-out outer subset exactly once.
  4. Combine the outer-fold results into the nested artifact.
  5. Run a fresh inner search on all 160 training samples and save the resulting full-cohort pipeline for reuse.

This costs approximately 3 × (12 × 2 + 1) + (12 × 2 + 1) = 100 pipeline fits. It is a teaching budget for a fairly wide space. BayesSearchCV in the pinned scikit-optimize version starts with 10 initialization candidates by default; 12 candidates allow a short adaptive phase. A two-candidate smoke run exercises the search workflow but cannot demonstrate a mature Bayesian optimization.

For 06_external_noisy_tuned, there is one inner search on the full training cohort, followed by evaluation of its fixed selected model on 80 external samples. External labels do not choose parameters, features, or thresholds.

Inspect what was actually chosen:

python scripts/summarize_results.py

Open reports/fitted_parameters.json for full-cohort pipeline settings, or:

from pathlib import Path
import joblib

path = Path("results/random-forest/03_noisy_tuned/training_random-forest_03_noisy_tuned_420.joblib")
model = joblib.load(path)["best_estimator"]
print(model.named_steps["estimator"].get_params())
selector = model.named_steps["preprocessor"].named_transformers_["peptides"].named_steps["feature_selection"]
print(selector.estimator_.get_params())

Outer-fold selected settings can differ from this full-cohort model; they are available from model_list in the nested artifact. The CLI saves the chosen models, not the complete BayesSearchCV.cv_results_ optimization history.

Implementation source: pipeline, selector and search. Parameter semantics were checked against the installed scikit-learn 1.5.2 and XGBoost 2.1.4 classes used for the reference run. Further reference: RandomForestClassifier, LogisticRegression, XGBoost parameters.

Plot configuration reference

Use phipml-plot --plot-config PATH.yaml for a saved recipe. Model fitting and plotting are separate commands: change colors or displayed annotations without fitting models again. Use phipml-heatmap with a CSV/TSV manifest to compare experiments. This guide describes the current public commands, not the older plotting modules retained in the repository.

Start with the right artifact

Result filename begins with Content Plot split
nested_ Outer-fold held-out predictions, SHAP, metrics, fold models train
validation_ Predictions/SHAP/metrics for an external or explicit hold-out cohort test
training_ Full-cohort fitted pipeline and input context Not an evaluation artifact; use it for model reuse, not performance plotting.

split: train is a storage key for the training cohort's out-of-fold evaluation. It does not mean the model is evaluated on samples it just fitted. For the internal examples the file contains one held-out prediction for each of the 160 training-cohort samples.

The files embed targets, feature names, SHAP values, metrics, and data-loading context. They do not contain a portable copy of every original input file. Use plotting.config to point to the local model YAML after moving the folder; the supplied recipes already do this. The original data are still required for colored feature-value SHAP displays and prevalence/clinical summaries.

A complete plotting recipe

This example is stored at configs/plots/random-forest/02_noisy.yaml.

plotting:
  results:
    - ../../../results/random-forest/02_noisy/nested_random-forest_02_noisy_420.joblib
  config: ../../models/random-forest/02_noisy.yaml
  split: train
  plots: [all]
  formats: [pdf, svg, png]
  dpi: 180
  output_dir: ../../../plots/random-forest/02_noisy
  output_prefix: 02_noisy
  title: "random-forest | 02 noisy"
  class_labels: [Control, Case]

  max_display: 10
  feature_ranking: mean-abs-shap
  ranking_top_k: 10
  min_top_k_frequency: 0
  shap_alignment: strict
  save_standalone: true
  reconstruct_data: true

  library_metadata: ../../../data/peptide_library.csv
  library_id_column: peptide_id
  validation_bootstraps: 0
  random_state: 420
  class_colors: ["#386CB0", "#D35D4D"]
  colors:
    roc: "#176B73"
    roc_band: "#B3DAD8"
    pr: "#8A5A44"
    pr_band: "#DFC5B8"
    confusion_cmap: Blues
    classification: "#52796F"
    shap_cmap: phipml_blue_gray_red
    shap_heatmap_cmap: phipml_purple_gray_orange
    shap_importance: "#6D597A"
  feature_table:
    annotation_columns: [Description, Role]
    extra_columns: []
    title: "Synthetic features: model importance and known roles"
    header_color: "#DCE9E9"
    row_colors: ["#F5F8F8", white]
    prevalence_cmap: phipml_prevalence
phipml-plot --plot-config configs/plots/random-forest/02_noisy.yaml

The figure files and feature table appear under plots/random-forest/02_noisy/. Relative YAML paths are resolved from the plotting YAML. Explicit CLI paths are resolved from the current terminal directory, and override YAML values.

Inputs, output, and figure selection

plotting key What to provide and what it controls
results List of evaluation joblib paths or quoted glob patterns. Multiple files are aggregated as repeated runs; they are not automatically drawn as distinct models in a comparison.
config Model YAML for the original input files/labels/annotations. Optional if embedded input paths still resolve; supplied here to make the examples portable. This is different from the command's --plot-config.
split train, test, or auto. auto infers an available evaluation split; explicit values are clearer.
plots List drawn from all, performance, roc, pr, confusion, classification, shap-beeswarm, shap-importance, shap-heatmap, feature-table.
formats One or more of pdf, svg, png. Vector SVG/PDF are useful for editing/sharing; PNG is convenient for slides and previews.
dpi Raster resolution; 180 for compact tutorial files. Increase to 300 or 600 if required. This does not change model results or the geometry of vector elements.
output_dir Plot destination. If omitted, the CLI considers an explicitly supplied model config's classification.plot_output_dir, then its classification.output_dir/plots; otherwise it uses plots/ next to the result. Explicit is easiest.
output_prefix Shared output filename stem, e.g. 02_noisy produces 02_noisy_performance.svg. Default phipml.
title Title for the combined performance summary. Some standalone panels have their own labels; do not expect this to rename every panel.
class_labels Display labels in [negative, positive] order. Defaults to model config labels if available. Relabeling does not change encoded targets.
save_standalone true saves requested individual ROC/PR/confusion/classification figures. false removes those standalone panel requests; include performance to retain the combined panel.
reconstruct_data true loads feature values from explicit/embedded model config. Set false for performance-only plotting without the original data.
validation_bootstraps Reconstruct external bootstrap intervals only if missing, for a single test artifact with targets available. 0 here because fitting already saved 500-resample intervals. This does not replace existing intervals.
random_state Seed for bootstrap reconstruction when used. It does not change predictions or retrain anything.

What each plot answers

plots choice Interpretation Typical filename suffix
performance Combined ROC, precision–recall, confusion matrix, and threshold-dependent metric panels. _performance
roc Sensitivity versus false-positive rate across score cutoffs. The legend reports ROC-AUC. _roc
pr Precision versus recall. The legend reports AP, not trapezoidal PR-AUC. Baseline precision depends on positive-class prevalence. _pr
confusion Counts of true/false positive/negative predictions at the saved threshold. _confusion
classification Metrics such as accuracy, sensitivity, specificity, F1 and MCC at that threshold. _classification
shap-beeswarm Distribution of signed contributions per feature; each point represents a sample. Color shows feature value. _shap_beeswarm
shap-importance Mean absolute contribution; magnitude without direction. _shap_importance
shap-heatmap Patterns of signed contributions across samples and selected features. _shap_heatmap
feature-table Compact importance table with class-wise prevalence/clinical summaries and annotations; also exports a detailed CSV. _feature_table, _feature_importance.csv

All nine figure types are generated by plots: [all]. The table CSV retains more features/annotations/audit fields than the compact rendered table.

Feature ranking and repeated runs

Key Meaning
max_display Maximum number of ranked features shown in SHAP figures and compact table; 10. This is a display limit, not model feature selection.
feature_ranking mean-abs-shap ranks mean absolute SHAP; top-k-frequency ranks how often a feature appears in each run's top K. auto chooses mean absolute SHAP for one result and top-K frequency for multiple results.
ranking_top_k Per-run top-K cutoff used for frequency ranking; 10. Separate from max_display.
min_top_k_frequency Minimum percentage of runs whose top K includes a feature; 50 means 50%, not 0.5. Set 0 for the single-file examples. With very few repeats this is only a demonstration of stability.
shap_alignment strict requires identical sample and feature sets across repeated files, reordered by ID as needed. intersection deliberately restricts to shared samples/features and changes the interpreted population.

Only aggregate runs with the same estimator, data, cohort, feature definition, class encoding, and threshold. Do not use a broad glob that mixes Random Forest, XGBoost, noisy/perfect data, training/external cohorts, or tuned/default experiments. Use separate plots or a heatmap manifest for those comparisons.

The repeated plot recipe uses a dedicated repeated_noisy folder, and three seeds on the same dataset. Signed SHAP values are averaged by sample/feature; ranking/stability statistics also retain run-level absolute importance. Top-K frequency differs from how often the elastic-net selector retained a feature. Both statistics can be inspected in the exported CSV.

Annotations and appearance

Key Meaning
library_metadata Optional annotation CSV, TSV/TXT, or Excel table used by the plotting CLI. Unlike the model loader, this explicit table reader does not accept pickle.
library_id_column Annotation column matching feature IDs; peptide_id. Set explicitly to ensure the join uses the right index.
class_colors Two colors in [Control, Case] order for prediction-direction labels. These do not encode biological class in the SHAP feature-value color scale.
colors.roc, colors.roc_band ROC curve and uncertainty shading.
colors.pr, colors.pr_band Precision–recall curve and shading.
colors.confusion_cmap Matplotlib colormap for confusion matrix counts.
colors.classification Threshold-metric bar color.
colors.shap_cmap Feature-value colormap for the beeswarm.
colors.shap_heatmap_cmap Diverging colormap for signed SHAP contributions.
colors.shap_importance Absolute-SHAP bar color.
feature_table.annotation_columns Annotation fields shown in the compact figure, e.g. [Description, Species, Protein]. Tutorial uses [Description, Role] to distinguish known synthetic signal/noise.
feature_table.extra_columns Optional audit columns to display, e.g. ["Top-k SHAP frequency (%)", "Selection frequency (%)"]. They remain in the CSV when omitted from the figure.
feature_table.title Compact-table title.
feature_table.header_color Header fill color.
feature_table.row_colors Alternating row colors, a two-item list.
feature_table.prevalence_cmap Colormap for prevalence cells.
feature_table.input Optional curated feature-importance CSV/TSV/Excel. The edited annotations and row order control feature-table rendering. Use an exported table as the starting point.

Quote hex colors because unquoted # begins a YAML comment. Use valid Matplotlib colormaps or the supplied phipml_blue_gray_red, phipml_purple_gray_orange, and phipml_prevalence maps. Annotation values in these examples are invented and have no biological meaning.

Supplying plotting data explicitly

These optional keys are useful for older artifacts or another controlled data source. They are not needed for the included recipes.

Key What to supply
features_table Sample-by-feature CSV/TSV/Excel, including sample_column. Clinical values must be numeric and use the same encodings as fitting. Include the original feature values, not SHAP values.
target_table Table containing sample IDs and the outcome column. Use already encoded 0/1 targets for the explicit table route.
sample_column Identifier column for these tables; default SampleName.
target_column Outcome column in target_table, or in features_table if no separate target table is supplied. Required with an explicit target table.
feature_importance_table Top-level alternative to feature_table.input. Use the exported CSV structure, including its Feature column.

Saved sample IDs determine alignment. The supplied model-config reconstruction route also checks reconstructed labels against saved targets. This check cannot prove all feature measurements are unchanged; preserve input data and hashes.

Common plotting commands

From the tutorial root:

# Change appearance/format without retraining.
phipml-plot --plot-config configs/plots/random-forest/02_noisy.yaml \
  --formats svg png --dpi 300 --max-display 8 \
  --output-dir plots/custom_noisy --output-prefix presentation

# Only performance panels; original feature data are not required.
phipml-plot results/random-forest/02_noisy/nested_random-forest_02_noisy_420.joblib \
  --split train --plots performance --no-reconstruct-data \
  --class-labels Control Case --output-dir plots/performance_only

# External model evaluation.
phipml-plot --plot-config configs/plots/xgboost/06_external_noisy_tuned.yaml

# Regenerate a compact table after manually editing a COPY of its exported CSV.
phipml-plot --plot-config configs/plots/random-forest/02_noisy.yaml \
  --plots feature-table \
  --feature-importance-table plots/random-forest/02_noisy/02_noisy_feature_importance.csv \
  --output-dir plots/curated_table --max-display 8

Editing a feature table changes presentation; it does not revise SHAP values or model predictions. Display a curated subset transparently.

Comparison heatmaps

phipml-heatmap takes a manifest, not a plotting YAML. Each row identifies an evaluation artifact:

training,validation,path,split
RF noisy default,Internal CV,../results/random-forest/02_noisy/nested_random-forest_02_noisy_420.joblib,train
RF noisy default,External,../results/random-forest/05_external_noisy/validation_random-forest_05_external_noisy_420.joblib,test

The path is relative to the manifest file. The training value labels a heatmap column and validation labels a row; labels can describe experiment conditions as in the supplied manifests. Multiple rows for one cell are aggregated as repeated run estimates. Use explicit files, not globs, in the manifest.

phipml-heatmap --manifest manifests/random-forest.csv \
  --metric roc.auc --palette viridis --vmin 0 --vmax 1 --dpi 180 \
  --title "Random Forest: internal and external performance" \
  --output plots/random-forest/roc_auc_overview

phipml-heatmap --manifest manifests/xgboost.csv \
  --metric pr.ap --palette viridis --vmin 0 --vmax 1 --dpi 180 \
  --output plots/xgboost/ap_overview
Heatmap argument Meaning
--manifest Required CSV/TSV/TXT with training, validation, path, and optional split.
--metric Dotted metric key, e.g. roc.auc, pr.ap, classification.f1, classification.balanced_accuracy, or classification.mcc.
--training-order Optional ordered list of column labels; must include all manifest training labels.
--validation-order Optional ordered list of row labels.
--order Legacy shared order for a square matrix; do not combine with the two independent orders. Usually unnecessary.
--title Figure title.
--palette Matplotlib/seaborn colormap name; viridis here.
--vmin, --vmax Color-scale bounds. Use 0–1 for this overview, −1–1 for MCC. The CLI defaults are 0.5–1.0, which can obscure below-0.5 values.
--output Required filename stem; requested format extensions are added/replaced.
--formats One or more of pdf svg png; defaults to all three.
--dpi Raster resolution.
--annotate-uncertainty / --no-annotate-uncertainty Show/hide available native intervals or variability. Default shows them.

The rectangular overview uses only conditions that exist. An internal-CV cell and an external cell summarize different evaluation procedures; their numbers are not a paired statistical test.

Interpreting uncertainty and SHAP honestly

Evaluation Displayed uncertainty What it describes
One nested-CV artifact Across-outer-fold variability; curve bands are mean ±1 sample SD. Sensitivity to the fold partition. It is not a formal 95% confidence interval.
Multiple seed artifacts 2.5–97.5% empirical interval across run estimates. Variation across repeated fits on reused participants; not independent-cohort confidence.
One external artifact with bootstrap Class-stratified paired bootstrap interval at the requested confidence level. Sampling variability of the external observations conditional on the fitted model; it excludes refitting/tuning uncertainty.

The perfect synthetic scenario can have a 1.0–1.0 bootstrap interval because every resampled set remains perfectly separated. This is a property of the constructed demonstration, not evidence that a clinical model has no uncertainty.

Positive SHAP contributions move the model output toward the positive class; negative values move it away. Random Forest tree SHAP is in its probability output scale here, while XGBoost's default tree SHAP explains the raw margin (log-odds for the binary logistic objective). Do not compare their numerical SHAP magnitudes directly or describe XGBoost values as probability-point changes. SHAP describes a fitted predictive relationship, not causation or a p-value. Correlated engineered peptides may share importance unevenly, and independent noise can acquire apparent importance in a finite sample.

Source: the pinned repository's plot CLI, result aggregation, and heatmap CLI.

Reference execution and observed results

Run date: 8 September 2026. Source: phipml 4.2.0, commit 932f14ffca41ba4f24f3cc570afea27265a5893c. The upstream source checkout was not modified.

Environment and budget

The run installed the pinned source into an isolated Linux Python 3.12.13 environment with the dependency ranges declared in pyproject.toml. Main versions: numpy 1.26.4, pandas 2.2.3, scipy 1.15.3, scikit-learn 1.5.2, scikit-optimize 0.10.2, XGBoost 2.1.4, SHAP 0.46.0, matplotlib 3.10.9, and seaborn 0.12.2. The full platform-specific inventory is in reports/requirements-reference.txt; it is distinct from the repository's Micromamba specification. The Micromamba, Docker, and Apptainer installation routes were inspected in upstream documentation, not separately executed here.

Data seed 20260908; model seed 420. Three outer folds, two inner folds, 12 candidates in each tuned search, one worker at each level, 500 external bootstrap draws, confidence level 0.95, classification threshold 0.5.

The complete twelve-scenario command, including per-scenario plots, ran in approximately five minutes on this environment, excluding dependency installation. This is an observation, not a runtime promise. The model-event log spans 4 minutes 36 seconds, with the final plot/heatmap/export work afterward. No outcome-dependent revisions were made to the generated data or tuning space.

What was run

python scripts/generate_examples.py
bash scripts/run_all.sh both
bash scripts/run_repeated.sh random-forest
phipml -c configs/models/random-forest/08_reuse_noisy.yaml
phipml -c configs/models/xgboost/08_reuse_noisy.yaml
python scripts/summarize_results.py

The main run covers 12 estimator/scenario combinations, each producing nine figure types in PDF/SVG/PNG plus the feature table CSV. Both estimator heatmaps were generated for ROC-AUC and AP. The optional repeated-run example was executed for Random Forest; the parallel XGBoost repeat recipe is supplied but was not separately executed. Pretrained model reuse was executed for both.

Metrics

Internal entries show mean ± sample SD across outer folds. External entries show the point estimate followed by the 95% stratified bootstrap interval. These are different uncertainty summaries and should not be read as equivalent 95% intervals. All participants here are synthetic.

Estimator Scenario N evaluated ROC-AUC Average precision
random-forest 01_perfect 160 1.000 ± 0.000 1.000 ± 0.000
random-forest 02_noisy 160 0.912 ± 0.028 0.908 ± 0.030
random-forest 03_noisy_tuned 160 0.911 ± 0.038 0.915 ± 0.044
random-forest 04_external_perfect 80 1.000 [1.000, 1.000] 1.000 [1.000, 1.000]
random-forest 05_external_noisy 80 0.810 [0.709, 0.898] 0.843 [0.769, 0.914]
random-forest 06_external_noisy_tuned 80 0.806 [0.705, 0.894] 0.805 [0.708, 0.900]
xgboost 01_perfect 160 1.000 ± 0.000 1.000 ± 0.000
xgboost 02_noisy 160 0.917 ± 0.038 0.921 ± 0.045
xgboost 03_noisy_tuned 160 0.926 ± 0.030 0.928 ± 0.037
xgboost 04_external_perfect 80 1.000 [1.000, 1.000] 1.000 [1.000, 1.000]
xgboost 05_external_noisy 80 0.827 [0.728, 0.920] 0.825 [0.730, 0.931]
xgboost 06_external_noisy_tuned 80 0.808 [0.706, 0.906] 0.804 [0.707, 0.919]

Tuning did not improve external AP for either estimator in this run. The external population has a deliberately weaker signal. These results illustrate why external data must stay out of model selection; they are not evidence that tuning is generally unhelpful or that one estimator is superior for real PhIP-seq data.

The CSV additionally records pooled-score ROC-AUC/AP, accuracy, and F1. For internal evaluation, pooled-score metrics are not necessarily equal to the mean of the fold metrics. Prediction CSVs use the pooled sample-level scores.

Verification performed

Detailed checks are recorded in reports/verification_checks.json. The normal training/plotting logs are in logs/. The model/config generator, inputs, search spaces, artifacts, plot recipes, exported predictions, and observed metrics are included so that the walkthrough can be independently rerun.

Scope of interpretation

The perfect data intentionally encode their target in engineered predictors. All annotations are simulated. The noisy scenario is a controlled learning exercise, not a biological benchmark. The current phipml prevalence filter is applied before outer CV on the training cohort; it is inactive here at 0–100%. The current CLI uses sample-level stratified folds, so participant-group cross-validation would require additional support for repeated-measure studies.

These notes describe observed execution and current limitations; optional longer-search, clinical-only, and per-sample-file recipes are explanatory examples rather than additional completed benchmark runs.