From peptide data
to model interpretation.
Install, configure, fit, validate, and explain. A complete command-line workflow with real synthetic-data outputs and a field-by-field reference.
phipml: from peptide data to model interpretation
A runnable tutorial for Random Forest, XGBoost, nested cross-validation, external validation, and publication-ready plots.
Built against phipml 4.2.0, commit
932f14f,
inspected on 8 September 2026. The commands below use the implemented
configuration schema and were run on the supplied synthetic data.
You will start with peptide and sample metadata CSVs, fit either estimator, compare fixed/default and tuned models, evaluate an independent synthetic cohort, and produce ROC, precision–recall, classification, SHAP, and annotated feature plots. Finished results are included so you can inspect them before rerunning anything.
Start here: follow sections 1–6 for your first model and plots. Then run the six scenarios, compare the results, and use the references when adapting your own configuration. For a workshop, start with Random Forest and add XGBoost after the workflow is familiar.
What is included
| Item | Location |
|---|---|
| This walkthrough | README.md |
| Every model-config argument | docs/MODEL_CONFIG.md |
| Shared selector, RF and XGBoost hyperparameters | docs/HYPERPARAMETERS.md |
| Every plot-config argument and heatmap command | docs/PLOTTING_CONFIG.md |
| Observed results and verification notes | reports/REFERENCE_RUN.md |
| Deterministic mock-data/config generator | scripts/generate_examples.py |
| Twelve ready-to-run model recipes | configs/models/{random-forest,xgboost}/01...06*.yaml |
| Corresponding plot recipes | configs/plots/{random-forest,xgboost}/ |
| CSV data, metadata, annotations, data hashes | data/ |
| Fitted models, predictions, and metrics | results/, reports/ |
| Actual plots, in PDF/SVG/PNG | plots/ |
| One-command workflow and repeated-run example | scripts/run_all.sh, scripts/run_repeated.sh |
The tutorial is a standalone folder. It can also be added under
tutorials/from_data_to_plots/ in the phipml repository: relative paths inside
the folder continue to work. Run its commands from that folder. It does not
require changes to phipml's source.
1. Install phipml
Recommended: the repository's Micromamba environment
Use Linux/macOS, or an equivalent Bash environment. Install Git and Micromamba first if needed; platform setup and Docker/Apptainer alternatives are in the repository's installation guide.
git clone https://github.com/csReynaB/phipml.git
cd phipml
# Pin the source used by this tutorial. This is a local checkout only.
git checkout 932f14ffca41ba4f24f3cc570afea27265a5893c
micromamba create --yes --name phipml --file ML_env.yml
micromamba activate phipml
python -m pip install --no-build-isolation --no-deps -e .
phipml --version
phipml -h
phipml-plot -h
phipml-heatmap -h
ML_env.yml supplies the tested scientific package pins with Python 3.10.16.
--no-deps prevents pip from replacing those packages. -e links the installed
package to the cloned source, so retain the checkout.
Alternative: pip in a virtual environment
In the same pinned repository checkout, use an installed Python 3.10–3.12:
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install .
phipml --version
Although the package declares Python >=3.10, its dependency upper bounds make
3.10–3.12 the practical choices for this tutorial; do not assume every newer
Python can resolve the pinned scientific stack. Pip resolves compatible ranges
and can choose versions different from ML_env.yml. The included run used
Python 3.12.13, sklearn 1.5.2, XGBoost 2.1.4, and SHAP 0.46.0; exact package
versions are recorded in reports/requirements-reference.txt.
The virtual environment must remain active while running the tutorial. The small examples use CPU execution; no GPU is required.
Open the tutorial folder
Extract phipml-tutorial-932f14f.zip and enter the extracted folder. For example,
if your downloaded archive is in the current directory:
unzip phipml-tutorial-932f14f.zip
cd phipml-tutorial-932f14f
All commands below assume this directory. Data/config files are already included. To recreate them deterministically:
python scripts/generate_examples.py
The generator rewrites its example data/configs with the documented defaults;
copy a config before customizing it. It does not delete old results. If you
change --data-seed or --n-iter, keep a separate tutorial copy/output directory
so old artifacts cannot be mistaken for the new experiment.
2. Understand the example data
There are two datasets, perfect and noisy. Each includes 160 training samples and 80 independent external samples, with balanced Control/Case labels. Every sample has 80 binary peptide measurements; four numeric clinical predictors are supplied through metadata.
| File | Rows represent | Required identifiers/content |
|---|---|---|
data/noisy/peptides.csv |
Samples | First column SampleName; remaining columns are named peptides containing 0/1. |
data/noisy/metadata.csv |
Samples | SampleName, group_test, cohort, Sex, Smoking, Age, BMI. |
data/peptide_library.csv |
Peptides | peptide_id, Description, Role, Species, Protein, is_demo_signal. Used for annotation, not as predictor columns. |
The matching files under data/perfect/ have the same schema. Inspect the
actual data with:
python - <<'PY'
import pandas as pd
X = pd.read_csv("data/noisy/peptides.csv", index_col=0)
metadata = pd.read_csv("data/noisy/metadata.csv").set_index("SampleName")
print("Peptide matrix:", X.shape)
print(X.iloc[:4, :6])
print(metadata.head(4))
print(pd.crosstab(metadata["cohort"], metadata["group_test"]))
assert X.index.is_unique and metadata.index.is_unique
assert set(X.index) == set(metadata.index)
PY
Expected counts:
| Cohort | Control | Case | Total |
|---|---|---|---|
| training | 80 | 80 | 160 |
| external | 40 | 40 | 80 |
group_tests: [Control, Case] encodes Control=0 and Case=1. cohort selects
training/external rows and is not included in the predictor matrix. The model
receives 80 peptides plus the four configured extras: 84 features before
pipeline feature selection.
What “perfect” and “noisy” mean
In the perfect data, six engineered peptides exactly encode the label or its inverse. This is an intentional software demonstration. The remaining 74 peptides and clinical features are independent of class in the data-generating process. Correlated perfect peptides can share model importance unevenly.
In the noisy data, six peptides have class-dependent presence probabilities: 0.70 versus 0.30 in training, and 0.65 versus 0.35 externally, with alternating directions. The external relationship is deliberately slightly weaker. The external age distribution is also shifted by four years, independently of class. There is no promise that tuning improves the result.
The generator draws training and external samples independently. Default and
tuned runs reuse exactly the same input files and model seed. Synthetic
annotations identify the intended signal/noise roles; they are not real taxa
or biomarkers. data/generation_manifest.json records the generation settings
and SHA-256 hashes.
Matrix orientation
The tutorial uses sample rows × peptide columns with transposed: false.
For a peptide-row matrix, use transposed: true. Equivalent
peptides_transposed.csv files are included for illustration. The first column
must contain the relevant row identifiers in either orientation.
3. Read the first model config
Open configs/models/random-forest/01_perfect.yaml. The main decisions are:
# Paths below are relative to configs/models/random-forest/.
data_input: ../../../data/perfect/peptides.csv
metadata_input: ../../../data/perfect/metadata.csv
data_input_mode: matrix
transposed: false
col_sample_name: SampleName
col_target: group_test
group_tests: [Control, Case]
peptide_prefixes: [agilent_, twist_, corona2_]
extra_features_to_include: [Sex, Smoking, Age, BMI]
fillna_value: null
lib_metadata_input: ../../../data/peptide_library.csv
lib_col_peptide_name: peptide_id
random_state: 420
classification:
model_type: random-forest
param_grid_name: random-forest
seed: 420
train_filters: {cohort: training}
run_nested_cv: true
only_train_model: true
use_pretrained: false
with_oligos: true
with_additional_features: true
prevalence_threshold_min: 0
prevalence_threshold_max: 100
outer_cv_splits: 3
inner_cv_splits: 2
n_iter: 1
n_jobs_outer: 1
n_jobs_inner: 1
classification_threshold: 0.5
validation_sets: []
output_dir: ../../../results/random-forest/01_perfect
output_name: 01_perfect
param_grid: {}
This is a complete minimal working configuration; the supplied file spells out additional options documented in the model reference.
param_grid: {} means no hyperparameter search. Constant-feature removal and
elastic-net peptide selection still fit inside each training pipeline.
run_nested_cv: true requests outer-CV evaluation; with an empty grid there is
no inner tuning. only_train_model: true then saves a model fitted on the whole
training cohort and skips external validation.
External rows are present in the input files but excluded from all training
and tuning by train_filters: {cohort: training}.
4. Fit a Random Forest and locate the outputs
phipml -c configs/models/random-forest/01_perfect.yaml
The log should report samples=160, features=84. This run creates:
Output under results/random-forest/01_perfect/ |
Use |
|---|---|
nested_random-forest_01_perfect_420.joblib |
Held-out outer-CV predictions, metrics, SHAP, selected features, and fold models. Use this to plot internal performance. |
training_random-forest_01_perfect_420.joblib |
Pipeline refitted on all training samples. Use this for later prediction/external validation. |
Filename pattern:
{nested|training}_{model_type}_{output_name}_{seed}.joblib.
Running the same configuration/seed again replaces the same filenames.
5. Generate the first plots
phipml-plot --plot-config configs/plots/random-forest/01_perfect.yaml
Open plots/random-forest/01_perfect/01_perfect_performance.png. It should show
near-perfect discrimination because the data were constructed to be perfectly
separable. The plotting recipe produces nine figure types in PDF, SVG, and PNG,
plus a detailed feature-importance CSV.
The essential plotting fields are:
plotting:
results:
- ../../../results/random-forest/01_perfect/nested_random-forest_01_perfect_420.joblib
config: ../../models/random-forest/01_perfect.yaml
split: train
plots: [all]
formats: [pdf, svg, png]
dpi: 180
output_dir: ../../../plots/random-forest/01_perfect
output_prefix: 01_perfect
class_labels: [Control, Case]
max_display: 10
library_metadata: ../../../data/peptide_library.csv
library_id_column: peptide_id
--plot-config reads the plotting recipe; its config: entry points to the
model/data YAML. split: train refers to the training cohort's out-of-fold
predictions. Use split: test for an external validation artifact. The separate
training_...joblib file is not an evaluation result to plot.
See the plotting reference for all fields, colors, annotation controls, feature ranking, and standalone-panel commands.
6. Move to noisy data, then tune
# Identical schema, imperfect signal, default hyperparameters.
phipml -c configs/models/random-forest/02_noisy.yaml
phipml-plot --plot-config configs/plots/random-forest/02_noisy.yaml
# The same noisy samples and seed, with inner-CV hyperparameter search.
phipml -c configs/models/random-forest/03_noisy_tuned.yaml
phipml-plot --plot-config configs/plots/random-forest/03_noisy_tuned.yaml
The tuned YAML adds search spaces beneath param_grid.random-forest and sets
n_iter: 12. It searches the peptide selector's regularization, forest tree
count/depth, split/leaf sizes, feature subsampling, and class weighting.
All candidate comparisons happen in training folds. The held-out outer samples
are evaluated only after a candidate has been selected.
For example, increasing min_samples_leaf makes leaves depend on more samples
and generally smooths predictions; increasing selector C weakens
regularization. Those are different parts of the pipeline. The
hyperparameter guide explains every search parameter,
the direction of its effect, and the exact YAML syntax.
What a real noisy result looks like
This figure is generated by the command above. ROC/PR summarize outer-fold performance. The confusion matrix pools held-out predictions. The classification bars in this version use outer-fold means and variability when present; these can differ slightly from metrics recomputed on all pooled predictions.
Interpret the learned features
The highest-ranked peptides largely recover the engineered signal in this example. Noise peptides and clinical variables can still receive nonzero importance. Class-wise prevalence/means and SHAP importance answer different questions: prevalence describes the samples; SHAP explains the fitted model. The CSV also includes column types and statistic labels omitted from the compact figure. Here peptide values are percentages and Age/BMI entries are means in their original units.
7. Evaluate the external cohort
Use run_nested_cv: false, only_train_model: false, and add the validation
cohort. Each model is fitted/tuned using the training rows and then evaluated
on the 80 independent external samples.
classification:
train_filters: {cohort: training}
run_nested_cv: false
only_train_model: false
validation_sets:
- name: 05_external_noisy
filters: {cohort: external}
classification_threshold: 0.5
bootstrap_validation: true
bootstrap_n_resamples: 500
bootstrap_confidence_level: 0.95
The fragment illustrates the change; run the complete supplied files:
phipml -c configs/models/random-forest/04_external_perfect.yaml
phipml-plot --plot-config configs/plots/random-forest/04_external_perfect.yaml
phipml -c configs/models/random-forest/05_external_noisy.yaml
phipml-plot --plot-config configs/plots/random-forest/05_external_noisy.yaml
phipml -c configs/models/random-forest/06_external_noisy_tuned.yaml
phipml-plot --plot-config configs/plots/random-forest/06_external_noisy_tuned.yaml
For example, the noisy external artifact is
results/random-forest/05_external_noisy/validation_random-forest_05_external_noisy_420.joblib.
Its name comes from the validation set's name, not output_name.
The current CLI selects cohorts from shared input files. Separate per-validation
matrix/metadata paths are not implemented in validation_sets at this commit.
For new files, see external-input guidance.
8. Run the same six scenarios with XGBoost
Use the XGBoost configurations supplied alongside Random Forest:
for scenario in 01_perfect 02_noisy 03_noisy_tuned 04_external_perfect 05_external_noisy 06_external_noisy_tuned; do
phipml -c "configs/models/xgboost/$scenario.yaml"
phipml-plot --plot-config "configs/plots/xgboost/$scenario.yaml"
done
The input/cohort definitions stay the same. XGBoost tuning instead searches tree count, learning rate, depth, child weight, row/column subsampling, and leaf/split regularization, along with the shared peptide selector.
Each dot is one external sample. Red/blue show high/low feature values; left/right show contributions away from/toward Case. XGBoost's SHAP values here are raw-margin/log-odds contributions, not probability-point differences. Mean absolute SHAP can rank features but does not give direction, statistical significance, or a causal interpretation.
9. One command for all scenarios
After installation, from the tutorial directory:
# Six models, all plots, heatmaps, and a metric/prediction export.
bash scripts/run_all.sh random-forest
# Or use XGBoost.
bash scripts/run_all.sh xgboost
# Or run all twelve estimator/scenario combinations.
bash scripts/run_all.sh both
| Scenario | What it evaluates | Tuning? | Evaluation samples |
|---|---|---|---|
01_perfect |
Strong, engineered signal via outer CV | No | 160 |
02_noisy |
Imperfect signal via outer CV | No | 160 |
03_noisy_tuned |
Same noisy data via nested CV | Yes | 160 |
04_external_perfect |
Full training model on independent perfect data | No | 80 |
05_external_noisy |
Full training model on weaker external signal | No | 80 |
06_external_noisy_tuned |
Training-selected tuned model on the same noisy external data | Yes | 80 |
Core recipes use three outer folds, two inner folds, one worker, and 12 search candidates for tuned runs. This is a short tutorial budget, not a recommended final analysis design. Hardware, installation startup, tuning, and plot export affect runtime; the reference execution timings are recorded separately.
10. Compare the observed results
python scripts/summarize_results.py
phipml-heatmap --manifest manifests/random-forest.csv \
--metric roc.auc --vmin 0 --vmax 1 --palette viridis --dpi 180 \
--output plots/random-forest/roc_auc_overview
reports/observed_metrics.csv contains actual metrics, uncertainties, and
artifact paths. reports/predictions/ contains sample IDs, encoded labels,
Case probabilities, and predictions at 0.5. reports/fitted_parameters.json
records the selected full-cohort pipeline settings.
Reference point estimates, seed 420:
| Scenario | RF ROC-AUC | RF AP | XGBoost ROC-AUC | XGBoost AP |
|---|---|---|---|---|
| Perfect internal CV | 1.000 | 1.000 | 1.000 | 1.000 |
| Noisy internal CV | 0.912 | 0.908 | 0.917 | 0.921 |
| Noisy tuned internal CV | 0.911 | 0.915 | 0.926 | 0.928 |
| Perfect external | 1.000 | 1.000 | 1.000 | 1.000 |
| Noisy external | 0.810 | 0.843 | 0.827 | 0.825 |
| Noisy tuned external | 0.806 | 0.805 | 0.808 | 0.804 |
Internal ROC-AUC/AP values are means across outer folds, not metrics calculated from pooled scores. The exported table reports both. External metrics evaluate one fixed full-cohort model. Tuning improved some internal estimates but did not improve external AP in this run. This is useful: tuning on training data cannot guarantee better generalization, especially after a distribution shift. Do not select the deployment model by repeatedly inspecting this external set.
The tutorial uses AP consistently: it is not interchangeable with trapezoidal area under a plotted precision–recall curve. The baseline AP of a random ranking is approximately the positive-class prevalence, here 0.5.
11. Repeat CV to examine stability
bash scripts/run_repeated.sh random-forest
This executes seeds 420, 421, and 422 on the same noisy dataset, stores them in
results/random-forest/repeated_noisy/, and aggregates them using
configs/plots/random-forest/07_repeated_noisy.yaml. It skips unnecessary
full-cohort model fitting for these repeat-only runs.
Equivalent model commands, showing the relevant overrides:
for seed in 420 421 422; do
phipml -c configs/models/random-forest/02_noisy.yaml --seed "$seed" \
--output-dir ../../../results/random-forest/repeated_noisy \
--no-only-train-model
done
phipml-plot --plot-config configs/plots/random-forest/07_repeated_noisy.yaml
The plot recipe keeps each run's top 10 features, then displays leading features that occur in at least 50% of runs. These three repeats demonstrate the mechanics; they are not enough to characterize feature stability precisely. More repeats reuse the same participants and must not be treated as independent studies.
Use one estimator/experiment per glob. To compare RF against XGBoost, use separate panels or a heatmap, not repeated-run aggregation.
12. Reuse the saved fitted model
After 03_noisy_tuned, validate its saved full-cohort model without another
search/refit:
phipml -c configs/models/random-forest/08_reuse_noisy.yaml
The key settings are:
classification:
run_nested_cv: false
only_train_model: false
use_pretrained: true
input_dir: ../../../results/random-forest/03_noisy_tuned
input_name: training_random-forest_03_noisy_tuned_420.joblib
The complete recipe includes the training input context and external filters,
and saves validation_random-forest_08_reuse_noisy_420.joblib under
results/random-forest/08_reuse_noisy/. Use the matching environment when loading
serialized fitted estimators. The XGBoost equivalent is supplied too.
13. Adapt this workflow to your own data
- Copy a supplied model YAML to another file in the same model-config folder. Replace the input paths, sample/target column names, class order, matrix orientation, and peptide prefixes. Check sample IDs and class counts.
- Choose numeric clinical extras. Handle genuine missingness and categorical
encodings deliberately;
fillna_value: 0is not a general solution. - Define disjoint training/validation cohort filters and choose the split design before inspecting external performance. This CLI uses stratified sample-level CV, not participant-group CV.
- Start with
param_grid: {}. Then copy the corresponding estimator's tuning space, revise it for your data size, and increase the search budget if useful. - Set distinct
output_dir,output_name, and validation names. Fit the model. - Copy the matching plot YAML; point
resultsto the new evaluation artifact, setconfig, selecttrain/test, and supply annotation column names that actually exist. Plot without retraining. - Preserve inputs, configs, source commit, dependency versions, predictions, and evaluation outputs together.
Troubleshooting
| Symptom | What to check |
|---|---|
phipml: command not found |
Activate the environment containing the installed package; run which python and which phipml. |
| File not found despite an existing CSV | Resolve model paths from the model YAML, plot paths from the plot YAML, and plot CLI overrides from the terminal directory. |
| No sample IDs in common, or unexpected sample count | Matrix orientation, first-column IDs, metadata SampleName, filters, and target spelling. |
| Missing columns/unsupported clinical data | Match extra_features_to_include exactly; encode nonnumeric categorical values before fitting. |
| No useful peptide features survive | Inspect prefixes, prevalence filters, constant columns, and selector C; verify binary measurements and labels before widening a search. |
| A parameter-grid key fails | Use random-forest/xgboost consistently and the exact pipeline path from the hyperparameter guide. |
| Validation did not run | Set only_train_model: false and add nonempty validation_sets. |
| Plotting says a metric/split is unavailable | Use a nested_ artifact with train, or a validation_ artifact with test, not a training_ model file. |
| Plotting after moving the folder fails | Set the plot recipe's config to the current local model YAML. Keep input files unchanged. |
| Repeated SHAP alignment fails | Check identical samples/features/experiment across matched files. Do not silently merge unrelated cohorts. |
| Tuned results are worse | This can be legitimate. Check the search objective, data size, uncertainty, and external distribution; do not search for a favorable seed. |
| Exact metrics differ from the table | Check source/dependency versions, generation seed/data hashes, model seed, and search budget. |
Further implementation details are linked from the three reference documents. This tutorial's observed results are synthetic examples for learning the tool, not estimates of diagnostic performance on patient data.
Model configuration reference
This reference describes the implemented CLI at phipml 4.2.0, commit
932f14f.
Start with a complete YAML in configs/models/; use this page when adapting it.
The three parts of a model YAML
| Section | Controls | Example |
|---|---|---|
| Top level | Files, sample IDs, target encoding, feature sources | metadata_input, group_tests, transposed |
classification: |
Estimator, cohorts, CV, output paths, decision threshold | model_type, train_filters, outer_cv_splits |
param_grid: |
Search spaces for pipeline parameters | random-forest:, xgboost: |
Use spaces for YAML indentation, native true/false, null for no value, and
lists in brackets or with hyphens. Quoted "false" is a string, not a Boolean.
Class and column names must match the input files, including capitalization.
Path rule: model YAML paths are relative to the YAML's own directory.
classification.output_dir and input_dir CLI overrides follow this same rule.
In a plotting YAML, paths are relative to that plotting YAML; explicit plotting
CLI paths are relative to your terminal's current directory. Absolute paths
work in either case.
Explicit model CLI values override the classification: section, which
overrides internal defaults. classification.seed overrides random_state.
Avoid adding unrecognized configuration keys: nested mappings can contain keys
that are not consumed by the CLI.
Inputs and target: top-level keys
| Key | What to provide and what it means |
|---|---|
data_input |
Required path to the peptide matrix, directory of enriched-peptide sample files, or sample-file manifest. In this tutorial: ../../../data/noisy/peptides.csv. |
metadata_input |
Required sample metadata CSV, TSV/TXT, XLSX, or XLS. One row per sample. Must contain the configured sample and target columns. |
data_input_mode |
matrix, sample-files, or auto. auto treats a directory as sample files and a file as a matrix; explicitly select sample-files for a manifest file. |
project |
Optional descriptive identifier; output filenames are controlled separately by classification.output_name. |
col_sample_name |
Sample-ID column in metadata, here SampleName. IDs must be unique and match the matrix. The combined matrix's first column is read as its index. |
col_target |
Metadata column containing the target classes, here group_test. The classification CLI is binary. |
group_tests |
Two target labels in encoding order: [Control, Case] gives Control=0 and Case=1. Predicted probabilities and positive SHAP direction refer to Case. Non-selected labels are excluded; check that the intended samples remain. |
col_predict |
Compatibility column name used by some workflows; default class1_proba. This does not change the target or the model's decision threshold. Omitted in the examples. |
transposed |
false for sample rows × peptide columns; true for peptide rows × sample columns. Sample-file mode constructs sample rows automatically. |
peptide_prefixes |
Prefixes distinguishing peptides from clinical features. Here [agilent_, twist_, corona2_]. A missing trailing underscore is added. All peptide IDs must use your configured prefixes. |
extra_features_to_include |
Metadata columns to append, here [Sex, Smoking, Age, BMI]; enabled by classification.with_additional_features. Do not include target, cohort, participant ID, or outcomes measured after the prediction time. |
fillna_value |
Whole-matrix missing-value fill after loading. null preserves missingness. A number such as 0 fills both peptide and clinical missingness, so it is not a clinical imputation strategy. |
random_state |
General seed; model/CV runs use classification.seed if supplied. This does not regenerate synthetic data; the generator has a separate --data-seed. |
lib_metadata_input |
Peptide annotation file: CSV, TSV/TXT, Excel, or a pandas pickle. Required for library-based peptide filters; also useful for annotated plots. Not required for unfiltered numeric modeling. |
lib_col_peptide_name |
Library column identifying peptides, here peptide_id. If null, table loading uses the first column; pickle input can already use a peptide-ID index. |
filters_metadata |
Base sample filter, e.g. {timepoint: baseline}. Conditions across keys use AND, a list of allowed values within one key uses OR. Used as the default when no training filter is set. Run-specific training/validation filters replace this mapping when provided, so include required conditions in each cohort filter. |
combined_filters_metadata |
List of condition mappings: OR between mappings, AND within each mapping. This stage still applies alongside run-specific cohort filters; keep it null for the tutorial. |
oligo_filters |
Default peptide-library filter mapping, such as {Species: Homo sapiens}. Use annotation columns, not sample metadata. The run can supply its own filters. |
oligo_filter_mode |
all means AND across library conditions; any means OR. Defaults to all. |
The loader intersects sample IDs between metadata and the matrix. A typo can
therefore remove samples without meaning that the cohort definition was right.
Check sample counts before fitting and the Final training data log message.
Alternative input: one enrichment file per sample
The main examples use combined matrices. For per-sample enriched-peptide lists, replace only the data-loading part of your YAML:
data_input: ../../../my_data/enriched_samples
data_input_mode: sample-files
sample_file_patterns: ["*.csv"]
sample_file_peptide_column: ID
sample_name_regex: null
| Key | Meaning |
|---|---|
sample_file_patterns |
Glob patterns within the directory; recursive patterns such as batch_*/**/*.csv are supported. |
sample_file_peptide_column |
Column containing enriched peptide IDs in every sample file, e.g. ID. These are presence lists, not raw sequencing-count matrices. |
sample_name_regex |
Optional regex to extract a sample ID from the filename; preferably a named group such as 'prefix_(?P<sample>.+)_enriched'. null uses the filename stem. |
A tab-delimited manifest can list sample_name and path columns, or one path
per line. Paths in a sample manifest are relative to that manifest. An explicit
sample name takes precedence over the filename-derived name.
Execution: classification: keys
Model and run type
| Key | Meaning and tutorial value |
|---|---|
model_type |
Exactly random-forest or xgboost; not random_forest, rf, or XGB. |
param_grid_name |
Which top-level param_grid subsection to read. Keep it equal to model_type in these examples. Switching models in a tuned run requires switching this too. |
seed |
Estimator/CV/split/bootstrap seed and filename suffix; 420 here. One seed creates one CV run; multiple repeats require multiple commands. |
run_nested_cv |
Evaluate the training cohort with outer stratified folds. Each outer model tunes on inner folds if a search space exists; with param_grid: {}, this is ordinary outer CV without tuning. |
only_train_model |
Fit/save the full-cohort model and stop before external evaluation. If nested CV is enabled, CV runs first. Internal examples set true; external examples set false. |
use_pretrained |
Load a fitted pipeline for full-model external evaluation. It does not stop an independently enabled nested-CV run from fitting fold models. Use run_nested_cv: false when only reusing a model. |
input_dir |
Directory containing a pretrained artifact, used with use_pretrained: true. |
input_name |
Direct .joblib filename or base name used to find standard saved-model names. An explicit filename is clearest. |
output_dir |
Destination for joblib artifacts; created automatically. Give each experiment a separate directory. |
output_name |
Base name for nested_... and training_... artifacts. External artifacts instead use each validation set's name. |
Features and preprocessing
| Key | Meaning and practical effect |
|---|---|
with_oligos |
Include peptide features. true here. At least this or with_additional_features must be true. |
with_additional_features |
Include columns in extra_features_to_include. true here. Adding Sex/Age makes them predictors; it does not by itself establish a causal adjustment. |
subgroup |
all retains the available library; another column name is a compatibility shortcut for selecting rows where that annotation is true. Use explicit oligo_filters for new analyses. |
oligo_filters |
Run-specific library filters. Example: {Species: Homo sapiens, is_PNP: true}. Use this with subgroup: all. Do not combine a named subgroup with non-null filters. |
oligo_filter_mode |
all/any across run-specific annotation conditions. |
prevalence_threshold_min |
Minimum percentage of training samples with a nonzero, nonmissing peptide value. Inclusive; tutorial 0. 1 means 1%, not a proportion of 1. |
prevalence_threshold_max |
Maximum percentage with a nonzero, nonmissing peptide value. Inclusive; tutorial 100. |
impute_extra_numeric |
Fit a SimpleImputer inside each pipeline for continuous numeric clinical features. Tutorial data are complete, so false. It does not impute peptides or numeric binary extras. |
extra_numeric_impute_strategy |
mean, median, most_frequent, or constant; median here. Active only when imputation is enabled. The CLI has no dedicated constant-fill-value option. |
fill_missing_peptides_with_zero |
Aligning external data: add absent peptide columns as zero when true, otherwise fail. We set false because all expected features are measured. Missing clinical columns always fail. This option handles absent columns, not arbitrary NaNs inside present columns. |
The peptide pipeline removes constant columns, then fits an elastic-net logistic-regression selector. Clinical variables bypass that peptide selector. Feature selection and optional continuous imputation are fitted inside CV pipelines. The current CLI applies prevalence filtering to the full training cohort before outer CV, although it excludes external samples. The tutorial uses 0–100% to make this filter inactive. Do not describe a restricted prevalence filter as being fitted separately inside every CV fold in this version.
Sex/Gender values F/Female/M/Male and 0/1 are recognized. Other categorical
clinical variables require numeric encoding before fitting. Define encodings
using training data and apply the same schema externally. Genuine missing
measurements should not automatically become biological absences.
CV, tuning, and uncertainty
| Key | Meaning and tutorial value |
|---|---|
outer_cv_splits |
Number of outer evaluation folds; 3 for the tutorial. Each training sample receives one held-out prediction per run. |
inner_cv_splits |
Number of folds used to rank hyperparameter candidates; 2 here. The final full-cohort model uses an additional inner search on the entire training cohort. Ignored for fitting when there is no search. |
n_iter |
Number of candidate configurations per Bayesian search, not number of CV repeats or trees. 12 in tuned recipes. It has no tuning effect when param_grid: {}. |
n_jobs_outer |
Concurrent outer-fold jobs; 1. |
n_jobs_inner |
Concurrent search/CV jobs; 1. Use a bounded positive number or -1 for available CPUs. Estimators themselves are created with n_jobs=1; avoid multiplying outer and inner parallelism on small machines. |
classification_threshold |
Preselected Case-probability cutoff for confusion matrix, accuracy, sensitivity, specificity, F1, etc.; 0.5. ROC-AUC and AP do not depend on this cutoff. |
bootstrap_validation |
Bootstrap the external predictions for uncertainty; true. This resamples observations with a fixed fitted model, without refitting or retuning. |
bootstrap_n_resamples |
External bootstrap draws; 500 for a short demonstration, implementation default 1000. More draws reduce Monte Carlo noise, not the uncertainty caused by limited sample size. |
bootstrap_confidence_level |
External interval level, as a fraction; 0.95 means 95%. It does not turn CV fold SD into a confidence interval. |
The search score is average precision in this commit. It is not a YAML
option. XGBoost's built-in eval_metric="auc" does not change the search scorer.
Class stratification preserves class proportions; it does not group repeated
samples from one participant. This CLI has no participant-group CV setting.
The mock data represent independent samples.
For a larger exploratory run, a reasonable starting budget is 5 outer folds, 3 inner folds, and 30 candidates, then repeated seeds if compute permits. Choose the split design and sample sizes for the scientific study; more repeats do not create more independent participants. Never select a seed or tuning recipe by its external-cohort result.
Cohorts and optional hold-out splitting
| Key | Meaning |
|---|---|
train_filters |
Metadata conditions selecting training samples, here {cohort: training}. Omitting this from a combined training/external dataset can include external samples in fitting. |
validation_sets |
A list of {name: ..., filters: {...}} entries, all read from the configured inputs. Names control validation filenames. Example below. |
split_filters |
Optional cohort to partition into train and hold-out portions; null in the six core examples. |
train_size |
Fraction of the split cohort assigned to training; default 0.7. Active only with split_filters. |
split_only |
true trains/evaluates only the split cohort. With false, the split's training portion is appended to train_filters, and its held-out portion is appended to each validation set. Use disjoint selections. |
classification:
train_filters: {cohort: training}
run_nested_cv: false
only_train_model: false
validation_sets:
- name: independent_site
filters: {cohort: external}
Keep participants and any repeated samples disjoint between cohorts. Labeling
rows external does not make them independent if they are copied from training.
What this version supports for external files
At the inspected commit, validation_sets accepts name and filters, not a
separate per-cohort data_input or metadata_input. The six examples therefore
use a combined matrix and metadata with disjoint training/external IDs. The
training filter excludes external rows from fitting and tuning.
If a new cohort arrives in separate files, combine compatible tables first
(with unique sample IDs, a cohort column, and checked feature/metadata schemas)
and then use the filters above. Or adapt the Python API explicitly. The current
pretrained CLI still prepares its training input context, so these reuse recipes
retain the training rows; an external-only CLI workflow is not documented here
as if it were already implemented.
Useful command overrides
Run these from the tutorial directory:
# Same noisy dataset, a different CV seed, dedicated output folder.
phipml -c configs/models/random-forest/02_noisy.yaml --seed 421 \
--output-dir ../../../results/random-forest/repeat_421
# A longer search. Output is kept separate from the reference examples.
phipml -c configs/models/xgboost/03_noisy_tuned.yaml \
--n-iter 30 --inner-cv-splits 3 --outer-cv-splits 5 \
--output-dir ../../../results/xgboost/noisy_longer_search
# Untuned clinical-only baseline, on the same target/cohort/CV seed.
phipml -c configs/models/random-forest/02_noisy.yaml \
--no-with-oligos --with-additional-features \
--output-name clinical_only --output-dir ../../../results/random-forest/clinical_only
# Untuned peptides-only comparison.
phipml -c configs/models/random-forest/02_noisy.yaml \
--with-oligos --no-with-additional-features \
--output-name peptides_only --output-dir ../../../results/random-forest/peptides_only
For a tuned estimator change, use the supplied estimator-specific YAML. Merely
passing --model-type xgboost to a Random Forest tuning YAML leaves its selected
grid pointing at Random Forest parameters unless it is also replaced.
Source: the pinned repository's data/config loader, run settings, and training CLI.
What the hyperparameters mean
phipml fits a pipeline: peptide preprocessing/selection, optional clinical imputation, then a Random Forest or XGBoost classifier. A hyperparameter is a setting chosen before fitting; training then estimates model coefficients or trees from the training samples. The tuned examples choose hyperparameters by inner cross-validation, using average precision (AP).
These are small teaching search spaces, not claimed optimal settings for PhIP-seq. The same noisy data, target, cohort filters, and seed are used for untuned/tuned pairs. Better training or inner-CV performance does not guarantee better outer-CV or external performance.
Read a parameter path
param_grid:
random-forest:
estimator__max_depth:
type: integer
low: 2
high: 10
classification.param_grid_name: random-forest selects this section.
estimator__max_depth means “the max_depth parameter of the pipeline step
called estimator.” Double underscores traverse nested pipeline objects.
The longer path
preprocessor__peptides__feature_selection__estimator__C changes the logistic
regression used to select peptides, not the final Random Forest/XGBoost.
| Search specification | Meaning | Valid example |
|---|---|---|
type: integer |
Whole-number choices between inclusive bounds | {type: integer, low: 2, high: 10} |
type: real |
Continuous values between bounds | {type: real, low: 0.2, high: 0.8} |
type: categorical |
Choose one of the listed alternatives | {type: categorical, categories: [sqrt, log2, 0.5]} |
prior: uniform |
Equal weight to equal-width intervals; default for reals | Useful for a fraction such as subsample. |
prior: log-uniform |
Equal weight across multiplicative scales | Useful for C, learning rate, or regularization; both bounds must be strictly positive. |
Do not use equal low/high to hold a value fixed. Use a one-element categorical
space, e.g. {type: categorical, categories: [100]}. This still invokes the
search machinery if param_grid is nonempty. The current CLI does not expose a
separate estimator_params block for fixed overrides. param_grid: {} uses
the built-in pipeline and estimator defaults; it does not disable feature
selection.
Shared peptide selector
Both models fit elastic-net logistic regression with solver: saga, then retain
peptides whose absolute coefficient meets the selection threshold.
| Full parameter key | Tutorial search | Meaning | Effect of increasing it |
|---|---|---|---|
preprocessor__peptides__feature_selection__estimator__l1_ratio |
0.2–0.8, uniform | Mix of L1 and L2 penalties: 0=pure L2, 1=pure L1. Fixed default 0.5. | Generally promotes sparser selection as the L1 share increases, though correlated predictors can change which peptide survives. |
preprocessor__peptides__feature_selection__estimator__C |
0.1–10, log-uniform | Inverse regularization strength. Fixed default 1.0. | Weaker shrinkage, often more selected peptides; a small value may eliminate nearly all peptide signal. |
Fixed implementation settings are VarianceThreshold(threshold=0.0), selector
threshold=1e-5, and logistic-regression max_iter=10000. The selector is
supervised and learns from class labels inside the fitted CV pipeline. Clinical
features do not pass through this peptide selector. Inactive peptide search
parameters are ignored when no peptide branch exists.
Feature selection can choose different peptides in different folds. The saved SHAP matrices retain original feature names and assign zero to unused features for the relevant fitted model/fold. A zero SHAP value is not proof of no biological association.
Random Forest
A Random Forest averages many randomized decision trees. Each tree can capture nonlinear patterns; averaging reduces the sensitivity to any one fitted tree.
Use the exact final-model prefix estimator__ for every parameter below.
| Parameter key | Tutorial search | Meaning | Practical tradeoff |
|---|---|---|---|
estimator__n_estimators |
Integer 60–160 | Number of trees. Untuned sklearn default: 100. | More trees usually stabilize predictions, but cost time/memory; tree count alone does not control overfitting as strongly as tree size. |
estimator__max_depth |
Integer 2–10 | Maximum depth of each tree. Untuned default: None (no explicit depth limit). |
Deeper trees fit finer interactions and noise; shallow trees may underfit. |
estimator__min_samples_split |
Integer 2–10 | Minimum observations required in a node before attempting a split. Untuned default: 2. | Larger values suppress splits on small groups. This is different from the size of each resulting leaf. |
estimator__min_samples_leaf |
Integer 1–6 | Minimum observations in each terminal leaf. Untuned default: 1. | Larger leaves smooth predicted probabilities and reduce small-sample fitting. |
estimator__max_features |
sqrt, log2, or 0.5 |
Number/fraction of available features considered at a split, after preprocessing. Untuned default: sqrt. |
Smaller subsets decorrelate trees, but can hide useful features at a split. 0.5 is numeric: half of available features, not the string "0.5". |
estimator__class_weight |
null or balanced |
Weight the contribution of classes to fitting. Untuned default: None. |
balanced weights inversely to class frequency. It can change decision behavior; it does not rebalance the evaluation cohort. |
Other supported Random Forest parameters can be added through a valid pipeline search path. Common optional choices:
| Parameter | How to specify | Meaning |
|---|---|---|
estimator__bootstrap |
{type: categorical, categories: [true, false]} |
Whether each tree uses a bootstrap sample. Default true. |
estimator__criterion |
{type: categorical, categories: [gini, entropy]} |
Criterion for scoring candidate splits; default gini. |
estimator__class_weight |
Add balanced_subsample to categories |
Recalculate inverse-frequency weights separately for each tree's bootstrap sample. |
The balanced mock classes are included to keep the demonstration easy to read; class weighting is present to teach its meaning, not because imbalance needs correction in these data.
XGBoost
XGBoost builds trees sequentially. Each boosting round adds a correction to the
current model. Smaller learning steps often need more trees. The phipml factory
sets objective="binary:logistic", tree_method="hist", eval_metric="auc",
n_jobs=1, and the configured seed. Other untuned settings come from the
installed XGBoost version. There is no early-stopping evaluation set supplied
by this CLI.
| Parameter key | Tutorial search | Meaning | Practical tradeoff |
|---|---|---|---|
estimator__n_estimators |
Integer 50–160 | Maximum number of boosting rounds/trees for this tree booster. | More rounds add capacity; combine with learning_rate. Without early stopping, the configured rounds are fitted. |
estimator__learning_rate |
0.02–0.25, log-uniform | Shrinkage applied to each added tree; also called eta. | Smaller steps often need more rounds and can improve generalization; too few rounds then underfit. |
estimator__max_depth |
Integer 2–5 | Maximum depth of a tree. | Larger values fit more complex interactions and require more data/memory. |
estimator__min_child_weight |
Integer 1–8 | Minimum sum of second-order loss weights (Hessians) required in a child node. | Larger values discourage splits. This is not a literal minimum sample count for logistic classification. |
estimator__subsample |
0.65–1.0, uniform | Fraction of training observations used in each boosting iteration. | Values below 1 introduce row subsampling; too little data per round can weaken learning. |
estimator__colsample_bytree |
0.65–1.0, uniform | Fraction of available features sampled for each tree. | Can reduce correlated tree behavior and fitting to noise; may miss useful predictors. |
estimator__reg_lambda |
0.1–10, log-uniform | L2 regularization on leaf weights. | Larger values shrink leaf contributions and make the model more conservative. |
estimator__reg_alpha |
0.0001–1.0, log-uniform | L1 regularization on leaf weights. | Larger values can set some leaf contributions to zero. The positive lower bound makes the log-uniform prior valid. |
estimator__gamma |
0–2, uniform | Minimum reduction in objective required to accept a split. | Larger values discourage low-benefit splits. |
Optional estimator__scale_pos_weight increases the training weight assigned to
the positive class. If exploring it, use training-fold information and validate
its effect. It is omitted here because the synthetic classes are balanced.
Do not add class_weight to an XGBoost grid or learning_rate to a Random
Forest grid.
How tuning is evaluated
For 03_noisy_tuned:
- Split the training cohort into three outer folds.
- For each outer fold, run 12 candidate configurations using two inner folds on the outer-training subset; refit the best candidate on that subset.
- Predict and explain the held-out outer subset exactly once.
- Combine the outer-fold results into the nested artifact.
- Run a fresh inner search on all 160 training samples and save the resulting full-cohort pipeline for reuse.
This costs approximately 3 × (12 × 2 + 1) + (12 × 2 + 1) = 100 pipeline fits.
It is a teaching budget for a fairly wide space. BayesSearchCV in the pinned
scikit-optimize version starts with 10 initialization candidates by default;
12 candidates allow a short adaptive phase. A two-candidate smoke run exercises
the search workflow but cannot demonstrate a mature Bayesian optimization.
For 06_external_noisy_tuned, there is one inner search on the full training
cohort, followed by evaluation of its fixed selected model on 80 external
samples. External labels do not choose parameters, features, or thresholds.
Inspect what was actually chosen:
python scripts/summarize_results.py
Open reports/fitted_parameters.json for full-cohort pipeline settings, or:
from pathlib import Path
import joblib
path = Path("results/random-forest/03_noisy_tuned/training_random-forest_03_noisy_tuned_420.joblib")
model = joblib.load(path)["best_estimator"]
print(model.named_steps["estimator"].get_params())
selector = model.named_steps["preprocessor"].named_transformers_["peptides"].named_steps["feature_selection"]
print(selector.estimator_.get_params())
Outer-fold selected settings can differ from this full-cohort model; they are
available from model_list in the nested artifact. The CLI saves the chosen
models, not the complete BayesSearchCV.cv_results_ optimization history.
Implementation source: pipeline, selector and search. Parameter semantics were checked against the installed scikit-learn 1.5.2 and XGBoost 2.1.4 classes used for the reference run. Further reference: RandomForestClassifier, LogisticRegression, XGBoost parameters.
Plot configuration reference
Use phipml-plot --plot-config PATH.yaml for a saved recipe. Model fitting
and plotting are separate commands: change colors or displayed annotations
without fitting models again. Use phipml-heatmap with a CSV/TSV manifest to
compare experiments. This guide describes the current public commands, not the
older plotting modules retained in the repository.
Start with the right artifact
| Result filename begins with | Content | Plot split |
|---|---|---|
nested_ |
Outer-fold held-out predictions, SHAP, metrics, fold models | train |
validation_ |
Predictions/SHAP/metrics for an external or explicit hold-out cohort | test |
training_ |
Full-cohort fitted pipeline and input context | Not an evaluation artifact; use it for model reuse, not performance plotting. |
split: train is a storage key for the training cohort's out-of-fold
evaluation. It does not mean the model is evaluated on samples it just fitted.
For the internal examples the file contains one held-out prediction for each
of the 160 training-cohort samples.
The files embed targets, feature names, SHAP values, metrics, and data-loading
context. They do not contain a portable copy of every original input file.
Use plotting.config to point to the local model YAML after moving the folder;
the supplied recipes already do this. The original data are still required
for colored feature-value SHAP displays and prevalence/clinical summaries.
A complete plotting recipe
This example is stored at configs/plots/random-forest/02_noisy.yaml.
plotting:
results:
- ../../../results/random-forest/02_noisy/nested_random-forest_02_noisy_420.joblib
config: ../../models/random-forest/02_noisy.yaml
split: train
plots: [all]
formats: [pdf, svg, png]
dpi: 180
output_dir: ../../../plots/random-forest/02_noisy
output_prefix: 02_noisy
title: "random-forest | 02 noisy"
class_labels: [Control, Case]
max_display: 10
feature_ranking: mean-abs-shap
ranking_top_k: 10
min_top_k_frequency: 0
shap_alignment: strict
save_standalone: true
reconstruct_data: true
library_metadata: ../../../data/peptide_library.csv
library_id_column: peptide_id
validation_bootstraps: 0
random_state: 420
class_colors: ["#386CB0", "#D35D4D"]
colors:
roc: "#176B73"
roc_band: "#B3DAD8"
pr: "#8A5A44"
pr_band: "#DFC5B8"
confusion_cmap: Blues
classification: "#52796F"
shap_cmap: phipml_blue_gray_red
shap_heatmap_cmap: phipml_purple_gray_orange
shap_importance: "#6D597A"
feature_table:
annotation_columns: [Description, Role]
extra_columns: []
title: "Synthetic features: model importance and known roles"
header_color: "#DCE9E9"
row_colors: ["#F5F8F8", white]
prevalence_cmap: phipml_prevalence
phipml-plot --plot-config configs/plots/random-forest/02_noisy.yaml
The figure files and feature table appear under plots/random-forest/02_noisy/.
Relative YAML paths are resolved from the plotting YAML. Explicit CLI paths are
resolved from the current terminal directory, and override YAML values.
Inputs, output, and figure selection
plotting key |
What to provide and what it controls |
|---|---|
results |
List of evaluation joblib paths or quoted glob patterns. Multiple files are aggregated as repeated runs; they are not automatically drawn as distinct models in a comparison. |
config |
Model YAML for the original input files/labels/annotations. Optional if embedded input paths still resolve; supplied here to make the examples portable. This is different from the command's --plot-config. |
split |
train, test, or auto. auto infers an available evaluation split; explicit values are clearer. |
plots |
List drawn from all, performance, roc, pr, confusion, classification, shap-beeswarm, shap-importance, shap-heatmap, feature-table. |
formats |
One or more of pdf, svg, png. Vector SVG/PDF are useful for editing/sharing; PNG is convenient for slides and previews. |
dpi |
Raster resolution; 180 for compact tutorial files. Increase to 300 or 600 if required. This does not change model results or the geometry of vector elements. |
output_dir |
Plot destination. If omitted, the CLI considers an explicitly supplied model config's classification.plot_output_dir, then its classification.output_dir/plots; otherwise it uses plots/ next to the result. Explicit is easiest. |
output_prefix |
Shared output filename stem, e.g. 02_noisy produces 02_noisy_performance.svg. Default phipml. |
title |
Title for the combined performance summary. Some standalone panels have their own labels; do not expect this to rename every panel. |
class_labels |
Display labels in [negative, positive] order. Defaults to model config labels if available. Relabeling does not change encoded targets. |
save_standalone |
true saves requested individual ROC/PR/confusion/classification figures. false removes those standalone panel requests; include performance to retain the combined panel. |
reconstruct_data |
true loads feature values from explicit/embedded model config. Set false for performance-only plotting without the original data. |
validation_bootstraps |
Reconstruct external bootstrap intervals only if missing, for a single test artifact with targets available. 0 here because fitting already saved 500-resample intervals. This does not replace existing intervals. |
random_state |
Seed for bootstrap reconstruction when used. It does not change predictions or retrain anything. |
What each plot answers
plots choice |
Interpretation | Typical filename suffix |
|---|---|---|
performance |
Combined ROC, precision–recall, confusion matrix, and threshold-dependent metric panels. | _performance |
roc |
Sensitivity versus false-positive rate across score cutoffs. The legend reports ROC-AUC. | _roc |
pr |
Precision versus recall. The legend reports AP, not trapezoidal PR-AUC. Baseline precision depends on positive-class prevalence. | _pr |
confusion |
Counts of true/false positive/negative predictions at the saved threshold. | _confusion |
classification |
Metrics such as accuracy, sensitivity, specificity, F1 and MCC at that threshold. | _classification |
shap-beeswarm |
Distribution of signed contributions per feature; each point represents a sample. Color shows feature value. | _shap_beeswarm |
shap-importance |
Mean absolute contribution; magnitude without direction. | _shap_importance |
shap-heatmap |
Patterns of signed contributions across samples and selected features. | _shap_heatmap |
feature-table |
Compact importance table with class-wise prevalence/clinical summaries and annotations; also exports a detailed CSV. | _feature_table, _feature_importance.csv |
All nine figure types are generated by plots: [all]. The table CSV retains
more features/annotations/audit fields than the compact rendered table.
Feature ranking and repeated runs
| Key | Meaning |
|---|---|
max_display |
Maximum number of ranked features shown in SHAP figures and compact table; 10. This is a display limit, not model feature selection. |
feature_ranking |
mean-abs-shap ranks mean absolute SHAP; top-k-frequency ranks how often a feature appears in each run's top K. auto chooses mean absolute SHAP for one result and top-K frequency for multiple results. |
ranking_top_k |
Per-run top-K cutoff used for frequency ranking; 10. Separate from max_display. |
min_top_k_frequency |
Minimum percentage of runs whose top K includes a feature; 50 means 50%, not 0.5. Set 0 for the single-file examples. With very few repeats this is only a demonstration of stability. |
shap_alignment |
strict requires identical sample and feature sets across repeated files, reordered by ID as needed. intersection deliberately restricts to shared samples/features and changes the interpreted population. |
Only aggregate runs with the same estimator, data, cohort, feature definition, class encoding, and threshold. Do not use a broad glob that mixes Random Forest, XGBoost, noisy/perfect data, training/external cohorts, or tuned/default experiments. Use separate plots or a heatmap manifest for those comparisons.
The repeated plot recipe uses a dedicated repeated_noisy folder, and three
seeds on the same dataset. Signed SHAP values are averaged by sample/feature;
ranking/stability statistics also retain run-level absolute importance.
Top-K frequency differs from how often the elastic-net selector retained a
feature. Both statistics can be inspected in the exported CSV.
Annotations and appearance
| Key | Meaning |
|---|---|
library_metadata |
Optional annotation CSV, TSV/TXT, or Excel table used by the plotting CLI. Unlike the model loader, this explicit table reader does not accept pickle. |
library_id_column |
Annotation column matching feature IDs; peptide_id. Set explicitly to ensure the join uses the right index. |
class_colors |
Two colors in [Control, Case] order for prediction-direction labels. These do not encode biological class in the SHAP feature-value color scale. |
colors.roc, colors.roc_band |
ROC curve and uncertainty shading. |
colors.pr, colors.pr_band |
Precision–recall curve and shading. |
colors.confusion_cmap |
Matplotlib colormap for confusion matrix counts. |
colors.classification |
Threshold-metric bar color. |
colors.shap_cmap |
Feature-value colormap for the beeswarm. |
colors.shap_heatmap_cmap |
Diverging colormap for signed SHAP contributions. |
colors.shap_importance |
Absolute-SHAP bar color. |
feature_table.annotation_columns |
Annotation fields shown in the compact figure, e.g. [Description, Species, Protein]. Tutorial uses [Description, Role] to distinguish known synthetic signal/noise. |
feature_table.extra_columns |
Optional audit columns to display, e.g. ["Top-k SHAP frequency (%)", "Selection frequency (%)"]. They remain in the CSV when omitted from the figure. |
feature_table.title |
Compact-table title. |
feature_table.header_color |
Header fill color. |
feature_table.row_colors |
Alternating row colors, a two-item list. |
feature_table.prevalence_cmap |
Colormap for prevalence cells. |
feature_table.input |
Optional curated feature-importance CSV/TSV/Excel. The edited annotations and row order control feature-table rendering. Use an exported table as the starting point. |
Quote hex colors because unquoted # begins a YAML comment. Use valid
Matplotlib colormaps or the supplied phipml_blue_gray_red,
phipml_purple_gray_orange, and phipml_prevalence maps. Annotation values in
these examples are invented and have no biological meaning.
Supplying plotting data explicitly
These optional keys are useful for older artifacts or another controlled data source. They are not needed for the included recipes.
| Key | What to supply |
|---|---|
features_table |
Sample-by-feature CSV/TSV/Excel, including sample_column. Clinical values must be numeric and use the same encodings as fitting. Include the original feature values, not SHAP values. |
target_table |
Table containing sample IDs and the outcome column. Use already encoded 0/1 targets for the explicit table route. |
sample_column |
Identifier column for these tables; default SampleName. |
target_column |
Outcome column in target_table, or in features_table if no separate target table is supplied. Required with an explicit target table. |
feature_importance_table |
Top-level alternative to feature_table.input. Use the exported CSV structure, including its Feature column. |
Saved sample IDs determine alignment. The supplied model-config reconstruction route also checks reconstructed labels against saved targets. This check cannot prove all feature measurements are unchanged; preserve input data and hashes.
Common plotting commands
From the tutorial root:
# Change appearance/format without retraining.
phipml-plot --plot-config configs/plots/random-forest/02_noisy.yaml \
--formats svg png --dpi 300 --max-display 8 \
--output-dir plots/custom_noisy --output-prefix presentation
# Only performance panels; original feature data are not required.
phipml-plot results/random-forest/02_noisy/nested_random-forest_02_noisy_420.joblib \
--split train --plots performance --no-reconstruct-data \
--class-labels Control Case --output-dir plots/performance_only
# External model evaluation.
phipml-plot --plot-config configs/plots/xgboost/06_external_noisy_tuned.yaml
# Regenerate a compact table after manually editing a COPY of its exported CSV.
phipml-plot --plot-config configs/plots/random-forest/02_noisy.yaml \
--plots feature-table \
--feature-importance-table plots/random-forest/02_noisy/02_noisy_feature_importance.csv \
--output-dir plots/curated_table --max-display 8
Editing a feature table changes presentation; it does not revise SHAP values or model predictions. Display a curated subset transparently.
Comparison heatmaps
phipml-heatmap takes a manifest, not a plotting YAML. Each row identifies
an evaluation artifact:
training,validation,path,split
RF noisy default,Internal CV,../results/random-forest/02_noisy/nested_random-forest_02_noisy_420.joblib,train
RF noisy default,External,../results/random-forest/05_external_noisy/validation_random-forest_05_external_noisy_420.joblib,test
The path is relative to the manifest file. The training value labels a heatmap
column and validation labels a row; labels can describe experiment conditions
as in the supplied manifests. Multiple rows for one cell are aggregated as
repeated run estimates. Use explicit files, not globs, in the manifest.
phipml-heatmap --manifest manifests/random-forest.csv \
--metric roc.auc --palette viridis --vmin 0 --vmax 1 --dpi 180 \
--title "Random Forest: internal and external performance" \
--output plots/random-forest/roc_auc_overview
phipml-heatmap --manifest manifests/xgboost.csv \
--metric pr.ap --palette viridis --vmin 0 --vmax 1 --dpi 180 \
--output plots/xgboost/ap_overview
| Heatmap argument | Meaning |
|---|---|
--manifest |
Required CSV/TSV/TXT with training, validation, path, and optional split. |
--metric |
Dotted metric key, e.g. roc.auc, pr.ap, classification.f1, classification.balanced_accuracy, or classification.mcc. |
--training-order |
Optional ordered list of column labels; must include all manifest training labels. |
--validation-order |
Optional ordered list of row labels. |
--order |
Legacy shared order for a square matrix; do not combine with the two independent orders. Usually unnecessary. |
--title |
Figure title. |
--palette |
Matplotlib/seaborn colormap name; viridis here. |
--vmin, --vmax |
Color-scale bounds. Use 0–1 for this overview, −1–1 for MCC. The CLI defaults are 0.5–1.0, which can obscure below-0.5 values. |
--output |
Required filename stem; requested format extensions are added/replaced. |
--formats |
One or more of pdf svg png; defaults to all three. |
--dpi |
Raster resolution. |
--annotate-uncertainty / --no-annotate-uncertainty |
Show/hide available native intervals or variability. Default shows them. |
The rectangular overview uses only conditions that exist. An internal-CV cell and an external cell summarize different evaluation procedures; their numbers are not a paired statistical test.
Interpreting uncertainty and SHAP honestly
| Evaluation | Displayed uncertainty | What it describes |
|---|---|---|
| One nested-CV artifact | Across-outer-fold variability; curve bands are mean ±1 sample SD. | Sensitivity to the fold partition. It is not a formal 95% confidence interval. |
| Multiple seed artifacts | 2.5–97.5% empirical interval across run estimates. | Variation across repeated fits on reused participants; not independent-cohort confidence. |
| One external artifact with bootstrap | Class-stratified paired bootstrap interval at the requested confidence level. | Sampling variability of the external observations conditional on the fitted model; it excludes refitting/tuning uncertainty. |
The perfect synthetic scenario can have a 1.0–1.0 bootstrap interval because every resampled set remains perfectly separated. This is a property of the constructed demonstration, not evidence that a clinical model has no uncertainty.
Positive SHAP contributions move the model output toward the positive class; negative values move it away. Random Forest tree SHAP is in its probability output scale here, while XGBoost's default tree SHAP explains the raw margin (log-odds for the binary logistic objective). Do not compare their numerical SHAP magnitudes directly or describe XGBoost values as probability-point changes. SHAP describes a fitted predictive relationship, not causation or a p-value. Correlated engineered peptides may share importance unevenly, and independent noise can acquire apparent importance in a finite sample.
Source: the pinned repository's plot CLI, result aggregation, and heatmap CLI.
Reference execution and observed results
Run date: 8 September 2026. Source: phipml 4.2.0, commit
932f14ffca41ba4f24f3cc570afea27265a5893c.
The upstream source checkout was not modified.
Environment and budget
The run installed the pinned source into an isolated Linux Python 3.12.13
environment with the dependency ranges declared in pyproject.toml. Main
versions: numpy 1.26.4, pandas 2.2.3, scipy 1.15.3, scikit-learn 1.5.2,
scikit-optimize 0.10.2, XGBoost 2.1.4, SHAP 0.46.0, matplotlib 3.10.9,
and seaborn 0.12.2. The full platform-specific inventory is in
reports/requirements-reference.txt; it is distinct from the repository's
Micromamba specification. The Micromamba, Docker, and Apptainer installation
routes were inspected in upstream documentation, not separately executed here.
Data seed 20260908; model seed 420. Three outer folds, two inner folds, 12 candidates in each tuned search, one worker at each level, 500 external bootstrap draws, confidence level 0.95, classification threshold 0.5.
The complete twelve-scenario command, including per-scenario plots, ran in approximately five minutes on this environment, excluding dependency installation. This is an observation, not a runtime promise. The model-event log spans 4 minutes 36 seconds, with the final plot/heatmap/export work afterward. No outcome-dependent revisions were made to the generated data or tuning space.
What was run
python scripts/generate_examples.py
bash scripts/run_all.sh both
bash scripts/run_repeated.sh random-forest
phipml -c configs/models/random-forest/08_reuse_noisy.yaml
phipml -c configs/models/xgboost/08_reuse_noisy.yaml
python scripts/summarize_results.py
The main run covers 12 estimator/scenario combinations, each producing nine figure types in PDF/SVG/PNG plus the feature table CSV. Both estimator heatmaps were generated for ROC-AUC and AP. The optional repeated-run example was executed for Random Forest; the parallel XGBoost repeat recipe is supplied but was not separately executed. Pretrained model reuse was executed for both.
Metrics
Internal entries show mean ± sample SD across outer folds. External entries show the point estimate followed by the 95% stratified bootstrap interval. These are different uncertainty summaries and should not be read as equivalent 95% intervals. All participants here are synthetic.
| Estimator | Scenario | N evaluated | ROC-AUC | Average precision |
|---|---|---|---|---|
| random-forest | 01_perfect | 160 | 1.000 ± 0.000 | 1.000 ± 0.000 |
| random-forest | 02_noisy | 160 | 0.912 ± 0.028 | 0.908 ± 0.030 |
| random-forest | 03_noisy_tuned | 160 | 0.911 ± 0.038 | 0.915 ± 0.044 |
| random-forest | 04_external_perfect | 80 | 1.000 [1.000, 1.000] | 1.000 [1.000, 1.000] |
| random-forest | 05_external_noisy | 80 | 0.810 [0.709, 0.898] | 0.843 [0.769, 0.914] |
| random-forest | 06_external_noisy_tuned | 80 | 0.806 [0.705, 0.894] | 0.805 [0.708, 0.900] |
| xgboost | 01_perfect | 160 | 1.000 ± 0.000 | 1.000 ± 0.000 |
| xgboost | 02_noisy | 160 | 0.917 ± 0.038 | 0.921 ± 0.045 |
| xgboost | 03_noisy_tuned | 160 | 0.926 ± 0.030 | 0.928 ± 0.037 |
| xgboost | 04_external_perfect | 80 | 1.000 [1.000, 1.000] | 1.000 [1.000, 1.000] |
| xgboost | 05_external_noisy | 80 | 0.827 [0.728, 0.920] | 0.825 [0.730, 0.931] |
| xgboost | 06_external_noisy_tuned | 80 | 0.808 [0.706, 0.906] | 0.804 [0.707, 0.919] |
Tuning did not improve external AP for either estimator in this run. The external population has a deliberately weaker signal. These results illustrate why external data must stay out of model selection; they are not evidence that tuning is generally unhelpful or that one estimator is superior for real PhIP-seq data.
The CSV additionally records pooled-score ROC-AUC/AP, accuracy, and F1. For internal evaluation, pooled-score metrics are not necessarily equal to the mean of the fold metrics. Prediction CSVs use the pooled sample-level scores.
Verification performed
- All generated CSVs match their recorded SHA-256 hashes.
- Training and external sample IDs are disjoint in both datasets.
- Every full-cohort artifact contains exactly the 160 intended training labels.
- Every nested artifact covers all 160 intended training samples, with each sample assigned to exactly one outer validation fold per run.
- Every external artifact covers exactly the 80 intended external samples.
- Saved labels agree with the input metadata, probabilities are finite/in range, and SHAP matrices contain finite values.
- Every core recipe generated all 27 requested figure files and its feature CSV.
- Representative performance, feature-table, and XGBoost SHAP figures were visually inspected for legibility and clipping.
- Reusing the saved full-cohort tuned model reproduced scenario 06 predictions and SHAP values for both estimators.
- Plotting was exercised from a relocated folder using an explicit local model YAML, with deliberately invalid original embedded paths in a temporary test artifact. Performance, SHAP beeswarm, and feature-table reconstruction passed.
Detailed checks are recorded in reports/verification_checks.json. The normal
training/plotting logs are in logs/. The model/config generator, inputs,
search spaces, artifacts, plot recipes, exported predictions, and observed
metrics are included so that the walkthrough can be independently rerun.
Scope of interpretation
The perfect data intentionally encode their target in engineered predictors. All annotations are simulated. The noisy scenario is a controlled learning exercise, not a biological benchmark. The current phipml prevalence filter is applied before outer CV on the training cohort; it is inactive here at 0–100%. The current CLI uses sample-level stratified folds, so participant-group cross-validation would require additional support for repeated-measure studies.
These notes describe observed execution and current limitations; optional longer-search, clinical-only, and per-sample-file recipes are explanatory examples rather than additional completed benchmark runs.