Skip to content

Downstream analysis: recommended defaults, clearer labels and decision plots - #35

Draft
t0mdavid-m wants to merge 4 commits into
mainfrom
claude/project-thread-mae3oz
Draft

t0mdavid-m wants to merge 4 commits into
mainfrom
claude/project-thread-mae3oz

Conversation

@t0mdavid-m

@t0mdavid-m t0mdavid-m commented Sep 27, 2026 •

Copy link
Copy Markdown
Member

Requested by Tom · project thread

Before: the sidebar had "Results" and "Differential Protein Analysis", and it was unclear which table a plot showed. The heatmaps plotted the raw abundance table while PCA used the processed one. Filtering, Imputation, Normalization and Statistical Inference opened on arbitrary first options. With those options, "log2FC" was a difference of raw intensities, because the statistics function subtracts group means and assumes log2 input. Choices reset on every visit, and re-applying a step left the later steps' tables in place.

After: the sidebar reads Workflow Results, Downstream Analysis (1. Filtering, 2. Imputation, 3. Normalization, 4. Statistics) and Downstream Plots. Every plot page states which step's output it shows, and the heatmaps use the same processed table as PCA. Each step opens on a recommended setting and has a "Help me choose" plot. Every plot, including Volcano, PCA and the heatmaps, has a one- to two-sentence explanation. Choices persist per workspace, and re-applying a step clears the steps after it.

Step Default Decision plot
1. Filtering Low Repeatability, ≤50% missing per group Proteins kept across the threshold range
2. Imputation MNAR, per-protein minimum Mean intensity of proteins with vs without gaps; where imputed values land
3. Normalization log2 + median, no scaling Per-sample box plots before/after (live preview)
4. Statistics limma-like + BH; warns if input is not log2 p-value histogram

Why per-protein minimum imputation. I ran a simulation with 500 proteins, 50 true 4-fold changes and 3 vs 3 samples. With 10% of values dropping out at random, the global minimum found 0 of 50 changes, because one imputed dropout inflates a protein's variance. The per-protein minimum found 36 of 50 with no false positives. When values were missing below a detection limit, it found 49 with 1 false positive.

How. Widgets are keyed on new postproc-* entries in default-parameters.json. postprocessing_param() seeds keys for older workspaces. clear_downstream_steps() drops stale step outputs. get_active_table() and show_pipeline_banner() give every plot page the same source and label. The plots live in src/common/postprocessing_plots.py, use Plotly and go through show_fig; OpenMS-Insight has no histogram or box plot component. Volcano, PCA and the heatmaps pass regenerate_cache=True, the same lines as #36. After #36, the only overlap is the sidebar entry in app.py.

Protein inference is unchanged: ProteomicsLFQ aggregation, picked protein FDR and unique peptides, and TMT ProteinInference with PEP/best/picked FDR. Both already match quantms' defaults.

Tested: I ran all four steps and the four plot pages in sequence with AppTest on a synthetic LFQ workspace, including an older params.json without the new keys, and checked each decision plot rendered from synthetic data. pytest tests passes (88 passed, 2 skipped). pylint --errors-only is clean apart from the pre-existing pyopenms import notices.

🤖 Generated with Claude Code

https://claude.ai/code/session_015XpSPmvsUELJs4GbbDXkrj

… normalization and statistics

- Filtering defaults to Low Repeatability, at most 50% missing per group.
- Imputation defaults to MNAR smallest value per protein (row scope).
- Normalization defaults to log2 + median, no row scaling; Statistical
  Inference warns when its input is not log2-scaled, since log2FC is the
  difference of group means.
- Choices are keyed widgets backed by default-parameters.json, so they
  persist per workspace; older workspaces fall back to the shipped defaults.
- Re-applying a step drops the stored outputs of the steps after it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015XpSPmvsUELJs4GbbDXkrj
@t0mdavid-m t0mdavid-m self-assigned this Sep 27, 2026
…on plots

- Sidebar: "Workflow Results", "Downstream Analysis" (numbered steps 1-4)
  and "Downstream Plots".
- Plot pages show which step's output they display; heatmaps now plot the
  same processed table as PCA instead of the raw abundance table.
- Volcano, PCA and heatmaps key their OpenMS-Insight cache on a hash of
  the plotted data, so re-applying a step does not show old plots.
- Each step gets a plot for choosing its setting: proteins kept across the
  filter threshold, missingness vs intensity and an imputation preview,
  per-sample distributions before/after normalization, p-value histogram.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015XpSPmvsUELJs4GbbDXkrj
@t0mdavid-m t0mdavid-m changed the title Recommended defaults for the postprocessing pipeline Downstream analysis: recommended defaults, clearer labels and decision plots Sep 27, 2026
…ylint E0606

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015XpSPmvsUELJs4GbbDXkrj
Hash-suffixed cache ids left a new cache directory in the app folder on
every re-applied step. Use the fixed ids with regenerate_cache=True, the
same change as the tables PR, so the two merge without conflicts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015XpSPmvsUELJs4GbbDXkrj
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants