Downstream analysis: recommended defaults, clearer labels and decision plots - #35
Draft
t0mdavid-m wants to merge 4 commits into
Draft
t0mdavid-m wants to merge 4 commits into
t0mdavid-m wants to merge 4 commits into
Conversation
… normalization and statistics - Filtering defaults to Low Repeatability, at most 50% missing per group. - Imputation defaults to MNAR smallest value per protein (row scope). - Normalization defaults to log2 + median, no row scaling; Statistical Inference warns when its input is not log2-scaled, since log2FC is the difference of group means. - Choices are keyed widgets backed by default-parameters.json, so they persist per workspace; older workspaces fall back to the shipped defaults. - Re-applying a step drops the stored outputs of the steps after it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015XpSPmvsUELJs4GbbDXkrj
…on plots - Sidebar: "Workflow Results", "Downstream Analysis" (numbered steps 1-4) and "Downstream Plots". - Plot pages show which step's output they display; heatmaps now plot the same processed table as PCA instead of the raw abundance table. - Volcano, PCA and heatmaps key their OpenMS-Insight cache on a hash of the plotted data, so re-applying a step does not show old plots. - Each step gets a plot for choosing its setting: proteins kept across the filter threshold, missingness vs intensity and an imputation preview, per-sample distributions before/after normalization, p-value histogram. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015XpSPmvsUELJs4GbbDXkrj
…ylint E0606 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015XpSPmvsUELJs4GbbDXkrj
Hash-suffixed cache ids left a new cache directory in the app folder on every re-applied step. Use the fixed ids with regenerate_cache=True, the same change as the tables PR, so the two merge without conflicts. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015XpSPmvsUELJs4GbbDXkrj
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Requested by Tom · project thread
Before: the sidebar had "Results" and "Differential Protein Analysis", and it was unclear which table a plot showed. The heatmaps plotted the raw abundance table while PCA used the processed one. Filtering, Imputation, Normalization and Statistical Inference opened on arbitrary first options. With those options, "log2FC" was a difference of raw intensities, because the statistics function subtracts group means and assumes log2 input. Choices reset on every visit, and re-applying a step left the later steps' tables in place.
After: the sidebar reads Workflow Results, Downstream Analysis (1. Filtering, 2. Imputation, 3. Normalization, 4. Statistics) and Downstream Plots. Every plot page states which step's output it shows, and the heatmaps use the same processed table as PCA. Each step opens on a recommended setting and has a "Help me choose" plot. Every plot, including Volcano, PCA and the heatmaps, has a one- to two-sentence explanation. Choices persist per workspace, and re-applying a step clears the steps after it.
Why per-protein minimum imputation. I ran a simulation with 500 proteins, 50 true 4-fold changes and 3 vs 3 samples. With 10% of values dropping out at random, the global minimum found 0 of 50 changes, because one imputed dropout inflates a protein's variance. The per-protein minimum found 36 of 50 with no false positives. When values were missing below a detection limit, it found 49 with 1 false positive.
How. Widgets are keyed on new
postproc-*entries indefault-parameters.json.postprocessing_param()seeds keys for older workspaces.clear_downstream_steps()drops stale step outputs.get_active_table()andshow_pipeline_banner()give every plot page the same source and label. The plots live insrc/common/postprocessing_plots.py, use Plotly and go throughshow_fig; OpenMS-Insight has no histogram or box plot component. Volcano, PCA and the heatmaps passregenerate_cache=True, the same lines as #36. After #36, the only overlap is the sidebar entry inapp.py.Protein inference is unchanged: ProteomicsLFQ aggregation, picked protein FDR and unique peptides, and TMT ProteinInference with PEP/best/picked FDR. Both already match quantms' defaults.
Tested: I ran all four steps and the four plot pages in sequence with
AppTeston a synthetic LFQ workspace, including an olderparams.jsonwithout the new keys, and checked each decision plot rendered from synthetic data.pytest testspasses (88 passed, 2 skipped).pylint --errors-onlyis clean apart from the pre-existing pyopenms import notices.🤖 Generated with Claude Code
https://claude.ai/code/session_015XpSPmvsUELJs4GbbDXkrj