Python script to dynamically generate the bioinfo production checklists for QC, Delivery and Close of NGI sequencing projects. The script uses Quarto to generate the checklists in HTML or markdown format. The checklists are based on templates that can be customized to fit the needs of different projects.
The templates are based on the following internal documents and versions:
- Bioinfo QC: 1617:6
- Delivery: 1286:23
- Close: 1262:18
- Python 3.10 or higher
- Python dependencies:
pip install -r requirements.txt - Quarto installed
Clone the repository and run the following minimal command in the root directory:
python generate_checklists.py --output-path . --format htmlThis will generate three checklists (i.e. QC, Delivery and Close) in HTML format in the current directory. It will also generate the corresponding .qmd files with the same base name, and place them in the qmds folder. A .qmd file is a Quarto document that can be edited and rendered to generate a new HTML file with the updated checklist.
To save an HTML checklist as a PDF while preserving the ticked checkboxes, open the HTML file in a web browser, tick the completed checklist items, and click the "Download PDF" button in the bottom-right corner (or use the browser's print dialog, e.g. Ctrl+P / Cmd+P), then choose "Save as PDF" as the destination.
To re-generate any of the checklists after having modified its .qmd file, run the following command:
quarto render qmds/<qmd_file> --to html --embed-resources --standaloneInstead, if you want to generate the checklists in markdown format, run:
quarto render qmds/<qmd_file> --to markdown --embed-resources --standalone--help: Show the help message and exit.--templates-path: The path to the template files. The default istemplates. The script will look for three files in this directory:QC_template.qmd,Delivery_template.qmd, andClose_template.qmd. These files are used to generate the QC checklist, delivery checklist, and close checklist, respectively.--format: The output format for the checklist:markdownorhtml. This is required, either via this option or theformatkey in the configuration file.--name: The project name (format:<username>_<year>_<index>). If not provided, the script will leave a generic placeholder (<project_name>) in the output file.--project: The project ID for which the QC checklist will be generated. If not provided, the script will leave a generic placeholder (<project_id>) in the output file.--flowcell: The flowcell ID for the project. If not provided, the script will leave a generic placeholder (<flowcell_id>) in the output file.--slide: The slide ID (used by the Visium best practice checklist). If not provided, the script will leave a generic placeholder (<slide_id>) in the output file.--genome-path: The path to the genome files. If not provided, the script will leave a generic placeholder (<genome_path>) in the output file.--transcriptome-path: The path to the transcriptome files. If not provided, the script will leave a generic placeholder (<transcriptome_path>) in the output file.--author: The full name of the author. If not provided, the script will not include the author in the output file.--signature: The author signature (initials) used in the running notes. If not provided, the signature placeholders are removed from the output file.--email: The email address of the author. If not provided, the script will not include the email in the output file.--instrument: The instrument type:illumina(default) oraviti.--best-practice: Set tovisiumto also generate the Visium data analysis checklist.--ngi-path: The path to the NGI folder on Miarka. This is used to generate the path to the project folders. If not provided, the script will leave a generic placeholder (<ngi_path>) in the output file.--incoming-path: The path to the incoming data folder. If not provided, the script will leave a generic placeholder (<incoming_path>) in the output file.--visium-base-path: The path to the Visium base folder (probe sets, feature references). If not provided, the script will leave a generic placeholder (<visium_base_path>) in the output file.--config-path: The path to the TACA configuration folder. This is used to generate the path to the project folders. If not provided, the script will leave a generic placeholder (<config_path>) in the output file.--genstat-url: The URL for the Genomics Status page. If not provided, the script will leave a generic placeholder (<genstat_url>) in the output file.--charon-url: The URL for the Charon page. If not provided, the script will leave a generic placeholder (<charon_url>) in the output file.--quarto-path: The path to the Quarto executable. If not provided, or if the executable is not found at the given path, the script will attempt to search forquartoin the system path.--output-path: The directory where the output files will be saved. The default is the current directory.--local-reports-path: The local path where the MultiQC and reports folders are saved. If not provided, the script will leave a generic placeholder (<local_reports_path>) in the output file.--timestamp: Whether to include a timestamp in the output file name. The default isFalse. IfTrue, the output files will be named<YYYYMMDD>_<project_id>_<Label>.{html,md,qmd}(e.g.20260127_P37871_QC.html).--output-structure: The structure of the output. The default isflat. The other option isnested. Ifnestedis selected, the output files will be saved in a subdirectory named after the timestamp and project ID.--script-assets-path: The path to the assets folder used by the templates. The default is theassetsfolder in this repository.--force: Force overwrite of existing files. If not provided and the output file already exists, the script will not overwrite it and will exit with an error message.--log-level: The logging level. The default isINFO. Other options areDEBUG,WARNING,ERROR, andCRITICAL. This can be set toDEBUGfor more detailed logging information.
If a config.json file is present in the same directory as the script, it will be used to set the default values for all or some of the options. The configuration file should be in JSON format and can include the following keys:
templates_path [string]format [string]project [string]flowcell [string]author [string]email [string]ngi_path [string]incoming_path [string]config_path [string]genstat_url [string]charon_url [string]quarto_path [string]output_path [string]timestamp [bool]output_structure [string]force [bool]log_level [string]slide [string]genome_path [string]transcriptome_path [string]instrument [string]visium_base_path [string]local_reports_path [string]script_assets_path [string]
Note:
config.jsonis gitignored and will not be tracked. To get started, copyconfig.json.exampletoconfig.jsonand fill in the values that apply to your environment. Any keys omitted fromconfig.jsonwill fall back to their command-line argument counterparts.
{
"author": "Your Name",
"email": "your.name@scilifelab.se",
"format": "markdown",
"output_path": "/path/to/output/directory",
"output_structure": "nested",
"quarto_path": "/usr/local/bin/quarto",
"ngi_path": "/path/to/NGI/folder",
"incoming_path": "/path/to/incoming/folder",
"genstat_url": "https://genomics-status.example.com",
"charon_url": "https://charon.example.com",
"config_path": "/path/to/conf/TACA",
"log_level": "INFO"
}