<b>Homo sapiens - Clinical-Grade Synthetic Cancer Mutation Panel of 55 Genes with Multi-Complexity Sequence Architecture Open Access Datas</b><b>et</b>
Providing a comprehensive synthetic genomic reference panel for clinical-grade diagnostic assay validation and research applications across 55 important cancer-associated genes, this project offers a full complement. The dataset consists of 100,000 pairs of normal and mutated sequences for each gene. The length of each sequence is 1,100 base pairs. All sequences are designed to represent human genomic DNA with clinically relevant cancer associated mutations, with perfect sequence identity between paired normal and mutated samples except at defined mutation sites. The panel includes 55 genes across multiple functional categories. Tumor suppressors include TP53 (missense, nonsense, frameshift, and deletion mutations), BRCA1 and BRCA2 (frameshift, nonsense, and deletion), RB1 (nonsense, frameshift, and deletion), PTEN (missense, nonsense, and frameshift), APC (nonsense and frameshift), NF1 and NF2 (nonsense, frameshift, and splice), VHL (missense, nonsense, and frameshift), CDKN2A (deletion, nonsense, and missense), TSC1 and TSC2 (nonsense and frameshift), and SMAD4 (missense, nonsense, and deletion). DNA repair and mismatch repair genes covered include MLH1, MSH2, MSH6, and PMS2 (nonsense, frameshift, and splice), along with MUTYH (missense and nonsense). Damage response genes include ATM (nonsense, frameshift, and missense), ATR (missense and nonsense), and CHEK2 (frameshift, nonsense, and missense). Chromatin modifiers and epigenetic regulators include ARID1A and ARID2 (frameshift and nonsense), SETD2 (nonsense, frameshift, and missense), and KMT2D (frameshift and nonsense). RAS pathway signaling genes include KRAS, NRAS, and HRAS with missense mutations at codons 12, 13, and 61. Receptor tyrosine kinases include EGFR (missense, deletion, and amplification marker), ERBB2 and ERBB3 (missense and amplification marker), ALK and ROS1 (missense and splice), RET (missense), NTRK1, NTRK2, and NTRK3 (missense and splice), MET (missense and splice), and FGFR1, FGFR2, and FGFR3 (missense and amplification marker). Other kinases include BRAF with V600E hotspot missense mutations, PIK3CA with E542, E545, and H1047 hotspot missense mutations, STK11 (nonsense, frameshift, and missense), and CDK4 and CDK6 (missense and amplification marker). Metabolic enzymes include IDH1 (R132 missense) and IDH2 (R140 and R172 missense). Transcription factors and signaling genes include MYC (missense and amplification marker), CTNNB1 (missense), and NOTCH1 (frameshift, nonsense, and missense). Cell cycle regulators include CCND1 (missense and amplification marker) and FBXW7 (missense and nonsense). Other critical genes included are POLE (missense) and KEAP1 (missense and nonsense). Key features of this dataset include 100,000 original and 100,000 mutated sequences per gene, totaling 11 million sequences across all 55 genes. Each sequence is 1,100 base pairs in length with greater than 99.99 percent accuracy. The dataset includes comprehensive metadata linking paired original and mutated sequences through CSV mapping files. GC content is stratified across Low, Mid, High, and VeryHigh levels ranging from 20 to 80 percent, with complexity levels varying from Low to VeryHigh. Mutations are biased 60 percent toward known cancer hotspots and 40 percent at random positions, with an Illumina-like sequencing error profile applied. The data is organized in per-gene TAR.GZ archives containing multi-FASTA files and CSV metadata, with a combined TAR.GZ file encompassing all 55 genes. Clinical applications include NGS assay validation across major cancer pathways, variant calling pipeline training and benchmarking, clinical laboratory proficiency testing, diagnostic kit development and quality control, and pan-cancer genomics research. This synthetic dataset serves as a valuable resource for bioinformatics pipeline development, machine learning model training, and clinical diagnostic assay validation.<br>Data Structure:<br><br>The dataset is provided as a single compressed TAR.GZ file (ishant.borse_sequences.genovaxis.tar.gz, 9.31 GB) containing:<br><br> 55 gene-specific directories (ALK through VHL)<br> Each gene directory contains:<br> - [GENE]_original_sequences.fasta (100,000 wild-type sequences)<br> - [GENE]_mutated_sequences.fasta (100,000 mutated sequences)<br> - [GENE]_mapping.csv (metadata linking original ↔ mutated pairs)<br> - [GENE].tar.gz (individual gene archive)<br> - master_index.csv (global metadata index)<br><br>File Format:<br> FASTA format with sequence headers containing unique identifiers<br> CSV format with fields: original_id, mutated_id.<br>The metadata is also registered as <b>BioProject</b> at <b>NCBI </b>having accession number <b>PRJNA1498227.</b>