Deep intronic variants are noncoding changes located more than 100 base pairs from the nearest exon–intron junction that can disrupt splicing or other regulatory functions, causing disease even when standard exome sequencing returns negative. When a patient's phenotype strongly suggests a genetic cause and exome sequencing (ES) is unrevealing, the immediate next step is whole-genome sequencing (WGS) paired with RNA-seq from an appropriate tissue or surrogate, with candidate variants prioritized using SpliceAI delta scores and PDIVAS ensemble scoring, filtered against gnomAD population frequencies, and cross-referenced against ClinVar for any prior clinical submissions. Long-read sequencing platforms, specifically PacBio and Oxford Nanopore, add resolution in repetitive or structurally complex loci where short reads fail. For cases requiring transcript-level proof or therapeutic translation, patient-derived iPSC models and antisense oligonucleotide (ASO) rescue experiments close the gap between a computational prediction and a clinically actionable finding.
Immediate triage checklist when you suspect a deep intronic cause:
- Confirm ES was performed with adequate intronic coverage; most exome captures extend only 10–20 bp into flanking introns.
- Order WGS if not already done; annotate all intronic variants with SpliceAI and PDIVAS, and filter to gnomAD allele frequency below 0.01% for rare Mendelian disease.
- Select RNA source before ordering: blood is accessible but misses tissue-restricted transcripts; fibroblasts, urinary epithelial cells, or disease-relevant iPSC-derived cells are often better surrogates.
- Flag candidates with SpliceAI delta scores above 0.2 (high sensitivity) or above 0.5 (higher specificity) for experimental follow-up.
- Submit validated findings to ClinVar with full evidence documentation to support community curation.
Pro Tip: Run SpliceAI and PDIVAS in parallel from the start. PDIVAS ensembles splice predictions with gene constraint metrics and typically narrows a genome-wide candidate list to roughly 27 actionable sites, which is a number your RNA team can realistically validate.
Key Takeaways
Deep intronic variants require a multi-modal detection and validation strategy: WGS provides the genomic data, RNA-seq confirms the transcript consequence, and functional assays establish pathogenicity with the rigor needed for clinical reporting and therapeutic translation.
| Point | Details |
|---|---|
| Definition and threshold | Deep intronic variants are noncoding changes more than 100 bp from the exon–intron junction; pseudo-exon inclusion is the most common disease mechanism. |
| Triage to roughly 27 candidates | PDIVAS ensemble scoring reduces a genome-wide candidate list to a manageable number of actionable sites per individual at high sensitivity. |
| RNA source selection matters | Blood RNA misses tissue-restricted transcripts; fibroblasts, urinary cells, or iPSC-derived cell types are often required for accurate splicing analysis. |
| ClinVar requires critical appraisal | ClinVar entries can be reclassified; always integrate gnomAD frequency, transcript evidence, and functional data before finalizing a pathogenicity call. |
| ASO rescue is dual-purpose | A successful ASO rescue experiment in patient iPSC-derived cells supports PS3 classification and serves as early evidence for a therapeutic program. |
Table of Contents
- What qualifies as a deep intronic variant and how does it cause disease?
- Clinical examples where deep intronic variants were the actual cause
- How to detect deep intronic variants: sequencing and RNA strategies
- Computational prediction: which tools to use and where they fall short
- Experimental validation: from RT-PCR to iPSC models
- How to interpret and report deep intronic variants clinically
- A stepwise diagnostic workflow for suspected deep intronic causes
- Common pitfalls and red flags when working with deep intronic variants
- Patient iPSC modeling, ASO screening, and the path to therapy
- Selected resources for tools, databases, and protocols
- The diagnostic gap is real, but the tools to close it now exist
- Sources
What qualifies as a deep intronic variant and how does it cause disease?
The term "deep intronic" lacks a single universal cutoff, and that ambiguity matters in practice. Published literature has used thresholds of more than 20 bp, more than 30 bp, and more than 100 bp from the exon–intron boundary. For clinical diagnostic workflows, the greater-than-100-bp threshold is the most defensible working definition: it excludes the canonical splice-site region (positions ±1 and ±2) and the extended splice region (roughly ±3 to ±8 for donors, ±3 to ±20 for acceptors) that standard ES annotation already captures, and it focuses attention on variants that require dedicated detection strategies.
A landmark review documented pathogenic deep intronic mutations across more than 75 disease-associated genes, with pseudo-exon inclusion as the dominant mechanism. The core molecular mechanisms fall into five categories:
- Pseudo-exon activation (exonification): A variant creates or strengthens a cryptic donor or acceptor site deep within an intron, causing the spliceosome to recognize a new exon. The inserted sequence typically introduces a frameshift or premature stop codon.
- Cryptic splice-site creation or strengthening: Single nucleotide changes, particularly GT or AG dinucleotide conversions at positions that mimic canonical splice signals, are the most common trigger. Pan-cancer WGS and RNA-seq analyses confirm that cryptic GT/AG creation and enhancer gain or silencer loss are the dominant molecular signatures of intronic mis-splicing mutations.
- Splicing enhancer or silencer disruption: Variants that destroy an intronic splicing silencer (ISS) or create an intronic splicing enhancer (ISE) shift the balance of exon recognition without necessarily touching a splice site directly.
- Branchpoint or polypyrimidine tract disruption: Variants in the branchpoint consensus (roughly 18–40 bp upstream of the acceptor) or the polypyrimidine tract impair U2 snRNP recognition and can cause exon skipping or intron retention.
- Noncoding RNA and regulatory element effects: Some deep intronic variants fall within intronic regulatory elements or noncoding RNA genes, altering transcription factor binding, chromatin accessibility, or miRNA processing rather than splicing per se.
The mechanism predicts the best validation assay. A predicted pseudo-exon inclusion is most directly confirmed by RT-PCR with primers flanking the expected insertion. A predicted enhancer disruption may require minigene assays with systematic mutagenesis to isolate the causal motif.
Clinical examples where deep intronic variants were the actual cause
Concrete disease contexts help map abstract mechanisms to realistic diagnostic scenarios.
- Duchenne muscular dystrophy (DMD): Multiple pathogenic deep intronic variants in DMD activate pseudo-exons within large introns, inserting out-of-frame sequence and truncating dystrophin. These are missed entirely by standard exome capture and require muscle RNA or surrogate cell RNA to detect the aberrant transcript.
- Fabry disease (GLA): Deep intronic variants create cryptic splice sites that insert intronic sequence into the GLA transcript, reducing alpha-galactosidase A activity. Because GLA is X-linked, even heterozygous females can present clinically, making detection in a negative exome a meaningful diagnostic gap.
- Androgen insensitivity syndrome (AR): Regulatory-element disruption deep within AR introns has been reported to reduce receptor expression without altering the coding sequence, a mechanism that standard sequencing panels miss entirely.
- Hereditary cancer syndromes: Long-read DNA and cDNA sequencing identified pathogenic deep intronic variants in BRCA1, PALB2, and ATM in previously unsolved families, with pseudo-exon inclusion confirmed at the transcript level. These findings would have remained invisible to exome-first or short-read WGS workflows without paired RNA analysis.
- Other Mendelian diseases: Pathogenic deep intronic variants have been reported in CFTR (cystic fibrosis), OPA1 (dominant optic atrophy), NF1 (neurofibromatosis type 1), and ABCA4 (Stargardt disease), among others, consistently via pseudo-exon or cryptic splice-site mechanisms.
One important caveat: pathogenic deep intronic variants are rare relative to the total number of intronic variants in any genome. Most intronic variants are benign. Confirmed pathogenic calls in the literature almost universally required transcript-level evidence or functional assay data to move beyond computational prediction.
How to detect deep intronic variants: sequencing and RNA strategies
Exome sequencing captures roughly 1–2% of the genome and typically extends only 10–20 bp into flanking intronic sequence. That design choice is deliberate and efficient for coding variants, but it means the vast majority of intronic sequence is simply not interrogated. WGS is the standard escalation when ES is negative and clinical suspicion remains high, because it provides uniform coverage across the full genome, including deep intronic regions. Combining multiple detection approaches including WGS, RNA-seq, and targeted functional assays substantially increases diagnostic yield over exome-first strategies alone.
RNA-seq from patient tissue adds a complementary dimension: instead of inferring splicing from DNA sequence, you observe the transcript directly. Paired DNA and RNA increases confidence because a variant that both alters a predicted splice site and produces an aberrant transcript in the expected tissue is far more convincing than a computational prediction alone. The critical constraint is tissue selection. Many disease-relevant genes are not expressed in blood at levels sufficient for reliable splicing analysis. Fibroblasts, urinary epithelial cells, and iPSC-derived cell types are commonly used surrogates, each with their own expression profile limitations.
Long-read sequencing on PacBio or Oxford Nanopore platforms resolves two problems that short reads cannot. First, reads spanning entire introns can phase a deep intronic variant with nearby coding variants on the same allele, which is essential for compound heterozygosity assessment. Second, long-read cDNA sequencing directly reveals pseudo-exon inclusion and complex splice isoforms that short-read RNA-seq may misassemble or miss entirely. Adaptive sampling on Oxford Nanopore allows targeted enrichment of specific loci without a separate capture step, which is useful for hard-to-sequence regions.
| Test | Key Benefits | Limitations | Typical Turnaround | Ideal Use Case |
|---|---|---|---|---|
| Exome sequencing (ES) | Cost-efficient, high coding coverage | Misses >100 bp intronic variants | 4–8 weeks | First-tier Mendelian workup |
| Whole-genome sequencing (WGS) | Full intronic coverage, structural variants | Higher cost, large data volume | weeks | Negative ES with strong phenotype |
| Short-read RNA-seq | Detects aberrant transcripts, quantitative | Tissue-restricted; misses complex isoforms | weeks | Paired with WGS for candidate confirmation |
| Long-read DNA (PacBio/ONT) | Phases variants, resolves complex loci | Higher cost, lower throughput | weeks | Repetitive regions, compound het phasing |
| Long-read cDNA (PacBio/ONT) | Full isoform resolution, pseudo-exon detection | Requires RNA; tissue constraints apply | 4–8 weeks | Complex splicing, short-read RNA-seq failure |
| Targeted intronic panel | Focused, cost-effective for known loci | Misses novel deep intronic sites | 3–5 weeks | Known disease gene with suspected intronic cause |
Pro Tip: For next-generation sequencing in ultra-rare disease, always collect and bank RNA at the time of initial sample collection. Retrospective RNA extraction from stored DNA is impossible, and a second biopsy or blood draw adds weeks of delay to an already urgent diagnostic process.
Computational prediction: which tools to use and where they fall short

A single human genome contains on the order of 1.5 million deep intronic variants. No laboratory can validate all of them. Computational prioritization is not optional; it is the only way to reduce that number to a testable set. PDIVAS reported an average precision of 0.92 and, at a recommended sensitivity threshold, extracts roughly 27 candidate pathogenic sites per genome, which is a number a molecular lab can realistically work through.
The main tools and their outputs:
- SpliceAI: A deep-learning model trained on GTEx splicing data that returns delta scores for donor gain, donor loss, acceptor gain, and acceptor loss at each position. Scores above 0.2 are commonly used as a high-sensitivity filter; scores above 0.5 as a higher-specificity threshold. SpliceAI is powerful but is a black-box model with limited mechanistic interpretability, meaning it tells you that splicing is likely affected but not always how or why.
- Pangolin: A related deep-learning model that predicts tissue-specific splicing changes across multiple GTEx tissues, adding tissue context that SpliceAI lacks.
- PDIVAS: An ensemble predictor that combines splice-prediction scores (including SpliceAI) with gene constraint metrics from gnomAD (LOEUF, pLI) and other features to estimate pathogenicity specifically for deep intronic variants causing aberrant splicing. Its practical value is in reducing the candidate list rather than in mechanistic explanation.
- MaxEntScan: A maximum entropy model for canonical splice-site strength. Less useful for deep intronic sites but valuable for scoring the cryptic site itself once a candidate pseudo-exon is identified.
- ConSpliceML: A constraint-based model that integrates regional splicing constraint to flag variants in intolerant intronic regions.
Interpretability limitations are a genuine concern with deep-learning splicing models. When a black-box model flags a variant, clinicians often want to know the predicted mechanism (cryptic donor creation vs. enhancer loss) to select the right validation assay. Ensemble models that output mechanism type alongside a pathogenicity score give labs a clearer experimental roadmap.
Population frequency from gnomAD is a non-negotiable filter at every stage. Gene constraint metrics (low LOEUF scores indicating intolerance to loss-of-function) add a second layer: a high-scoring SpliceAI variant in a highly constrained gene deserves more urgent follow-up than the same score in a tolerant gene.
Statistic to anchor your prioritization: PDIVAS reduces a genome-wide deep intronic candidate list to a small, manageable number of sites per individual at a threshold that maintains high sensitivity, compared with hundreds or thousands of candidates from SpliceAI alone at a delta score cutoff of 0.2.
Experimental validation: from RT-PCR to iPSC models
Computational prediction establishes biological plausibility. Experimental validation establishes evidence. The validation ladder runs from simple and fast to complex and definitive:
- Allele-specific RT-PCR: Design primers flanking the predicted pseudo-exon or cryptic splice site. A new band of the expected size in patient RNA but not in controls is strong evidence of aberrant splicing. This is the fastest and cheapest first step, typically achievable in days from a banked RNA sample.
- Targeted RNA-seq: Quantitative and unbiased, RNA-seq confirms the aberrant transcript, measures its abundance relative to the canonical transcript, and detects unexpected secondary splicing changes. It is the preferred method when the predicted mechanism is uncertain or when multiple candidates need simultaneous evaluation.
- Minigene assays: A genomic fragment containing the variant (and flanking exons as splice-site anchors) is cloned into an expression vector and transfected into a cell line. Splicing of the minigene is then compared between wild-type and mutant constructs. Minigenes are particularly useful when patient RNA is unavailable or when you need to confirm that the variant itself, rather than a linked variant, drives the splicing change.
- Patient-derived primary cells or iPSCs: For tissue-restricted genes, iPSC differentiation into the relevant cell type (cardiomyocytes, neurons, hepatocytes) provides the most physiologically relevant splicing context. This is the gold standard for tissue-specific splicing validation and is required when blood or fibroblast RNA gives a false negative.
- ASO rescue experiments: Antisense oligonucleotides designed to block the cryptic splice site or pseudo-exon inclusion restore canonical splicing. A rescue experiment that corrects the aberrant transcript is both functional proof that the variant drives the splicing change and an early translational signal that an ASO-based therapy could work. This is the most compelling single piece of functional evidence for pathogenicity.
Each assay has a cost-benefit profile. RT-PCR is fast (days) and cheap but requires RNA from an expressing tissue. Minigene assays take 2–4 weeks and can be done in any cell line but may not replicate tissue-specific splicing regulation. iPSC differentiation takes months and is expensive but provides the most disease-relevant context. ASO rescue adds 4–8 weeks on top of the iPSC work but generates therapeutic-grade evidence.
For ACMG/ClinGen classification, transcript-level evidence (RT-PCR or RNA-seq showing the predicted aberrant transcript) typically supports PS3 (functional evidence of pathogenicity) when combined with appropriate controls. A minigene result alone is generally coded as PM3 or PS3 at a lower weight without patient-tissue confirmation.

Pro Tip: Always include a positive control (a known pathogenic splice-site variant in the same gene, if available) and a negative control (a synonymous intronic variant with no predicted splice effect) in every RT-PCR and minigene experiment. Without controls, reviewers and curators cannot assess assay sensitivity.
How to interpret and report deep intronic variants clinically
Clinical interpretation of deep intronic variants requires integrating multiple evidence streams, none of which is sufficient alone. The ACMG/AMP framework provides the structure; the challenge is mapping non-coding evidence onto evidence codes designed primarily for coding variants.
Key integration principles:
- Computational evidence (PP3/BP4): SpliceAI or PDIVAS scores above validated thresholds support PP3 (pathogenic supporting). Scores below threshold support BP4 (benign supporting). Neither is sufficient alone.
- Population frequency (BA1/BS1/PM2): A gnomAD allele frequency above 5% is BA1 (benign stand-alone). Frequency above the disease-specific threshold is BS1. Absent or ultra-rare frequency supports PM2 (pathogenic moderate) when combined with other evidence.
- Transcript evidence (PS3): RT-PCR or RNA-seq demonstrating the predicted aberrant transcript in patient tissue, with appropriate controls, supports PS3 at moderate to strong weight depending on assay rigor.
- Functional assay (PS3): Minigene confirmation or ASO rescue elevates PS3 weight further, particularly when patient tissue RNA is unavailable.
- Segregation (PP1/BS4): Co-segregation with disease in affected family members supports PP1; absence in affected members supports BS4.
ClinVar is an essential cross-reference, but treat its entries as supporting evidence rather than ground truth. Submission heterogeneity is real: the same variant may carry conflicting interpretations from different submitters, and reclassification rates are non-trivial as new population data or functional evidence accumulates. Periodic review of ClinVar entries for variants you have classified is good laboratory practice, not optional housekeeping.
Reporting language for a suspected splice-altering deep intronic variant should state: the variant's location and predicted mechanism, the computational tools and scores used, the population frequency from gnomAD, any available transcript or functional evidence, the resulting ACMG classification, and the recommended follow-up (RNA testing in a specific tissue, family segregation, or functional assay). Avoid language that implies certainty when transcript evidence is absent.
Practical lab checklist before submitting a variant to ClinVar:
- Document sample provenance: tissue type, collection method, RNA quality metrics (RIN score).
- State tissue selection justification: why this tissue was chosen and what its expression limitations are.
- Include assay details: primer sequences, RT-PCR conditions, RNA-seq pipeline and version.
- Link raw evidence: VCF, BAM files, gel images, or RNA-seq output.
- State the ACMG classification with explicit evidence codes and their weights.
- Recommend clinical actions: confirmatory testing, family studies, or specialist referral.
A stepwise diagnostic workflow for suspected deep intronic causes
The workflow below is designed for a clinical or research laboratory that has already obtained a negative or inconclusive ES result and has a patient with a phenotype that strongly implicates a specific gene or pathway. For a broader variant assessment framework, the stepwise logic applies across variant classes.
- Phenotype-driven gene review: Confirm the clinical diagnosis and generate a ranked gene list. Review ES data for variants of uncertain significance in candidate genes, including synonymous and near-splice-site variants that may have been deprioritized.
- WGS ordering and annotation: Order WGS if not already performed. Annotate all intronic variants in candidate genes with SpliceAI and PDIVAS. Apply gnomAD frequency filter (allele frequency below 0.01% for rare Mendelian disease). Flag variants with SpliceAI delta score above 0.2 or PDIVAS score above the recommended sensitivity threshold.
- RNA-seq or targeted RT-PCR: Select the most appropriate tissue or surrogate. Order RNA-seq or design RT-PCR primers for top candidates. Interpret results in the context of expected transcript abundance and tissue-specific splicing patterns.
- Minigene or iPSC functional workup: For candidates with positive RNA evidence, proceed to minigene confirmation if patient tissue RNA is limited or if the mechanism needs isolation from linked variants. Escalate to iPSC differentiation when tissue-specific splicing context is required or when therapeutic translation is under consideration.
- ASO rescue and therapeutic discussion: If iPSC data confirms aberrant splicing, design ASOs targeting the cryptic splice site or pseudo-exon. A successful rescue experiment supports pathogenic classification and opens a conversation about compassionate-use ASO programs or gene therapy feasibility.
| Workflow step | Approximate timeline | Relative cost | Common bottleneck |
|---|---|---|---|
| WGS + bioinformatic annotation | weeks | Moderate | Sequencing queue, data storage |
| RNA-seq (short-read) | weeks | Low–moderate | Tissue availability, RNA quality |
| RT-PCR validation | 1–2 weeks | Low | Primer design, RNA access |
| Minigene assay | weeks | Moderate | Cloning, cell line selection |
| iPSC generation and differentiation | 4–9 months | High | Reprogramming efficiency, differentiation protocol |
| ASO design and rescue | 4–8 weeks (after iPSC) | Moderate–high | ASO synthesis lead time |
Sample handling notes: Collect PAXgene or Tempus RNA tubes at the time of the clinical visit and store at minus 80°C immediately. RNA degrades rapidly at room temperature; a sample collected but not stabilized within 30 minutes is often unusable for splicing analysis. Obtain informed consent for functional assays and data sharing at the time of initial sample collection, not retrospectively. Institutional review board approval is required for iPSC generation and for sharing cell lines or sequence data through public repositories.
Common pitfalls and red flags when working with deep intronic variants
The failure modes in this field are predictable, and most are avoidable with the right controls.
- Mapping artifacts in repetitive regions: Short-read aligners frequently misplace reads in low-complexity or segmentally duplicated regions, generating false-positive variant calls that do not exist in the genome. Any deep intronic candidate in a known "dark" region of the genome should be confirmed by long-read sequencing before proceeding to RNA validation. Long-read sequencing is specifically designed to resolve these loci.
- Tissue expression mismatch: Blood RNA is convenient but misses tissue-restricted transcripts. A gene expressed primarily in neurons or cardiomyocytes will show little or no transcript in blood, producing a false-negative RT-PCR result that could lead you to abandon a real pathogenic variant. Always check GTEx expression data for the candidate gene before selecting your RNA source.
- gnomAD coverage gaps: gnomAD coverage is not uniform across intronic regions. Some deep intronic positions have low read depth in the gnomAD dataset, meaning an apparently absent variant may simply be uncovered rather than genuinely rare. Check the gnomAD coverage track before interpreting a variant as absent from the population.
- Population stratification: A variant absent from gnomAD overall may be present at moderate frequency in a specific population not well represented in gnomAD. Always check population-stratified frequencies, not just the global allele frequency.
- ClinVar submission heterogeneity: A ClinVar entry labeled "pathogenic" from a single submitter without functional evidence is not the same as a consensus classification with multiple independent submissions and RNA data. Reclassification from pathogenic to VUS or benign occurs when larger population datasets or new functional evidence contradicts the original call. Never use a single ClinVar submission as the sole basis for a clinical report.
- Implausible pseudo-exon size: A predicted pseudo-exon smaller than 50 bp or larger than 300 bp should raise suspicion. Most genuine pseudo-exons fall within a biologically plausible size range; extreme predictions often reflect model artifacts rather than real splicing events.
- Strong prediction in a phenotypically mismatched gene: A high SpliceAI score in a gene with no established connection to the patient's phenotype is almost always a false lead. Phenotype-gene concordance is a prerequisite for experimental follow-up, not an afterthought.
Statistic to keep in mind: PDIVAS reduces a genome-wide candidate list to a small, manageable number of sites per individual, but even at that scale, most candidates will be benign. Experimental validation is the only way to distinguish a real pathogenic variant from a high-scoring false positive.
Patient iPSC modeling, ASO screening, and the path to therapy
Transcript-level evidence from patient blood or fibroblasts is often the first validation milestone, but it is rarely the last one needed for therapeutic translation. Patient-derived iPSC models close the gap between a splicing observation in a surrogate tissue and a mechanistic proof in a disease-relevant cell type.
The workflow for iPSC-based validation:
- iPSC generation: Reprogram patient somatic cells (typically peripheral blood mononuclear cells or fibroblasts) into iPSCs using episomal or mRNA-based reprogramming factors. Quality control includes karyotyping, pluripotency marker confirmation, and mycoplasma testing.
- Isogenic CRISPR controls: Correct the deep intronic variant in the patient iPSC line using CRISPR-Cas9 with a single-stranded oligonucleotide donor. The corrected isogenic line serves as the ideal control because it is genetically identical to the patient line except at the variant of interest, eliminating background genetic noise.
- Differentiation into disease-relevant cell types: Differentiate iPSCs into the cell type where the gene is expressed and the disease mechanism operates (cardiomyocytes for cardiac genes, motor neurons for neuromuscular disease, hepatocytes for metabolic disorders). Splicing assays in differentiated cells reflect the tissue-specific regulatory environment that blood or fibroblast RNA cannot replicate.
- High-throughput ASO screening: Design a panel of ASOs targeting the cryptic splice site, the pseudo-exon itself, or nearby splicing regulatory elements. Screen the panel in iPSC-derived cells, measuring rescue of canonical splicing by RT-PCR or RNA-seq. A lead ASO that restores normal transcript ratios is both functional validation and an early therapeutic candidate.
How these results inform clinical actionability: a successful ASO rescue experiment in patient iPSC-derived cells supports PS3 at strong weight for ACMG classification, provides the mechanistic evidence needed for a compassionate-use IND application, and informs gene therapy feasibility assessment by confirming the aberrant transcript is the primary disease driver.
Practical constraints are real. iPSC generation and differentiation takes 4–9 months and requires specialized infrastructure. ASO synthesis and screening adds cost and lead time. Ethical consent for iPSC generation, cell-line banking, and data sharing must be obtained prospectively. For families and physicians who need this level of translational validation but lack in-house capacity, Hopeatrarelabs provides patient-derived iPSC modeling, CRISPR isogenic controls, high-throughput ASO screening, and gene therapy feasibility assessment as contracted services, designed specifically for ultra-rare and undiagnosed genetic diseases where no approved therapy exists.
Pro Tip: When designing ASOs for a pseudo-exon target, screen at least 5–10 different ASO sequences spanning the cryptic splice sites and the pseudo-exon body. The most effective blocking position is not always the one closest to the splice site; a systematic screen identifies lead candidates faster than single-ASO testing.

Selected resources for tools, databases, and protocols
These are the resources you will actually open during a deep intronic variant workup, organized by the question they answer.
- Population frequency and gene constraint: gnomAD provides allele frequencies, coverage tracks, and LOEUF/pLI constraint scores. Check coverage depth at your candidate position before interpreting absence as rarity.
- Variant classification and prior submissions: ClinVar aggregates clinical interpretations with supporting evidence. Filter by submitter type and evidence level; single-submitter entries without functional data carry less weight.
- Splice prediction: SpliceAI Lookup provides precomputed delta scores for SNVs genome-wide. For indels or novel variants, run the SpliceAI Python package locally.
- Ensemble pathogenicity scoring: PDIVAS is described in its publication with code and threshold guidance; apply it after SpliceAI annotation to reduce candidates to a manageable set.
- Long-read sequencing protocols: The Genome Research study on long-read cDNA sequencing in cancer-predisposing genes provides a practical protocol template for multiplexed long-read DNA and cDNA sequencing in unsolved families.
- Functional validation protocols: The pan-cancer intronic mis-splicing study describes experimental validation approaches for cryptic splice-site and enhancer/silencer variants, including minigene and RNA-seq confirmation methods.
- Integrative detection review: Leveraging multiple approaches for deep intronic variant detection summarizes the evidence for combining WGS, RNA-seq, and functional assays and is a useful reference for justifying multi-modal testing to payers or institutional review boards.
- Variant interpretation frameworks: The Hopeatrarelabs gene variant interpretation guide covers non-coding variant curation workflows and evidence integration in clinical practice.
When submitting variant evidence to ClinVar or an institutional database, always link primary literature that supports the functional assay design and interpretation. A submission with raw evidence files, assay details, and cited methodology is far more durable than one that relies on computational prediction alone.
The diagnostic gap is real, but the tools to close it now exist
The conventional advice in clinical genetics has long been to start with exome sequencing and escalate only when coding variants fail to explain the phenotype. That logic made sense when WGS was prohibitively expensive and RNA-seq required specialized infrastructure. Neither constraint holds the same way today, yet the field still underutilizes paired DNA and RNA analysis in the first diagnostic tier for patients with strong phenotype-gene concordance and a negative exome.
The deeper problem is not sequencing access. It is the interpretive bottleneck. A genome contains roughly 1.5 million deep intronic variants, and without computational triage, that number is paralyzing. The practical contribution of tools like PDIVAS is not that they replace experimental validation; it is that they make experimental validation feasible by reducing the candidate list to a size a laboratory can actually work through. That reduction, from millions to roughly 27 candidates per genome, is what turns WGS from a data-generation exercise into a diagnostic tool.
What the field still underestimates is the tissue-selection problem. Clinicians order RNA-seq from blood because blood is easy to collect, and then interpret a negative result as evidence against a splice-altering variant. For any gene with restricted expression, that interpretation is wrong. A negative blood RNA result for a neuronal or cardiac gene means nothing about splicing in the relevant tissue. The solution is not more sequencing; it is better sample planning at the time of initial collection, before the diagnostic window closes.
The translational opportunity is also underappreciated. An ASO rescue experiment in patient iPSC-derived cells is not just a validation assay. It is a proof-of-concept therapeutic experiment. For ultra-rare diseases where no approved therapy exists, that experiment can be the first step toward a compassionate-use program. The gap between "we found the variant" and "we have a therapeutic lead" is shorter than most clinicians realize, provided the right experimental infrastructure is in place from the start.
This article is general information, not a substitute for advice from a qualified doctor. Consult a qualified healthcare professional about your own circumstances before acting on anything here.
Sources
- Deep intronic mutations and human disease
- Long-read DNA and cDNA sequencing identify cancer-predisposing deep intronic variation in tumor-suppressor genes
- ClinVar
- gnomAD
