readread定序儀產生的一段序列觀測。一條 read 通常來自一個 DNA 分子的片段,但不等同於一個細胞;仍需考慮重複、嵌合與定序錯誤。A sequence observation produced by the sequencer, usually a fragment from one DNA molecule. A read is molecular evidence, not a cell; duplicates, chimeras and sequencing errors still need to be considered.完整條目 →
定序儀對單一 DNA 分子局部序列的觀測
VAFVAF 變異等位基因頻率在某個位點上,支持 alt allele 的 read 佔全部 read 的比例。VAF 不等於帶有這個突變的細胞比例。Variant Allele Frequency: the fraction of reads supporting the alternate allele. Not the same as the fraction of cells carrying it.完整條目 →
在特定位點上,支持變異的 read 佔全部涵蓋該位點 read 的比例
教學用簡化區段模型
左側呈現樣本中的四個細胞群,右側呈現各群對應的 DNA 組成;下方則表示標準 bulk 定序後可觀測的 VAF。橫線區隔細胞層級組成與定序後的分子層級資料。
圖中元素的意義如下:
圖中元素
代表意義
灰色細胞
正常細胞,佔 50%
紅/藍/橘色細胞
三個腫瘤細胞群;最早形成者為 founder cloneclone 克隆源自共同祖先細胞,並共享一組可辨識 somatic mutations 的細胞群。A population of tumour cells descended from a common ancestor and sharing the same set of somatic mutations.完整條目 →,由其分出的後續細胞群為 subclonesubclone 次克隆clone 內取得額外變異並擴增的子群;治療後可能富集具有抗性的 subclone。A subset of a clone that acquired further mutations and expanded. Treatment-resistant populations are often subclones.完整條目 →
藍色長條 HP1/橘色長條 HP2
同一段染色體的兩份 DNA 拷貝;跨多個位置的一組序列差異構成 haplotypehaplotype 單倍型同一條實體染色體拷貝上,具有一致相位的一組 alleles 或 variants。HP1 與 HP2 是任意的相對標籤。A set of variants that lie on the same physical chromosome copy and are inherited together.完整條目 →
綠點
由親代遺傳、存在於多數體細胞的 DNA 差異germline variant 生殖系變異在生殖系形成、通常存在於多數細胞的變異;可遺傳給子代,但不代表必然傳遞。A variant arising in the germline and therefore present in essentially every cell of the individual. It may be inherited or arise de novo, and may be passed to offspring.完整條目 →;常見的單一核苷酸差異稱為 SNPSNP 單核苷酸多型性族群中以多型形式存在的單一鹼基差異;本教材主要以可作為 phasing 錨點的 germline SNP 為例。A single-nucleotide polymorphism observed as a population variant. This tutorial mainly uses germline SNPs as phasing anchors.完整條目 →,可協助區分 HP1 與 HP2
紅點 S1/S2/S3
受精後新產生、僅存在於部分細胞譜系的 三個體細胞變異somatic variant 體細胞變異在非生殖系細胞譜系中、通常於受精後取得的變異;可只存在於部分細胞,通常不由親代遺傳給子代。癌症基因體學常分析此類變異。A variant acquired outside the germline, usually after conception. It may be restricted to a subset of cells and is generally not passed to offspring; it is a central focus of cancer genomics.完整條目 →
locuslocus 位點基因體上的一個特定位置。講「這個 locus」時通常指某個座標,例如 chr1:162,404,043。A specific position in the genome, usually given as a chromosome and coordinate.完整條目 →
一個位置
chr1:162,404,043
alleleallele 等位基因同一個 locus 上可能出現的其中一種序列版本。在一般二倍體常染色體位點,個體通常有兩份 allele;性染色體、拷貝數變化與腫瘤基因體可能不符合這個簡化模型。One of the alternative sequence versions at a locus. At a typical diploid autosomal locus, an individual has two allele copies; sex chromosomes, copy-number changes and tumour genomes may differ.完整條目 →
該位置上其中一種序列
這個 locus 可以是 A 或 G,各是一個 allele
variant / SNVSNV 單核苷酸變異單一鹼基的改變。與 SNP 的差別在於 SNV 不預設它在族群中常見 —— 腫瘤裡新產生的單點突變叫 somatic SNV(sSNV),不叫 SNP。A single-nucleotide change. Unlike SNP, it carries no implication of population frequency; new tumour mutations are sSNVs.完整條目 →
variant 是相對參考序列的差異;SNV 是單一鹼基替換的 variant
參考為 A、樣本為 G → 一個 SNV
genotypegenotype 基因型某個 locus 上觀察到的 allele 組合;對一般二倍體常染色體位點,通常是兩份 allele。VCF 裡寫成 0/1(未定相)或 0|1(已定相)。The allele combination observed at a locus. At a typical diploid autosomal locus it has two copies; VCF writes these as 0/1 (unphased) or 0|1 (phased).完整條目 →
該位置上兩份 allele 的組合
0/1:一份跟參考一樣、一份不一樣
haplotypehaplotype 單倍型同一條實體染色體拷貝上,具有一致相位的一組 alleles 或 variants。HP1 與 HP2 是任意的相對標籤。A set of variants that lie on the same physical chromosome copy and are inherited together.完整條目 →
同一條染色體拷貝上具有一致相位的一組 alleles
HP1 = A…C…G,HP2 = T…A…C
上表以正常二倍體常染色體為模型,因此同一個 locus 通常有兩份 allele。
這種具有兩份拷貝的狀態稱為 diploiddiploid 二倍體正常人多數常染色體區段通常有兩份拷貝的狀態;性染色體、拷貝數變化與腫瘤細胞可能不同。A state in which a typical autosomal locus has two copies. Sex chromosomes, copy-number-altered regions and tumour cells may differ.完整條目 →;性染色體與腫瘤拷貝數異常將於後續模組另行處理。
關鍵差別在最後兩個:genotype 是單一位置的,haplotype 是跨越多個位置的。
genotype 說「這裡有一份 A、一份 G」,但沒說 A 跟隔壁位置的 C 是不是在同一條染色體上。
補上這個資訊,正是 phasingphasing 定相為可判定的 variants 建立其位於不同實體染色體拷貝上的相位關係;VCF 常以 0|1 等形式表示。長 read 可提供跨位點的分子層級觀測證據。Determining which variants lie on the same physical chromosome, turning 0/1 into 0|1. Long reads provide direct molecular evidence.完整條目 → 要推定的關係之一。
一個 locus 若兩份 allele 不同(heterozygousheterozygous 異型合子在一般二倍體位點上,兩份 allele 不同(例如一份 A、一份 G)。此類位點可作為區分 haplotype 的資訊錨點;複雜拷貝數情況需另行解讀。At a typical diploid locus, carrying two different allele copies (for example A and G). These sites can anchor phasing; copy-number changes require separate interpretation.完整條目 →,例如 0/1),
在讀取品質足夠的簡化模型中,可用來區分兩條 haplotype —— 看到 A 的 read 可歸到一條,看到 G 的可歸到另一條。
若兩份相同(homozygoushomozygous 同型合子在一般二倍體位點上,兩份 allele 相同。此位點本身通常無法用來區分兩條 haplotype。At a typical diploid locus, carrying two identical allele copies. The site itself usually cannot distinguish haplotypes.完整條目 →),該位點本身通常不能區分兩條 haplotype。
het 位點密度是影響 phasing 的因素之一;讀長、coverage 與錯誤率也會影響可建立的相位範圍。
腫瘤中的 LOHLOH 異型合子性喪失原本 heterozygous 的區域變成只剩一種 allele。LOH 不等於缺失 —— 也可能是一條 haplotype 遺失後另一條被複製(copy-neutral LOH)。Loss of heterozygosity: a formerly het region retains only one allele. Not necessarily a deletion.完整條目 →(loss of heterozygosity,異型合子性喪失)可能使一個區段失去其中一種 allele,
使原本可用的 heterozygous 位點減少,因而降低可用的 phasing 錨點。
germline variantgermline variant 生殖系變異在生殖系形成、通常存在於多數細胞的變異;可遺傳給子代,但不代表必然傳遞。A variant arising in the germline and therefore present in essentially every cell of the individual. It may be inherited or arise de novo, and may be passed to offspring.完整條目 →
somatic variantsomatic variant 體細胞變異在非生殖系細胞譜系中、通常於受精後取得的變異;可只存在於部分細胞,通常不由親代遺傳給子代。癌症基因體學常分析此類變異。A variant acquired outside the germline, usually after conception. It may be restricted to a subset of cells and is generally not passed to offspring; it is a central focus of cancer genomics.完整條目 →
來源
生殖系;可能由親代遺傳或 de novo 形成
非生殖系細胞譜系
分布
通常存在於多數細胞(視嵌合而定)
可限於部分細胞,也可能為 clonal
是否可傳給子代
可能,但不代表必然傳遞
通常不以遺傳方式傳遞
典型 VAFVAF 變異等位基因頻率在某個位點上,支持 alt allele 的 read 佔全部 read 的比例。VAF 不等於帶有這個突變的細胞比例。Variant Allele Frequency: the fraction of reads supporting the alternate allele. Not the same as the fraction of cells carrying it.完整條目 →
在二倍體且無污染/CNV 時,het 常接近 0.5
受 purity、copy number、multiplicity 與 clonality 影響
在本課程示例中的角色
工具 —— 可作為 phasing 的錨點
分析目標 —— 需依流程與證據判定
在特定 somatic calling 流程中,germline calls 常被排除;然而 heterozygous germline SNPSNP 單核苷酸多型性族群中以多型形式存在的單一鹼基差異;本教材主要以可作為 phasing 錨點的 germline SNP 為例。A single-nucleotide polymorphism observed as a population variant. This tutorial mainly uses germline SNPs as phasing anchors.完整條目 →
也可作為建立 haplotype 骨架的錨點之一。錨點不足時,區分 HP1/HP2 及連結 somatic 變異會更困難。
clone 與 subclone
回到 M1 的簡化區段模型。一個細胞取得突變後增生,形成一群共享可辨識突變的細胞,稱為 cloneclone 克隆源自共同祖先細胞,並共享一組可辨識 somatic mutations 的細胞群。A population of tumour cells descended from a common ancestor and sharing the same set of somatic mutations.完整條目 →。
其中某些細胞又取得新突變、再擴增,就形成 subclonesubclone 次克隆clone 內取得額外變異並擴增的子群;治療後可能富集具有抗性的 subclone。A subset of a clone that acquired further mutations and expanded. Treatment-resistant populations are often subclones.完整條目 →。
許多 somatic 變異可能是 passenger;可提高生長或存活優勢的候選變異稱為 driver mutationdriver mutation 驅動突變可提升腫瘤細胞生長或存活優勢的候選變異;passenger mutation 通常不具有已知的選擇優勢。A mutation that confers a growth advantage, as opposed to the far more numerous passenger mutations.完整條目 →,
其判定需依癌別與功能證據。
因此,除列出腫瘤中的變異外,也需評估這些變異可能分布於哪些細胞群。
這種細胞組成的不均勻,就叫 intratumour heterogeneityintratumour heterogeneity 腫瘤內異質性同一腫瘤內不同細胞可帶有不同突變組合,因而同時存在多種基因型。Different cells within one tumour carrying different mutation combinations.完整條目 →。
真實證據
評估方法時需有可供對照的高可信度參考標註。
truth set 通常由多種技術與證據建立,但仍可能包含不確定區域;benchmark 還包含評估範圍與程序。
HCC1395 將於後續示例中再次出現;可比較兩份資料:
HCC1395_HKU 與 HCC1395_NYGC,同一個細胞株、不同的 basecallerbasecalling 鹼基判讀把定序儀的原始訊號(ONT 是電流)轉成 A/C/G/T 字母的步驟。不同 basecaller 版本會產生不同的錯誤特性 —— 所以「同一個細胞株、不同 basecaller」是兩份不同的資料集。Converting raw sequencer signal into A/C/G/T. Different basecallers yield different error profiles.完整條目 → 與分析流程。
兩份資料的組成相似度(0.909)可作為跨流程一致性的觀察,但不能單獨證明方法穩健性。
因此,provenanceprovenance 資料來源履歷記錄資料來源、細胞株、定序平台、basecaller、reference build、caller、benchmark 與混樣方式。在本實驗室中,這些資訊是重要實驗變數,而非附註。The full origin record of a dataset. In this lab it is a first-class experimental variable, not metadata.完整條目 → 應視為重要實驗變數,而非附註。
長讀long read 長讀單條可達數千至數萬鹼基的定序片段(例如 ONT、PacBio)。若可靠地同時覆蓋多個 variant,可提供它們位於同一 DNA 分子上的直接觀測證據。A sequencing read thousands to tens of thousands of bases long. Its length lets one molecule link multiple variants directly.完整條目 →提供的長距離資訊
兩個變異位點相距 3 kb。上:在本示意中,約 100 bp 的 Illumina reads 可分別觀察兩個位點,但沒有單條 read 同時涵蓋兩點。下:一條較長的 ONT read 涵蓋兩點,並可在完成修飾辨識時關聯中間的甲基化訊號。
ONTONT 奈米孔定序Oxford Nanopore Technologies。以電流訊號讀取 DNA,可產生長 read;在 native-DNA 流程與適當模型下,可由訊號推定 5mC,通常不需 bisulfite 轉換。Oxford Nanopore Technologies. Reads DNA through ionic-current signals, producing long reads. With native DNA and a suitable model, 5mC can be inferred without bisulfite conversion.完整條目 →
讀長
常見約 100–150 bp,依平台與建庫而異
可達數千至數萬 bp,分布依流程而異
單鹼基準確度
通常較高
依化學、basecaller 與變異類型而異;indelindel 插入/缺失短的插入(insertion)或缺失(deletion)。在 nanopore 資料上通常比 SNV 更難可靠判讀,因為 homopolymer 與 alignment 位置漂移都可能製造假訊號。A short insertion or deletion. It is often harder to call reliably than an SNV in nanopore data because of homopolymers and alignment ambiguity.完整條目 → 常較具挑戰
跨越多個變異
限於片段長度與建庫設計
較能建立長距離相位與分子連結
甲基化
bisulfite-seq 是常見方法之一
native-DNA 流程可搭配模型由訊號推定修飾
常見優勢
高準確度的局部變異偵測
長距離 phasing 與分子連結
長讀 indel 判讀的主要限制
ONT 與 Illumina 的錯誤特性不同。homopolymerhomopolymer 同聚物同一個鹼基連續重複的區段,例如 AAAAAA。nanopore 在此類區段較容易發生長度判讀錯誤,是 indel 假陽性的常見來源之一。A run of identical bases. Nanopore miscounts their length, a major source of false indel calls.完整條目 → 是 ONT indel 錯誤的常見高風險區域,
例如 AAAAAA。
下表提供效能示例,用於比較 SNV 與 indel 的相對差異。
數值解讀需同時記錄 caller 名稱與版本、資料集、truth set 與指標定義:
Caller
SNV F1
INDEL F1
ClairS(示例)
0.71
0.32
DeepSomatic(示例)
0.81
0.53
在此示例中,indel F1 低於 SNV F1,表示 somatic-indel 的整體偵測效能仍有改善空間;
後續方法會利用更多 read-level 特徵處理此問題。
補充:重複序列中的等價比對位置
除了 homopolymer,另一個原因是 alignment 位置的模糊性:
同一個 deletion 在重複序列中可有多種等價表示位置,使不同 reads 的訊號分散。
這些差異會反映在 CIGARCIGAR描述一條 read 如何對上參考基因體的緊湊字串。例如 5M1I5M 表示兩端各有 5 個對齊欄位,中間有 1 個相對於參考序列的插入;M 可能代表吻合,也可能代表錯配。A compact string describing how read bases and reference positions are consumed by an alignment. For example, 5M1I5M has five alignment columns, one insertion, then five alignment columns; M may contain matches or mismatches.完整條目 →;CIGAR 是描述 read 與參考序列對齊操作的字串。
三項核心數值
量
意思
本教材示例值
coveragecoverage 覆蓋度/深度某個位置被多少條 read 覆蓋。50× 表示平均每個位置約有 50 條 read 覆蓋,不代表它們均支持同一 allele。How many reads span a position. 50x means on average 50 reads per position.完整條目 →
某個位置被多少條 read 覆蓋
tumor 50×,normal 25×
VAFVAF 變異等位基因頻率在某個位點上,支持 alt allele 的 read 佔全部 read 的比例。VAF 不等於帶有這個突變的細胞比例。Variant Allele Frequency: the fraction of reads supporting the alternate allele. Not the same as the fraction of cells carrying it.完整條目 →
支持 alt allele 的 read 比例
germline het 在簡化二倍體模型中 ;somatic 依 purity、copy number 與 clonality 而異
puritytumour purity 腫瘤純度樣本中腫瘤細胞所佔的比例。purity 越低,somatic 訊號被正常細胞稀釋得越嚴重,偵測越困難。The fraction of cells in a sample that are tumour cells. Low purity dilutes somatic signal.完整條目 →
在二倍體、單拷貝、clonal、無偏取樣且 tumor DNA fraction 為 0.2 的簡化模型中,
期望支持數為 條。訊號數量不足是主要限制之一;背景錯誤、品質與演算法也會影響可靠性。
basecaller 是實驗變數
basecallingbasecalling 鹼基判讀把定序儀的原始訊號(ONT 是電流)轉成 A/C/G/T 字母的步驟。不同 basecaller 版本會產生不同的錯誤特性 —— 所以「同一個細胞株、不同 basecaller」是兩份不同的資料集。Converting raw sequencer signal into A/C/G/T. Different basecallers yield different error profiles.完整條目 → 是把電流訊號轉成 A/C/G/T 的那一步。不同版本的 basecaller
會產生不同的錯誤特性,所以:
同一個細胞株,不同 basecaller,可視為兩份不同的資料集
此記錄並非次要細節。跨 basecaller 的同株資料可作為穩健性評估的一部分;例如
HCC1395_HKU 與 HCC1395_NYGC 來自同一細胞株,但流程不同。
因此分析時應將 provenanceprovenance 資料來源履歷記錄資料來源、細胞株、定序平台、basecaller、reference build、caller、benchmark 與混樣方式。在本實驗室中,這些資訊是重要實驗變數,而非附註。The full origin record of a dataset. In this lab it is a first-class experimental variable, not metadata.完整條目 →(細胞株、平台、basecaller、reference buildreference genome 參考基因體用於比對的標準序列(人類資料常用 GRCh38)。所有座標均相對於指定版本;更換 build 會改變座標系,座標不可直接比較。The standard sequence everything is aligned to. All coordinates are relative to it, so changing build changes coordinates.完整條目 →、
caller 版本、benchmark 來源與混合方式)與結果一併記錄。
真實證據
合成不同 purity 的樣本
若要評估方法在不同 tumor DNA fraction 下的效能,可依計算後的比例混合 tumor 與 normal BAM,
建立預先設定的合成梯度:
但請注意:列為低可信候選不等於可以直接刪除。
低 purity 樣本中的 genuine somatic 變異也可能只有少數 reads 支持。
僅依低支持度刪除候選可能降低 recallrecall 召回率在指定評估範圍內,所有真陽性中被找出的比例:TP / (TP + FN)。若只對固定的 caller 候選集做後處理,recall 只能維持或下降,不能恢復 caller 從未輸出的真陽性;報告時應說清楚這個候選集範圍。Within a stated evaluation scope, the fraction of true positives that were found: TP / (TP + FN). With a fixed caller candidate set, post-filtering can preserve or lower recall but cannot recover variants the caller never emitted; report the candidate-set scope explicitly.完整條目 →。
本節會用一個很小的 toy locus 貫穿範例。它不是可直接拿去分析的完整檔案,而是把同一個位置在不同資料層級的樣子並排,讓你能回答「這一列在記錄什麼」以及「我要回哪個檔案找證據」。
概念與互動
一、先建立檔案地圖
格式/檔案
它保存什麼
一筆 record 通常代表什麼
它回答的問題
FASTAFASTA以文字保存 DNA 或 RNA 序列的格式;在本模組中,FASTA 主要用來提供比對所需的參考基因體序列。A text format for DNA or RNA sequences; here it mainly supplies the reference genome used for alignment.完整條目 →
參考基因體reference genome 參考基因體用於比對的標準序列(人類資料常用 GRCh38)。所有座標均相對於指定版本;更換 build 會改變座標系,座標不可直接比較。The standard sequence everything is aligned to. All coordinates are relative to it, so changing build changes coordinates.完整條目 →序列
一條序列或 contig 的 header 與鹼基
參考座標上的標準序列是什麼?
FASTQFASTQ存放 read 序列與每個鹼基品質分數的文字格式;每筆紀錄由四行組成。Text format holding read sequences plus per-base quality scores; four lines per record.完整條目 →
定序儀產生的 read 序列與品質
四行 read record:名稱、序列、分隔線、品質
定序儀觀察到了哪些片段?
BAMBAM已比對到參考基因體的 read 的二進位格式(SAM 的壓縮版)。每一列是一條 read,帶著它的位置、CIGAR、品質與各種 tag。Binary format for reads aligned to the reference. One record per read, carrying position, CIGAR, quality and tags.完整條目 →
已對齊至參考的 read;它是 SAM 文字格式的二進位版本
一筆 alignment record,通常對應一條 read
這條 read 被放在參考的哪裡?對齊得怎樣?
.bai
BAM 的索引,不是另一批生物學觀測
由座標建立的查詢索引
如何快速取出某個區間的 BAM records?
VCFVCF存放 variant 的文字格式。每一列是一個位點,記錄位置、ref/alt allele、品質、FILTER 與各種 FORMAT 欄位。Text format for variants: one line per site with position, ref/alt alleles, quality, FILTER and FORMAT fields.完整條目 →
PSPS tag phase set標示某群 variants 屬於同一 phase block 的編號。同一 PS 內的相位方向可相互比較;不同 PS 的 HP1/HP2 方向彼此獨立。Phase set identifier grouping variants phased together. Haplotype orientation is only consistent within one PS.完整條目 →
讀 BAM 時,可以把一筆 alignment record 想成一列「read 如何被放回參考」的資料。POS 告訴你從哪個參考座標開始,MAPQ 是比對位置的信心摘要,SEQ 是 read 序列;CIGARCIGAR描述一條 read 如何對上參考基因體的緊湊字串。例如 5M1I5M 表示兩端各有 5 個對齊欄位,中間有 1 個相對於參考序列的插入;M 可能代表吻合,也可能代表錯配。A compact string describing how read bases and reference positions are consumed by an alignment. For example, 5M1I5M has five alignment columns, one insertion, then five alignment columns; M may contain matches or mismatches.完整條目 → 則用一串數字與字母,描述 read 與參考序列之間各段如何對齊。
soft clipsoft clippingread 端部未與參考序列對齊,但序列仍保留在 BAM 中(CIGAR 為 S)。大量 clipping 集中於同一位置時,可能提示結構變異斷點,仍需其他證據確認。A read end that failed to align but is retained in the BAM (CIGAR S). Clipping clustered at one position may suggest a structural breakpoint and needs corroborating evidence.完整條目 →:read 保留,但這段沒有納入對齊
✗
✓
H
hard clip:該段序列不保留在這筆 alignment record 中
✗
✗
M 不等於完全吻合。如果要知道 M 中哪些鹼基真的不同,要一起看 read 序列、參考序列或 MD tag;不能只靠 CIGAR 的 M 判斷。
六、soft clip:BAM 對齊資訊的一個判讀例子
當 read 的一端無法可靠地對上參考時,比對器可能把那一端標成 S。soft-clipped 的序列仍保留在 BAM,所以它和 hard clip 不同;只是該段不消耗參考座標。
HPHP tagBAM 裡標示某條 read 屬於哪一條 haplotype 的 tag。LongPhase-S 的 germline haplotag 寫成整數 HP:i:1;somatic haplotag 與 LongPhase-TO 則寫成字串 HP:Z:1-1。BAM tag assigning a read to a haplotype. Integer HP:i:1 for germline haplotag; string HP:Z:1-1 for somatic.完整條目 →
工具對這條 read 的 haplotype assignment,例如 HP:i:1 或其他工具版本的 HP 值
HP1/HP2 是工具標籤,不自動代表父系/母系;沒有 HP 不代表 read 支持相反 allele。
HP3
在 somatic haplotypesomatic haplotype在 germline haplotype 之下再細分出來的層級。某條 HP1 的 read 若帶有 somatic 突變,就屬於 HP1-1;HP2 的則是 HP2-1。無法判定來源的歸為 HP3。A sub-level beneath the germline haplotype: HP1-1 and HP2-1 carry somatic mutations derived from HP1 and HP2; HP3 is unassignable.完整條目 → 分層中,帶有 somatic signal,但工具尚未可靠判定它源自哪一條 germline haplotype
HP3 不是「沒有 HP」。一般 untagged read 只是沒有 HP assignment;只有同時有 somatic signal 且來源未定,才使用 HP3 這個概念。
MM/MLMM/ML tagBAM 中儲存每條 read 甲基化判讀的兩個 tag:MM 記錄位置、ML 記錄機率。在轉換或重新比對流程中可能遺失,需確認 tags 是否保留。BAM tags carrying per-read methylation calls: MM for positions, ML for probabilities. Easily lost during processing.完整條目 →
前面那張畫面還有一個沒用上的線索:這些 read 分別來自哪一條染色體。
一個人的每個位置有兩條染色體,一條來自父親、一條來自母親。
把 read 依這兩條分成兩群(也就是 phasingphasing 定相為可判定的 variants 建立其位於不同實體染色體拷貝上的相位關係;VCF 常以 0|1 等形式表示。長 read 可提供跨位點的分子層級觀測證據。Determining which variants lie on the same physical chromosome, turning 0/1 into 0|1. Long reads provide direct molecular evidence.完整條目 →),同一個畫面會變得好判斷得多:
這也解釋了一個看起來矛盾的地方:明明要先有變異才能 phasing,怎麼又說 phasing 幫助 calling?
因為兩者是來回兩次:先用最有把握的位點做一次 calling,用它們把 read 分群,
再把 HP 標籤HP tagBAM 裡標示某條 read 屬於哪一條 haplotype 的 tag。LongPhase-S 的 germline haplotag 寫成整數 HP:i:1;somatic haplotag 與 LongPhase-TO 則寫成字串 HP:Z:1-1。BAM tag assigning a read to a haplotype. Integer HP:i:1 for germline haplotag; string HP:Z:1-1 for somatic.完整條目 →當成模型的額外輸入重判一次。Clair3 的 full-alignment 階段與
PEPPER-Margin-DeepVariant 都採用這個順序。
同一個「集中或分散」的線索在腫瘤資料裡會再出現一次,只是那時要判斷的不只是真假,
還包括這個變異長在哪一條 haplotypehaplotype 單倍型同一條實體染色體拷貝上,具有一致相位的一組 alleles 或 variants。HP1 與 HP2 是任意的相對標籤。A set of variants that lie on the same physical chromosome copy and are inherited together.完整條目 → 上 —— 需要的證據也更多。以下這張圖先把最重要的陷阱說清楚:
兩種情況的支持 read 總數可以幾乎一樣,所以不能靠數量判斷。
系統性錯誤、mapping bias,以及腫瘤裡的 copy-number 變化或
LOHLOH 異型合子性喪失原本 heterozygous 的區域變成只剩一種 allele。LOH 不等於缺失 —— 也可能是一條 haplotype 遺失後另一條被複製(copy-neutral LOH)。Loss of heterozygosity: a formerly het region retains only one allele. Not necessarily a deletion.完整條目 →,都可能讓分布不符合上述的簡化模式。它是重要線索,不是單獨的判定依據。
真實證據
本教材示例使用 Clair3Clair3以深度學習做 germline variant calling 的工具,長 read 上常用。A deep-learning germline variant caller widely used for long reads.完整條目 →;部分流程也會使用 PEPPER-Margin-DeepVariant。
以下提供一組示範指令:
最後的 --ont_r9_guppy5_sup 是與特定 ONT 化學版本/basecaller 資料相容的模型選項。
R10 與 PacBio HiFi 應依所用版本文件選擇相容模型(例如 --ont_r10_q20 或 --hifi)。
模型不匹配可能降低效能,且流程未必報錯;請將平台、basecaller 與模型版本記錄於 provenanceprovenance 資料來源履歷記錄資料來源、細胞株、定序平台、basecaller、reference build、caller、benchmark 與混樣方式。在本實驗室中,這些資訊是重要實驗變數,而非附註。The full origin record of a dataset. In this lab it is a first-class experimental variable, not metadata.完整條目 →。
Clair3:Zheng Z 等,Symphonizing pileup and full-alignment for deep learning-based
long-read variant calling,Nature Computational Science,2022。程式碼在
github.com/HKU-BAL/Clair3。
DeepVariant:Poplin R 等,A universal SNP and small-indel variant caller using deep
neural networks,Nature Biotechnology,2018(doi 10.1038/nbt.4235)。程式碼在
github.com/google/deepvariant。
本模組術語
Clair3
以深度學習做 germline variant calling 的工具,長 read 上常用。
DeepVariant
Google 開發的 germline variant caller,把 pileup 轉成影像再用 CNN 分類。
phasing 可能因缺乏重疊 read 或 informative 位點而中止,形成新的 phase blockphase block一段可建立連續相位關係的區域。read 長度不足、缺少 informative heterozygous 位點或證據不一致時,可能形成不同 phase blocks。A contiguous stretch over which phasing is consistent. It can break when reads are too short, informative heterozygous sites are absent or evidence conflicts.完整條目 →。
VCF 用 PS 欄位標示每個 block。
switch errorswitch errorphasing 結果自某一位置起將兩條 haplotype 的方向對調,是常見的 phasing 錯誤類型。A phasing error in which the two haplotype labels swap from some point onward; a common phasing error mode.完整條目 → 是常見的 phasing 錯誤類型:自某一位置起,演算法將兩條 haplotype 的標籤對調。
它可能發生在 het 位點稀疏、read 重疊不足或證據歧義較高的區域。
LOHLOH 異型合子性喪失原本 heterozygous 的區域變成只剩一種 allele。LOH 不等於缺失 —— 也可能是一條 haplotype 遺失後另一條被複製(copy-neutral LOH)。Loss of heterozygosity: a formerly het region retains only one allele. Not necessarily a deletion.完整條目 → 可能減少可用的 heterozygous 錨點,使 phasing 變弱或分段;LOH 可為部分、copy-neutral 或 subclonal,
未必使所有位點都變成 homozygous。缺乏 informative 位點時,跨區域連結會變得不可靠。
phasingphasing 定相為可判定的 variants 建立其位於不同實體染色體拷貝上的相位關係;VCF 常以 0|1 等形式表示。長 read 可提供跨位點的分子層級觀測證據。Determining which variants lie on the same physical chromosome, turning 0/1 into 0|1. Long reads provide direct molecular evidence.完整條目 → 將 variant 的 haplotype assignment 寫入 VCF;
haplotagginghaplotagging根據已 phase 好的 variants,把每一條 read 指派到 HP1 或 HP2,並把結果寫回 BAM 的 HP tag。Assigning each read to HP1 or HP2 using phased variants, and writing the result back as a BAM HP tag.完整條目 → 則將每條 read 的 assignment 寫入 BAM 的 HPHP tagBAM 裡標示某條 read 屬於哪一條 haplotype 的 tag。LongPhase-S 的 germline haplotag 寫成整數 HP:i:1;somatic haplotag 與 LongPhase-TO 則寫成字串 HP:Z:1-1。BAM tag assigning a read to a haplotype. Integer HP:i:1 for germline haplotag; string HP:Z:1-1 for somatic.完整條目 → tag。
之後可在 IGV 中依 haplotype 分組或著色顯示。
完成後,可在 IGV 中依 HP tag 分組,觀察 read 是否呈現 haplotype 分層。
判讀練習
在一段區域完成 haplotag 後,若有 30% 的 read 沒有 HP tag,請列出至少三個可能原因,
並說明可如何區分這些原因。
展開答案
可從以下方向逐一排查:
read 未覆蓋 informative het 位點 —— 沒有足夠的判斷依據,可能保持未標記。
這段區域可能是 LOHLOH 異型合子性喪失原本 heterozygous 的區域變成只剩一種 allele。LOH 不等於缺失 —— 也可能是一條 haplotype 遺失後另一條被複製(copy-neutral LOH)。Loss of heterozygosity: a formerly het region retains only one allele. Not necessarily a deletion.完整條目 → —— 若區段缺乏可用的 het 位點,read 的 haplotype assignment 會變得困難。
若未標記 read 集中於連續區段,應優先檢查 LOH、低複雜度或 coverage。
上排用這個病人自己的正常組織,下排改查公開族群資料庫。輸出的差別只有位置 ③ —— 那是這個人自己的 germline 變異,族群裡幾乎沒人有,所以資料庫查不到,於是被留下來變成假陽性。位置 ② 兩排都靠 read 的分布排除,跟有沒有正常樣本無關。圖上的顏色是給讀者看的答案,程式看不到。
請注意這張圖真正的重點:位置 ① 與 ③ 都是這個人天生就有的 germline variantgermline variant 生殖系變異在生殖系形成、通常存在於多數細胞的變異;可遺傳給子代,但不代表必然傳遞。A variant arising in the germline and therefore present in essentially every cell of the individual. It may be inherited or arise de novo, and may be passed to offspring.完整條目 →,
在腫瘤樣本裡長得一模一樣(各三條支持、VAF 都是 0.5)。它們的命運不同,
不是因為變異本身有什麼差別,而是因為資料庫知不知道它。
位置 ④ 兩種模式都留得住,那是我們要的 somatic variantsomatic variant 體細胞變異在非生殖系細胞譜系中、通常於受精後取得的變異;可只存在於部分細胞,通常不由親代遺傳給子代。癌症基因體學常分析此類變異。A variant acquired outside the germline, usually after conception. It may be restricted to a subset of cells and is generally not passed to offspring; it is a central focus of cancer genomics.完整條目 →。
caller 產生的是 候選變異candidate 候選變異caller 提出、但尚未被確認的變異位點。本實驗室多數工具位於 caller 之後,進行再校正,處理的就是這些 candidates。A proposed but unconfirmed variant. Most tools in this lab post-process the caller's candidate set.完整條目 →,還不是最終答案;本教材以 ClairSClairS配對 tumor–normal 的 somatic variant caller。tumor-only 版本叫 ClairS-TO。A somatic variant caller for matched tumour–normal data; ClairS-TO is the tumour-only version.完整條目 →、DeepSomaticDeepSomaticGoogle 的 somatic variant caller,同樣有 tumor-only 版本。Google's somatic variant caller, also available in a tumour-only mode.完整條目 →
作為 tumour-normaltumour-normal 配對腫瘤/正常同時定序病人的腫瘤與正常組織;正常樣本提供個體 germline 背景,有助於辨識 somatic 候選。Sequencing a patient's tumour alongside their normal tissue, which supplies the germline background.完整條目 →(TN)calling 的範例,tumour-onlytumour-only 僅腫瘤只有腫瘤樣本、沒有配對正常樣本。在成本或檢體受限時很常見,但少了 germline 對照,分析難度會大幅提高。Having only the tumour sample. Common for cost or specimen reasons, but much harder without a germline baseline.完整條目 →(TO)則以 ClairS-TO 為例,
實際輸入與過濾流程請依版本文件確認。
沒有正常組織時,改用什麼當對照
在部分臨床或研究情境,matched normal 可能因成本、檢體量或取得條件而缺失。
替代方案是公開的族群資料庫:利用 allele frequency 與已知變異紀錄提供 germline 線索。
本教材在 LongPhase-TO/ClairS-TO 情境中將這組集合稱為 PONPON 正常樣本面板/族群資料庫(依流程而異)PON 依分析流程有兩種用法,兩者不可互換:本教材的 tumor-only 流程以它指稱 population germline database set(如 1000G、CoLoRSdb、dbSNP、gnomAD 的聯集);GATK Mutect2 的 Panel of Normals(PoN)則由多個正常樣本建立,用來標記反覆出現的技術性 artifact。兩者都不能完全取代同一病人的 matched normal。The abbreviation is workflow-dependent and the two uses are not interchangeable. In this tutorial's tumour-only workflows, PON denotes a population germline database set; in GATK Mutect2, a Panel of Normals (PoN) is built from multiple normal samples to flag recurrent technical artifacts. Either can miss a private germline variant or an unseen artifact, so neither replaces a patient's matched normal.完整條目 →(population database set),
它包含 1000 Genomes、gnomAD、dbSNP、CoLoRSdb 等來源。
上圖下排的橫軸就是這件事的關鍵。資料庫有一個收錄門檻——
實際下載的檔名裡寫著 af-ge-0.001,意思是只收頻率 0.001 以上的位點。
位置 ① 的族群頻率 0.31,遠高於門檻,查得到;位置 ③ 只有 0.0002,落在門檻左邊,
資料庫裡根本沒有這一筆。「族群常見」與「這個人有」不是同一件事,缺口就從這裡來。
回到那張 PONPON 正常樣本面板/族群資料庫(依流程而異)PON 依分析流程有兩種用法,兩者不可互換:本教材的 tumor-only 流程以它指稱 population germline database set(如 1000G、CoLoRSdb、dbSNP、gnomAD 的聯集);GATK Mutect2 的 Panel of Normals(PoN)則由多個正常樣本建立,用來標記反覆出現的技術性 artifact。兩者都不能完全取代同一病人的 matched normal。The abbreviation is workflow-dependent and the two uses are not interchangeable. In this tutorial's tumour-only workflows, PON denotes a population germline database set; in GATK Mutect2, a Panel of Normals (PoN) is built from multiple normal samples to flag recurrent technical artifacts. Either can miss a private germline variant or an unseen artifact, so neither replaces a patient's matched normal.完整條目 → 涵蓋不到的左下角。要把住在同一格裡的三種東西分開,
必須換一個問法:不再問「這個位置別人有沒有」,而是問「支持它的那些 read,彼此的結構像不像一次真的突變」。
真的變異一定是從某一條既有的 haplotypehaplotype 單倍型同一條實體染色體拷貝上,具有一致相位的一組 alleles 或 variants。HP1 與 HP2 是任意的相對標籤。A set of variants that lie on the same physical chromosome copy and are inherited together.完整條目 → 上長出來的,所以支持它的 read
應該集中在一條 haplotype 上。如果同一個「變異」的支持 read 分散在兩條染色體上,
那等於同一個位置各獨立突變了一次 —— 機率極低,更像比對或定序造成的問題:
這個想法在兩個工具裡有不同的實作:LongPhase-S 做成 read-level 與 haplotype origin
兩個過濾器(四個過濾器與各自的門檻見 M12);LongPhase-TO 把它擴展成三個位點的
triplet graphtriplet graphLongPhase-TO 判斷候選變異真假時用的最小結構:候選位置加上左右各一個變異,共三個位置。把 read 上看到的 allele 組合畫成路徑後,真的 somatic 變異會形成一條「跟某一條 germline haplotype 只差候選這一格」的新路徑;定序錯誤則湊不出一致的路徑。因為左右兩側都要對得上,所以三個位置是最小單位。LongPhase-TO's minimal unit for judging a candidate: the candidate plus one flanking variant on each side. A true somatic allele forms a third path differing from one parental haplotype at the candidate alone; artifacts show no consistent path.完整條目 →,用路徑結構判斷(見 M13)。
要注意這個判準本身是理想化的:實務資料不會這麼乾淨,
系統性 artifact 或 phasing error 都會造成例外,所以它是提高或降低疑慮,不是定案。
前面提到的三位點路徑結構(triplet graphtriplet graphLongPhase-TO 判斷候選變異真假時用的最小結構:候選位置加上左右各一個變異,共三個位置。把 read 上看到的 allele 組合畫成路徑後,真的 somatic 變異會形成一條「跟某一條 germline haplotype 只差候選這一格」的新路徑;定序錯誤則湊不出一致的路徑。因為左右兩側都要對得上,所以三個位置是最小單位。LongPhase-TO's minimal unit for judging a candidate: the candidate plus one flanking variant on each side. A true somatic allele forms a third path differing from one parental haplotype at the candidate alone; artifacts show no consistent path.完整條目 →)就是為了讓這個判準更穩健:
左右鄰居都要對得上,才算得出「跟某一條既有 haplotype 只差候選這一格」的那條路徑。
ClairS(tumour-normal):ClairS: a deep-learning method for long-read somatic small
variant calling,bioRxiv,2023。程式碼在
github.com/HKU-BAL/ClairS。
ClairS-TO(tumour-only):ClairS-TO: a deep-learning method for long-read tumor-only
somatic small variant calling,Nature Communications,2025。程式碼在
github.com/HKU-BAL/ClairS-TO。
DeepSomatic:Accurate somatic small variant discovery for multiple sequencing
technologies with DeepSomatic,Nature Biotechnology,2025。程式碼在
github.com/google/deepsomatic。
這些名詞看起來很多,但它們其實只做一件事:在不同層級上數東西。
puritytumour purity 腫瘤純度樣本中腫瘤細胞所佔的比例。purity 越低,somatic 訊號被正常細胞稀釋得越嚴重,偵測越困難。The fraction of cells in a sample that are tumour cells. Low purity dilutes somatic signal.完整條目 → 與 CCFcancer cell fraction 癌細胞比例帶有某個特定突變的腫瘤細胞佔全部腫瘤細胞的比例。用來區分 clonal()與 subclonal()突變。不等於 VAF。The fraction of tumour cells carrying a given mutation; distinguishes clonal from subclonal. Not the same as VAF.完整條目 → 數的是細胞;
copy numbercopy number 拷貝數某段基因體在細胞內的拷貝數。多數正常常染色體區段為 2,可再分為 major 與 minor allele copy number。How many copies of a genomic segment a cell carries; a typical diploid autosomal segment has two. It can be split into major- and minor-allele copy numbers.完整條目 → 與 multiplicitymutation multiplicity在帶有某突變的細胞中,該突變所占的拷貝數;是 VAF 與 cancer cell fraction 換算時的重要參數。How many mutant-allele copies are present in a cell carrying the mutation; a necessary term when converting VAF to cancer cell fraction.完整條目 → 數的是一個細胞裡有幾份 DNA;
VAFVAF 變異等位基因頻率在某個位點上,支持 alt allele 的 read 佔全部 read 的比例。VAF 不等於帶有這個突變的細胞比例。Variant Allele Frequency: the fraction of reads supporting the alternate allele. Not the same as the fraction of cells carrying it.完整條目 → 數的是read。搞混它們,通常就是搞混了分母。
中間那層最容易被跳過,但它是後面所有推論的地基。三個詞的差別很小、影響很大:
copy number 是「這一段在一個細胞裡有幾份」,multiplicity 是「其中幾份帶著這個突變」,
ploidyploidy 倍性細胞內染色體套數。多數正常常染色體為 2(diploid);腫瘤可呈現 3、4 或其他非整倍體狀態。The number of chromosome sets in a cell. Most normal autosomal regions are diploid (2); tumours may be triploid, tetraploid or otherwise non-diploid.完整條目 → 是「把整個基因體的 copy number 平均起來」。
四格都假設檢體全是腫瘤細胞、而且每個細胞都帶有這個突變 ——
唯一的差別是拷貝狀態。VAF 從 0.50 變成 0.67 或 1.00,不是因為突變變多了,
是因為它所在的那份 DNA 被複製了,或另一份被丟掉了(那正是 copy-loss LOHLOH 異型合子性喪失原本 heterozygous 的區域變成只剩一種 allele。LOH 不等於缺失 —— 也可能是一條 haplotype 遺失後另一條被複製(copy-neutral LOH)。Loss of heterozygosity: a formerly het region retains only one allele. Not necessarily a deletion.完整條目 →)。
下半部區分 copy number(某一段的份數)與 ploidy(整個基因體的平均)。
腫瘤細胞比例tumour purity 腫瘤純度樣本中腫瘤細胞所佔的比例。purity 越低,somatic 訊號被正常細胞稀釋得越嚴重,偵測越困難。The fraction of cells in a sample that are tumour cells. Low purity dilutes somatic signal.完整條目 →
通常需估計
totalCN
該位置的 總拷貝數copy number 拷貝數某段基因體在細胞內的拷貝數。多數正常常染色體區段為 2,可再分為 major 與 minor allele copy number。How many copies of a genomic segment a cell carries; a typical diploid autosomal segment has two. It can be split into major- and minor-allele copy numbers.完整條目 →(整體平均值即 ploidyploidy 倍性細胞內染色體套數。多數正常常染色體為 2(diploid);腫瘤可呈現 3、4 或其他非整倍體狀態。The number of chromosome sets in a cell. Most normal autosomal regions are diploid (2); tumours may be triploid, tetraploid or otherwise non-diploid.完整條目 →)
通常需估計
multiplicity
突變佔了幾份拷貝mutation multiplicity在帶有某突變的細胞中,該突變所占的拷貝數;是 VAF 與 cancer cell fraction 換算時的重要參數。How many mutant-allele copies are present in a cell carrying the mutation; a necessary term when converting VAF to cancer cell fraction.完整條目 →
通常需推論
CCF
帶有此突變的腫瘤細胞比例cancer cell fraction 癌細胞比例帶有某個特定突變的腫瘤細胞佔全部腫瘤細胞的比例。用來區分 clonal()與 subclonal()突變。不等於 VAF。The fraction of tumour cells carrying a given mutation; distinguishes clonal from subclonal. Not the same as VAF.完整條目 →
通常是分析所要估計的目標
VAF
支持 alt 的 read 比例
可由 reads 估計
圖上方拆解分子與分母:分母是此位置讀到的 DNA 總量,由腫瘤與正常細胞共同貢獻;分子是其中帶有變異的份數。下方示範同一個 VAF 0.20 可由多種生物狀態產生,僅第三例具有 subclonal CCF。
LOHLOH 異型合子性喪失原本 heterozygous 的區域變成只剩一種 allele。LOH 不等於缺失 —— 也可能是一條 haplotype 遺失後另一條被複製(copy-neutral LOH)。Loss of heterozygosity: a formerly het region retains only one allele. Not necessarily a deletion.完整條目 → 是「原本 heterozygous 的區域只剩下一種 allele」。
但其形成機制不只一種;以下比較兩類常見情形:
可呼叫區域的 het 密度受樣本、族群、coverage、caller 與過濾條件影響,應以配對 normal 或鄰近區段建立基線。
若某區段密度顯著低於基線,可列為 LOH 候選;purity 混合可能保留部分 het reads,仍需結合 allele balance、coverage 與其他證據。
5mC5mC 5-甲基胞嘧啶胞嘧啶第 5 個碳上的甲基化修飾。這是常見的 DNA 甲基化形式,主要見於 CpG;可影響轉錄調控,效應依基因組位置與細胞類型而異。5-methylcytosine, the commonest DNA methylation mark, usually at CpG sites. Affects expression without changing sequence.完整條目 → 是在胞嘧啶(C)的 5 位碳上加上甲基。在序列字母表示上,它不改變 C 的識別;
5mC 可與基因調控及表現相關,但作用取決於基因組脈絡。
在哺乳類,5mC 主要集中於 CpGCpG序列上一個 C 後接一個 G(p 代表兩者之間的磷酸鍵)。哺乳類多數 5mC 位於 CpG,特定細胞或情況亦可見非 CpG 甲基化。A cytosine followed by a guanine. Most mammalian 5mC occurs at CpG sites, although non-CpG methylation is also observed in particular cells and contexts.完整條目 → 位點(一個 C 後面接著一個 G),但也可能出現在非 CpG 脈絡。
在本節的分析中,基本單位是「某個 CpG 位點上,一條 read 的 methylated 或 unmethylated 狀態」。
bisulfite-seq 是測量甲基化的常用方法之一,會將未甲基化的 C 轉換為 U(在定序中通常讀為 T),流程較為繁複。
ONTONT 奈米孔定序Oxford Nanopore Technologies。以電流訊號讀取 DNA,可產生長 read;在 native-DNA 流程與適當模型下,可由訊號推定 5mC,通常不需 bisulfite 轉換。Oxford Nanopore Technologies. Reads DNA through ionic-current signals, producing long reads. With native DNA and a suitable model, 5mC can be inferred without bisulfite conversion.完整條目 → 可利用電流訊號搭配修飾辨識模型推定 5mC,通常不需 bisulfite 前處理;
在完成修飾辨識且保留 MM/ML 後,同一條 read 可同時提供序列變異與甲基化狀態。
上方區分序列與甲基化資訊;示意中的修飾標記位於 C 後接 G 的 CpG 位點,且不改變序列字母。下方示範同一分子可同時關聯變異與修飾狀態。
至少可比較配對正常樣本的同一區域,並納入獨立重複或 allele-aware 分析:
若 normal 與 tumor REF reads 的甲基化模式相似,可降低區域背景差異的解釋,但不能完全排除。
因此,該方法將甲基化定位為事後註記而非推論證據;在缺乏獨立對照時,這可避免過度的因果解讀。
實作練習
甲基化資訊存在 BAM 的哪裡
前面那張 0/1 表不是憑空來的,它從 BAM 的兩個 tag 讀出來:
tag
內容
MMMM/ML tagBAM 中儲存每條 read 甲基化判讀的兩個 tag:MM 記錄位置、ML 記錄機率。在轉換或重新比對流程中可能遺失,需確認 tags 是否保留。BAM tags carrying per-read methylation calls: MM for positions, ML for probabilities. Easily lost during processing.完整條目 →
修飾類型及其相對於 read 上標準鹼基的位置與跳過規則
ML
與 MM 項目對應的修飾機率或信心,以 0–255 的位元組值編碼
資料格式:ML 是機率,不是二元標籤
ML 是機率而不是 0/1:要得到前面那張表,還需要一個門檻把機率二值化,
或直接保留機率值。門檻怎麼定會影響分群結果,屬於分析選擇而非資料本身。
重新比對可能遺失 MM/ML 標籤
若未明確傳遞 SAM tags,將帶有甲基化的 BAM 轉成 FASTQ 再重新比對時,通常不會保留 MM/ML。
結果可能是格式正常、但修飾欄位空白的 BAM;工具是否發出警告取決於實作與版本。
在支援 base-modification 且 BAM 含有效 MM/ML、參考序列與相容版本時,IGV 可顯示 ONT 甲基化。載入 BAM 之後右鍵 →
Color alignments by → base modification (5mC)。
搭配前面的 Group by phase,即可同時檢視 haplotype 與甲基化資訊。
學習檢核
本模組術語
5mC(5-甲基胞嘧啶)
胞嘧啶第 5 個碳上的甲基化修飾。這是常見的 DNA 甲基化形式,主要見於 CpG;可影響轉錄調控,效應依基因組位置與細胞類型而異。
CpG
序列上一個 C 後接一個 G(p 代表兩者之間的磷酸鍵)。哺乳類多數 5mC 位於 CpG,特定細胞或情況亦可見非 CpG 甲基化。
precisionprecision 精確率所有 calls 中實際為真陽性的比例:TP / (TP + FP)。後處理常以提升此指標為目標。Of the calls you made, the fraction that are correct: TP / (TP + FP).完整條目 →
TP ÷(TP+FP)
所報變異中屬於真陽性的比例
通常上升(移除 FP)
recallrecall 召回率在指定評估範圍內,所有真陽性中被找出的比例:TP / (TP + FN)。若只對固定的 caller 候選集做後處理,recall 只能維持或下降,不能恢復 caller 從未輸出的真陽性;報告時應說清楚這個候選集範圍。Within a stated evaluation scope, the fraction of true positives that were found: TP / (TP + FN). With a fixed caller candidate set, post-filtering can preserve or lower recall but cannot recover variants the caller never emitted; report the candidate-set scope explicitly.完整條目 →
TP ÷(TP+FN)
所有真陽性中被找出的比例
至多維持,可能下降
F1F1precision 與 recall 的調和平均。當資料裡 FN 遠多於 FP 時,precision 的改善對 F1 的影響有限 —— 見 M10。The harmonic mean of precision and recall. A precision gain has little effect when false negatives dominate the evaluation.完整條目 →
precision 與 recall 的調和平均
兩項指標的綜合表現
依 TP、FP、FN 的變化而定
上方定義三項指標;下方示範在固定候選集合、且後處理僅移除假陽性的情況下,FP 減少而 FN 維持不變。因此 precision 上升 0.080,而 F1 僅上升 0.024。
truth settruth set 標準答案集評估用的參考標註。可靠度依資料來源與建立方法而異,困難區域可能包含標註錯誤。The reference labels used for evaluation. A truth set is not infallible truth: reliability varies, and hard regions can contain label errors.完整條目 → 並非無誤的標準
評估需要參考答案;不同資料來源的標註可靠度可能存在差異:
來源
形成方式
可靠度評估
SEQC2
多平台、多中心的共識計畫
需依變異類型與區域評估
NYGC
單一機構產生
需依驗證設計評估
Google(orthogonal tools)
多個工具交叉比對的共識
需確認工具獨立性與共識形成方式
LongPhase-S 在搜尋過濾參數時,
依不同 cell line 的 benchmark 可靠度乘上相對應的權重。
此設計可能減少 benchmark 差異的影響,但比較不同論文的數字時,仍應確認 truth set 與加權方式是否一致。
評估邊界:benchmark 只涵蓋高信心區域
benchmark 通常附一個 high-confidence regionhigh-confidence regionbenchmark 建立者指定的高可信度區間,通常以 BED 檔表示;區間外不宜視為具有相同標註可靠度。Intervals that benchmark authors consider reliable, usually supplied as a BED file. Results outside those intervals should not be assumed to have the same label reliability.完整條目 → BED 檔,標示「我們有把握的區間」。
區間外的標註信心通常較低,結果不宜直接解讀,除非有其他獨立驗證。
三項可能造成評估偏差的因素
因素
可能徵兆
本例的處理
data leakagedata leakage 資料洩漏訓練或模型選擇階段取得了評估資料的資訊,讓效能被高估。實驗室的做法是按染色體切分,確保同一個位點不會同時出現在訓練、驗證與最終測試中。Evaluation information entering training or model selection and inflating performance. The lab splits by chromosome so a locus cannot appear in training, validation and final test partitions at once.完整條目 →
distribution shiftdistribution shift 分布偏移測試資料的組成與訓練資料存在顯著差異(例如正負樣本比例不同),使得模型表現不如預期。Test data differing in composition from training data, so measured performance does not transfer.完整條目 →
訓練與測試的類別或特徵分布不同
indel 精修模型的訓練集 FP 佔 4%、測試集佔 19%,因此不能只看 loss
label noiselabel noise 標籤雜訊訓練或評估標籤與真實狀態不一致的情況。在 homopolymer 與重複區域可能較常見;特徵相近的位點也可能得到相反標籤。Errors in the labels themselves. They can be more frequent in homopolymers and repeats, where near-identical sites may receive different labels.完整條目 →
兩條曲線共用同一條 epoch 橫軸。上面的 validation loss 一路降到第 22 個 epoch 才最低;下面的任務分數第 12 個 epoch 就到頂,之後開始下滑。兩個最佳點相差十個 epoch —— 取 loss 最低的那一個,拿到的是一個任務目標已經變差的模型。
核心原則
模型選擇的標準應與任務目標對齊,而不必與 loss 最小化完全相同。
此原則適用於 variant calling 以外的應用 ML 任務。
本例交叉驗證以樣本為分割單位
LongPhase-S 用 leave-k-out 交叉驗證評估腫瘤 DNA 比例的迴歸模型,
每次留下 1 或 2 個 cell line 當測試集。結果 MAE 約 4%、 約 0.95。
這種做法叫 cross-validationcross-validation 交叉驗證輪流保留資料子集作為驗證折,用來估計模型的泛化能力與選擇設定;這些折是 validation,不是最終 test。最終測試集應在模型與設定確定後才使用,並另行保留。Repeatedly holding out validation folds to estimate generalisation and choose model settings. The folds are validation data, not the final test; an untouched final test set is used only after the model is fixed.完整條目 →。此處保留 cell line 作測試,有助於評估跨樣本的泛化能力。
若同一 cell line 的高度相關位點同時出現在訓練與測試中,可能造成 data leakagedata leakage 資料洩漏訓練或模型選擇階段取得了評估資料的資訊,讓效能被高估。實驗室的做法是按染色體切分,確保同一個位點不會同時出現在訓練、驗證與最終測試中。Evaluation information entering training or model selection and inflating performance. The lab splits by chromosome so a locus cannot appear in training, validation and final test partitions at once.完整條目 → 或效能高估。
在哪個資料集、用哪一份 truth set?Google 的 orthogonal-tools 共識與 SEQC2
的標註範圍與難度可能不同;是否限制在 high-confidence regionhigh-confidence regionbenchmark 建立者指定的高可信度區間,通常以 BED 檔表示;區間外不宜視為具有相同標註可靠度。Intervals that benchmark authors consider reliable, usually supplied as a BED file. Results outside those intervals should not be assumed to have the same label reliability.完整條目 → 內,也可能造成明顯差異。
LongPhaseLongPhase實驗室開發的長 read phasing 工具,是後續所有工具的共同基礎。輸出 germline haplotype。The lab's long-read phasing tool and the common foundation for everything else. Produces germline haplotypes.完整條目 → 是本教材後續工具共用的基礎,處理的是一般(非腫瘤)樣本的 phasing。
開始分析腫瘤資料前,先掌握這個基本流程。
三種失敗模式及其可觀測指標。① 分段破碎:檢查 phase set 的數量。② 方向對調:與 truth 比對時,錯誤從某點起系統性出現。③ 大量 read 無法指派:檢查 HP 標記率。
失敗
對應到上面的哪一步
徵兆
phase blockphase block一段可建立連續相位關係的區域。read 長度不足、缺少 informative heterozygous 位點或證據不一致時,可能形成不同 phase blocks。A contiguous stretch over which phasing is consistent. It can break when reads are too short, informative heterozygous sites are absent or evidence conflicts.完整條目 → 斷裂
投票平手,或沒有 read 跨過(沒有邊)
同一條染色體上出現很多不同的 PS 值
switch errorswitch errorphasing 結果自某一位置起將兩條 haplotype 的方向對調,是常見的 phasing 錯誤類型。A phasing error in which the two haplotype labels swap from some point onward; a common phasing error mode.完整條目 →
某條邊的 ESR 偏高卻仍勉強過關,方向選錯
與 truth 比對時,錯誤從某一點之後系統性地發生
大量 read 未標記
沒過 --readConfidence(0.65)
haplotag 後很多 read 沒有 HP
第一項與第三項在腫瘤樣本上可能更明顯,因為 LOHLOH 異型合子性喪失原本 heterozygous 的區域變成只剩一種 allele。LOH 不等於缺失 —— 也可能是一條 haplotype 遺失後另一條被複製(copy-neutral LOH)。Loss of heterozygosity: a formerly het region retains only one allele. Not necessarily a deletion.完整條目 → 與 copy number 異常會減少可用的 phasing 錨點。
LongPhase-S 與 LongPhase-TO 都針對這些腫瘤資料情境提供額外處理。
可投票的位點太稀疏。如果 VCF 裡的 het 位點本來就少(樣本雜合度低,
或這條染色體有大段 LOHLOH 異型合子性喪失原本 heterozygous 的區域變成只剩一種 allele。LOH 不等於缺失 —— 也可能是一條 haplotype 遺失後另一條被複製(copy-neutral LOH)。Loss of heterozygosity: a formerly het region retains only one allele. Not necessarily a deletion.完整條目 →),可用錨點就會減少。
LongPhase:Lin JH、Chen LC、Yu SQ、Huang YT,
LongPhase: an ultra-fast chromosome-scale phasing algorithm for small and large
variants,Bioinformatics,2022。程式碼在
github.com/twolinin/LongPhase。
LongPhase-SLongPhase-S配對 tumor–normal 版本。做腫瘤 DNA 比例估計(輸出欄位名為 purity)、比例感知的 somatic variant 過濾,以及 somatic haplotagging。The matched tumour–normal version: tumour DNA fraction estimation (the output field is named purity), fraction-aware somatic filtering, and somatic haplotagging.完整條目 → 的做法是多做一層:除了兩條 germline haplotype,
再把「帶著 somatic 變異的那些 read」單獨拉成它們的後代分支。
這一層做出來之後有三個連帶好處 —— 可以直接估腫瘤來源 DNA 佔了多少、
可以用來重新判斷變異真假、也可以看出哪些 read 屬於同一個癌細胞群。
第一項要先把量講清楚,因為工具的參數名稱與輸出欄位都叫 purity,
但它估的是 腫瘤 DNA 比例tumour DNA fraction 腫瘤 DNA 比例樣本 DNA 中源自腫瘤的比例。與 tumor purity(細胞比例)在 aneuploid 或 WGD 的情況下會不一樣。The fraction of DNA originating from tumour cells; diverges from cellular purity under aneuploidy or WGD.完整條目 →(分母是分子),
不是 cellular puritytumour purity 腫瘤純度樣本中腫瘤細胞所佔的比例。purity 越低,somatic 訊號被正常細胞稀釋得越嚴重,偵測越困難。The fraction of cells in a sample that are tumour cells. Low purity dilutes somatic signal.完整條目 →(分母是細胞)。
M3 與 M8 的重要區分 #2 已建立這兩個量在非二倍體之下並不相等;
本章說明 LongPhase-S 落在哪一邊,以及為什麼它不可能落在另一邊。
概念與互動
先看懂 somatic haplotagging 本身
整章的其他東西都建立在這一步上:somatic haplotaggingsomatic haplotagging把帶有 somatic 突變的 read 指派到它們源自的 germline haplotype。這是 LongPhase-S 的核心產出。Assigning somatic-mutation-bearing reads to the germline haplotype they arose from. The core output of LongPhase-S.完整條目 → 就是給每條 tumor read 貼一個標籤。
判斷只用兩件事 —— 這條 read 上的 germline allele 屬於哪一條,以及它有沒有帶 somatic allele。
算全基因體的 GHIRGHIR 生殖系單倍型失衡比Germline Haplotype Imbalance Ratio:在候選 somatic 位點上,取標為 HP1 與 HP2 的 read 數中較大者除以兩者之和,值域為 0.5 至 1。須注意兩件事:分母只含這兩個 germline 計數,HP1-1/HP2-1/HP3 皆不在內;且這些標籤是整條 read 的判定(該 read 任一處帶 somatic 等位即離開 germline 計數),並非該位點的等位計數。它不是直接的 purity 讀數:拷貝數變異、LOH、read 跨距內的突變密度、標記錯誤與抽樣不足都會使它偏移。Germline Haplotype Imbalance Ratio: at a candidate somatic locus, the larger of the HP1 and HP2 tagged read counts divided by their sum, ranging from 0.5 to 1. Two caveats: the denominator contains only those two germline counts, excluding HP1-1/HP2-1/HP3; and those tags are whole-read decisions (a read carrying a somatic allele anywhere leaves the germline counts), not per-locus allele counts. It is not a direct purity readout — copy number, LOH, mutation density within the read span, tagging error and sparse sampling all shift it.完整條目 → 分布、過濾低信心位點、迴歸估腫瘤 DNA 比例
一個比例數值
③
依該比例調整門檻,跑四個過濾器重新校準候選變異
高信心的 somatic VCF
④
用校準後的變異重新 haplotag,並處理 HP3
最終 HP 標籤
變異呼叫本身不在這四個階段裡:tumor 的候選變異由 ClairSClairS配對 tumor–normal 的 somatic variant caller。tumor-only 版本叫 ClairS-TO。A somatic variant caller for matched tumour–normal data; ClairS-TO is the tumour-only version.完整條目 → 或 DeepSomatic 提供,
normal 的 germline 變異由 Clair3Clair3以深度學習做 germline variant calling 的工具,長 read 上常用。A deep-learning germline variant caller widely used for long reads.完整條目 → 或 DeepVariant 提供,並先 phasephasing 定相為可判定的 variants 建立其位於不同實體染色體拷貝上的相位關係;VCF 常以 0|1 等形式表示。長 read 可提供跨位點的分子層級觀測證據。Determining which variants lie on the same physical chromosome, turning 0/1 into 0|1. Long reads provide direct molecular evidence.完整條目 → 成 HP1/HP2 骨架。
LongPhase-S 是接在它們後面的一層。
若套用適合純腫瘤的嚴格門檻,真變異可能被當成雜訊而漏掉,recallrecall 召回率在指定評估範圍內,所有真陽性中被找出的比例:TP / (TP + FN)。若只對固定的 caller 候選集做後處理,recall 只能維持或下降,不能恢復 caller 從未輸出的真陽性;報告時應說清楚這個候選集範圍。Within a stated evaluation scope, the fraction of true positives that were found: TP / (TP + FN). With a fixed caller candidate set, post-filtering can preserve or lower recall but cannot recover variants the caller never emitted; report the candidate-set scope explicitly.完整條目 → 下降。
若一律放寬,高比例樣本的偽陽性則可能增加。
因此門檻不是一組固定數字:流程先估比例,再依比例級距選用對應的那組參數。
這一步只需要「支持 read 大概會有幾條」,而那正是 DNA 比例決定的量 ——
過濾這件事本來就不需要細胞比例。
GHIR:兩條 germline haplotype 的 read 有多偏
GHIRGHIR 生殖系單倍型失衡比Germline Haplotype Imbalance Ratio:在候選 somatic 位點上,取標為 HP1 與 HP2 的 read 數中較大者除以兩者之和,值域為 0.5 至 1。須注意兩件事:分母只含這兩個 germline 計數,HP1-1/HP2-1/HP3 皆不在內;且這些標籤是整條 read 的判定(該 read 任一處帶 somatic 等位即離開 germline 計數),並非該位點的等位計數。它不是直接的 purity 讀數:拷貝數變異、LOH、read 跨距內的突變密度、標記錯誤與抽樣不足都會使它偏移。Germline Haplotype Imbalance Ratio: at a candidate somatic locus, the larger of the HP1 and HP2 tagged read counts divided by their sum, ranging from 0.5 to 1. Two caveats: the denominator contains only those two germline counts, excluding HP1-1/HP2-1/HP3; and those tags are whole-read decisions (a read carrying a somatic allele anywhere leaves the germline counts), not per-locus allele counts. It is not a direct purity readout — copy number, LOH, mutation density within the read span, tagging error and sparse sampling all shift it.完整條目 →(Germline Haplotype Imbalance Ratio)在每一個候選 somatic 位點上各算一次:
本來就已經失衡的位點(多為 germline 結構變異或 CNVcopy number 拷貝數某段基因體在細胞內的拷貝數。多數正常常染色體區段為 2,可再分為 major 與 minor allele copy number。How many copies of a genomic segment a cell carries; a typical diploid autosomal segment has two. It can be split into major- and minor-allele copy numbers.完整條目 →)
上圖那五個「已知比例」的樣本,是把純腫瘤細胞株的 BAM 與正常樣本的 BAM
依 read 數混合出來的。所以標籤 0.2 的意思是「這堆 read 裡有兩成來自腫瘤細胞株」——
它控制的是分子的比例,也就是 tumour DNA fractiontumour DNA fraction 腫瘤 DNA 比例樣本 DNA 中源自腫瘤的比例。與 tumor purity(細胞比例)在 aneuploid 或 WGD 的情況下會不一樣。The fraction of DNA originating from tumour cells; diverges from cellular purity under aneuploidy or WGD.完整條目 →,一般記作 。
迴歸學到的是「GHIR 的中位數與四分位距 → 這個標籤」。
既然標籤是 ,學到的曲線輸出的當然也是 。
要把它轉成 cellular puritytumour purity 腫瘤純度樣本中腫瘤細胞所佔的比例。purity 越低,somatic 訊號被正常細胞稀釋得越嚴重,偵測越困難。The fraction of cells in a sample that are tumour cells. Low purity dilutes somatic signal.完整條目 → ,需要腫瘤的 倍體ploidy 倍性細胞內染色體套數。多數正常常染色體為 2(diploid);腫瘤可呈現 3、4 或其他非整倍體狀態。The number of chromosome sets in a cell. Most normal autosomal regions are diploid (2); tumours may be triploid, tetraploid or otherwise non-diploid.完整條目 → :
該位點在 normal 樣本的 VAFVAF 變異等位基因頻率在某個位點上,支持 alt allele 的 read 佔全部 read 的比例。VAF 不等於帶有這個突變的細胞比例。Variant Allele Frequency: the fraction of reads supporting the alternate allele. Not the same as the fraction of cells carrying it.完整條目 →
大範圍的 LOHLOH 異型合子性喪失原本 heterozygous 的區域變成只剩一種 allele。LOH 不等於缺失 —— 也可能是一條 haplotype 遺失後另一條被複製(copy-neutral LOH)。Loss of heterozygosity: a formerly het region retains only one allele. Not necessarily a deletion.完整條目 → 或 CNVcopy number 拷貝數某段基因體在細胞內的拷貝數。多數正常常染色體區段為 2,可再分為 major 與 minor allele copy number。How many copies of a genomic segment a cell carries; a typical diploid autosomal segment has two. It can be split into major- and minor-allele copy numbers.完整條目 →。GHIR 量的是 haplotype 失衡,
而 LOH 與 aneuploidy 造成的失衡跟腫瘤佔多少無關 —— 若沒被 LCVF 濾掉,就會被誤讀成高比例。
值得注意的是:研究實際觀察到的結果沒有出現這個問題,
即使資料集中有帶染色體規模 LOH 的細胞株(例如 HCC1395),估計仍未受影響;
論文推測原因是這個估計本身是 phase-aware 的,而以等位比例或拷貝數為基礎的方法更容易被這類事件帶偏。
所以這一項要當成「先檢查 LCVF 有沒有正常工作」,而不是預設它一定出錯。
TINCTINC 腫瘤污染正常樣本Tumour-in-normal contamination:配對正常樣本含有腫瘤來源訊號,使 genuine somatic 變異也可能在 normal 中被觀察到,因而可能被錯誤過濾。Tumour-in-normal contamination: tumour cells present in the matched normal, causing true somatic variants to produce signal there as well.完整條目 →。如果 normal 樣本含有腫瘤 DNA 污染,germline 基準線可能發生偏移。
模型外推。迴歸是在六個細胞株的合成樣本上訓練的。
如果你的樣本在某個維度上不像那批訓練資料(不同癌別、不同 basecallerbasecalling 鹼基判讀把定序儀的原始訊號(ONT 是電流)轉成 A/C/G/T 字母的步驟。不同 basecaller 版本會產生不同的錯誤特性 —— 所以「同一個細胞株、不同 basecaller」是兩份不同的資料集。Converting raw sequencer signal into A/C/G/T. Different basecallers yield different error profiles.完整條目 →、不同 coverage),
預測可能不可靠。
由上而下就是程式實際跑的順序,不是隨便排的。
③ 需要 ① 算出來的 LOH 區段才做得下去;④ 需要 ③ 的定相結果才算得出來。
而 ④ 算出來的數字如果很高,還會回頭把 ③ 重跑一次 —— 這條回饋線是這個流程唯一的循環。
第一件事:哪些區段只剩一條染色體
人的每個位置本來有兩份,一份來自父親、一份來自母親。
腫瘤細胞常常整段整段地弄丟其中一份,這叫 LOHLOH 異型合子性喪失原本 heterozygous 的區域變成只剩一種 allele。LOH 不等於缺失 —— 也可能是一條 haplotype 遺失後另一條被複製(copy-neutral LOH)。Loss of heterozygosity: a formerly het region retains only one allele. Not necessarily a deletion.完整條目 →。
LOH 區段對後面每一步都有影響,所以要先找出來。
怎麼找?看一個很簡單的數字:這個區段裡,「兩份不一樣」的位置佔多少。
兩份都在的時候,一個人身上大約每一千個鹼基就有一個位置是一邊 A、一邊 G,
這種位置叫 雜合heterozygous 異型合子在一般二倍體位點上,兩份 allele 不同(例如一份 A、一份 G)。此類位點可作為區分 haplotype 的資訊錨點;複雜拷貝數情況需另行解讀。At a typical diploid locus, carrying two different allele copies (for example A and G). These sites can anchor phasing; copy-number changes require separate interpretation.完整條目 →。少了一份之後,這些位置只剩一種 allele,比例就會塌下來。
三層。① 區段的邊界不是憑空切的:read 的一端對不上參考序列時會被剪掉(soft clippingsoft clippingread 端部未與參考序列對齊,但序列仍保留在 BAM 中(CIGAR 為 S)。大量 clipping 集中於同一位置時,可能提示結構變異斷點,仍需其他證據確認。A read end that failed to align but is retained in the BAM (CIGAR S). Clipping clustered at one position may suggest a structural breakpoint and needs corroborating evidence.完整條目 →,M4 的 CIGAR S),
而斷點附近會有一整排 read 在同一個位置被剪。方向相反又靠得很近的兩個剪痕會被配成一對,代表中間夾著一個小事件;配不到對的落單剪痕就當成大區段的邊界。
② 每個區段算一次雜合比例,低於門檻的標成 LOH。③ 關鍵的一步:算比例時先把小事件區間裡的變異扣掉,否則一段連續的 LOH 會被切成好幾塊。
候選變異不是 LongPhase-TO 自己叫出來的 —— 那是 ClairSClairS配對 tumor–normal 的 somatic variant caller。tumor-only 版本叫 ClairS-TO。A somatic variant caller for matched tumour–normal data; ClairS-TO is the tumour-only version.完整條目 →-TO 或 DeepSomatic-TO 的工作。
LongPhase-TO 做的是再校正(recalibrationrecalibration 再校正以額外證據重新評估 caller 已輸出的 candidates。若僅處理既有候選,可移除 false positives,但不能恢復 caller 未輸出的變異。Re-judging a caller's candidates with extra evidence. When the candidate set is fixed, it can remove false positives but cannot recover variants the caller never emitted.完整條目 →):把明顯不合理的候選挑掉。
第一關很直接:拿去對 PONPON 正常樣本面板/族群資料庫(依流程而異)PON 依分析流程有兩種用法,兩者不可互換:本教材的 tumor-only 流程以它指稱 population germline database set(如 1000G、CoLoRSdb、dbSNP、gnomAD 的聯集);GATK Mutect2 的 Panel of Normals(PoN)則由多個正常樣本建立,用來標記反覆出現的技術性 artifact。兩者都不能完全取代同一病人的 matched normal。The abbreviation is workflow-dependent and the two uses are not interchangeable. In this tutorial's tumour-only workflows, PON denotes a population germline database set; in GATK Mutect2, a Panel of Normals (PoN) is built from multiple normal samples to flag recurrent technical artifacts. Either can miss a private germline variant or an unseen artifact, so neither replaces a patient's matched normal.完整條目 → 這種族群資料庫比。
大家都有的變異,多半是這個人與生俱來的,不是癌症造成的,先排除掉。
但族群資料庫收不到每個人的私有變異,所以過了這一關還不夠。
第二關才是重點,而且它用的證據只有 long read 給得起:
看支持這個候選的 read,左右兩邊的鄰居是什麼。
兩邊帶 T 的 read 刻意都是三條 —— 光數數量分不出真假,這正是重點。
差別在鄰居:左邊三條的鄰居都是 C 與 A,也就是三條全落在同一條染色體上;右邊三條各有各的鄰居,湊不出一致的來源。
下半部把同一批資料畫成路徑圖後更清楚:左邊多出一條完整的新路徑,而且它跟 HP2 只差候選這一格;右邊則沒有任何一條新路徑被穩定支持。
但這裡要問的是另一個問題:候選 allele 在左右兩側是不是都指向同一條 haplotype。
只有一個鄰居時,「跟左邊一致」很容易碰巧發生;要有左鄰居和右鄰居,
才看得出「這條新路徑從頭到尾只跟某一條差一格」。所以最小單位是三個位置,
這個結構就叫 triplet graphtriplet graphLongPhase-TO 判斷候選變異真假時用的最小結構:候選位置加上左右各一個變異,共三個位置。把 read 上看到的 allele 組合畫成路徑後,真的 somatic 變異會形成一條「跟某一條 germline haplotype 只差候選這一格」的新路徑;定序錯誤則湊不出一致的路徑。因為左右兩側都要對得上,所以三個位置是最小單位。LongPhase-TO's minimal unit for judging a candidate: the candidate plus one flanking variant on each side. A true somatic allele forms a third path differing from one parental haplotype at the candidate alone; artifacts show no consistent path.完整條目 →。
檢體不會是純的癌細胞,裡面一定混著正常細胞。混得多寡直接影響每個變異的訊號強度,
所以這個比例要估出來。M12 用的是 GHIRGHIR 生殖系單倍型失衡比Germline Haplotype Imbalance Ratio:在候選 somatic 位點上,取標為 HP1 與 HP2 的 read 數中較大者除以兩者之和,值域為 0.5 至 1。須注意兩件事:分母只含這兩個 germline 計數,HP1-1/HP2-1/HP3 皆不在內;且這些標籤是整條 read 的判定(該 read 任一處帶 somatic 等位即離開 germline 計數),並非該位點的等位計數。它不是直接的 purity 讀數:拷貝數變異、LOH、read 跨距內的突變密度、標記錯誤與抽樣不足都會使它偏移。Germline Haplotype Imbalance Ratio: at a candidate somatic locus, the larger of the HP1 and HP2 tagged read counts divided by their sum, ranging from 0.5 to 1. Two caveats: the denominator contains only those two germline counts, excluding HP1-1/HP2-1/HP3; and those tags are whole-read decisions (a read carrying a somatic allele anywhere leaves the germline counts), not per-locus allele counts. It is not a direct purity readout — copy number, LOH, mutation density within the read span, tagging error and sparse sampling all shift it.完整條目 →;這裡的比值算法一樣,差別只在數的是帶 reference allele 的 read。
程式把它寫成 ##tumor_purity=,但它量到的其實是
腫瘤 DNA 佔的比例tumour DNA fraction 腫瘤 DNA 比例樣本 DNA 中源自腫瘤的比例。與 tumor purity(細胞比例)在 aneuploid 或 WGD 的情況下會不一樣。The fraction of DNA originating from tumour cells; diverges from cellular purity under aneuploidy or WGD.完整條目 →,不是腫瘤細胞佔的比例。
在實驗室依覆蓋度混出來的合成樣本裡,兩者恰好相等,所以平常不容易察覺差別。
但真實腫瘤如果是 非整倍體aneuploidy 非整倍體染色體數目異常。腫瘤裡非常普遍,也可能使 cellular purity 與 DNA fraction 不一致。An abnormal chromosome count. It is common in tumours and can make cellular purity differ from tumour DNA fraction.完整條目 →或發生過全基因體加倍,一個癌細胞貢獻的 DNA 比正常細胞多,
兩個數字就會分開。報告數字時要說清楚是哪一個。
估計 puritytumour purity 腫瘤純度樣本中腫瘤細胞所佔的比例。purity 越低,somatic 訊號被正常細胞稀釋得越嚴重,偵測越困難。The fraction of cells in a sample that are tumour cells. Low purity dilutes somatic signal.完整條目 →
0.6(本題簡化模型中視為 tumor DNA fraction)
該區段 copy numbercopy number 拷貝數某段基因體在細胞內的拷貝數。多數正常常染色體區段為 2,可再分為 major 與 minor allele copy number。How many copies of a genomic segment a cell carries; a typical diploid autosomal segment has two. It can be split into major- and minor-allele copy numbers.完整條目 →
2(無 CNV,無 LOHLOH 異型合子性喪失原本 heterozygous 的區域變成只剩一種 allele。LOH 不等於缺失 —— 也可能是一條 haplotype 遺失後另一條被複製(copy-neutral LOH)。Loss of heterozygosity: a formerly het region retains only one allele. Not necessarily a deletion.完整條目 →)
周圍 200 bp 內的 het germline SNP
2 個,皆已 phase
支持 ALT 的 11 條 read 的 HP tagHP tagBAM 裡標示某條 read 屬於哪一條 haplotype 的 tag。LongPhase-S 的 germline haplotag 寫成整數 HP:i:1;somatic haplotag 與 LongPhase-TO 則寫成字串 HP:Z:1-1。BAM tag assigning a read to a haplotype. Integer HP:i:1 for germline haplotag; string HP:Z:1-1 for somatic.完整條目 →
PONPON 正常樣本面板/族群資料庫(依流程而異)PON 依分析流程有兩種用法,兩者不可互換:本教材的 tumor-only 流程以它指稱 population germline database set(如 1000G、CoLoRSdb、dbSNP、gnomAD 的聯集);GATK Mutect2 的 Panel of Normals(PoN)則由多個正常樣本建立,用來標記反覆出現的技術性 artifact。兩者都不能完全取代同一病人的 matched normal。The abbreviation is workflow-dependent and the two uses are not interchangeable. In this tutorial's tumour-only workflows, PON denotes a population germline database set; in GATK Mutect2, a Panel of Normals (PoN) is built from multiple normal samples to flag recurrent technical artifacts. Either can miss a private germline variant or an unseen artifact, so neither replaces a patient's matched normal.完整條目 → 查詢 —— 不在這個族群資料庫裡,會降低常見 germline 的可能性,
但排除不了病人私有的 germline 變異(M7 的核心限制)。
haplotype 結構 —— 這是 LongPhase-TO 主要利用的證據。
用左右兩個 germline SNP 建 triplet graphtriplet graphLongPhase-TO 判斷候選變異真假時用的最小結構:候選位置加上左右各一個變異,共三個位置。把 read 上看到的 allele 組合畫成路徑後,真的 somatic 變異會形成一條「跟某一條 germline haplotype 只差候選這一格」的新路徑;定序錯誤則湊不出一致的路徑。因為左右兩側都要對得上,所以三個位置是最小單位。LongPhase-TO's minimal unit for judging a candidate: the candidate plus one flanking variant on each side. A true somatic allele forms a third path differing from one parental haplotype at the candidate alone; artifacts show no consistent path.完整條目 →,看支持這個候選的 read
是不是都落在同一條 haplotype 上(M13 的第二件事)。本例 10:1 偏向 HP1,符合這個樣態。
假設你在這個 chr9 評估資料集上跑完流程,並與所採用的 SEQC2 truth set 比對後得到:
先將結果分為 TP、FP、FN,再說明 precision、recall 與 F1 各自使用哪些計數。以下數值僅適用於本例的評估設定:若 FP 減半且未誤刪 TP,precision 增加 0.080,而 F1 增加 0.024;recall 不變。
算出 precisionprecision 精確率所有 calls 中實際為真陽性的比例:TP / (TP + FP)。後處理常以提升此指標為目標。Of the calls you made, the fraction that are correct: TP / (TP + FP).完整條目 →、recallrecall 召回率在指定評估範圍內,所有真陽性中被找出的比例:TP / (TP + FN)。若只對固定的 caller 候選集做後處理,recall 只能維持或下降,不能恢復 caller 從未輸出的真陽性;報告時應說清楚這個候選集範圍。Within a stated evaluation scope, the fraction of true positives that were found: TP / (TP + FN). With a fixed caller candidate set, post-filtering can preserve or lower recall but cannot recover variants the caller never emitted; report the candidate-set scope explicitly.完整條目 →、F1F1precision 與 recall 的調和平均。當資料裡 FN 遠多於 FP 時,precision 的改善對 F1 的影響有限 —— 見 M10。The harmonic mean of precision and recall. A precision gain has little effect when false negatives dominate the evaluation.完整條目 →。然後回答:
如果你加一個後處理過濾器,把 FP 從 88 降到 44,F1 會變成多少?這個改善值得嗎?
展開答案
目前:
把 FP 減半之後(假設未誤刪 TP,屬於理想化情境):
precision = 412/456 = 0.904(+0.080)
recall = 0.551(不變)
F1 = 0.685(只有 +0.024)
值得注意的是這兩個數字動的幅度差很多。precision 增加 8 個百分點,
F1 只動了 2.4 個百分點 —— 因為 FN(335)遠多於 FP(88),
F1 被 recall 那一側綁住了。只要 FN 一直是大宗,再怎麼壓 FP,F1 都不會有大改變。
腫瘤並非由單一均質的細胞群體構成。它自單一細胞起始,
於分裂過程中持續累積 somatic 變異somatic variant 體細胞變異在非生殖系細胞譜系中、通常於受精後取得的變異;可只存在於部分細胞,通常不由親代遺傳給子代。癌症基因體學常分析此類變異。A variant acquired outside the germline, usually after conception. It may be restricted to a subset of cells and is generally not passed to offspring; it is a central focus of cancer genomics.完整條目 →,因而分化為數個群體,
各自帶有不同的變異組合;M2 稱這些群體為 cloneclone 克隆源自共同祖先細胞,並共享一組可辨識 somatic mutations 的細胞群。A population of tumour cells descended from a common ancestor and sharing the same set of somatic mutations.完整條目 → 與 subclonesubclone 次克隆clone 內取得額外變異並擴增的子群;治療後可能富集具有抗性的 subclone。A subset of a clone that acquired further mutations and expanded. Treatment-resistant populations are often subclones.完整條目 →。
這些群體之間存在先後關係:某一群衍生自另一群。
將「何者衍生自何者」表示為圖,即得一棵樹 ——
節點為一群細胞,邊代表「在祖先既有的變異之上再取得一個」。
此樹稱為 clone treeclone tree 克隆演化樹描述腫瘤內各群細胞祖先關係的樹:節點是一群帶有相同變異組合的細胞,邊代表在祖先之上又多拿到變異。要注意同一組群集常常有多棵樹同時相容。A tree describing ancestral relationships among cell populations in a tumour: nodes are groups of cells sharing a mutation set, edges represent additional mutations acquired on top of an ancestor. Multiple trees are often compatible with the same clusters.完整條目 →,即為以下三頁所要重建的對象。
四個轉折並非彼此取代,而是各自補足前一階段所不可見者。
① 以頻率將變異分群,首次使「腫瘤含幾群細胞」成為可計算的問題。
② 發現頻率分布本身具有形狀,可用於判定是否存在天擇。
③ 大規模評比顯露前兩者共同的上限:答案不唯一。
④ 兩種新的資訊來源 —— 甲基化提供解析度更高的時鐘,長讀提供分子層級的共現關係。
第一步:將一批變異的頻率繪為直方圖 —— mutation frequency spectrum
設一份腫瘤樣本含數千個 somatic 變異,各自具有其 VAFVAF 變異等位基因頻率在某個位點上,支持 alt allele 的 read 佔全部 read 的比例。VAF 不等於帶有這個突變的細胞比例。Variant Allele Frequency: the fraction of reads supporting the alternate allele. Not the same as the fraction of cells carrying it.完整條目 →。
將這數千個 VAF 繪為直方圖,可得一張具有結構的圖,
稱為 mutation frequency spectrummutation frequency spectrum 突變頻率譜把一份樣本裡所有 somatic 變異的 VAF 畫成直方圖後得到的分布。分布上的峰與肩對應不同大小的細胞群,最低頻端的尾巴斜率則被用來判斷有沒有天擇。The distribution obtained by histogramming the VAFs of all somatic variants in a sample. Peaks and shoulders correspond to cell populations of different sizes; the slope of the low-frequency tail is used to test for selection.完整條目 →。
三塊結構各有其意義。
① 最右側的峰對應全部腫瘤細胞皆帶有的變異,其位置由 純度tumour purity 腫瘤純度樣本中腫瘤細胞所佔的比例。purity 越低,somatic 訊號被正常細胞稀釋得越嚴重,偵測越困難。The fraction of cells in a sample that are tumour cells. Low purity dilutes somatic signal.完整條目 →決定(拷貝數正常、變異僅位於其中一份拷貝上時,恰為 purity ÷ 2)。
② 左側的肩對應僅一部分腫瘤細胞帶有的變異 —— 每一個肩對應一群 subclonesubclone 次克隆clone 內取得額外變異並擴增的子群;治療後可能富集具有抗性的 subclone。A subset of a clone that acquired further mutations and expanded. Treatment-resistant populations are often subclones.完整條目 →。
③ 中性演化的檢定所用者為中段(VAF 0.12–0.24),且其量測對象並非此直方圖 ——
須先將變異數累積後對 作圖,該線方為直線。
最左側三根標記叉號的長條為假陽性區:深度不足時定序錯誤亦落於該處,不可用於擬合。
區分 ① 與 ② 即為分群;③ 所問者為是否存在天擇。
2014 至 2015 年間出現三項影響最大的做法,其思路均為對此圖分群:
SciClone(2014)以 beta 混合模型(以數個鐘形分布擬合一張直方圖)擬合 VAF,
且主要僅採用拷貝數正常、無 LOHLOH 異型合子性喪失原本 heterozygous 的區域變成只剩一種 allele。LOH 不等於缺失 —— 也可能是一條 haplotype 遺失後另一條被複製(copy-neutral LOH)。Loss of heterozygosity: a formerly het region retains only one allele. Not necessarily a deletion.完整條目 → 的區段,因為僅該處的 VAF 易於解釋。
PyClone(2014)改以 Dirichlet process(不需預先固定群數的分群模型)
對「帶有該變異的細胞比例」分群,並將 拷貝數copy number 拷貝數某段基因體在細胞內的拷貝數。多數正常常染色體區段為 2,可再分為 major 與 minor allele copy number。How many copies of a genomic segment a cell carries; a typical diploid autosomal segment has two. It can be split into major- and minor-allele copy numbers.完整條目 →與正常細胞的混入一併納入模型。
PhyloWGS(2015)進一步將分群與建樹合為同一個推論。
前述分群的對象並非 VAF,而是 CCFcancer cell fraction 癌細胞比例帶有某個特定突變的腫瘤細胞佔全部腫瘤細胞的比例。用來區分 clonal()與 subclonal()突變。不等於 VAF。The fraction of tumour cells carrying a given mutation; distinguishes clonal from subclonal. Not the same as VAF.完整條目 →。
理由已見於 M8:VAF 為「帶變異的 read 所佔比例」,CCF 為「帶變異的癌細胞所佔比例」,
兩者相差一段換算。
換算式的分母並非 2:
由 VAF 至 CCF 須經過三個估計值。
純度與拷貝數來自其他程式,各自帶有誤差;
multiplicitymutation multiplicity在帶有某突變的細胞中,該突變所占的拷貝數;是 VAF 與 cancer cell fraction 換算時的重要參數。How many mutant-allele copies are present in a cell carrying the mutation; a necessary term when converting VAF to cancer cell fraction.完整條目 →(變異位於幾份腫瘤拷貝上)通常僅能在 1 與 2 之間推定。
推定偏差一格,CCF 即差兩倍 —— 而 CCF 相差兩倍已足以將 clonal 變異誤讀為 subclonal。
下方三條軸所繪為同一個 VAF = 0.30 在不同假設下的落點。
因此頻率路線的第一層不確定性源於換算:CCF 為三個估計值的商,
而三者皆可能偏誤。此層已於 M8 標記為「重要區分 #3:VAF 不等於 cancer cell fraction」。
設換算完全正確,得到三群變異,CCF 分別為 1.0、0.6、0.3。
其次須將此三群排為一棵 clone treeclone tree 克隆演化樹描述腫瘤內各群細胞祖先關係的樹:節點是一群帶有相同變異組合的細胞,邊代表在祖先之上又多拿到變異。要注意同一組群集常常有多棵樹同時相容。A tree describing ancestral relationships among cell populations in a tumour: nodes are groups of cells sharing a mutation set, edges represent additional mutations acquired on top of an ancestor. Multiple trees are often compatible with the same clusters.完整條目 →,即確定何者為何者的祖先。
在此情形下,方法必須引入額外的偏好。最常用者為 簡約parsimony 簡約法在所有與資料相容的解裡,選步數(或成本)最少的那一個。它是一個偏好而不是證據 —— 當多個解並列時,簡約法決定選哪一個,但資料本身沒有排除其他解。Choosing, among all solutions compatible with the data, the one requiring the fewest steps or lowest cost. It is a preference rather than evidence: when several solutions tie, parsimony picks one, but the data has not ruled the others out.完整條目 →:
於所有相容的樹中選取步數最少者。此為合理的偏好,但它是偏好,而非證據。
此外,為使樹得以連通,有時須補入未獲任何資料直接支持的中間狀態,
此類節點稱為 latent nodelatent node 潛在節點建樹時為了讓圖連得起來而補進的中間狀態,沒有被任何 read 直接觀測到。它是模型的產物,不能當成「還沒觀察到的細胞」。An intermediate state added during tree construction to keep the graph connected, not directly observed in any read. It is a product of the model and must not be read as an unobserved cell population.完整條目 →。
成因:頻率為 marginal distribution,樹所需者為 joint distribution
前述兩層不確定性表面上為兩個獨立的問題,實則同一。以兩個位點即可完整說明。
設一個連鎖視窗內有兩個 somatic 位點 A、B。頻率給出兩個數字:
A 出現於 50% 的細胞、B 出現於 30% 的細胞。
此二數字各自描述單一位點,此類分布稱為 marginal distributionmarginal distribution 邊際分布只描述單一變數的分布。VAF 與甲基化 β 值都是邊際統計:它們各自只講一個位點有多少比例帶有標記,不講兩個位點在同一個分子或同一個細胞上的搭配情形。A distribution over a single variable. VAF and methylation beta values are both marginal statistics: each describes one site in isolation and says nothing about how two sites co-occur on the same molecule or cell.完整條目 →。
而樹所要回答的問題為:同時帶有 A 與 B 的細胞佔多少。
此為兩個位點聯合的分布,稱為 joint distributionjoint distribution 聯合分布多個變數一起看的分布,也就是「哪些組合各出現多少」。演化樹的形狀取決於聯合分布,而不是邊際分布;長讀的價值在於一條分子上的組合是直接觀測到的。A distribution over several variables jointly, i.e. how often each combination occurs. Tree topology depends on the joint distribution, not the marginals; the value of long reads is that combinations on one molecule are observed directly.完整條目 →。
由邊際分布無法推得聯合分布 —— 此非統計方法不足,而是該資訊確實不存在於資料中。
既然序列上的頻率不足,更換標記是否有所改善?甲基化5mC 5-甲基胞嘧啶胞嘧啶第 5 個碳上的甲基化修飾。這是常見的 DNA 甲基化形式,主要見於 CpG;可影響轉錄調控,效應依基因組位置與細胞類型而異。5-methylcytosine, the commonest DNA methylation mark, usually at CpG sites. Affects expression without changing sequence.完整條目 →為自然的候選。
其吸引力在於速率。epimutationepimutation 表觀突變細胞分裂時甲基化狀態的隨機翻轉。文獻常引的量級是每個 CpG、每次分裂 到 ,而 DNA 突變約每鹼基 到 ,同量級對比約差五個數量級,所以甲基化是解析度更高的譜系時鐘。要注意這個速率隨位點與量測方式差異很大(有估到 的),應當成量級而非定值。A stochastic flip of methylation state during cell division. Commonly cited at – per CpG per division versus – per base for DNA mutation — roughly five orders of magnitude faster like-for-like — which makes it a higher-resolution lineage clock. Treat it as an order of magnitude, not a constant: estimates vary by site and assay, up to ~.完整條目 →(甲基化狀態的隨機翻轉)的發生速率,
文獻常引的量級為每個 CpGCpG序列上一個 C 後接一個 G(p 代表兩者之間的磷酸鍵)。哺乳類多數 5mC 位於 CpG,特定細胞或情況亦可見非 CpG 甲基化。A cytosine followed by a guanine. Most mammalian 5mC occurs at CpG sites, although non-CpG methylation is also observed in particular cells and contexts.完整條目 →、每次細胞分裂 至 ;DNA 突變則為每個鹼基 至 。
同量級相比,快約五個數量級。此即表示甲基化為解析度高得多的時鐘 ——
兩群細胞分離僅數十代時,序列上尚無任何差異,甲基化上已可觀測。
一段區域內的四個 CpG,左右兩側各位點的 β 值皆為 0.50,
逐位點統計完全無法區分。
左:每個分子皆為雜亂的半甲基化,對應同一群細胞的隨機起伏。
右:分子分為兩組,一組全甲基化、一組全不甲基化。
兩者的差異同樣僅在同一分子上的四個 CpG 聯合觀測時出現。
將單一分子上的甲基化組合視為一個單位,該單位稱為 epialleleepiallele 表觀等位型同一條分子上一組相鄰 CpG 的甲基化組合,例如四個 CpG 的 1101。逐位點的 β 值看不到它 —— 要算 epiallele 的組成,一條 read 至少得跨過 4 個 CpG。The combination of methylation states across a set of neighbouring CpGs on a single molecule, e.g. 1101 over four CpGs. Per-site beta values cannot resolve it; computing epiallele composition requires reads spanning at least four CpGs.完整條目 →。
惟右側的兩組未必對應兩群細胞,見下方警告。
至此轉折已明確。一條 long readlong read 長讀單條可達數千至數萬鹼基的定序片段(例如 ONT、PacBio)。若可靠地同時覆蓋多個 variant,可提供它們位於同一 DNA 分子上的直接觀測證據。A sequencing read thousands to tens of thousands of bases long. Its length lets one molecule link multiple variants directly.完整條目 → 對應一個分子,
其在多個位點上的 allele 為一次觀測所得的組合,而非由兩個邊際推得。
亦即:聯合分布不需推估,可直接計數。
尚有第二層結構須先區分。somatic 變異somatic variant 體細胞變異在非生殖系細胞譜系中、通常於受精後取得的變異;可只存在於部分細胞,通常不由親代遺傳給子代。癌症基因體學常分析此類變異。A variant acquired outside the germline, usually after conception. It may be restricted to a subset of cells and is generally not passed to offspring; it is a central focus of cancer genomics.完整條目 →發生於某一細胞的某一條染色體上,
故同一個 somatic 位點的變異僅會出現於其中一條 haplotypehaplotype 單倍型同一條實體染色體拷貝上,具有一致相位的一組 alleles 或 variants。HP1 與 HP2 是任意的相對標籤。A set of variants that lie on the same physical chromosome copy and are inherited together.完整條目 → 上。
先將 read 依 單倍型家族單倍型家族 一條 germline 單倍型,加上由它衍生的 somatic 單倍型把 read 依 HP tag 分成的兩組之一。家族一是 HP1 與從它長出來的 HP1-1,家族二是 HP2 與 HP2-1;歸不到任何一條 germline 單倍型的 HP3 不屬於任何一族。叫「家族」是因為它把一條 germline 單倍型與由它衍生的 somatic 單倍型收在同一組裡 —— 分組看的是 germline 那一層,不是有沒有帶 somatic 突變。
在同一個 phase block 內,一個家族對應一條染色體拷貝,所以「兩族」就是那個位置上的兩條同源染色體。但兩件事不成立:其一,軟體判定不出哪一族來自父親、哪一族來自母親(那需要另外定序父母);其二,標號只在該 phase block 內有定義,跨 block 的「家族一」並非同一條染色體。One of the two groups reads are split into by HP tag. Family 1 is HP1 plus the somatic haplotype HP1-1 derived from it; family 2 is HP2 plus HP2-1; HP3, which cannot be assigned to either germline haplotype, belongs to neither. It is called a family because it groups a germline haplotype together with the somatic haplotypes descended from it — the split is by the germline layer, not by whether a read carries a somatic mutation. Within one phase block a family corresponds to one chromosome copy, so the two families are the two homologous chromosomes at that locus. Two things do not follow: which family is paternal cannot be determined without sequencing the parents, and the labels are defined only within that phase block.完整條目 →分為兩組再各自建樹,
一方面符合生物學,一方面使各組的狀態空間減半。
分析單位的切分方式。
① 先須落於同一個 phase setphase block一段可建立連續相位關係的區域。read 長度不足、缺少 informative heterozygous 位點或證據不一致時,可能形成不同 phase blocks。A contiguous stretch over which phasing is consistent. It can break when reads are too short, informative heterozygous sites are absent or evidence conflicts.完整條目 → 之內 —— 跨 phase set 的相位關係本即無定義。
② 再依 read 實際可跨越的距離切為 連鎖視窗連鎖視窗同一個 phase block 之內,由 read 連鎖的傳遞閉包所界定的一段:凡有某條 read 同時覆蓋至少兩個 somatic 位點,該段即成一個視窗;另一條 read 疊到已連鎖的位點又碰到新位點時,兩段合併成更長的視窗。不是固定寬度,也不跨 phase block。早期估算比例時所用的「20 kb 視窗」是固定寬度的近似,兩者的計數不可互換。Within a single phase block, the span defined by the transitive closure of read linkage: any read covering at least two somatic sites opens a window, and overlapping reads merge windows. It is not a fixed width and never crosses a phase block. The fixed-width 20 kb window used in early proportion estimates is an approximation; the two counts are not interchangeable.完整條目 →,每個連鎖視窗含 k 個 somatic 位點(k 通常為 2 或 3)。
③每個連鎖視窗之內再依單倍型家族分為兩組。
故最小分析單位為 單倍型連鎖區段單倍型連鎖區段一個連鎖視窗再限定到一條單倍型家族 —— 也就是「一個連鎖視窗、一條染色體拷貝」。這是局部共現分析的最小單位:其中的 read 全部來自同一條染色體拷貝,另一條拷貝屬於另一個連鎖區段。因此一個連鎖視窗至多切出兩個連鎖區段,連鎖區段數大於連鎖視窗數,兩者不可混用為同一個分母。單倍型標號只在該區段內有定義,跨區段的同名標號並非同一條染色體。A linkage window restricted to one haplotype family — one window, one chromosome copy. It is the smallest unit of local co-occurrence analysis: all its reads come from the same chromosome copy, the other copy forming a separate segment. One window yields at most two segments, so segments outnumber windows and the two must not share a denominator. Haplotype labels are defined only within a segment.完整條目 →:一個連鎖視窗的一條單倍型家族,而非整條染色體。
上下兩層的 x 座標刻意對齊:兩者為同一批 read,僅重新分組。
M8 已提及此量,稱為 read-AFread-AF read 層級的 ALT 比例在同一個分析區域、同一個單倍型家族的 read 之中,某個 somatic 位點帶 ALT 的比例。分母已限縮到同一條 haplotype 的同一段區域,所以不必經過純度與拷貝數換算;用途是在步數並列的候選樹之間排序。The fraction of reads carrying the ALT allele at a somatic site, computed within one analysis window and one haplotype family. Because the denominator is already restricted, no purity or copy-number correction is needed; it is used to rank candidate trees of equal cost.完整條目 →:於同一連鎖視窗、同一單倍型家族的 read 之中,
某個 somatic 位點的 ALT 所佔比例。它與全基因體 VAF 的差別在於分母受限縮 ——
僅計入同一連鎖視窗、同一條 haplotype 的 read。
SciClone:Miller CA 等,SciClone: inferring clonal architecture and tracking the
spatial and temporal patterns of tumor evolution,PLoS Computational Biology 10(8):e1003665,2014。
PyClone:Roth A 等,PyClone: statistical inference of clonal population structure
in cancer,Nature Methods 11(4):396–398,2014。
methclone:Li S 等,Dynamic evolution of clonal epialleles revealed by methclone,
Genome Biology 15:472,2014。
PhyloWGS:Deshwar AG 等,PhyloWGS: reconstructing subclonal composition and
evolution from whole-genome sequencing of tumors,Genome Biology 16:35,2015。
甲基化與突變所建之樹互相吻合:Mazor T 等,DNA methylation and somatic mutations
converge on the cell cycle and define similar evolutionary histories in brain tumors,
Cancer Cell 28(3):307–317,2015。
1/f 中性演化檢定:Williams MJ 等,Identification of neutral tumor evolution
across cancer types,Nature Genetics 48(3):238–244,2016。
對中性檢定的質疑:Tarabichi M 等,Neutral tumor evolution?,
Nature Genetics 50(12):1630–1633,2018;以及 Bozic I 等,
On measuring selection in cancer from subclonal mutation frequencies,
PLoS Computational Biology 15(9):e1007368,2019。
epimutation 速率的整理:Chen S、Wu J、Gaiti F,Methylation-based lineage tracing
in cancer,Blood,2026(doi:10.1182/blood.2024028196)。文中的速率寫為「up to」,屬上界而非定值。
大規模評比:Salcedo A 等,Crowd-sourced benchmarking of single-sample tumor
subclonal reconstruction,Nature Biotechnology 43(4):581–592(2024 年線上發表,2025 年刊出)。
EVOFLUx:Gabbutt C、Duran-Ferrer M 等,Fluctuating DNA methylation tracks cancer
evolution at clinical scale,Nature 645:764–773,2025。程式碼在
github.com/Duran-FerrerM/evoflux。
本模組術語
5mC(5-甲基胞嘧啶)
胞嘧啶第 5 個碳上的甲基化修飾。這是常見的 DNA 甲基化形式,主要見於 CpG;可影響轉錄調控,效應依基因組位置與細胞類型而異。
CpG
序列上一個 C 後接一個 G(p 代表兩者之間的磷酸鍵)。哺乳類多數 5mC 位於 CpG,特定細胞或情況亦可見非 CpG 甲基化。
其餘 5% 的連鎖視窗所提供的資訊性質不同:它們給出的不是更多資料點,
而是資料點之間的關係 —— 哪兩個變異不可能屬於同一群、哪一群是哪一群的祖先。
而這類關係不需要大量:一棵 clone treeclone tree 克隆演化樹描述腫瘤內各群細胞祖先關係的樹:節點是一群帶有相同變異組合的細胞,邊代表在祖先之上又多拿到變異。要注意同一組群集常常有多棵樹同時相容。A tree describing ancestral relationships among cell populations in a tumour: nodes are groups of cells sharing a mutation set, edges represent additional mutations acquired on top of an ancestor. Multiple trees are often compatible with the same clusters.完整條目 → 所要確定的是群與群之間的少數幾個關係,
而非一萬五千個變異各自的歸屬。
在同一個 phase block 內,一個家族對應一條染色體拷貝,所以「兩族」就是那個位置上的兩條同源染色體。但兩件事不成立:其一,軟體判定不出哪一族來自父親、哪一族來自母親(那需要另外定序父母);其二,標號只在該 phase block 內有定義,跨 block 的「家族一」並非同一條染色體。One of the two groups reads are split into by HP tag. Family 1 is HP1 plus the somatic haplotype HP1-1 derived from it; family 2 is HP2 plus HP2-1; HP3, which cannot be assigned to either germline haplotype, belongs to neither. It is called a family because it groups a germline haplotype together with the somatic haplotypes descended from it — the split is by the germline layer, not by whether a read carries a somatic mutation. Within one phase block a family corresponds to one chromosome copy, so the two families are the two homologous chromosomes at that locus. Two things do not follow: which family is paternal cannot be determined without sequencing the parents, and the labels are defined only within that phase block.完整條目 →,以下簡稱連鎖區段)
都算過那一批 read 的共現狀態,並判定過屬於串接或分岔;
所需的原料因此已經在手上,但目前停留在連鎖視窗層級,未被傳遞至全基因體尺度。
此即 pigeonhole 原則pigeonhole 鴿籠原理由父代與子代的細胞比例限制樹形的算術規則:一個細胞至多屬於父節點底下的一個子節點,所以各子節點的細胞比例加起來不得超過父節點。名稱來自鴿籠原理 —— 東西放進籠子,總量不會憑空變多。在 subclone 重建的文獻中也稱為 sum rule 或 crossing rule。它只能排除樹,不能挑出樹:通過檢查的候選通常仍不只一棵,其餘要靠 parsimony 之類的偏好決定。The arithmetic constraint that limits tree shape from parent and child cell fractions: a cell belongs to at most one child of a given parent, so the children's cell fractions cannot sum to more than the parent's. The name comes from the pigeonhole principle. Also called the sum rule or crossing rule in the subclonal reconstruction literature. It can only rule trees out, never select one: the surviving candidates are usually more than one, and the choice among them falls to a preference such as parsimony.完整條目 →(鴿籠原理:東西放進籠子,總量不會憑空變多)。
subclone 重建的文獻中亦稱 sum rule 或 crossing rule,本頁一律用前者。
關鍵在於它只能排除樹,不能挑出樹:
通過此檢查的候選通常不只一棵,其餘由簡約性parsimony 簡約法在所有與資料相容的解裡,選步數(或成本)最少的那一個。它是一個偏好而不是證據 —— 當多個解並列時,簡約法決定選哪一個,但資料本身沒有排除其他解。Choosing, among all solutions compatible with the data, the one requiring the fewest steps or lowest cost. It is a preference rather than evidence: when several solutions tie, parsimony picks one, but the data has not ruled the others out.完整條目 →等偏好決定。
上篇的測驗即針對此點:通過相容性檢查與被資料證成是兩回事。
第 群的 CCFcancer cell fraction 癌細胞比例帶有某個特定突變的腫瘤細胞佔全部腫瘤細胞的比例。用來區分 clonal()與 subclonal()突變。不等於 VAF。The fraction of tumour cells carrying a given mutation; distinguishes clonal from subclonal. Not the same as VAF.完整條目 →:帶有該群突變的腫瘤細胞比例
由 clone 比例、tree 與指派共同決定
樣本純度tumour purity 腫瘤純度樣本中腫瘤細胞所佔的比例。purity 越低,somatic 訊號被正常細胞稀釋得越嚴重,偵測越困難。The fraction of cells in a sample that are tumour cells. Low purity dilutes somatic signal.完整條目 →:樣本中腫瘤細胞的比例
推導量:,其中 是正常細胞的比例(見第二節)
變異 所在腫瘤區段的 total copy number
通常由 copy-number 分析提供;不是 read depth
帶原腫瘤細胞中,幾份拷貝帶有此變異(multiplicitymutation multiplicity在帶有某突變的細胞中,該突變所占的拷貝數;是 VAF 與 cancer cell fraction 換算時的重要參數。How many mutant-allele copies are present in a cell carrying the mutation; a necessary term when converting VAF to cancer cell fraction.完整條目 →)
在同一個 phase block 內,一個家族對應一條染色體拷貝,所以「兩族」就是那個位置上的兩條同源染色體。但兩件事不成立:其一,軟體判定不出哪一族來自父親、哪一族來自母親(那需要另外定序父母);其二,標號只在該 phase block 內有定義,跨 block 的「家族一」並非同一條染色體。One of the two groups reads are split into by HP tag. Family 1 is HP1 plus the somatic haplotype HP1-1 derived from it; family 2 is HP2 plus HP2-1; HP3, which cannot be assigned to either germline haplotype, belongs to neither. It is called a family because it groups a germline haplotype together with the somatic haplotypes descended from it — the split is by the germline layer, not by whether a read carries a somatic mutation. Within one phase block a family corresponds to one chromosome copy, so the two families are the two homologous chromosomes at that locus. Two things do not follow: which family is paternal cannot be determined without sequencing the parents, and the labels are defined only within that phase block.完整條目 →就是依 HP tag 分出來的兩組 read 之一:
更完整地說, 包含細胞比例 、clone treeclone tree 克隆演化樹描述腫瘤內各群細胞祖先關係的樹:節點是一群帶有相同變異組合的細胞,邊代表在祖先之上又多拿到變異。要注意同一組群集常常有多棵樹同時相容。A tree describing ancestral relationships among cell populations in a tumour: nodes are groups of cells sharing a mutation set, edges represent additional mutations acquired on top of an ancestor. Multiple trees are often compatible with the same clusters.完整條目 → 、
突變指派 、逐拷貝基因型 ,以及觀測錯誤參數 。
由 、 與 決定, 由 決定。
不是一個錯誤率,而是一組觀測通道參數的縮寫,包括稍後的逐位點錯誤 與
家族誤標率 。
外部提供的 copy number 與校準後的 是條件輸入,為了避免式子過長沒有逐項寫在條件線右側。
在同一個 phase block 內,一個家族對應一條染色體拷貝,所以「兩族」就是那個位置上的兩條同源染色體。但兩件事不成立:其一,軟體判定不出哪一族來自父親、哪一族來自母親(那需要另外定序父母);其二,標號只在該 phase block 內有定義,跨 block 的「家族一」並非同一條染色體。One of the two groups reads are split into by HP tag. Family 1 is HP1 plus the somatic haplotype HP1-1 derived from it; family 2 is HP2 plus HP2-1; HP3, which cannot be assigned to either germline haplotype, belongs to neither. It is called a family because it groups a germline haplotype together with the somatic haplotypes descended from it — the split is by the germline layer, not by whether a read carries a somatic mutation. Within one phase block a family corresponds to one chromosome copy, so the two families are the two homologous chromosomes at that locus. Two things do not follow: which family is paternal cannot be determined without sequencing the parents, and the labels are defined only within that phase block.完整條目 →。
這個切法很重要:它表示一個連鎖區段裡的 read 全部來自同一條染色體拷貝,
另一條拷貝屬於同一個視窗的另一條連鎖區段。因此
不可以在連鎖區段內部再把兩條拷貝混起來寫成各佔一半 ——
那是尚未依家族分組時的寫法,用在已分組的資料上會把細胞比例整體算錯將近一倍。
基準值確實由平台與 basecaller 決定,可以事先校準一次。但它逐脈絡變動 ——
homopolymerhomopolymer 同聚物同一個鹼基連續重複的區段,例如 AAAAAA。nanopore 在此類區段較容易發生長度判讀錯誤,是 indel 假陽性的常見來源之一。A run of identical bases. Nanopore miscounts their length, a major source of false indel calls.完整條目 → 區、低複雜度序列與 mapping/參考偏誤都會把它抬高數倍,
所以進模型的是一張依脈絡分層的表,不是一個全域常數
局部超立方體用的是「一步一個突變」的簡約假設,
但全域的 把多個突變指派到同一個 clone —— 在全域這一層,
完全可以是一條邊,中間不需要任何節點。
所以那個補進來的節點是局部表示法的產物,不是「還沒找到的 subclone」。
這正是 latent nodelatent node 潛在節點建樹時為了讓圖連得起來而補進的中間狀態,沒有被任何 read 直接觀測到。它是模型的產物,不能當成「還沒觀察到的細胞」。An intermediate state added during tree construction to keep the graph connected, not directly observed in any read. It is a product of the model and must not be read as an unobserved cell population.完整條目 → 那條警告的具體形式:潛在節點數 仍然是個有用的統計量
(它量的是這份重建有多少比例來自推論),但它是記帳,不是證據。
延續中篇第三節那個例子(ρ、A、B、C 都不變),只加上第四群 D 與兩個連鎖區段。
上排是三份證據各自說了什麼 —— 注意頻率譜給的是一棵部分決定的樹,不是一棵完整的樹。
中排說明目標函數的兩項各自回答什麼、又各自不回答什麼。
下排是三者合起來時,相容的候選樹如何被砍到只剩一棵。
成因是每個連鎖區段只裝得下落在該連鎖視窗內的那幾個變異。
上圖的連鎖區段甲裝到的是 B 群與 C 群的變異,兩群互不包含,故其局部形狀是分岔;
連鎖區段乙裝到的是 B 群與 D 群的變異,後者巢狀於前者,故其局部形狀是一條線。
兩者皆為同一棵全域 clone tree 在不同變異子集上的投影 ——
形狀相異是因為子集相異,而非證據衝突。
合起來為什麼更準、解析度更高
關鍵在於認清頻率譜的輸出並不是一棵樹,而是一棵部分決定的樹。
本例中它換算出三群(),而這三群在 pigeonholepigeonhole 鴿籠原理由父代與子代的細胞比例限制樹形的算術規則:一個細胞至多屬於父節點底下的一個子節點,所以各子節點的細胞比例加起來不得超過父節點。名稱來自鴿籠原理 —— 東西放進籠子,總量不會憑空變多。在 subclone 重建的文獻中也稱為 sum rule 或 crossing rule。它只能排除樹,不能挑出樹:通過檢查的候選通常仍不只一棵,其餘要靠 parsimony 之類的偏好決定。The arithmetic constraint that limits tree shape from parent and child cell fractions: a cell belongs to at most one child of a given parent, so the children's cell fractions cannot sum to more than the parent's. The name comes from the pigeonhole principle. Also called the sum rule or crossing rule in the subclonal reconstruction literature. It can only rule trees out, never select one: the surviving candidates are usually more than one, and the choice among them falls to a preference such as parsimony.完整條目 → 之下
有兩棵樹同時通過:線性的 A → X → D,與分支的 A → 。
取兩棵的交集,才是頻率譜真正確定下來的東西:
for K = 1 ... Kmax:
列舉/搜尋這個 K 下的離散結構候選
for each 候選:
固定結構 → 在 simplex 上最大化 log-likelihood ← 內層,直接解
記下這個候選的最佳 log-likelihood
取這個 K 的最佳值
比較各 K 的 Score(K, Θ̂_K),取最大者為 K̂
但中間那個「列舉/搜尋」不是一個真的 for 迴圈。
可用的做法包括啟發式搜尋、分支定界、以取樣取代列舉,以及最重要的一項:
用局部證據先把候選砍掉。第八節那個候選集合正是為此而存在 ——
每個連鎖視窗算出來的局部候選整組帶進外層,
不相容的全域結構因而在列舉之前就被排除,
這比列舉完再逐一評分便宜得多。
for each unit u:
for each read r spanning ≥1 site in u:
emit (u, phase_set, sites(u), mask(r), pattern(r), 1)
aggregate → sparse rows
report: 每個遮罩的 read 數、連鎖區段的最大跨越深度、
以及「若某狀態真佔 5%,這個連鎖區段看得到它的機率」
可跨越範圍受限的成因,來自以下三層結構:
read 跨距限制的是能否計數共現,
phase blockphase block一段可建立連續相位關係的區域。read 長度不足、缺少 informative heterozygous 位點或證據不一致時,可能形成不同 phase blocks。A contiguous stretch over which phasing is consistent. It can break when reads are too short, informative heterozygous sites are absent or evidence conflicts.完整條目 → 限制的是能否確定兩個位點是否位於同一條染色體。
局部狀態表僅能自最上層產生,而該層亦為三者中範圍最小者。
其二:候選排序的分數要換成似然
最小成本候選常不只一個,現行流程以 read-AFread-AF read 層級的 ALT 比例在同一個分析區域、同一個單倍型家族的 read 之中,某個 somatic 位點帶 ALT 的比例。分母已限縮到同一條 haplotype 的同一段區域,所以不必經過純度與拷貝數換算;用途是在步數並列的候選樹之間排序。The fraction of reads carrying the ALT allele at a somatic site, computed within one analysis window and one haplotype family. Because the denominator is already restricted, no purity or copy-number correction is needed; it is used to rank candidate trees of equal cost.完整條目 → 的差值總和排序;
問題在於相加的項數由拓撲形狀決定:
本頁「現行實作的量測結果」一節的數值均出自:
廖子游,Subclonal reconstruction using somatic haplotagging and methylation profiles with
Nanopore sequencing,國立中正大學資訊工程學系碩士學位口試,2026 年 7 月 30 日。
以 read 上的多變異狀態計數(而非單點 VAF)直接建模 subclone,
最接近的既有做法是 PairClone 與 TreeClone:Zhou T, Sengupta S, Müller P, Ji Y,
PairClone: a Bayesian subclone caller based on mutation pairs,
J R Stat Soc Ser C 2019;68:705–725
(academic.oup.com);
以及 Zhou T et al., TreeClone: Reconstruction of tumor subclone phylogeny based on genotype pairs,
arXiv:1703.03853。
兩者處理的是成對的位點與部分觀測的 read;本頁規格與其相異之處在於
可變的 、逐 phase set 的方向邊際化,以及與全基因體單變異連鎖視窗的相乘接合。
此一相異性尚未經文獻檢索確認,成文前應先查證。
前一篇把兩類觀測寫成兩個相乘的因子,接上同一組全域參數
,目標仍然是一棵全基因體、細胞層級的 clone treeclone tree 克隆演化樹描述腫瘤內各群細胞祖先關係的樹:節點是一群帶有相同變異組合的細胞,邊代表在祖先之上又多拿到變異。要注意同一組群集常常有多棵樹同時相容。A tree describing ancestral relationships among cell populations in a tumour: nodes are groups of cells sharing a mutation set, edges represent additional mutations acquired on top of an ancestor. Multiple trees are often compatible with the same clusters.完整條目 →。
這個目標有一個現實問題:以邊際 VAF 與 等位特異拷貝數allele-specific copy number 等位特異拷貝數把一個區段的總拷貝數拆成兩個親源等位各自的份數,通常記為 major 與 minor。總拷貝數相同而兩個等位不同的狀態(例如 2+0 與 1+1)在深度上完全一致,只有 B-allele frequency 分得開 —— copy-neutral LOH 即為此類。The total copy number of a segment split into the counts contributed by each parental allele, usually reported as major and minor. States with equal total but different split (2+0 versus 1+1) are identical in depth and separable only by B-allele frequency; copy-neutral LOH is exactly such a case.完整條目 →
為觀測的既有方法已經成熟且被系統性評比過,
要在同一個目標上勝過它們,需要的證據量遠大於多變異連鎖視窗能提供的。
中篇的計算已經指出,可用的多變異連鎖區段在低突變負荷樣本只有數百個。
本篇因此換一個估計目標:不重建全域的細胞層級系統發生,
而是重建每一個連鎖區段內部的分子系譜 —— 哪幾種局部狀態真的存在、
各佔多少分子、以及它們之間的先後關係。這個目標的解析度是單倍型連鎖區段 —— 一個連鎖視窗(read 連鎖起來的一段,不跨 phase setphase block一段可建立連續相位關係的區域。read 長度不足、缺少 informative heterozygous 位點或證據不一致時,可能形成不同 phase blocks。A contiguous stretch over which phasing is consistent. It can break when reads are too short, informative heterozygous sites are absent or evidence conflicts.完整條目 →)的一條染色體拷貝,不是「全基因體共用的一個 CCF 位置」。
在同一個 phase block 內,一個家族對應一條染色體拷貝,所以「兩族」就是那個位置上的兩條同源染色體。但兩件事不成立:其一,軟體判定不出哪一族來自父親、哪一族來自母親(那需要另外定序父母);其二,標號只在該 phase block 內有定義,跨 block 的「家族一」並非同一條染色體。One of the two groups reads are split into by HP tag. Family 1 is HP1 plus the somatic haplotype HP1-1 derived from it; family 2 is HP2 plus HP2-1; HP3, which cannot be assigned to either germline haplotype, belongs to neither. It is called a family because it groups a germline haplotype together with the somatic haplotypes descended from it — the split is by the germline layer, not by whether a read carries a somatic mutation. Within one phase block a family corresponds to one chromosome copy, so the two families are the two homologous chromosomes at that locus. Two things do not follow: which family is paternal cannot be determined without sequencing the parents, and the labels are defined only within that phase block.完整條目 →,
其位點集合為 、位點數 。
可以含沒有被任何分子觀測到的狀態。
觀測到 與 而 沒看到時,
兩者若不共用一個帶第一個突變的祖先,第一個突變就得發生兩次。
這種狀態只有兩種身分 —— 現存但沒抽到(),
或已被後代取代的歷史狀態()。
前一篇還有第三種:純粹因為「一步一個突變」的表示法而被迫補進來的記帳節點;
本規格允許一條邊帶多個突變,那一種就隨表示法一起消失了。
latent nodelatent node 潛在節點建樹時為了讓圖連得起來而補進的中間狀態,沒有被任何 read 直接觀測到。它是模型的產物,不能當成「還沒觀察到的細胞」。An intermediate state added during tree construction to keep the graph connected, not directly observed in any read. It is a product of the model and must not be read as an unobserved cell population.完整條目 → 那條警告仍然適用,但適用範圍窄了一半。
各節點的 CCFcancer cell fraction 癌細胞比例帶有某個特定突變的腫瘤細胞佔全部腫瘤細胞的比例。用來區分 clonal()與 subclonal()突變。不等於 VAF。The fraction of tumour cells carrying a given mutation; distinguishes clonal from subclonal. Not the same as VAF.完整條目 →
純度與該處 拷貝數copy number 拷貝數某段基因體在細胞內的拷貝數。多數正常常染色體區段為 2,可再分為 major 與 minor allele copy number。How many copies of a genomic segment a cell carries; a typical diploid autosomal segment has two. It can be split into major- and minor-allele copy numbers.完整條目 →,由全域基準提供
四配子相容性與 perfect phylogeny 都預設突變只增不減,所以
一段發生 LOHLOH 異型合子性喪失原本 heterozygous 的區域變成只剩一種 allele。LOH 不等於缺失 —— 也可能是一條 haplotype 遺失後另一條被複製(copy-neutral LOH)。Loss of heterozygosity: a formerly het region retains only one allele. Not necessarily a deletion.完整條目 → 而失去某個突變的區域,在模型眼中會被描述成
「那個狀態從來沒有取得過」,而不是「取得後又失去」。
家族標號在每個 phase setphase block一段可建立連續相位關係的區域。read 長度不足、缺少 informative heterozygous 位點或證據不一致時,可能形成不同 phase blocks。A contiguous stretch over which phasing is consistent. It can break when reads are too short, informative heterozygous sites are absent or evidence conflicts.完整條目 → 內獨立決定,
故仍引入 並邊際化。但與前一篇不同的是,
這裡有一件可以放心的事:局部系譜的形狀完全不依賴 ——
連鎖區段內部的狀態集合與包含關係與家族標號無關,所以 不進核心推論,只進註解層。
全域基準是一個標準的 頻率譜mutation frequency spectrum 突變頻率譜把一份樣本裡所有 somatic 變異的 VAF 畫成直方圖後得到的分布。分布上的峰與肩對應不同大小的細胞群,最低頻端的尾巴斜率則被用來判斷有沒有天擇。The distribution obtained by histogramming the VAFs of all somatic variants in a sample. Peaks and shoulders correspond to cell populations of different sizes; the slope of the low-frequency tail is used to test for selection.完整條目 →分群,
用哪一個既有實作都可以,它不是本規格的貢獻。它必須輸出四樣東西:
純度 ;CCF 原子的位置 與權重 ;
各原子的後驗不確定度;以及分層估到的 overdispersion 。
先後關係的先驗:鴿籠pigeonhole 鴿籠原理由父代與子代的細胞比例限制樹形的算術規則:一個細胞至多屬於父節點底下的一個子節點,所以各子節點的細胞比例加起來不得超過父節點。名稱來自鴿籠原理 —— 東西放進籠子,總量不會憑空變多。在 subclone 重建的文獻中也稱為 sum rule 或 crossing rule。它只能排除樹,不能挑出樹:通過檢查的候選通常仍不只一棵,其餘要靠 parsimony 之類的偏好決定。The arithmetic constraint that limits tree shape from parent and child cell fractions: a cell belongs to at most one child of a given parent, so the children's cell fractions cannot sum to more than the parent's. The name comes from the pigeonhole principle. Also called the sum rule or crossing rule in the subclonal reconstruction literature. It can only rule trees out, never select one: the surviving candidates are usually more than one, and the choice among them falls to a preference such as parsimony.完整條目 →只能軟用
read 跨距限制的是能否計數共現,
phase blockphase block一段可建立連續相位關係的區域。read 長度不足、缺少 informative heterozygous 位點或證據不一致時,可能形成不同 phase blocks。A contiguous stretch over which phasing is consistent. It can break when reads are too short, informative heterozygous sites are absent or evidence conflicts.完整條目 → 限制的是能否確定兩個位點是否位於同一條染色體。
本規格的每一列都只能自最上層產生,而該層亦為三者中範圍最小者 ——
這就是「不可拼接成一棵大樹」的物理來源,不是保守的選擇。
體細胞標記haplotagging根據已 phase 好的 variants,把每一條 read 指派到 HP1 或 HP2,並把結果寫回 BAM 的 HP tag。Assigning each read to HP1 or HP2 using phased variants, and writing the result back as a BAM HP tag.完整條目 →輸出的詞彙恰好是
1、2、1-1、2-1、3 五個值。
1-1 表示這條分子屬於 germline 單倍型 1,且帶有可歸因於它的體細胞改變;
3 表示體細胞改變無法歸因於任何一條 germline 單倍型。
沒有更長的後綴,所以後綴不是一條獲得序列 ——
不可以把它讀成「先 1 再加一步」。不認得的標籤一律排除在所有分組之外。
既有實作的連鎖視窗定義比「固定寬度 20 kb 視窗」精確得多,也更該沿用:
在同一個 phase setphase block一段可建立連續相位關係的區域。read 長度不足、缺少 informative heterozygous 位點或證據不一致時,可能形成不同 phase blocks。A contiguous stretch over which phasing is consistent. It can break when reads are too short, informative heterozygous sites are absent or evidence conflicts.完整條目 → 內,凡有某條 read 同時覆蓋至少兩個 sSNV,該段即為一個連鎖視窗;
若另一條 read 疊到已連鎖的位點又碰到新的位點,兩段合併成更長的連鎖視窗。三個推論:
因此每個候選 的特徵,都是對把它拿掉之後建成的局部系譜
算出來的。這條留一法不是保險,是必要條件:
否則特徵會直接把「這個候選有幾條 ALT read」抄一遍,
在訓練時看起來很有效,在真正的低頻變異上完全沒有幫助。
錨點也不得取自 truth set —— 那是最典型的資料洩漏data leakage 資料洩漏訓練或模型選擇階段取得了評估資料的資訊,讓效能被高估。實驗室的做法是按染色體切分,確保同一個位點不會同時出現在訓練、驗證與最終測試中。Evaluation information entering training or model selection and inflating performance. The lab splits by chromosome so a locus cannot appear in training, validation and final test partitions at once.完整條目 →。
上排是 caller 本來的流程,三個箭頭是三種接法的接入位置 ——
位置決定成本與風險。越靠前上限越高,也越不可逆:
A 接在給分之後,只改排序,碰不到已經被丟掉的候選;
B 接進張量後重訓,上限最高但難以歸因;
C 接在產生候選之前,被它砍掉的候選不會出現在任何後續指標裡,
所以它傷害 recall 的方式是安靜的。
建議順序是 A 先做。它便宜、可歸因(每一組特徵可以單獨加減)、
而且如果 A 沒有效果,B 幾乎不可能有 —— 因為 A 的失敗代表這些特徵
在 caller 已有的資訊之外沒有增量。
錨點必須由 caller 自己產生,不得取自 truth set 或 高信心區域high-confidence regionbenchmark 建立者指定的高可信度區間,通常以 BED 檔表示;區間外不宜視為具有相同標註可靠度。Intervals that benchmark authors consider reliable, usually supplied as a BED file. Results outside those intervals should not be assumed to have the same label reliability.完整條目 →的標註。
驗收須跨純度、突變負荷、平台與 basecaller 分層報告,
以檢查分布偏移distribution shift 分布偏移測試資料的組成與訓練資料存在顯著差異(例如正負樣本比例不同),使得模型表現不如預期。Test data differing in composition from training data, so measured performance does not transfer.完整條目 →。
本頁「輸入與前處理契約」、「推論演算法與計算預算」的上限與成本數字,
以及「既有實作的資料規模」的全部計數,
均出自本實驗室既有實作的紀錄:廖子游,
Subclonal reconstruction using somatic haplotagging and methylation profiles with
Nanopore sequencing,國立中正大學資訊工程學系碩士論文,2026 年 7 月。
該實作與本規格的目標一致(局部、分子層級、不重建全域樹),
但推論方式不同:它以 Camin–Sokal 簡約的最小潛在頂點搜尋產生候選,
再以 read-AF 分數縮減候選集合;本規格把後者換成邊際似然,並補上錯誤地板與全域先驗。
凡本頁引用其數值之處,均為該實作實測到的量,不是本規格已驗證的結果。
以 read 上的多位點狀態計數直接建模 subclone,最接近的既有做法是
Zhou T, Sengupta S, Müller P, Ji Y, PairClone: a Bayesian subclone caller based on mutation pairs,
J R Stat Soc Ser C 2019;68:705–725
(academic.oup.com);
以及 Zhou T et al., TreeClone: Reconstruction of tumor subclone phylogeny based on genotype pairs,
arXiv:1703.03853。
M8 已建立 VAFVAF 變異等位基因頻率在某個位點上,支持 alt allele 的 read 佔全部 read 的比例。VAF 不等於帶有這個突變的細胞比例。Variant Allele Frequency: the fraction of reads supporting the alternate allele. Not the same as the fraction of cells carrying it.完整條目 → 的完整關係式,並示範同一個 VAF = 0.20 可由三種不同的細胞狀態產生。
該章將「估計腫瘤細胞比例」列為一個獨立的研究題目,本頁即為該題目的標準解法。
討論之前須先分開三個量,其分母各不相同:
符號
名稱
分母
cellular puritytumour purity 腫瘤純度樣本中腫瘤細胞所佔的比例。purity 越低,somatic 訊號被正常細胞稀釋得越嚴重,偵測越困難。The fraction of cells in a sample that are tumour cells. Low purity dilutes somatic signal.完整條目 →
全部細胞
腫瘤的 倍體ploidy 倍性細胞內染色體套數。多數正常常染色體為 2(diploid);腫瘤可呈現 3、4 或其他非整倍體狀態。The number of chromosome sets in a cell. Most normal autosomal regions are diploid (2); tumours may be triploid, tetraploid or otherwise non-diploid.完整條目 →
(每個腫瘤細胞的平均拷貝數)
tumour DNA fractiontumour DNA fraction 腫瘤 DNA 比例樣本 DNA 中源自腫瘤的比例。與 tumor purity(細胞比例)在 aneuploid 或 WGD 的情況下會不一樣。The fraction of DNA originating from tumour cells; diverges from cellular purity under aneuploidy or WGD.完整條目 →
全部分子
分母的兩項即為兩種細胞各自貢獻的 DNA 量:腫瘤細胞每個 份,正常細胞每個 2 份。
時 ,此即三個量在二倍體之下不易被察覺分家的原因。
此換算式沒有任何一篇文獻將其列為編號式。它是 ABSOLUTE 所定義的
(樣本平均拷貝數)與該文
「the smallest possible value is 2(1−α)/D … corresponds to the fraction of DNA from normal cells」
一語的直接推論:正常細胞所貢獻的 DNA 比例為 ,其補數即為 。
同一個 在 ASCAT 記為 ,在 PURPLE 為 normFactor 的倒數的兩倍。
以下以 ASCAT 與 PURPLE 為對照基準,兩者皆自 B-allele frequencyB-allele frequency B 等位頻率在 germline heterozygous 位點上,兩個等位其中一個所佔的 read 比例。腫瘤樣本中混入的正常細胞在此值上恆為 0.5,該常數即為反推純度所依據的錨。ASCAT 與 PURPLE 皆刻意不做相位,故其標號在各位點之間獨立。The fraction of reads carrying one of the two alleles at a germline heterozygous site. Contaminating normal cells contribute exactly 0.5, and that constant is the anchor from which purity is recovered. ASCAT and PURPLE deliberately do not phase, so the A/B label is independent per site.完整條目 → 與定序深度
同時估計純度、倍體與 等位特異拷貝數allele-specific copy number 等位特異拷貝數把一個區段的總拷貝數拆成兩個親源等位各自的份數,通常記為 major 與 minor。總拷貝數相同而兩個等位不同的狀態(例如 2+0 與 1+1)在深度上完全一致,只有 B-allele frequency 分得開 —— copy-neutral LOH 即為此類。The total copy number of a segment split into the counts contributed by each parental allele, usually reported as major and minor. States with equal total but different split (2+0 versus 1+1) are identical in depth and separable only by B-allele frequency; copy-neutral LOH is exactly such a case.完整條目 →。
所要建立的並非其操作流程,而是四件事:前向式的形式、反解式、
解為何不唯一,以及兩者各以什麼手段選出其中一個。
深度比logR 深度比對數同一位點上腫瘤與正常樣本的 read 深度比取以 2 為底的對數。ASCAT 再以全基因體比值的平均重新置中,因此 0 並不代表兩份拷貝 —— 倍體須由擬合結果另行推算。單獨的 logR 只能決定拷貝數輪廓至一個仿射變換。The base-2 logarithm of the tumour-to-normal read-depth ratio at a locus, re-centred on the genome-wide mean ratio. Zero therefore does not mean two copies; ploidy has to be recomputed from the fitted profile. logR alone determines the copy-number profile only up to an affine map.完整條目 →為同一位點上腫瘤與正常樣本的 read 深度比取以 2 為底的對數。
ASCAT 的實作(ascat.prepareHTS.R 的 ascat.getBAFsAndLogRs)
在取對數之後,再以全基因體比值的算術平均重新置中。
B-allele frequencyB-allele frequency B 等位頻率在 germline heterozygous 位點上,兩個等位其中一個所佔的 read 比例。腫瘤樣本中混入的正常細胞在此值上恆為 0.5,該常數即為反推純度所依據的錨。ASCAT 與 PURPLE 皆刻意不做相位,故其標號在各位點之間獨立。The fraction of reads carrying one of the two alleles at a germline heterozygous site. Contaminating normal cells contribute exactly 0.5, and that constant is the anchor from which purity is recovered. ASCAT and PURPLE deliberately do not phase, so the A/B label is independent per site.完整條目 →僅在 germline heterozygous 位點計算。
germline homozygous 位點在腫瘤與正常樣本中皆為 0 或 1,對純度零資訊,故予以剔除。
兩個等位的 A/B 標號逐位點隨機決定(selector = round(runif(len)))——
ASCAT 刻意不做相位,因此其等位軌道對稱於 0.5,PURPLE 的 AMBER 則取兩者較大值,值域為
。此設計選擇在下篇會再度出現。
沿染色體的三個區段。1+1 與 2+0 的深度完全相同 ——
兩者的總拷貝數皆為 2,故深度比軌道上看不出任何差異;只有等位比例軌道分得開。
此即 M8 所述 copy-neutral LOHLOH 異型合子性喪失原本 heterozygous 的區域變成只剩一種 allele。LOH 不等於缺失 —— 也可能是一條 haplotype 遺失後另一條被複製(copy-neutral LOH)。Loss of heterozygosity: a formerly het region retains only one allele. Not necessarily a deletion.完整條目 → 僅依 coverage 難以辨識的代數成因:
深度只承載總拷貝數,等位的拆分方式落在另一條軌道上。
全基因體加倍whole-genome doubling 全基因體加倍腫瘤演化過程中整套基因體複製一次的事件,使各處拷貝數同時加倍。其重要後果是可辨識性:若加倍後所有等位特異拷貝數皆為偶數,則該解與未加倍的解對深度與等位比例給出完全相同的觀測;純度為 1 時兩者逐點相同。An event doubling the entire tumour genome, multiplying every copy number by two. Its consequence is identifiability: if all allele-specific copy numbers are even afterwards, the doubled solution reproduces the observed depth and allele ratios exactly, and at purity one the two models are pointwise identical.完整條目 →
此即 M8 的 VAF 關係式反解為 。PURPLE 回報的是 本身,
而非 CCFcancer cell fraction 癌細胞比例帶有某個特定突變的腫瘤細胞佔全部腫瘤細胞的比例。用來區分 clonal()與 subclonal()突變。不等於 VAF。The fraction of tumour cells carrying a given mutation; distinguishes clonal from subclonal. Not the same as VAF.完整條目 → —— 兩者相差一個 multiplicity。
此處須區分兩個概念。精度指兩種情形已可區分之後,量測是否準確,
它隨資料量改善。可辨識性identifiability 可辨識性資料在原則上能否區分兩組不同的參數值。若兩組參數對任何可能的觀測給出相同的機率,則兩者不可辨識,增加深度、變異數目或更換工具皆無作用 —— 須改變的是可行域,例如加入外部約束或錨。此性質與精度不同:精度指可區分之後量得準不準。Whether the data can in principle distinguish two parameter values. If both assign the same probability to every possible observation they are unidentifiable, and more depth, more variants or a different tool cannot help; only changing the feasible set — an external constraint or anchor — can. This differs from precision, which concerns accuracy once two states are already distinguishable.完整條目 →指資料原則上能否區分兩種情形,
它不隨資料量改善。上述偏差屬於後者,另加一道硬性下限所造成的截斷。
ASCAT:Van Loo P et al., Allele-specific copy number analysis of tumors,
PNAS 2010;107:16910–16915(doi 10.1073/pnas.1009843107);
多樣本分割見 Ross EM et al., Allele-specific multi-sample copy number segmentation in ASCAT,
Bioinformatics 2021;37:1909–1911(doi 10.1093/bioinformatics/btaa538)。
程式碼在 github.com/VanLoo-lab/ascat。
本頁的前向式、反解式、距離函數與四道硬性條件的常數,均取自 ASCAT/R/ascat.runAscat.R
與 ASCAT/R/ascat.prepareHTS.R。
PURPLE:Priestley P et al., Pan-cancer whole-genome analyses of metastatic solid tumours,
Nature 2019;575:210–216(doi 10.1038/s41586-019-1689-y)。
程式碼與說明文件在 github.com/hartwigmedical/hmftools
的 purple 之下。本頁的擬合分數、事件懲罰、偏離懲罰與各項常數,取自
PurityAdjuster.java、PloidyDeviation.java、FittedPurityFactory.java
與 PurpleConstants.java。
的換算式源自 ABSOLUTE:Carter SL et al.,
Absolute quantification of somatic DNA alterations in human cancer,
Nat Biotechnol 2012;30:413–421(doi 10.1038/nbt.2203)。
Battenberg 見 Nik-Zainal S et al., The life history of 21 breast cancers,
Cell 2012;149:994–1007(doi 10.1016/j.cell.2012.04.023)。
M12 已完整教過 LongPhase-S 的四個階段。此處只需要第二階段的一句話摘要:
把全基因體的 GHIRGHIR 生殖系單倍型失衡比Germline Haplotype Imbalance Ratio:在候選 somatic 位點上,取標為 HP1 與 HP2 的 read 數中較大者除以兩者之和,值域為 0.5 至 1。須注意兩件事:分母只含這兩個 germline 計數,HP1-1/HP2-1/HP3 皆不在內;且這些標籤是整條 read 的判定(該 read 任一處帶 somatic 等位即離開 germline 計數),並非該位點的等位計數。它不是直接的 purity 讀數:拷貝數變異、LOH、read 跨距內的突變密度、標記錯誤與抽樣不足都會使它偏移。Germline Haplotype Imbalance Ratio: at a candidate somatic locus, the larger of the HP1 and HP2 tagged read counts divided by their sum, ranging from 0.5 to 1. Two caveats: the denominator contains only those two germline counts, excluding HP1-1/HP2-1/HP3; and those tags are whole-read decisions (a read carrying a somatic allele anywhere leaves the germline counts), not per-locus allele counts. It is not a direct purity readout — copy number, LOH, mutation density within the read span, tagging error and sparse sampling all shift it.完整條目 → 分布收成中位數與四分位距兩個數字,
再以一條二次多項式迴歸換成一個純量。
那個純量是腫瘤 DNA 比例tumour DNA fraction 腫瘤 DNA 比例樣本 DNA 中源自腫瘤的比例。與 tumor purity(細胞比例)在 aneuploid 或 WGD 的情況下會不一樣。The fraction of DNA originating from tumour cells; diverges from cellular purity under aneuploidy or WGD.完整條目 →,
不是 cellular puritytumour purity 腫瘤純度樣本中腫瘤細胞所佔的比例。purity 越低,somatic 訊號被正常細胞稀釋得越嚴重,偵測越困難。The fraction of cells in a sample that are tumour cells. Low purity dilutes somatic signal.完整條目 → 。理由不在實作,在模型形式:
上排為現行流程:GHIR 分布經中位數與四分位距兩個特徵,
以迴歸給出 tumour DNA fraction 一個純量;純度與倍體不在輸出之中。
下排為主題二所要建立者:三個觀測通道分別量在 germline 位點、視窗全體與 somatic 位點上,
聯合估計純度與每段的等位特異拷貝數,倍體與 DNA 比例由兩者導出。
本頁只處理三個通道各自量到什麼,合成一條估計式屬下篇。
右側對照上篇:ASCAT 與 PURPLE 輸出純度與倍體之後,須以換算式才進入同一個尺度。
現行實作已有的原料,以及還缺的一格
現行實作在每個 somatic 位點上,已將覆蓋該位點的 read 依標籤分入
H1、H2、H1-1、H2-1、H3 五個計數
(SomaticVarCaller.cpp 的 ReadHpCount),
並已算出三個比值:germlineHaplotypeImbalanceRatio(即 GHIRGHIR 生殖系單倍型失衡比Germline Haplotype Imbalance Ratio:在候選 somatic 位點上,取標為 HP1 與 HP2 的 read 數中較大者除以兩者之和,值域為 0.5 至 1。須注意兩件事:分母只含這兩個 germline 計數,HP1-1/HP2-1/HP3 皆不在內;且這些標籤是整條 read 的判定(該 read 任一處帶 somatic 等位即離開 germline 計數),並非該位點的等位計數。它不是直接的 purity 讀數:拷貝數變異、LOH、read 跨距內的突變密度、標記錯誤與抽樣不足都會使它偏移。Germline Haplotype Imbalance Ratio: at a candidate somatic locus, the larger of the HP1 and HP2 tagged read counts divided by their sum, ranging from 0.5 to 1. Two caveats: the denominator contains only those two germline counts, excluding HP1-1/HP2-1/HP3; and those tags are whole-read decisions (a read carrying a somatic allele anywhere leaves the germline counts), not per-locus allele counts. It is not a direct purity readout — copy number, LOH, mutation density within the read span, tagging error and sparse sampling all shift it.完整條目 →)、
allelicImbalanceRatio 與 somaticHaplotypeImbalanceRatio。
而單一個 GHIR 值無法分開兩者。
時,二倍體 clonal 的位點給出 ;
而一個該處無 somatic 變異、但等位特異拷貝數為 的位點,
其 GHIR 同樣為 。這正是 單倍型失衡haplotype imbalance 單倍型失衡指派至兩條親源單倍型的 read 數不相等。其成因有二:該處兩條單倍型的拷貝數不同,或其中一條的分子被改標至 somatic 子單倍型。兩者在數值上形式相同,僅憑一個失衡值無法區分。Unequal read counts assigned to the two parental haplotypes. Two mechanisms produce it: unequal allele-specific copy number, or molecules of one haplotype being relabelled to a somatic sub-haplotype. The two are indistinguishable from a single imbalance value.完整條目 →
一詞所涵蓋的兩種成因。
前身工具見 Lin J-H et al., LongPhase: an ultra-fast chromosome-scale phasing
algorithm for small and large variants, Bioinformatics 2022;38:1816–1822
(doi 10.1093/bioinformatics/btac058)。
分段不能沿用 phase blockphase block一段可建立連續相位關係的區域。read 長度不足、缺少 informative heterozygous 位點或證據不一致時,可能形成不同 phase blocks。A contiguous stretch over which phasing is consistent. It can break when reads are too short, informative heterozygous sites are absent or evidence conflicts.完整條目 →
第三節:兩件不同的工作
偵測與標記誤差尚未進入似然
第二節的第四層與第六節
本頁的位階:規格,不是已驗證的方法
以下所寫的是一個可實作、可反駁的規格 ——
每一層的參數、每一個通道的分布形式、每一期可單獨驗證什麼。
它尚未在任何真實或模擬資料上執行過,因此不宣稱優於現行的 DNA 比例迴歸;
後者在其所針對的量上已有 MAE 0.03 的實測成績。
中篇曾以 phase blockphase block一段可建立連續相位關係的區域。read 長度不足、缺少 informative heterozygous 位點或證據不一致時,可能形成不同 phase blocks。A contiguous stretch over which phasing is consistent. It can break when reads are too short, informative heterozygous sites are absent or evidence conflicts.完整條目 → 的 N50 與拷貝數區段的尺度相近為由,
主張可直接以 block 分段。該論證不成立:尺度相近不蘊含斷點對應。
TINCTINC 腫瘤污染正常樣本Tumour-in-normal contamination:配對正常樣本含有腫瘤來源訊號,使 genuine somatic 變異也可能在 normal 中被觀察到,因而可能被錯誤過濾。Tumour-in-normal contamination: tumour cells present in the matched normal, causing true somatic variants to produce signal there as well.完整條目 →
掃參數空間的模擬:純度、倍體、加倍時序、subclonal 拷貝數變異、突變負荷、phase switch 錯誤率、深度、讀長、TINCTINC 腫瘤污染正常樣本Tumour-in-normal contamination:配對正常樣本含有腫瘤來源訊號,使 genuine somatic 變異也可能在 normal 中被觀察到,因而可能被錯誤過濾。Tumour-in-normal contamination: tumour cells present in the matched normal, causing true somatic variants to produce signal there as well.完整條目 →、兩個平台