Choose an analysis mode
Watterson’s theta
aₙ = Σᵢ₌₁ⁿ⁻¹ (1 / i) θW per locus = S / aₙ θW per site = S / (aₙ × L) Variable coverage: θW per site = S / Σ(callable sites in group × aₙ for that group) Effective population size: Ne = θW per site / (c × μ)
The symbol S is the segregating-site count. The value n is sampled chromosomes or sequences. The value L is the callable site count.
Enter comparable sequence information
- Count segregating sites after applying consistent quality filters.
- Enter chromosomes directly or select an individual-based sample type.
- Provide the callable sequence length using the correct unit.
- Use callable groups when sample size changes across sequence regions.
- Add mutation rate only when estimating effective population size.
- Review assumptions, cautions, and calculation steps before comparison.
Worked single-locus example
| Input | Example value | Meaning |
|---|---|---|
| Segregating sites | 18 | Eighteen variable sites passed filtering. |
| Sample chromosomes | 20 | Twenty haplotypes contribute to the estimate. |
| Callable length | 10,000 bp | Ten thousand sites were reliably assessed. |
| Mutation rate | 1 × 10−8 | An optional per-site, per-generation rate. |
Interpret estimates carefully
Watterson’s estimator is commonly interpreted under neutral infinite-sites assumptions. Selection, structure, recombination, and demographic history can change interpretation. Sequencing errors may inflate observed segregating-site counts.
Callable regions should use consistent filtering across compared datasets. Missing genotypes can change the harmonic correction across sites. Grouped sample sizes offer a practical correction here.
The analytical interval uses a plug-in variance approximation. It is omitted for variable callable sample sizes. Batch bootstrap intervals describe variation across supplied loci.
Theta calculator FAQs
1. What is Watterson’s theta?
Watterson’s theta estimates scaled genetic diversity from segregating sites. It adjusts the observed count using sampled chromosomes. Per-site normalisation supports comparisons across sequence lengths.
2. What is a segregating site?
A segregating site contains more than one allele. The variation must occur among sampled sequences. Quality filtering should happen before counting sites.
3. Should sample size mean individuals or chromosomes?
The formula uses sampled chromosomes or independent sequences. Diploid individuals usually contribute two autosomal chromosomes. Select the correct sample-count type before calculating.
4. How are diploid individuals converted?
The calculator multiplies diploid individuals by two. Twenty diploid individuals therefore represent forty chromosomes. Missing genotypes may require callable sample-size groups.
5. What is the harmonic correction?
The correction sums reciprocal integers through n minus one. It accounts for the effect of sample size. Larger samples increase the expected segregating-site count.
6. What is theta per site?
The per-site estimate divides locus theta by callable length. It supports comparison across differently sized regions. Comparable filtering remains essential for valid interpretation.
7. Can theta estimate effective population size?
Yes, when a suitable mutation rate is available. The scaling factor depends on the inheritance model. The result remains sensitive to biological assumptions.
8. How should missing sequence data be handled?
Use only sites that pass callable-data requirements. Group sites sharing the same chromosome count. The calculator then sums their denominator contributions.
9. Why can sequencing errors increase theta?
Errors may appear as false rare variants. These false variants increase segregating-site counts. Apply quality, depth, and genotype filters consistently.
10. Can estimates from different loci be compared?
They can be compared after per-site normalisation. Regions should use similar filtering and callability rules. Mutation rates and selective pressures may still differ.
11. What assumptions does the estimator make?
The classic interpretation assumes neutral, randomly sampled variation. It also uses an infinite-sites mutation framework. Real populations may depart from these assumptions.