DNA forengineers.
DNA is a polymer that stores information in a 4-symbol alphabet. Cells read it, copy it, and execute it. Engineers now write it too.
A T C G 2 bits per base. ~3.1 Gbp per haploid human genome ≈ 770 MB.
Chemistry, and why 5′→3′ is not a convention
A nucleotide is a sugar with a phosphate on one side and a base on the other. The sugar's carbons are numbered 1′ to 5′ ("prime"). Two of them matter.
This is the "input" end of a nucleotide.
The next nucleotide's phosphate bonds here. That link is a phosphodiester bond.
One end has a free 5′ phosphate, the other a free 3′ –OH. Polymerases can only add to the 3′ end. Everything downstream (replication, transcription, sequencing chemistry) follows from this.
DNA's sugar lacks the 2′-OH that RNA has. That makes DNA far more stable: good for storage. RNA's 2′-OH makes it reactive and short-lived: good for a working copy.
Bases and pairing
Two ring-classes. A big base always pairs with a small one, so every rung is the same width (~1.08 nm between C1′ atoms). Click a base.
Helix geometry
Two antiparallel strands, bases inside, sugar-phosphate backbone outside. The twist creates two grooves of unequal width.
One strand runs 5′→3′ top to bottom, the other 3′→5′. Required for the bases to face each other.
The strands are not 180° apart around the axis. That leaves a wide groove and a narrow one. Proteins (transcription factors, Cas9) read the sequence from the major groove without unzipping the helix: each base pair exposes a distinct pattern of H-bond donors and acceptors there.
The form in cells at normal hydration. A-form (RNA duplexes, dehydrated DNA) is wider and squatter; Z-form is left-handed and rare.
Every phosphate carries a −1 charge. That's why DNA moves toward the + electrode in gel electrophoresis, and why Mg²⁺ is in every enzyme buffer.
helix diameter
rise per base pair
per full turn (3.4 nm pitch)
major / minor groove width
Packing 2 metres into a 6 µm nucleus
A hierarchy of coiling. Each level multiplies compaction. The final chromosome is ~10,000× shorter than the naked DNA.
The genome as data
The human reference genome, in numbers an engineer can use.
haploid genome. 2 bits/base ≈ 770 MB raw. Compressible far below that; most of it repeats.
chromosomes: 22 autosome pairs + XX or XY. Each cell holds ~6.2 Gbp total.
protein-coding genes, only ~1.5% of the sequence. Roughly the count in a worm.
transposable elements: ancient copy-paste sequences. Mostly inert.
difference between two people: ~3–4 million single-nucleotide variants (SNVs).
to 250 Mbp per chromosome. One continuous double helix each.
Anatomy of a gene
A gene is not just the coding letters. It's a region with control elements, and (in eukaryotes) the coding parts are split by introns that get cut out.
Upstream sequence where RNA polymerase and transcription factors bind. In eukaryotes often a TATA box ~25 bp before the start. Engineers swap promoters to control how strongly a gene is expressed.
Exons are kept; introns are spliced out. Alternative splicing lets one gene yield many proteins: ~20k genes, >100k proteins.
Untranslated regions flank the coding sequence in the mRNA. They regulate stability, localisation, and translation rate.
Control elements that can sit thousands of bases away and loop over in 3D. Gene regulation is a graph, not a line.
Transcription and translation
DNA → RNA → protein. The one diagram that clears up most confusion: which strand gets read, and which one the mRNA looks like.
RNA polymerase unwinds ~15 bp, reads the template strand 3′→5′, and builds mRNA 5′→3′. The mRNA therefore matches the coding strand, with U for T. Then: 5′ cap, splicing, poly-A tail. Export to cytoplasm.
The ribosome scans for the first AUG, then reads 3 bases at a time. Each tRNA carries an anticodon and one amino acid. Stops at UAA, UAG, or UGA. Chain folds into a protein.
RNA pol: ~40 nt/s (human), ~50–90 nt/s (E. coli). Ribosome: ~5 aa/s human, ~20 aa/s bacteria. Bacteria translate while still transcribing; eukaryotes cannot (nuclear membrane).
Retroviruses (HIV) run RNA → DNA with reverse transcriptase. Labs use the same enzyme to turn mRNA into cDNA for sequencing (RNA-seq).
The codon table
64 codons map onto 20 amino acids plus stop. The code is degenerate: most amino acids have several codons, usually differing at the 3rd position ("wobble"). Click any cell.
Rows: 1st base. Columns: 2nd base. Inside each cell, top to bottom: 3rd base U, C, A, G. Standard code (NCBI table 1). Mitochondria and some microbes use slight variants. Shown in RNA (U); in DNA read T for U.
Replication
Semi-conservative: each daughter helix keeps one old strand. The machinery is constrained by one rule: polymerase only extends a 3′ end.
Unzips the parent duplex at the fork. Single-strand binding proteins keep it open; topoisomerase relieves the overwinding ahead.
Polymerase cannot start from nothing. Primase lays a ~10-nt RNA primer with a 3′-OH to extend from.
Adds nucleotides 5′→3′. Also proofreads: a 3′→5′ exonuclease removes a wrong base before moving on. ~1 error per 10⁷; after mismatch repair, ~1 per 10⁹.
Runs toward the fork, so it is synthesised continuously.
Runs away from the fork, so it is made backwards in Okazaki fragments (~150 nt eukaryote, ~1500 nt bacteria). Primers are removed, gaps filled, and ligase seals the nicks.
The last primer on the lagging strand leaves an unfilled gap. Chromosomes shorten each division unless telomerase extends the repeat cap (TTAGGG)ₙ.
Mutations
A change in sequence. What it does depends on where it lands and whether it shifts the reading frame. Pick a base, then edit it.
Silent: same amino acid (thanks to wobble). Missense: different amino acid (sickle cell: one A→T in β-globin). Nonsense: new stop codon, truncated protein.
Insertion or deletion. If the length is not a multiple of 3, every downstream codon shifts: a frameshift, usually fatal to the protein. Multiples of 3 add or drop whole amino acids.
Replication errors, UV (thymine dimers), chemicals, radiation, transposons. Most are repaired. Germline mutations are inherited; somatic ones are not.
Copy-number variants, inversions, translocations, whole-gene duplication (the raw material for new genes).
Reading DNA: PCR and sequencing
You cannot see a single molecule. So first you copy it a billion times, then you read it.
PCR: exponential copying
Three temperatures, cycled ~30 times. Two short primers (18–25 nt) define which region gets copied. A heat-stable polymerase (Taq) survives the cycling.
Sequencing: three generations
| Method | Read length | Error | Use |
|---|---|---|---|
| Sanger (1977) | ~800 bp | ~0.1% | Verifying a plasmid or a single gene |
| Illumina (short-read) | 2 × 150 bp | ~0.1–1% | Whole genomes, RNA-seq. Massive parallel; cheap per base |
| Nanopore / PacBio (long-read) | 10 kb – 1 Mb | ~1–5% raw, corrected by depth | Assembly, structural variants, direct methylation |
All produce short strings with per-base confidence. Coverage (how many reads overlap each base, e.g. 30×) is what makes the consensus reliable.
The file you'll actually get: FASTQ
@SRR001666.1 071112_SLXA-EAS1_s_7:5:1:817:345 length=36 GGGTGATGGCCGCTGCCGATGGCGTCAAATCCCACC + IIIIIIIIIIIIIIIIIIIIIIIIIIIIII9IG9IC
Header, sequence, +, quality string of equal length.
Q = −10·log₁₀(perror), stored as ASCII (char − 33). I = 40 → 1 error in 10,000. 9 = 24 → 1 in 250.
FASTA (sequence only, >header), SAM/BAM (reads aligned to a reference), VCF (variants vs. reference), GenBank/GFF (annotated features).
Writing and editing DNA
Three techniques cover most of bioengineering: cut with restriction enzymes, carry on plasmids, and edit in place with CRISPR.
Restriction enzymes
Bacterial scissors that cut at a specific 4–8 bp site, usually a palindrome (reads the same 5′→3′ on both strands). Staggered cuts leave sticky ends that base-pair with any matching end, so fragments from different sources can be joined by ligase.
Plasmids
Small circular DNA (3–10 kb) that bacteria copy on their own. The standard vehicle: ori for replication, a promoter driving your gene, a selection marker (antibiotic resistance) so only cells that took it up survive. Modern assembly (Gibson, Golden Gate) joins many parts in one reaction.
CRISPR-Cas9
A programmable nuclease. A 20-nt guide RNA base-pairs with the target; Cas9 cuts both strands 3 bp upstream of a PAM (NGG). The cell repairs the break: sloppily by NHEJ (knockouts via indels) or precisely by HDR if you supply a template (knock-ins). Base and prime editors avoid the double-strand break entirely.
Sequence toolkit
Paste raw bases or a FASTA record. Everything below is computed in the page; view source to see the algorithms.
bases
GC content
Tm estimate (°C)
bits · MW (kDa, dsDNA)
Tm: Wallace rule (2·AT + 4·GC) under 14 nt, else 64.9 + 41·(GC − 16.4)/N. Good enough for primer sanity checks; use nearest-neighbour models for real work.
CS ↔ Bio map
The concepts you already have, and what they're called on the other side.