GC% variation among species
varies among species: bacteria, plants, invertebrates
little variation among vertebrates
high veriation/ GC heterogenity within vertabrates:
vertebrate genome can be divided into isochores = long stretches (100s of kb) of DNA with uniform %GC.
explain isochores
long stretches of DNA with uniform %GC in vertabrates
100s of kb, sometimes >10 Mb
High %GC = heavy isochores (H)
H1 (42-47%)
H2(47-52%)
H3(>52%)
Low %GC = light isochores (L)
L1 (<37% GC)
L2 (37-42%)
which vertabrates have heavy isochores
Heavy isochores are found only in warm-blooded vertebrates (mammals, birds)
not in cold-blooded vertebrates (fish).
explain Selectionist hypothesis and Mutationist hypothesis for isochores
Selectionist hypothesis
GC-pairing is stronger than AT-pairing (3 vs. 2 hydrogen bonds)
-> may stabilize DNA at higher temperatures.
Support: heavy isochores found in warm-blooded vertebrates
Furthermore, heavy isochores are gene-rich.
Mutationist hypothesis
The pool of available nucleotides changes over replication
(which takes 8 or more hours for mammals).
-> There is more GC available early in replication, so mutations will be biased towards G or C
support: Over time, regions of the genome that replicate early
non-random bio-chemical errors during replication and repair
how much of human genome is organised in isochores?
Analysis 2005: By doing a sliding window analysis of the human genome DNA sequence and defining Isochores as segments >300 kb with distinct %GC and low heterogeneity:
isochores covered 41% of the human genome
most had low %GC
four-family model with mean GC contents of
35%, 38%, 41%, and 48%
GC% of functional parts of genome
genomic regions are GC rich
Coding regions > introns > 5’ flanking regions > 3’ flanking regions
explain codon bias
all of the synonymous codons for a particular amino acid are not used with equal frequency as would be expected at random
The preferred codons correspond to the most abundant tRNA in each species, suggesting that selection favors the use of codons that increase the level of gene expression
example codon bias Leu
Leucine: very high codon bias:
Leu can be encoded by 6 different codons:
CTG, CTA, CTC, CTT, TTG, TTA.
at radom: 1/6 (17%) of the time/ per codon
In highly expressed E. coli genes, CTG is used ≈90% of the time.
In yeast, TTG is used ≈90% of the time.
how to meassure codon bias
ENC = effective number of codons
the average number of codons that are used to encode the
20 amino acids
lies btw 20 and 61
Low ENC = high codon bias.
Fop = frequency of optimal codons
the frequency with which the “optimal” codon is used for
each amino acid
Optimal codons def: the one used with the highest frequency in highly expressed genes.
High Fop = high codon bias
do i need prior knowledge to use FOP or ENC?
Fop is species-specific
-> requires that optimal codons are known. For this, one must have many gene sequences and expression information.
ENC
can be applied to any species without prior knowledge of expression or codon usage.
where is codon bias highest?
Some observed patterns of codon bias:
in highly-expressed genes
in short genes compared to long genes
in female-expressed compared to male-expressed genes
laveal expressed genes
what might explain codon bias?
Selection
Mutation (neutral)
how might selection explain codon bias? what is eidence for this theory?
use of optimal codons (those that correspond to the most abundant tRNA) make translation faster and more accurate.
Codon bias could be used as a way to regulate gene expression post-transcriptionally
evidence:
Highly expressed genes have higher codon bias
Conserved protein motifs, such as DNA binding domains, have higher bias than other protein regions.
(This suggests selection for accuracy of translation)
Experimental replacement of optimal codons with non-optimal codons reduces the level of protein.
Replacing sub-optimal leucine codons in the Adh gene with optimal codons increases ADH enzymatic activity in larvae, but decreases it in adults.
-> This suggests that there may be a trade-off between optimal codon usage (or other factors) in different developmental stages. Codon bias is highest for larval-expressed genes.
give an example for:
Drosophila Alcohol dehydrogenase (ADH)
optimal leucine codons -> non-optimal codons
=> leads to a lower level of ADH protein in vivo and reduces ethanol tolerance in adult flies
Wa-F = wild-type, optimal leucine codons
1 leu = 1 luecine codon changed from optimal to non-optimal
6 leu = 6 leucine codons changed from optimal to non-optimal
10 leu = 10 luecine codons changed from optimal to non-optimal
In a comparison of ADH protein concentration, it was found that:
Wa-F > 1 leu > 6 leu > 10 leu
wild-type flies: more tolerant to ethanol that the mutant flies.
The LD50 = the ethanol concentration at which 50% of the flies were killed within 24 hours
The LD50 wild-type flies: 9%
for 10 leu flies: It was only 7.5%
how might neutral mutation explain codon bias?
Mutation bias specific bases?
eg:
if mutations from A or T to G or C are more frequent and no selection on synonymous sites:
-> should become GC-rich.
The reverse bias in mutation would lead to synonymous sites being AT rich
= this would lead to non-random codon usage
does mutation theory have evidence?
Most optimal codons end in G or C.
->hypothesis requires mutational bias towards G or C.
But most observations: bias is towards A or T,
But biased mismatch repair in favor of G or C.
= strongest in regions of high recombination (where most of the highly codon-biased genes are located).
could explain at least part of the observed codon bias
The GC content at third codon positions is correlated with local genomic GC content, suggesting an effect of mutation of codon usage.
how/ where/ can we distinguisch selection and mutation explaination for codon bias
The strength of selection acting on a particular codon is expected to be very weak (N*s=1 where N pop size, s selection coefficient )
-> often very difficult to distinguish between selective and neutral explanations for codon bias.
Typically, there is evidence for selection affecting codon usage in organisms with small genomes and large population sizes (bacteria, yeast, Drosophila).
evidence is much weaker in humans and other vertebrates, where it appears that mutational/repair biases can explain the observed codon usage.
Zuletzt geändertvor 8 Tagen