explain the C-Value paradox
There is not a strong correlation between organism complexity and genome size (C-value)
Large genomes may not necessarily gain DNA faster; they may simply lose unnecessary DNA much more slowly
inverse correlation:
large genome ~ 1/spontanious deletion rate
range of genome size prokaryotes vs eukaryotes
bacteria genomes: 580 Kb to 13 Mb,
-> 20–30-fold size variation within prokaryotes
A few eukaryotic genomes fall in the size range of bacteria (e.g. yeast), but most are much larger.
eukaryotic genomes: 8.8 Mb to ≈700 Gb
-> 80,000-fold size
relation genome size and gene number
in bacteria correlated
in eukaryotes not really, but a bit more than for genome size and complexity:
variation in eukaryotic gene number is estimated to be <50 fold
-> most of the C-value variation is due to the amount of non-coding DNA.
In general, there is a steep decline in the fraction of genic DNA (coding DNA) as genomes become larger.
what is heterochromatin
large regions of the genome with no (or very few) genes. It is difficult to clone and usually not sequenced in genome projects
what types of repetetive DNA are the?
Satellite DNA (2-2000bp, 1000’s x rep)
Minisatellites (9-100bp, 10-100 x rep)
Microsatellites (1-5bp, 10-100 x rep)
what is Satellite DNA how first identified?
These are repeated sequences that have either
high GC (heavy) or high AT (light) content
2–2000 repeated 1000’s of times in a row
in heterochromatic regions and around centromeres
FIRST IDENTIFIED with density centrifugation:
-> distinct bands of DNA that are heavier or lighter than the majority of genomic DNA
what are Minisatellites?
9–100 bp repeated 10–100 times.
In subtelomeric regions and (rarely) dispersed throughout chromosomes
what are Microsatellites, what are they also called?
= SRS “short repetitive sequences”
= STR “short tandem repeats”
= SSR “simple sequence repeats”
1-5 bp repeated 10–100 times
dispersed throughout chromosomes, often in and around genes
very high mutation rates (changing the repeat number, polymerase “slips“) -> useful in pop. genetics and DNA fingerprinting
eg: dinucleotide short tandem repeat CA is very common in the human genome (≈50,000 copies)
what is expansion of tri-nucleotide repeats (increase in repeat number) in or near genes is often associated with?
inherited diseases. Some examples include:
Fragile-X syndrome (CCG)
Huntington’s disease (CAG)
Schizophrenia? (CAG)
Myotonic Dystrophy (CTG)
… plus many other neuro-muscular disorders
What are Transposable Elements
= interspersed repetitive elements
= jumping genes
pieces of DNA that can move within the genome and increase in number
About 50% of the human genome is made up of TEs and remnants of TEs
which types of TEs are there?
transposons
2,500-7,000 bp long, DNA -> DNA
can be autonomous or non-outonomous (need helper element)
Conservative transposition – cut-and-paste”
does not necessarily lead to an increase in copy number/ Copy number can be increased through recombination between elements at different chromosomal locations.
-> equal number of gains and losses.
Replicative transposition – “copy-and-paste”
original element remains at donor site, while a new copy inserts into a new site
-> copy number is increased
retrotransposons
active or dead
Retrotransposition – “copy-and-paste”
the TE is transcribed into RNA, then reverse transcribed into cDNA, then inserts into new chromosomal location. The copy number increases.
typically the most abundant in a genome
explain autonomous vs non-autonomous Transposons and helper elements
autonomous
have terminal repeats at the ends, encode a single gene (transposase)
-> can move by themselves.
non-autonomous
have terminal repeats at ends, but no transposase gene. -> -> Cannot move by themselves,
-> but can move if there is another element in the genome producing transposase: Helper element
“Helper element”
does not have inverted repeats, but does have transposase gene. Cannot move, but can cause non-autonomous elements to move.
-> These are very useful for experiments in organisms like Drosophila.
explain how Helper genes are usefull for experiments
transposase gene of a TE can be replaced with any gene
a helper element can be used to make transposase and insert this gene into the Drosophila genome
then the helper element is removed so the new gene becomes a stable part of the genome
explain active vs dead retrotransposons
Active
have intact promoter, are transcribed, and can retrotranspose
“Dead” or “Dead On Arrival (DOA)”
retroelements are often truncated at the 5’ end when inserting into DNA.
-> they lose their promoter and no longer can be transcribed or retrotransposed.
= “junk DNA” that is under no selective constraint
-> accumulate mutations at random
what are Pseudogenes
Previously functional genes
Most cases: gene duplication
have lost their function due to mutation
eg:
mutation that introduces a stop codon into the ORF
InDel that disrupts the reading frame
In rare cases: LoF due to parasitic or symbiotic relationship with their host
-> genes are not needed and can be lost through mutations (ex. pathogenic bacteria, M. tuberculosis).
what types of pseudogenes are there
unprocessed
through tandem duplication: an entire section of DNA is duplicated during replication
-> two copies of a gene (adjacent in the genome)
If only one copy is required, the other copy may accumulate mutations and become a non-functional pseudogene
processed = or retrotransposed genes
through reverse transcription of mRNA of a nuclear gene into cDNA,
then re-inserts into the genome
Most likely this uses the reverse transcriptase and integrase enzymes encoded by a retroelement.
features of a processed pseudogene
Does not have introns present in the “parental” gene
If recent, may have a poly(A) sequence at 3’ end
lacks promoter sequences
(thus “Dead on arrival” = not expressed)
do all genes retrotransposase equally often? explain your answer
some more, depends on:
Expression level – highly-expressed genes have more mRNA and thus have a greater chance of being reverse transcribed.
Gene size – short mRNAs may retrotranspose better than long mRNAs.
Sequence specific – the primary sequence of some genes may be better for retrotransposition.
why is there so great variantion in genome size?
2 ideas
adaptive – the non-coding DNA is functionally important to the organism
junk DNA – most of the non-coding DNA serves no purpose. It may even be parasitic or “selfish DNA”.
what is the relation between genome size and rate odf sponaeous DNA deletion?
large genome ~ low deletion rate
order by genome size, deletion rate and deletion size:
Drosophila (fly)
Podisma (grasshopper )
Laupala (kcrickets)
genome size: Pod (20Gbp) > Lau (1.9Gbp) > Dro (168Mbp)
deletion rate: Dro > Lau (1.9Gbp) > Pod
deletion size: Dro > Lau (1.9Gbp) > Pod
what are NUMTs
Nuclear copies of mitochondrial genes
= pseudogenes in the nuclear DNA that are derived from mitochondrial genes
grasshopper has many of them (Podisma)
Last changed8 days ago