What is the C-value paradox
size of the genome = C-Value
no strong correlation between organism complexity and genome size
essentially constant within species, but varies widely among species
Which is the range of genome size in bacteria
580 kb to 13 Mb
-> 20-30 fold change variation within prokayotes
Which is the range of genome size in eukayotic
700 GB to 8.8 Gb
fold change of 80000
Is there a correlation between genome size and gene number?
In bacteria, yes.
In eukaryotes, no. Of course, there is some correlation in the genomes of the model organisms that have been sequenced (e.g., yeast < Drosophila < human), but the range of variation in eukaryotic gene number is estimated to be <50 fold. It may actually be much less.
What is the genome size of yeast and how many genes can be found
Yeast, 15 Mb total, 6,000 genes
What is the genome size of human and how many genes can be found
Human, 3,000 Mb (3 Gb) total, 25,000 genes
What is the C and G ration from human to yeast
C ratio = 3,000/15 = 200
G ratio = 25,000/6,000 = 4.2
-> most of the C-value variation is due to the amount of non-coding DNA
Heterochromatin
large regions of the genome with no (or very few) genes. It is difficult to clone and usually not sequenced in genome projects.
thight packed
How much of the human genome is protein coding? and what is the connection between coding DNA and genome size
2% coding
there is a steep decline in the fraction of genic DNA (coding DNA) as genomes become larger
Which 3 DNA elements you know that are repetitive
satellite
minisatellites
microsatellites
Satellite
first identified as distinct bands of DNA that are heavier or lighter than the majority of genomic DNA by density centrifugation.
repeated sequences that have either high GC (heavy) or high AT (light) content.
fairly short sequences (2–2000 bp) repeated 1000’s of times in a row
found in heterochromatic regions and around centromeres
Minisatellites
sequences of 9–100 bp repeated 10–100 times.
Found in subtelomeric regions and (rarely) dispersed throughout chromosomes.
Microsatellites & functions
very short sequences of 1-5 bp repeated 10–100 times
Found dispersed throughout chromosomes, often in and around genes.
e.g. CA is very common in human genome
in general high mutation rates
often variable within a population & useful for population genetocs
useful for “fingerprinting”
Which diseases you know are correlated with the expansion of tri-nucelotide repeats
Fragile-X syndrome (CCG)
Huntington’s disease (CAG)
Schizophrenia? (CAG)
Myotonic Dystrophy (CTG)
What are transposable elements and how often do they occur
known as interspersed repetitive elements or “jumping genes”
TEs are pieces of DNA that can move within the genome and increase in number.
About 50% of the human genome is made up of TEs and remnants of TEs.
Two sorts of transposable elements
transposons and retrotransposons
classifies by their mechanism of transposition
Conservative transposition
TE moves from one place in the genome to another.
does not necessarily lead to an increase in copy number.
Copy number can be increased through recombination between elements at different chromosomal locations.
However this should lead to an equal number of gains and losses.
involves only dna
Replicative transposition
copy number is increased because the original element remains at donor site, while a new copy inserts into a new site.
Retrotransposition
TE is transcribed into RNA, then reverse transcribed into
cDNA, then inserts into new chromosomal location
copy number increases
most abundant in a genome
requires RNA
How big are Transposons
2500 - 7000 bp long
autonomous vs non-autonomous
autonomous – have terminal repeats at the ends, encode a single gene (transposase), can move by themselves
non-autonomous – have terminal repeats at ends, but no transposase gene. Cannot move by themselves, but can move if there is another element in the genome producing transposase
What do we know about helper element from transposable elements
does not have inverted repeats, but does have transposase gene.
Cannot move, but can cause non-autonomous elements to move.
very useful for experiments in organisms like Drosophila.
For example, the transposase gene of a TE can be replaced with any gene, then a helper element can be used to make transposase and insert this gene into the Drosophila genome. Then the helper element is removed so the new gene becomes a stable part of the genome.
What is the structural and functional difference between an "active" and a "dead" (DOA) retrotransposon?
Active Retrotransposons: They possess an intact promoter, meaning they can be actively transcribed and continue to replicate/retrotranspose throughout the genome.
"Dead On Arrival" (DOA) Retrotransposons: They are often truncated (cut short) at the 5' end during the insertion process. Because of this, they lose their promoter and can no longer be transcribed or retrotransposed.
Evolutionary Status: DOA elements are considered "junk DNA" that is completely free from selective constraints, causing them to accumulate neutral mutations entirely at random.
What is a pseudogene, and what are the primary evolutionary causes for a gene to lose its function?
Definition: A pseudogene is a previously functional gene that has lost its biological function due to mutations.
Disruptive Mutations: This usually occurs via a mutation that introduces a premature stop codon into the Open Reading Frame (ORF), or through an insertion/deletion (indel) that disrupts the reading frame.
Host-Driven Decay: In rare cases, a parasitic or symbiotic relationship with a host makes certain genes redundant, allowing them to decay neutrally (e.g., pathogenic bacteria like M. tuberculosis or M. leprae).
The Most Common Origin: Most pseudogenes, however, do not come from simple host-driven decay but involve some type of gene duplication event followed by the inactivation of one copy.
How do unprocessed pseudogenes typically arise, and where are they located relative to the original gene?
Origin: They arise through tandem duplication, where an entire section of DNA is accidentally duplicated during replication, creating two copies of the same gene.
Location: The two copies are usually adjacent (right next to each other) in the genome.
Evolutionary Fate: Since only one copy is required to perform the biological task, the duplicate copy is free from selective constraints; it accumulates mutations over time and degrades into a non-functional pseudogene.
What is the molecular mechanism behind the creation of a processed pseudogene?
The Mechanism: An active nuclear gene is transcribed into mRNA, which is then reverse transcribed back into cDNA.
Integration: This newly formed cDNA patch is then re-inserted into a completely new, random location in the genome.
Enzymatic Hijacking: This process typically steals the reverse transcriptase and integrase enzymes encoded by a nearby mobile retroelement.
What are the 3 distinct structural features used to identify a processed pseudogene in the genome?
1. Lacks Introns: It contains no introns because it was copied directly from a mature, already-spliced mRNA template.
2. Poly(A) Tail: If the retrotransposition event occurred recently in evolutionary history, it will feature a distinct poly(A) sequence at its 3' end.
3. No Promoter Switch (DOA): It almost always lacks the necessary promoter sequence, rendering it "Dead on arrival" and unable to be expressed by the cell.
Why do some gene appear to retrotranspose more than others
Expression level – highly-expressed genes have more mRNA and thus have a greater chance of being reverse transcribed.
Gene size – short mRNAs may retrotranspose better than long mRNAs.
Sequence specific – the primary sequence of some genes may be better for retrotransposition.
What are the two classes of explanations for the variation in genome size
a) adaptive – the non-coding DNA is functionally important to the organism.
b) junk DNA – most of the non-coding DNA serves no purpose. It may even be parasitic or “selfish DNA”.
Can you order Laupala, Drosophila and Podisma based on deltion rate, deletion size and genome size
deletion rate: dros > lau > pod
deletion size: dros > lau > pod
genome size: pod > lau > dros
What evolutionary mechanism has been proposed to explain the massive variation in genome size (C-value) among different species?
Variations in the rate of spontaneous DNA deletion.
Organisms that delete unneeded DNA rapidly maintain small, streamlined genomes.
Organisms with slow deletion rates accumulate genomic "baggage" over time, leading to massive genome expansion.
How do the genomes and spontaneous deletion rates of Drosophila and Laupala (Hawaiian crickets) compare, and how were they measured?
Genomes: Laupala has a genome 11 times larger than Drosophila, while Drosophila has a very small genome with almost no pseudogenes.
Measurement: Deletion rates were estimated by tracking mutations in "Dead On Arrival" (DOA) transposable elements (which function like pseudogenes).
Result: Spontaneous DNA loss is much faster in Drosophila. Pseudogenes in Drosophila vanish so quickly due to deletion mutations that they rapidly become undetectable.
What did studying the grasshopper genus Podisma reveal about the inverse correlation between genome size and DNA deletion rates?
The Scale: Podisma has a massive genome (≈20 Gb), which is 10x larger than Laupala and 100x larger than Drosophila.
The Result: Podisma was found to have a very low rate of DNA loss (lower than both Laupala and Drosophila).
The Rule: This confirmed a strict inverse correlation across these three insect groups:
Higher Deletion Rate⟹Smaller Genome Size
What are NUMTs, and how are they utilized by evolutionary geneticists to calculate mutation and deletion rates?
Definition: NUMTs (Nuclear copies of miandtochondrial genes) are fragments of mitochondrial DNA that accidentally escaped into the cell nucleus.
Utility: Because they enter the nucleus as non-functional, "Dead-On-Arrival" sequences, they are completely free from selective constraints. They act as perfect neutral markers to observe raw, undisturbed mutation and deletion rates over evolutionary time.
What are the 3 distinct biological reasons why a NUMT cannot function once it is integrated into the nuclear genome?
1. Different Genetic Code: The translational code used inside the mitochondria does not match the universal genetic code of the nucleus.
2. No Promoter Switch: They typically migrate without a proper promoter sequence to initiate nuclear transcription.
3. Missing Target Signal: They lack the specific cellular signal sequence required to route any translated protein products back into the mitochondria.
what is NUMTs
NUMTs = Nuclear copies of mitochondrial genes = “new mites”
What is Gene Ontology?
A consortium that develops a common vocabulary to
describe the function of genes/proteins across taxa
What are the three ontologies?
Molecular Function, Cellular Component, Biological Process
white gene of Drosophila melanogaster
(also indexed by the FlyBase number FB:FBgn0003996).
How many GO terms are associated with this gene?
What biological processes is it involved in?
What is its molecular function?
What cell components is it found in?
50
eye pigment precursor transport, etc.
transmembrane transporter, etc.
plasma membrane, etc.
BRCA1 gene
331
double-strand break repair, etc.
nucleus, etc.
DNA binding, DNA repair, etc.
Zuletzt geändertvor 11 Tagen