proteinsamino acidspolypeptidesprotein foldingpeptide bonds

Proteins: The Molecular Machinery of Life

Proteins: The Molecular Machinery of Life Proteins are large biomolecules and macromolecules essential to every living organism. Composed of one or more long chains of amino acid residues...

Proteins: The Molecular Machinery of Life

Proteins are large biomolecules and macromolecules essential to every living organism. Composed of one or more long chains of amino acid residues, these versatile molecules act as the primary functional units of the cell. From catalyzing metabolic reactions and replicating DNA to providing structural integrity and transporting vital molecules, proteins perform a vast array of functions that sustain life.

The diversity of protein function is driven by the specific sequence of amino acids, which is dictated by the nucleotide sequence of an organism's genes. This sequence determines how a protein folds into a specific three-dimensional (3D) structure, and this structure, in turn, defines the protein's biological activity.

Key Facts

The enzyme hexokinase is shown as a conventional ball-and-stick molecular model. To scale in the top right-hand corner are two of its substrates, ATP and glucose.
The enzyme hexokinase is shown as a conventional ball-and-stick molecular model. To scale in the top right-hand corner are two of its substrates, ATP and glucose.
  • Composition: Proteins consist of amino acid residues linked by peptide bonds.
  • Genetic Control: The amino acid sequence is encoded by the nucleotide sequence of genes.
  • Structure: A protein's function is determined by its unique 3D folding pattern.
  • Lifespan: Protein half-lives vary from minutes to years, averaging 1–2 days in mammalian cells.
  • Diversity: While 20 standard amino acids are common, some organisms use selenocysteine or pyrrolysine.

The Building Blocks: Polypeptides and Peptides

Ribbon diagram of a mouse antibody against cholera that binds a carbohydrate antigen
Ribbon diagram of a mouse antibody against cholera that binds a carbohydrate antigen

A linear chain of amino acid residues is known as a polypeptide. While a protein must contain at least one long polypeptide, shorter chains containing fewer than 20–30 residues are typically referred to as peptides rather than proteins.

Polypeptide
Polypeptide

These residues are joined by peptide bonds, which are chemical links formed between adjacent amino acids. The specific order of these residues is defined by the genetic code. In most organisms, this code specifies 20 standard amino acids, though certain archaea and other organisms may incorporate selenocysteine or pyrrolysine.

Chemical structure of the peptide bond (bottom) and the three-dimensional structure of a peptide bond between an alanine and an adjacent amino acid (top/inset). The bond itself is made of the CHON elements.
Chemical structure of the peptide bond (bottom) and the three-dimensional structure of a peptide bond between an alanine and an adjacent amino acid (top/inset). The bond itself is made of the CHON elements.

The stability and nature of these bonds are further explained by their resonance structures, which contribute to the overall architecture of the protein polymer.

Resonance structures of the peptide bond that links individual amino acids to form a protein polymer
Resonance structures of the peptide bond that links individual amino acids to form a protein polymer

Protein Synthesis and Modification

Protein structure
Protein structure

Protein production begins with the DNA sequence of a gene, which encodes the necessary amino acid sequence. This information is transcribed and then translated by a ribosome, the cellular machine that assembles the polypeptide chain using mRNA as a template.

The DNA sequence of a gene encodes the amino acid sequence of a protein
The DNA sequence of a gene encodes the amino acid sequence of a protein

A ribosome produces a protein using mRNA as template.
A ribosome produces a protein using mRNA as template.

Once the chain is synthesized, proteins often undergo post-translational modification. These chemical changes occur during or shortly after synthesis and can significantly alter the protein's stability, folding, activity, and overall function. Some proteins also require cofactors—non-protein chemical compounds or ions—to become biologically active. In many cases, individual proteins associate to form stable protein complexes to achieve complex biological tasks.

Peptide synthesis
Peptide synthesis

Structural Hierarchy and Folding

Proteins in various cellular compartments and structures tagged with green fluorescent protein (here, white)
Proteins in various cellular compartments and structures tagged with green fluorescent protein (here, white)

The transition from a linear polypeptide to a functional protein requires precise folding. This process is sometimes assisted by chaperonins, which are large protein complexes that ensure other proteins fold correctly.

The crystal structure of the chaperonin, a huge protein complex. A single protein subunit is highlighted. Chaperonins assist protein folding.
The crystal structure of the chaperonin, a huge protein complex. A single protein subunit is highlighted. Chaperonins assist protein folding.

Proteins are often analyzed by their structural components. Protein domains are functional units that fold into defined 3D structures, whereas motifs are shorter sequences that may serve as binding sites but lack a stable 3D structure on their own.

Protein domains vs. motifs. Protein domains (such as the EVH1 domain) are functional units within proteins that fold into defined 3D structures. Motifs are usually short sequences with specific functions but without a stable 3D structure. Many motifs are binding sites for other proteins (such as the red and green bars shown here in the context of a VASP protein).[56]
Protein domains vs. motifs. Protein domains (such as the EVH1 domain) are functional units within proteins that fold into defined 3D structures. Motifs are usually short sequences with specific functions but without a stable 3D structure. Many motifs are binding sites for other proteins (such as the red and green bars shown here in the context of a VASP protein).[56]

The complexity of these structures can be visualized in various ways, from all-atom representations to solvent-accessible surfaces that highlight the distribution of acidic, basic, polar, and nonpolar residues.

Three possible representations of the three-dimensional structure of the protein triose phosphate isomerase. Left: All-atom representation colored by atom type. Middle: Simplified representation illustrating the backbone conformation, colored by secondary structure. Right: Solvent-accessible surface representation colored by residue type (acidic residues red, basic residues blue, polar residues green, nonpolar residues white).
Three possible representations of the three-dimensional structure of the protein triose phosphate isomerase. Left: All-atom representation colored by atom type. Middle: Simplified representation illustrating the backbone conformation, colored by secondary structure. Right: Solvent-accessible surface representation colored by residue type (acidic residues red, basic residues blue, polar residues green, nonpolar residues white).

Historically, myoglobin was the first protein to have its structure solved using X-ray crystallography, revealing the importance of α-helices and the heme group for oxygen binding.

A representation of the 3D structure of the protein myoglobin showing turquoise α-helices. This protein was the first to have its structure solved by X-ray crystallography. Toward the right-center among the coils, a prosthetic group called a heme group (shown in gray) with a bound oxygen molecule (red).
A representation of the 3D structure of the protein myoglobin showing turquoise α-helices. This protein was the first to have its structure solved by X-ray crystallography. Toward the right-center among the coils, a prosthetic group called a heme group (shown in gray) with a bound oxygen molecule (red).

John Kendrew with model of myoglobin in progress
John Kendrew with model of myoglobin in progress

Protein Turnover and Degradation

Constituent amino-acids can be analyzed to predict secondary, tertiary and quaternary protein structure, in this case hemoglobin containing heme units.
Constituent amino-acids can be analyzed to predict secondary, tertiary and quaternary protein structure, in this case hemoglobin containing heme units.

Proteins do not last indefinitely. Through a process called protein turnover, the cell degrades and recycles proteins. While some proteins last for years, the average lifespan in mammalian cells is 1–2 days. Proteins that are damaged, unstable, or misfolded are targeted for rapid destruction, often by the proteasome—a large protein assembly that breaks down proteins, sometimes after they have been marked for destruction by ubiquitin ligases.

Hydrolysis of protein. X = HCl and heat for industrial proteolysis. X = protease for biological proteolysis
Hydrolysis of protein. X = HCl and heat for industrial proteolysis. X = protease for biological proteolysis

Mechanical Properties and Classification

Proteins vary wildly in their physical properties depending on their class (fibrous, globular, or membrane). The Young's modulus, a measure of stiffness, varies significantly across different protein types.

Mechanical Properties of Selected Proteins
Protein Protein Class Young's Modulus
Keratin (cross-linked) Fibrous 1.5–10 GPa
Collagen (cross-linked) Fibrous 5–7.5 GPa
β-barrel outer membrane proteins Membrane 20–45 GPa
Fibrin (cross-linked) Fibrous 1–10 MPa
Elastin (cross-linked) Fibrous 1 MPa
Bovine serum albumin (cross-linked) Globular 2.5–15 kPa

Molecular surface of several proteins showing their comparative sizes. From left to right are: immunoglobulin G (IgG, an antibody), hemoglobin, insulin (a hormone), adenylate kinase (an enzyme), and glutamine synthetase (an enzyme).
Molecular surface of several proteins showing their comparative sizes. From left to right are: immunoglobulin G (IgG, an antibody), hemoglobin, insulin (a hormone), adenylate kinase (an enzyme), and glutamine synthetase (an enzyme).

Frequently Asked Questions

What is the difference between a peptide and a protein?

A peptide is a short chain of amino acids, typically containing fewer than 20–30 residues. A protein is a larger biomolecule consisting of at least one long polypeptide chain.

How is a protein's function determined?

A protein's function is determined by its specific three-dimensional structure, which is a result of the unique sequence of amino acids encoded by its gene.

What happens to misfolded proteins?

Misfolded or abnormal proteins are typically degraded more rapidly than healthy proteins, often through the action of the proteasome.

What are cofactors in protein activity?

Cofactors are non-protein chemical compounds or ions that certain proteins require to achieve biological activity.

What is post-translational modification?

This is a process where amino acid residues in a protein are chemically modified after or during synthesis, altering the protein's physical properties, stability, and function.

References

  1. Osborne TB (1909). "History". The Vegetable Proteins. pp. 1–6.
  2. Reynolds JA, Tanford C (2003). Nature's Robots: A History of Proteins (Oxford Paperbacks). New York, New York: Oxford University Press. p. 15. ISBN 978-0-19-860694-9.
  3. Tanford C (2001). Nature's robots: a history of proteins. Internet Archive. Oxford; Toronto: Oxford University Press. ISBN 978-0-19-850466-5.
  4. Mulder GJ (1838). "Sur la composition de quelques substances animales". Bulletin des Sciences Physiques et Naturelles en Néerlande: 104.
  5. Hartley H (August 1951). "Origin of the word 'protein'". Nature. 168 (4267): 244. Bibcode:1951Natur.168..244H. doi:10.1038/168244a0. PMID 14875059. S2CID 4271525.