1. Introduction
Up to the late 1970s, biochemistry had only dealt with molecular ensembles for which the laws of thermodynamics are readily applicable.
Ensemble methods have played and are still playing central roles in biological research, and have produced wonderful results such as the geometric principle of protein functioning and proteins having native structures, [1] and [2]. Anfinsen’s Thermodynamic Hypothesis was inferred from experiments and observations on molecular ensembles, which claims that “This hypothesis states that the three-dimensional structure of a native protein in its normal physiological milieu (solvent, pH, ionic strength, presence of other components such as metal ions or prosthetic groups, temperature, and other) is the one in which the Gibbs free energy of the whole system is lowest; that is, that the native conformation is determined by the totality of interatomic interactions and hence by the amino acid sequence, in a given environment.” [3].
In ensemble experiments, single molecule behaviour can only be inferred from the average of the ensemble. Averaging methods cannot really tell us the individual molecule’s dynamics.
But the reality in biology is that in natural biological environment, only a few molecules of the same biological macromolecules are involved in any particular interaction. For example, according to [4], for some yeast proteins, there are only 60 molecules per cell. Thus to follow nature, experiments, measuring, and observing of biological macromolecules should take single molecule methods. Indeed, the introduction of [5] is entitled as “Molecular biophysics at the twenty-first century: from ensemble measurements to single-molecule detection”.
Single molecule experiments are well developed now, single molecule theory lags behind. To solve the protein folding problem, the author has tried to develop a single molecule theory of protein dynamics based on two concepts, single molecule conformational Gibbs free energy function (CGF) and single molecule thermodynamic hypothesis (STH) that claims that all stable conformations are (local or global) minimizers of CGF, [6]-[13].
2. Method
Although a cell is crowded with various molecules, ions, etc., in protein folding process at any moment there are just a few molecules of the same protein, as stated in [14], “Folding generally involves only one molecule at a time, working, at least in most cases, without the aid of any other molecular actors except a suitable solvent. So no fancy biology needs to be invoked―chaperones, which after all consume valuable ATP, are actually used quite sparingly in vivo.” Thus, in dealing with protein dynamics we may imaging that protein molecules fold in their physiological environment independently to one another.
We can make a mental experiment that a single protein molecule folds in certain environment, its conformation changes from the initial conformation to the final, native structure, leaving a folding path consisting of a series conformations. Each of the conformations in the folding path possesses a Gibbs free energy whose value also depends on the environment such that the final, native structure has the minimum Gibbs free energy among all conformations in the folding path. These are the ideas of conformational Gibbs free energy function and single molecule thermodynamic hypothesis. The formal description is as follows.
Suppose that
is a molecule consisting of n atoms
. A conformation of
is a point in the 3n-dimensional Euclidean space
and denoted as
, where
is the nuclear centre of the atom
. When talking conformations we assume all covalent bonds and bond angles are correctly formed. There are standard bond lengths and angles for a molecule
which should be respected in any conformation
of
. Thus, we require that a conformation
of
satisfying the steric conditions below.
(1)
where
is the distance between
and
,
and
are the van der Waals radii of
and
,
and
are small positive constants,
are the standard bond lengths between
and
,
is the standard bond angle if there are covalent bonds between
and
and also between
and
,
are bond angles measured in
. The set of all conformations of
is denoted as
.
Let a molecule
be situated in an environment
(solvent, pH, ionic strength, temperature, pressure, etc.). Each conformation
occupies a space
in
. Adding one layer of environment particles surrounding
, we get a tiny open thermodynamic system
tailor made for the conformation
(“open” means that particles of
can enter and leave the system). The Gibbs free energy
of
can be denoted as the value of a function
at
.
is a single molecule conformational Gibbs free energy function (CGF), a function whose variables are the conformations Rs.
also has two groups of parameters
and
, the molecule and the environment.
Note that for different environments
and
,
and
are different functions. For example, if
and
are two different conformations of
, comparing
and
does not make sense. But comparing
and
does make sense, it shows the environment influence to the same conformation
. Similarly, let
and
be two molecules, in general we should not compare
and
, since to begin with,
and
are different.
The single molecule thermodynamic hypothesis (STH), is a single molecule version of Anfinsen’s Thermodynamic Hypothesis (ATH). It claims that in the environment
, all stable conformations of
must be (local or global) minimizers of the CGF
.
For small monomeric globular proteins (100 to 400 residues) many believe that the native structure
of
is the global minimizer of
where
is the aqueous solvent (water) environment, i.e.,
(2)
In 2011, the concept of CGF was thought as absurd, see [15] in which it was argued that such a function cannot exist and is the biggest pitfall of ATH. In fact, formula of the CGF,
, has been derived via quantum statistics,
(3)
where
is the intra-conformational potential energy depending only on the conformation
;
is the mean number of electrons in the tiny open thermodynamic system
and
is the chemical potential of an electron;
’s are mean numbers of first layer water molecules surrounding
which are positioned nearby moieties of
that having hydrophobicity level i,
, such water molecules then have chemical potentials
’s. Thus, the chemical potential of a water molecule in
depends on the water molecule’s position. All hydrophobic moieties have positive
’s, all hydrophilic moieties have negative
’s.
in general, if
is a protein then
since proteins are amphiphiles having at least two hydrophobicity classes. Occasionally, depending on conformation
, moieties may change from hydrophilic to hydrophobic. For example, when two polar or charged moieties formed a hydrogen bond or neutralized their charges due to their close up in
, then the two moieties should be denoted as hydrophobic in
. These intramolecular hydrogen bonds and charge neutralizations are important elements of the intra-conformational potential function
.
A geometric approximation of (3) is by an interface
between
and water molecules in
.
is a closed surface (no boundary) of genus zero (homeomorphic to a sphere).
is the boundary of a bounded domain
such that
, where the boundary
of
is just
. All water molecules in
are contained in
except that
has interior holes (contained in
) big enough to contain some water molecules, in this case, these interior water molecules are not counted in the one layer water molecules.
The geometric approximation of Formula (3) is as follows:
(4)
and
are volume and area,
the electron chemical potential per volume;
the diameter of a water molecule;
the area of the subsurface
covering all moieties of hydrophobicity level i that are exposed to water in
,
,
’s the chemical potentials per area.
The advantage of
in (4) is that it is differentiable, thus we can use its gradient
to discuss such as stable conformations and folding dynamics. For example, the negative gradient
will be the main folding force (Anfinsen stated in [3] that “This process (protein folding) is driven entirely by the free energy of conformation that is gained in going to stable, native structure.”). The advantage of Formula (3) is that it is easy to grasp and apply to formulations, such as the docking Gibbs free energy difference in §3.4.
Formulas (3) and (4) are derived via quantum statistics in [6]-[9] and [11]. Note that it is always deriving Formula (3) first via quantum statistics, then approaching it by Formula (4). To fill in the argument in limited pages, preparation materials are just treated as well known. Some readers may feel difficult to understand the derivation. Detailed derivation, including all necessary preparations in thermodynamics, statistic mechanics, and quantum chemistry, etc., will appear in a coming monograph, the preparation and derivation will occupy over 70 pages.
In previous derivations, the intra-conformational potential
was presented only as the Coulomb’s potential of nuclei
. This should be modified to also consider potentials involving electron cloud.
The term
in (3) comes purely from quantum mechanics. Applying classical statistics will result in the same formula as (3) but without this electron quantum term. Term
in (4) comes from translating
to
via a quantum mechanics argument, then approximate
by
.
The term
in (3) and its counterpart
in (4) come from the surrounding water molecules, thus can be called the Gibbs free energy of aqueous solvent. It reduces the heavy calculations involving surrounding water molecules in molecular dynamic simulation to an easier to calculate boundary energy formula. It is called solvation free energy in literature. Molecular dynamics simulators pursued it but unsuccessful, only resulting in a no calculable integral formula that is not recognized as the solvation free energy but is called the potential of mean force or effective energy, see for example, [16] [17].
The occupied space
of a conformation
is determined by an electron wave function Φ in quantum mechanics. It is enveloped by the largest connected branch of the boundary
, where
is a small number.
To make the interface
easy to calculate, we replace
by a very good approximation, a bunch of overlapping balls
, where
is the round ball centred at
with the van der Waals radius
of the atom
. Figures on page 4 of [18], where
,
au, 1 au = a0 = 0.529188 Å, show that
is so like a bunch of overlapping balls.
With
replacing
, there are many choices of the interface
, such as the van der Waals surface (it is just the outmost component of
); the solvent accessible surface defined by Lee and Richards in [19]; and the molecular surface defined by Richards in [20]. The first two are both bunches of overlapping spheres. The latter two surfaces are both generated by rolling a sphere of diameter
(the diameter of a water molecule) over
(or
, only that we do not know exactly what is
), see [21] and [22], if in case they are not connected, we take the outmost component of them as
.
A question was often asked, what are the use of these concepts and formulations? This article discusses the applications of this single molecule theory, predictions and explanations to various aspects of protein dynamics and structures. Mechanisms of dynamic protein structure, in particular protein folding, binding, allostery, denaturation, all are predicted and explained by CGF and STH.
3. Discussions
A theory should be able to make verifiable predictions and explain various known phenomena, the more the better.
With the shift of point of view from ensemble to single molecule, the concept of CGF together with STH can explain phenomena such as allostery, protein denaturation, and intrinsically disordered proteins.
In §3.1, STH predicts that after binding there will be conformational change, thus allostery is result of post-binding deformation, i.e., after binding conformational change, either they have allosteric effect or not. Thus the mechanism of the second secret of life-allostery, is nothing but the dynamic second law of thermodynamics. Life follows the dynamic second law of thermodynamics.
In §3.2, we argue that by CGF and STH, denaturation and folding are both the process of reducing conformational Gibbs free energy to achieve a stable conformation, the only difference is that they happen in different environments.
In §3.3 we discuss intrinsically disordered proteins, pointing out they are caused by the ensemble point of view. In fact, by CGF and STH, each individual molecule will find a stable conformation. If in an ensemble of molecules of the same protein there are multiple stable conformations, we get intrinsically disordered proteins.
Even not knowing the accurate values of the chemical potentials prevents us to do the most important verifiable prediction, ab initio prediction of proteins’ native structures, there are still other applications of the Formulas (3) and (4).
In §3.4, we derive a docking free energy difference formula
via (3). Unlike the quantitative task of structure prediction that needs accurate values of the chemical potentials, the docking free energy difference Formula (5) provides a qualitative but practical way to search would be docking sites that will be useful in finding new drugs. Finally we define the single molecule binding affinity such that the bigger the affinity the stronger the binding.
In §3.5, we apply (4) to predict the global geometric feathers of monomeric globular proteins, these predictions also explain the physics behind the well known global geometric features of globular proteins, they look like globules and are compactly packed, often, with hydrophobic cores in the structure interiors.
In §3.6 we argue that the dominate folding force is quantum force and hydrophobic force (aqueous force). We will show only a vanisingly tiny portion of polypeptide chains can be proteins. To be a monomeric globular protein, the entire polypeptide chain must be able to be arranged a delicate and balanced spacial arrangement to lower the intra-conformational potential
. We also argue that
cannot be dominate folding force.
In §3.7, we suggest a wholistic, top-down point of view of monomeric globular protein structures and list evidences supporting this view.
3.1. Demystify the Second Secret of Life: Mechanism of Allostery
Proteins work through binding to ligand.
3.1.1. Post-Binding Deformation
A dynamic version of STH, also a prediction, is: If a conformation
of
is not stable in environment
, then it will spontaneous (or be forced by
) fold to a stable conformation
which must be a (local or global) minimizer of
.
STH predicts that after binding there should be conformational change, or post-binding deformation, let
and
be stable conformations of
and
in environment
, it is very unlikely that
is also a stable conformation of the complex
, therefore, STH predicts that it will fold until it achies a stable conformation
that is necessarily a (local or global) minimizer of
. In particular,
.
3.1.2. Allostery
Post-binding deformation was first noted in allostery, thus our prediction of it was already virified.
In ([23], p. 434) allostery is described as “In both enzymes and receptors, the inherent function of the protein (catalysis, binding), which occurs at a certain location (e.g., the active site in enzymes), can be modulated by binding of a ligand at a different location. This phenomenon is called ‘allostery’ and the site to which the modulating ligand binds is called allosteric site.”
Allostery is the second secret of life, “Monod’s recognition of this important biological role led to the historical description of allostery as ‘the second secret of life’, second only to the genetic code.” [24].
3.1.3. Mechanism of Allostery
The mechanism of allostery is the mechanism of post-binding deformation. There are no allostery proteins, any protein after binding will have conformational change. For some proteins, post-binding deformations have the results of allostery. Post-binding deformations also plays important roles in proteins’ functions, for example, catalysis.
3.1.4. Dynamic Second Law of Thermodynamics
The second law of thermodynamics says that in an open thermodynamic system
surrounded by a heat bath of constant temperature and pressure, the Gibbs free energy will achieve the minimum at equilibrium. But our
is not fixed, changing conformation to reduce the Gibbs free energy will also change the system
, i.e., we are pursuing a system with minimum Gibbs free energy among all systems
, instead of pursuing minimum Gibbs free energy for a fixed system
. We will call this kind of minimization of Gibbs free energy the dynamic second law of thermodynamics. Protein folding and binding follow this dynamic second law of thermodynamics.
3.1.5. Life Follows Dynamic Second Law of Thermodynamics
Proteins’ functionalities rely on specific binding, the specificity comes from proteins’ native structures. Native structures and post-binding deformations come from spontaneous folding whose mechanism is the dynamic second law of thermodynamics. Since protein folding and allostery play central roles in protein functioning and regulations in almost all cellular processes [25], we may safely say that life follows the dynamic second law of thermodynamics.
3.2. Mechanism of Denaturation
All known denature phenomena are caused by changing an environment
in which a protein
has a native structure to a denaturation environment
. Many ways can change environment. “Acids and alkalies, salts of heavy metals, alcohol, ether and other organic solvents, concentrated urea and related compounds, heat, ultraviolet light, high pressure, shaking, supersonic, waves and even drying: all these can induce denaturation of the protein.” [2].
Roughly speaking, before the change of environment, an ensemble of protein molecules was in a kind of equilibrium such that most of these molecules are in native structure in the environment
. When the environment finally become the denaturation environment
, according to the STH, all molecules in this ensemble will fold to stable conformations, these stable conformations are minimizers of
.
The difference of folding and denaturation is that after folding almost all molecules take the same stable conformation
, a minimizer of
; but in denaturation molecules will take many different stable conformations,
may or may not be a minimizer of
.
For example, consider
is the native structure of a protein
in its physiological environment
that has constant temperature T. Now we change the environment by lifting T. In terms specific (per unit mass) internal energy
and entropy
, the per unit area chemical potentials
in (4) can be expressed as
, where m is the mass of a water molecule. Since it is always
, if we lift T sufficiently high, eventually all
and
will become negative. In particular, hydrophobic area
, hydrophilic area
. Once all chemical potentials become negative, according (4), the larger the volume
and area
, the lower the Gibbs free energy
, the conformation will become an extended random coil. Continue lifing T will break the molecule. Stop lifting T before breaking, many extended random coils will become local minimizers of
.
This is actually the observed phenomenon as stated in ([26]: p. 13), “Indeed the biological macromolecules, proteins as well as the nucleic acids attain ordered structure that are important and perform their functions in the cells as structures that are stable under certain circumstances, in particular low temperatures but can be disintegrated, ‘denaturised’ to flexible, non-ordered randomised coils at higher temperatures.”
Why in
, folding from a single initial conformation
can result in different stable conformations? Are stable conformations independent of initial conformations?
Actually, the sudden changes of the same natives structure to many different structures in the initial denaturation process of an ensemble of protein molecules is a phenomenon often happenes in nature, mathematical description of it is the catastrophe theory developed by Rene Thom [27], a definition of it is given in [28]: “Catastrophe theory is concerned with the mathematical modelling of sudden changes―so called ‘catastrophes’―in the behaviour of natural systems, which can appear as a consequence of continuous changes of the system parameters.” Indeed, from the physiological environment
to the denaturation environment
, there must be a one family of environments
connecting them. Although these environments are continues parameters in the CGFs
, catastrophe does happen so that various copies of the same native structure
suddenly changed to different conformations for different molecules and when the environment parameter reached the denaturation environment
, these changed conformations become different initial conformations.
3.3. Intrinsically Disordered Proteins
A large proportion of proteins are intrinsically disordered, “After half a century of structural studies on beautifully folded globular proteins, it is perhaps a shock to discover that up to some 40% of the proteins in the human proteome are estimated to be intrinsically disordered and become fully or partly structured on binding to binding partners in the cell.”, [29]. A recent literature analysis showed that there are approximately 1150 non-redundant proteins in the list of validated intrinsically disordered proteins, [30].
By CGF and STH, in the level of single molecule, any molecule always has a stable conformation and it must be a (local or global) minimizer of the CGF. For proteins having native structures, these molecules’ stable conformations are the same. Repeat testing in the same environment always get the same stable conformation. For intrinsically disordered proteins, there are multiple stable conformations, therefore, repeat testing in the same environment will get different stable conformations, or none at all. Thus, looking in an ensemble, they seem to have no structure at all in an averaging observation such as X-ray crystollgraphy.
The concepts of ordered and disordered proteins were introduced in the level of ensemble of molecules. It is because that all protein structure determination methods working on ensembles of proteins, except the cryo electron microscopy. For example, the X-ray crystallography experiments are performed to crystals of ensembles of protein molecules. In the crystal, protein molecules are in cells, the whole structure is obtained by adding up informations of individual cells together to get an average conformation. If all cells have the same conformation, for example, their structures are parallel motions (congruence) to each other, then we obtain a native structure by adding up all cells’ information. But if in different cells the molecules’ structures are different (no longer congruent), then the adding up of these structures will be blurred, showing no structure at all. Therefore, the situation for disordered proteins is: looking at individual molecules in the ensemble, each has a stable conformation, collectively as an ensemble, no average structure at all.
3.4. Docking and Binding
Formulas in (3) and (4) open the door of investigating all protein dynamic phenomena by first principle. Although quantitative investigations have to wait the relative accuracy chemical potential values to be determined, we still can do qualitative work like the following.
3.4.1. Docking Free Energy Difference
Here we just consider binding in the aqueous solvent
. Binding starts with docking, a process of two molecules come to each other and remove water molecules between surfaces of molecules (note that conformations are invariant under orientational preserving congruence of
, so we use the same notations
and
to express the moving conformations). Let
and
be the docking sets, i.e., water molecules between them are removed such that
and
(almost) touching each other along S and T to form a new combined conformation
of the would be complex
.
Let
be the open thermodynamic system tailor made for the conformation
, i.e., it is composed by
plus one layer of water molecules. The interface
inside
is
Starting from
along ST, whether or not a complex
can be formed and be further folded to a stable conformation will depend on the docking Gibbs free energy difference
If
,
has the chance to further fold to a stable conformation
of the complex
; if
, the conformation
will not stand, the two separate conformations
and
will be more stable than it. In other words, if
, the complex
has no chance to stabilize along the site ST.
3.4.2. Docking Free Energy Difference Formula
Suppose that there are
water molecules in
,
molecules in
, etc. Docking along ST means that
water molecules covering S and T are removed from the open thermodynamic system
. Denote
as the Gibbs free energy of
, remembering that if
along ST is possible, it is just
. By the formula in (3),
Let the total number of water molecules in
,
, and
be
then
Since
we have
(5)
3.4.3. Searching for Potential Docking Sites
To make
negative, the best chance is that we make all terms in (5) negative or zero. Since
,
. To make
negative, there should be more
than
among all hydrophobic classes
. Similarly, to make
negative, there should be more
than
among all hydrophilic classes
. Thus the best chance of
is that
,
. Indeed, in that case,
for each
, hence
; and
for each
; furthermore, since in that case it must be
,
.
Let
(
) be the set of atoms in
(
) such that these atoms (almost) touching S (T). Let
(
) be the subset corresponding to
(
). The potential difference
is essentially the potential between
and
,
. To make
negative or small positive, we want that S and T are complementary both in geometry and in electrostatic distribution, such that the distances between atoms of
and
in the conformation
are small, and positive and negative charges can be matched to produce negative or small positive electrostatic potentials, and polar moieties may match to form hydrogen bonds to further reduce docking free energy.
In practice, we may find that connected pieces
and
almost contained in
and
except containing some tiny hydrophilic pieces (hydrophilic islands) at which a hydrophilic moiety has a hydrogen bond with water, or a charge neutralized by water. We will call such S and T as hydrophobic pieces with hydrophilic islands. So we should search the largest such hydrophobic pieces with hydrophilic islands as candidates of docking site S and T, if they are complementary both in geometry and in electrostatic distribution. Indeed, though hydrophilic islands means water molecules with
being removed, geometric complementary and matching hydrophilic islands in S and T to either neutralize charges or form hydrogen bonds can reduce
to compensate or at least partially compensate the increased
, and keep
.
3.4.4. Single Molecule Binding Affinity
If
, STH predicts that
will spontaneously fold to a stable conformation
of the complex
.
is a (local or global) minimizer of
, so
. Thus, we have
(6)
The single molecule binding affinity is defined as
where
is the Boltzmann constant and T the temperature in K˚. Because of
,
, the bigger the K, the stronger the binding.
3.5. Global Geometric Featuers of Monomeric Globular Proteins
Based on the assumption (2), applying the formula in (4), we can predict that the native structure of a monomeric globular protein will have the following global geometric features:
1) It is compactly packed or it has a high packing density.
2) It has a globular shape, i.e., looks like an ellipsoid even sphere and its surface area is small comparing with other conformations.
3) Hydrophobic moieties trend to hide from water as much as possible.
These global geometric features of monomeric globular protein native structures have been well known, only via Formula (4) the physical laws behind these global geometric features become clear, they are all the results of minimization of the CGF
. And, as will be shown, even nobody knows them, we can infer them from Formula (4).
3.5.1. Small Volume
Means High Packing Density
Our first prediction of the global geometric feature of the native structure of a monomeric globular protein is that it must have a high packing density.
Let us clarify the relationship between density and volume by global geometry. Suppose that some quantity Q is distributed inside some domain Ω, i.e., out of Ω the quantity Q is zero. Then density is a measure of quantity Q per volume unit expressed as
.
Therefore, to measure density globally, the quantity Q and domain Ω are necessarily required. We have talked that a conformation
occupies a space
, here the quantity Q is the volume
. We also need a packing container to put
in. As stated in introduction,
, thus, the container is
.
Physically
is packed into the tailor made container
, how effectively we packed the molecule
then is measured by the packing density,
(7)
Note that
, and
if and only if
, thus,
is always less than 1 unless
. Because the van der Waals force, in nature a water molecular cannot really touching
, so the interface should not be selected as
. Thus,
is always less than 1.
We will replace
by its good approximation
as explained in Method. Note that since
,
For all conformations,
is a constant. By the steric condition (1), two balls
and
is overlapping, if and only if the atoms
and
are covalently bonded, i.e., the distance
for not covalently bonded atoms
and
,
. For covalently bonded
and
, since
is always around the standard bond length as required in (1),
will be almost a constant in all conformations, thus the overlapping volume
will almost be the same in all conformations. In rare cases that there are extra covalent bonds such as the disulfide bond in some conformations, the overlapping volume will increase a little. We conclude that the overlapping volume is almost the same for all conformations. Therefore, the volume
is almost the same for all conformations and equation (7) shows that the volume
alone determines the packing density
. Therefore, shrinking
is the only way to increase the packing density
.
Since
, shrinking
will reduce
and simultaneously enlarge the packing density. Let
be the native structure of
, then by (2), for any other
,
. Due to the effects of other terms in the Formula (4),
may not be the minimum volume, but the packing density surely is as high as possible.
Indeed, high packing density was observed half century ago. In 1974, Richards observed that protein interiors are closely packed and he showed the interior of lysozyme and ribonuclease having a packing density of 0.75 ([20]). But the definition of packing density is not the global one given in (7), it was via complicated local calculations to determine each atom occupy how much space. Only much later the relations between packing density and volume was studied, [31].
The global definition of packing density (7) shows that shrinking
must result the packing density becoming uniformly or homogenously high such that water molecules should not appear inside
since that will lower the packing density. This also was noted long ago, for example, it was observed that in general, water molecules are excluded from the interior of globular proteins, [20] [32] [33]. Unfilled cavities large enough to accommodate a water molecule only appear in very few cases, [20] [33]-[35]. In the literature, phrases such as “knobs in holes” and “ridges in grooves” are often used to describe the compactness (high packing density) of the packing, [32] [34] [36] [37], see Figure 1.
![]()
Figure 1. Left: Polyproline helices PPI (a right hand screw) and PPII (a left hand screw). Hydrogen bond is not necessary for a helix. Taking from [46]. Right: Taking from [47]. These screws make the packing very tight by occupying smaller space.
3.5.2. Small Surface Area Resulting Globular Shape
Our second prediction of the global geometric features of a monomeric globular protein’s native structure is that it is shaped like a globule because that if volume
is the conformation packer, the surface area
is the conformation shaper. This is because the Isoperimetric Inequality: Among all regions
with finite volume
and boundary area
,
(8)
Alternative statements of the isopremetric inequality is that: 1) among all regions with the same volume
and finite boundary surface area, only the round ball of radius
has the smallest boundary surface area
; 2) among all closed surfaces with the same surface area
, only the sphere with radius
bounds the largest volume
.
The isoperimetric inequality gives us a theoretical method to make a perfect sphere. Take a dough and change its shape (the volume V of the dough is preserved in changing), measure the boundary surface area A and calculate
, it is always
. If for some shape we have
, without looking at the dough, we know that we have got a perfect sphere by shaping the dough to a round ball.
Of course, in practice, nobody will make ball or sphere with such a clumsy procedure. But for protein structure, since
, by (4), shrinking
will reduce
. Since
and
is almost a constant for any conformations, the isoperimetric inequality implies that to make
smaller and smaller, the
must go towards to a round ball, otherwise,
would not keep shrinking as (2) required.
On the other hand, isoperimetric inequality also shows that
. Since
is almost a constant, one cannot make
too small.
Like the high packing density, it is observed long ago that the native structure of a globular protein has smaller (solvent accessible) surface area than that of other conformations ([19] [32] [33]). Janin ([38], an appendix to [39]) observed the need of “globular proteins to achieve a minimum accessible surface area compatible with their mass.” During the folding of a fully extended polypeptide chain to give the compact native structure for proteins with a molecular weight of 15,000, the (solvent accessible surface) area decreases to about one-third of its maximum value. The
is shrunk so much, it is observed that even the accessible area of charged groups (hydrophilic surface area
in our case) is substantially lower in the native protein (40% - 60% of
) than in the extended chain, see [32]. This decreasing of
will enlarge the
according to (4), one more example that the shrinking of
has to be cohesive. The
is decreased so much, thus
has to be shrunk too, the decreased Gibbs free energy via decreasing
and
seems more than enough to compensate the increased Gibbs free energy via shrinking
.
Novotny et al. [40] and [41] constructed incorrect conformations by putting the side chains from the N-end to the C-end of a globular protein (sea-worm hemerythrin) into the backbone of the native structure of another globular protein (variable domain of mouse immunoglobulin κ-chain) of the same length (113 residues), and vice versa, to get two incorrect conformations of proteins with known native structures. Comparing the incorrect conformations with native structures, they found that the solvent accessible surface areas of the incorrect conformations are significantly larger than that of the native structures. Many others also observed that smaller surface area phenomenon and by using various different models, concluded that it is a structural feature of the native structures of globular proteins, see, for example, [33] [42]-[44].
3.5.3. Hydrophobic Moieties Tend to Avoid Contacting with Water
By Formula (4), besides shringking the intra-conformational potential
and the electron quantum terms
, shrinking the hydrophobic area
and enlarging the hydrophilic area
also contribute to reducing the Gibbs free energy
. Thus, by (2), besides small volume and surface area, we can predict that the native structure of a monomeric globular protein should have a small hydrophobic portion and large hydrophilic portion in its surface. Or, in other words, hydrophobic moieties tend to avoid contacting with water.
3.5.4. Hydrophobic Core
In fact, further mathematical argument not involving (4) can show that if the polypeptide chain length N is between 40 and 250, there will be hydrophobic core that is untouched by water molecules. For
the structure will be composed by domains connected by loops, each domain having a globular shape and tightly packed hydrophobic core, we will leave the long geometric argument to elsewhere. Here we only point out that
(9)
is an indicator of how well a hydrophobic core is formed, Core(R) = 1 means all water exposed moieties are hydrophobic, Core(R) = 0 means a perfect hydrophobic core, i.e., not even one hydrophobic moiety is exposed to water.
3.6. Dominate Folding Force
3.6.1. Hyrophobic Force Comes From Aqueous Environment
In §3.5 we have shown that the native structure of a monomeric globular protein should be compactly packed and shaped like a globule with hydrophobic moiety hiding from water as much as possible. By formula of
in (4), shrinking
and enlarging
have the same result, reducing the Gibbs free energy. Shrinking
means that hydrophobic moieties having less chance to expose to water and indicating a hydrophobic effect. But enlarging
means that hydrophilic moieties will have more chance to expose to water. Because that
and
is also shrinking to reduce the Gibbs free energy, enlarging
will indirectly cause hydrophobic moieties having less chance to contact water, the same effect as the hydrophobic effect. Thus, Formula (4) shows the idea that there is a hydrophilic force or hydrophilic effect, [45], has a reason.
We can consider them together as the hydrophobic force, since they comes from the aqueous environment, we can also call them the aqueous force, or in general, environment force. In this sense, we may say that hydrophobic force is a force of protein folding. Is it the dominate folding force as claimed in [42]? It certainly is an important folding force.
3.6.2. Quantum Force
In §3.5 we have shown that the volume term
in (4) is the structure packer and the area term
is the structure shaper, without them we can hardly imaging the native structures of monomeric globular proteins are tightly packed and look like a glouble. Thus
and
are also major folding forces. As before mentioned, they come from quantum effect, let us call them quantum force.
3.6.3. Weak Forces
In ([25], pp. 110-114), there are four weak folding forces, or noncovalent bonds: “The folding of a protein chain is also determined by many different sets of weak noncovalent bonds that form between one part of the chain and another. These involve atoms in the polypeptide backbone, as well as atoms in the amino acid side chains. There are three types of these weak bonds: hydrogen bonds, electrostatic attractions, and van der Waals attractions, …” These weak forces are contained in
in formulas in (3) and (4). The fourth weak force in [25] is hydrophobic clustering force, just our hydrophobic force.
3.6.4. Dominate Folding Force
We claim that the quantum force and the aqueous force together is the dominate folding force. It shapes the native structure of a monomeric globular proteins such that it looks like a globule, it packed the native structure compactly and hiding the hdyrophobic moieties inside the interior as a hydrophobic core.
There are strong believes that hydrogen bonds should be dominate folding force. Hydrogen bonds essentially comes from electrostatic distribution, therefore, Intra-molecular hydrogen bonds, electrostatic attractions, and van der Waals attractions are countered in the intra-conformational potential
. Intermolecular hydrogen bonds are counted in aqueous force.
Since intra or inter molecular bydrogen bonds reduce similar amount of energy,
cannot be the dominate folding force. But it is necessary in stabilizing the folded native structure due to the fact that proteins are special among polypeptide chains.
3.6.5. Proteins Are Special
Due to the existence of main chain hydrophilic moieties (the NH and CO groups), even Core(R) = 0, there are still hydrophilic portion inside the interior of the protein, the hydrophobic core is never pure hydrophobic. These main chain interior hydrophilic moieties have to form intramolecular hydrogen bonds, otherwise the intra-conformational potential
will be too large such that the conformation is not stable. Secondary structures such as helices, strands, and sheets are not only packed tightly, but also supply spacial arrangements of the entire polypeptide chain that are able to form regular hydrogen bonding patterns associated with secondary structures. Besides the main chain hydrophilic moieties, if there are hydrophilic side chains in the interior of the protein, they have to be arranged in space such that all polar and charged interior side chains would be nearby enough to each other to neutralize electron charges, or to form hydrogen bonds, to keep
low. The longer the polypeptide chain, the more difficult to arrange such a delicate and balanced spacial arrangement of the entire polypeptide chain.
Because of this delicate and balanced spacial arrangement of the polypeptide chain for native structures of monomeric globular proteins, the polypeptide chains of them must be very special, and perhaps is selected by natural selection for just this advantage.
Not every polypeptide chain can have this kind of delicate and balanced spacial arrangement, i.e., polypeptide chains of proteins are very special, they are selected by nature selection. It is estimated in [6] that for a random polypeptide chain
of 400 residues, the probability
of
is a protein’s polypeptide chain is
is practically zero. The monomeric globular proteins are even more special.
3.7. A Wholistic View of Native Structures of Monomeric Globular Proteins
By assumption (2), the global minimizer of
is the native structure
. The dominate folding force, quantum force and aqueous force, are global ones, i.e., only taking the protein molecule as a whole we can measure them. Shrinking or enlarging global measures such as volume and area and hydrophobic and hydrophilic areas cohesively (in synergy) is a global folding force. As shown above, under these global forces the entire polypeptide chain has to be arranged in a delicate and balanced spacial arrangement, only a vanishingly tiny portion of polypeptide chains can be polypeptide chains of monomeric globular proteins.
As discussed above, the delicate and balanced spacial arrangement of the entire polypeptide chain prefers secondary structures such as α helices and β strands, and semi-local structures such as the pleated β sheets, with their regular hydrogen bond patterns.
Because of these global forces, we suggest a point of view of native structures of monomeric globular proteins, a wholistic top-down point of view: The entire molecule as a whole determines the local structures, i.e., global folding forces produce a compactly packed tertiary structure that looks like a globule. Secondary structures appear as by-products of this global folding force pushing high packing density and keep
as low as possible via the speciality of monomeric globular protein polypeptide chains that can be arranged a delicate and balanced spacial arrangement. Moreover, under this global folding force, this delicate and balanced spacial arrangement also hides hydrophobic moieties away from water as much as possible.
On the contrary, a bottom-up point of view of native structures of proteins believes that secondary structures appear due to the propensities of short peptide chains independent of the whole polypeptide chain, and the tertiary structure is formed by a suitable arrangement of these secondary structures connected by turns and loops.
In fact, a wholistic point of view was already implied in Anfinsen’s Thermodynamic Hypothesis, “the native conformation is determined by the totality of interatomic interactions and hence by the amino acid sequence, in a given environment.” [3].
Allosterey can make long distance conformational change, and other facts in the following subsections, PPI and PPII helices without hydrogen bonding, simulations by reducing hydrophobic surface alone can produce secondary structures and hydrogen bonds, chameleon sequences, point mutation causes dramatical conformation change, all of them support the wholistic top-down point of view. The bottom-up point of view cannot explain these phenomena.
3.7.1. Secondary Structures Are Due to Compact Packing
Regular intramolecular hydrogen bond patterns associated with α helix and pleated β sheet lowered the intra-conformational potential
in (3) and (4), thus
or
is lowered. These intramolecular hydrogen bond patterns give a reason to support of the bottom-up point of view of protein structures.
The wholistic top-down point of view suggests that it is the compact packing that produces the secondary structures, hydrogen bond patterns associated with secondary structures are only by-products of this compact packing via a delicate and balanced spacial arrangement of the entire polypeptide chain, and only a vanishing portion of polypeptide chains can afford this delicate and balanced spacial arrangement. While screws (helices) and strands being formed to get compact packing, regular hydrogen bond patterns can be simultaneously realized due to this delitcate and balanced spacial arrangement of these naturally selected special polypeptide chains.
From the geometric point of view, there are only very few methods to compactly pack a uneven rope (the extended, long and thin wire conformation of the polypeptide chain), one of them is to form some kind of screw shape, i.e., a helix, another is to form straight strands and arrange these strand as sheets; these correspond to secondary structures. Turns and loops are necessary connecting portions of the wire between consecutive secondary structures, see [42]. Compactly pack a random polypeptide chain, we will find a lot of unmatched hydrogen bond donors and accepters, as well as unmatched charged side chains, leading to a large intra-confomational potential
, therefore, a large Gibbs free energy. Due to unable to arrange a delicate and balanced spacial arrangement of the random polypeptide, the only way of lowering CGF is make all those hydrogen bond donors and acceptors to form hydrogen bonds with water molecules, as well as to neutralize charges with water. This will increase surface areas
,
, and
, but the lowered
will be more than enough to compensate. That is why these random polypeptide chains are not proteins. Instead of natives structures, they will have many extended stable conformations, just like water soluble polymers. On the other hand, without main chain hydrophilic moieties, a polyproline peptide in a protein can form helices such as PPI and PPII as shown in Figure 1. Of course, these PPI and PPII are formed by compact packing and do not have any hydrogen bond.
The common geometric characteristic of PPI and PPII, as well as the regular α helices, is that their screw pattern makes the structure denser, or packed more tightly to reduce the occupied spaces to as small as possible. Once the concept of helices be modified accordingly, very simplified models such as the HP model [42] and tube model [48] obtained “secondary structures” like helices and sheets, of course no hydrogen bonds exist because of the simplicity of these models, but indeed the packing density is high. Thus, intramolecular hydrogen bond patterns are not necessary of secondary structures, instead, they are by-products of a global packing force to the specially selected protein polypeptide chains.
3.7.2. An ab initio Simulation of Shrinking Hydrophobic Area Alone Produces Secondary Structures and Hydrogen Bonds
An ab initio simulation of shrinking the hydrophobic surface area
alone has produced secondary structures such as α helices and β strands and their associated intramolecular hydrogen bonds, with statistical significance, [49]. This ab initio simulation shows that a global force alone can really produce secondary structures, supporting the wholistic top-down point of view of monomeric globular protein structures.
It is admitted in [50], hydrogen bonding must be explicitly modelled for helix formation, otherwise molecular dynamics simulation based on pairwise potential energy cannot produce secondary structures.
3.7.3. Chameleon Sequences
A chameleon sequence is a short peptide sequence that in native structures of different proteins may appear as different secondary structurss, for example, α helix in protein
, and β strand in protein
.
Note that in the discussion of mechanism of allostery above, although
is a stable conformation of
, but
and
may not be a stable conformation of
and
. Global (
) stability does not guarantee local (
and
) stabilities. Even though
and
are not stable as parts of
, the compensation is that
is stable. On the other hand,
and
are both stable, together
is not stable. Thus, local stabilities do not guarantee global stability either.
These observations suggest that local conformations has to adjust to one another in space to achieve global stability, therefore, suggest that the folding force must be a global one, the entire polypeptide chain(s) must be considered together to achieve a stable conformation. Indeed, in the water environment
, look at the CGF Formula (4), volume and area, and hydrophobic and hydrophilic areas, all are truly global, one cannot define them without onsidering the entire polypeptide chain.
A chameleon sequence has to take a α helix secondary structure to help achieving a stable conformation
of
; but the same chameleon sequence has to take a β strand secondary structure to help to achieve a stable
of
, even thought the β strand may less stable than α helix, the compensation is that
is stable. Stability of the conformation as a whoe requires the same chameleon sequence taking different secondary structures in different proteins.
3.7.4. Point Mutation
Point mutation, or single mutation, is the substitution of just one residue of a polypeptide chain. Although just one residue is substituted, point mutation may cause dramatic conformational change. The reference [51] is entitled a “One sequence plus one mutation equals to two folds.” The reference [52] is entitled as “Structure of the Hydrophobic Core Determines the 3D Protein Structure-Verification by Single Mutation Proteins”, its abstract states: “Four de novo proteins differing in single mutation positions, with a chain length of 56 amino acids, represent diverse 3D structures: monomeric
and
folds. The reason for this diversity is seen in the different structure of the hydrophobic core as a result of synergy leading to the generation of a system in which the polypeptide chain as a whole participates.”
4. Conclusions
We have demonstrated that CGF and STH can make predictions and explanations, but only qualitatively. Formulas of CGF in (3) and (4) have opened the door to quantitative investigations to resolve old problems in life science, such as the protein foliding problem, ab initio prediction of the native structure of a water soluble protein. Quantitative calculating the conformational changes in allostery is essentially also a folding problem. Resolutions of these problems will make not only theoretical progresses but also will have many practical applications. But, a obstacle remains.
Looking at Formula (4), because
, one way to shrink
is to change
such that the volume
and surface area
decrease. Similarly, we want the hydrophobic surface area
to get as small as possible and the hydrophilic surface area
to get as large as possible, to shrink
. There are some restrictions to these decreases and increases, for example, since
,
cannot become too large while
is shrinking with changing conformations. Indeed, to make
as small as possible, we have to balance between volume, area, hydrophobic area, and hydrophilic area; as well as the intra-conformational potential
. Thus, we should cohesively simultaneously shrinking the volume
, the area
, and the hydrophobic surface area
; as well as to shrink the intra-conformational potential
to make more intramolecular hydrogen bonds, neutralised charges, and to avoid overlapping of no covalent bonded atoms. If
and
both become smaller,
will increasing or decreasing depending on the degrees of shrinking of
and
.
How to cohesively (in synergy) do all of these? We need the accurate values of
,
,
,
,
. Once we know these accurate values, one can readily do ab initio structure prediction of a globular protein via Formula (3) or (4). Thus, the ab initio structure predicting problem for globular proteins and accurately predict post-binding deformations will become pure mathematical minimization problems.
Unfortunately, so far, the obstacle is that there is no accurate determination of these chemical potentials. Determining these chemical potentials is the next task towards resolving these famous problems by first principle. The determination methods will be both theoretical and experimental. Heavy computations in training on a set of native structures of monomeric globular proteins is necessary.