Full Text
GIANT MOLECULES
W. L. Bragg*)
I shall deal here chiefly with investigations of the structure of crystalline proteins by X-ray methods. This is a grandiose task. X-ray analysis is now being applied to more and more complex molecules; but in proteins we encounter a degree of complexity several orders higher than in the most complex organic molecules that have so far been successfully deciphered. The following table of molecular weights illustrates what has been said:
| Naphthalene | 128 | Tobacco-seed globulin | 300,000 |
| Penicillin | 302 | Hemocyanin | 6,800,000 |
| Insulin | 12,000 | Tomato bushy stunt virus | 10,000,000 |
| Myoglobin | 17,000 | ||
| Pepsin | 37,000 | ||
| Hemoglobin | 66,700 |
The naphthalene molecule was the first organic molecule successfully analyzed by X-ray methods about 25 years ago. Penicillin is one of the latest triumphs of X-ray analysis. Insulin, which is one of the simplest proteins, has a molecular weight almost 40 times greater than that of penicillin, and after it we pass to the colossal molecules of viruses and other complex proteins. All these substances form quite perfect crystals and give X-ray photographs. These X-ray photographs show that proteins are built very regularly down to atomic distances. The task of X-ray analysis is to find a way of interpreting these diagrams. It is not difficult to measure several thousand reflections from each crystal, and we are faced with the task of reading an enciphered letter without the key.
*) Nature 164, 7–10 (1949) (cf. UFN XXII, 98–104 (1939). Translated by E. N. Belova).
Proteins, as was shown by Emil Fischer in 1906, are built of long chains or rings composed of amino-acid residues. A typical amino acid has the formula
\[ \begin{array}{c} \mathrm{H}\quad R\\[-2mm] \diagdown\ \diagup\\[-1mm] \mathrm{C}\\[-1mm] \diagup\ \diagdown\\[-1mm] \mathrm{H_2N}\quad \mathrm{COOH} \end{array} \qquad \text{or} \qquad \begin{array}{c} \mathrm{H}\quad R\\[-2mm] \diagdown\ \diagup\\[-1mm] \mathrm{C}\\[-1mm] \diagup\ \diagdown\\[-1mm] {}^{+}\mathrm{H_3N}\quad \mathrm{COO}^{-} \end{array} \tag{1} \]
Here \(R\) denotes a monovalent side chain, specific for each characteristic amino acid. The acidic group of any amino acid may be linked with the basic group of another amino acid with the elimination of a particle of water; as a result of repeated repetition of such a reaction, a chain of residues arises, in which the different side chains \(R\) are like variously colored beads strung on a necklace. The formula of such a peptide chain will be:
\[ <\cdots \mathrm{CO} \begin{array}{c} \\[-1.2em] \end{array} \mathrm{NH} \begin{array}{c} R'\\[-0.3em] |\\[-0.3em] \mathrm{CH} \end{array} \mathrm{CO} \mathrm{NH} \begin{array}{c} \\[-1.2em] \end{array} \mathrm{CH} \begin{array}{c} \\[-0.3em] |\\[-0.3em] R'' \end{array} \mathrm{CO} \mathrm{NH} \begin{array}{c} R'''\\[-0.3em] |\\[-0.3em] \mathrm{CH} \end{array} \cdots > \tag{2} \]
We now know 23 different amino acids, whose complexity increases from glycine, in which \(R\) is a hydrogen atom, to phenylalanine, which contains a benzene ring, and tryptophan, which contains fused five- and six-membered rings. Most amino acids consist exclusively of carbon, oxygen, nitrogen, and hydrogen. But some also contain other elements, for example sulfur. In most amino acids the side chain \(R\) is neutral; for example, in alanine—\(\mathrm{CH_3}\), in valine—\(\mathrm{CH(CH_3)_2}\), in phenylalanine—\(\mathrm{CH_2 \cdot C_6H_5}\). The side chain may be basic, as, for example, in lysine—\(\mathrm{CH_2 \cdot CH_2 \cdot CH_2 \cdot CH_2 \cdot NH_2}\), or acidic, as in aspartic acid—\(\mathrm{CH_2 \cdot COOH}\). Cystine is an amino acid with two ends, \(\mathrm{-CH_2 \cdot S \cdot S \cdot CH_2-}\), and forms bridges between one chain and another. Analysis of the amino-acid residues in any protein is a very long and difficult task, and nevertheless many proteins have been analyzed with a high degree of completeness. Insulin (molecular weight 12,000), according to Kibanall and Sanger, is composed of 106 residues joined in four chains, between which six cystine disulfide bridges are thrown. Myoglobin (17,000) contains 140 residues, and hemoglobin (67,000)—about 540. If the 23 types of amino-acid residues are likened to the letters of an alphabet, then the chain of these residues in myoglobin will be like a phrase of 20–30 words, whereas hemoglobin will have to be character-
write a whole paragraph of about 10 sentences. For reasons still unclear, nature has chosen this simple method of forming protoplasm molecules in both the animal and the plant world. The complex and specific functions that molecules must perform turn out to be possible not by the formation of any especially complex organic molecules, but by stringing these relatively simple groups in various orders and in different quantities. Thus, with the help of one and the same small number of letters of the alphabet, both a Miltonic poem and a page from a telephone directory can be written.
Some characteristic features of the structure of proteins deserve special attention:
a) Almost all known amino acids enter into almost all known proteins.
b) The average molecular weight of the residues is almost the same in all proteins and lies between 110 and 120. Since the molecular weight of that part of the atoms which enters into the backbone of the chain, \(\mathrm{CO \cdot CH \cdot NH}\), is 56, it follows, therefore, that half of the molecular weight in proteins is concentrated in this backbone (“spine”) and the other half in the side groups.
c) It can be asserted almost with certainty that all amino acids occurring in nature are characterized by one and the same steric configuration around the central carbon atom, which was shown above (1). Since the central carbon atom is tetrahedrally surrounded by four different groups, all amino acids are possible either in the right-handed form or in the left-handed form. Both are mirror images of one another, just as the right hand is the reflection of the left. All amino acids belong to the left-handed forms, with the exception, of course, of glycine, in which the residue \(R\) is a second hydrogen atom, and therefore for glycine there can be no distinction between right- and left-handed forms. Evidently it is a simple matter of chance that all forms of life are connected with an identical left-handed chain. If our world were reflected in Alice’s mirror, then it would evidently function with the same success, and only some accident determined the “left-handedness” of all forms of living matter.
d) The crystalline structure of several simple amino acids, and specifically of dipeptides composed of two residues linked to one another, has already been determined by means of X-ray analysis (for example, glycine). The amino acids deciphered are characterized by two essential features. First, they rather strictly obey the rule of constancy of distances between atoms, as well as the constancy of angles between bonds, which were established in other organic compounds, in particular in those in which there are no large-
ing stresses or tensions. Hettig gave the following summary of constants for the principal constituent of the chain:
\[ \begin{array}{ccccc} & \mathrm{H} & R & \mathrm{O} & \\ & \cdot & \cdot & \cdot & \\ -\mathrm{N}-\mathrm{C}_1-\mathrm{C}_2- & & & & \\ & & \mathrm{H} & & \end{array} \]
(they are averaged over many structures):
\[ \begin{aligned} \mathrm{N}-\mathrm{C}_1 &= 1.41\,\text{\AA}; \qquad & \mathrm{C}_1-\mathrm{C}_2 &= 1.52\,\text{\AA};\\ \mathrm{C}_2-\mathrm{O} &= 1.25\,\text{\AA}; \qquad & \mathrm{C}_2-\mathrm{N} &= 1.33\,\text{\AA};\\ \text{angle } \mathrm{NC}_1\mathrm{C}_2 &= 112^\circ; \qquad & \text{angle } \mathrm{C}_1\mathrm{C}_2\mathrm{N} &= 118^\circ;\\ & & \text{angle } \mathrm{C}_2\mathrm{NC}_1 &= 118^\circ. \end{aligned} \tag{3} \]
Secondly, all protein structures contain numerous hydrogen bonds between nitrogen and oxygen, \(\mathrm{N}-\mathrm{H}-\mathrm{O}\). Apparently, precisely this bond is of very great importance for the question of the form of the structure; its length is \(2.65\,\text{\AA}\). These numerical data are very important in constructing possible models of polypeptide chains of any appreciable length.
The difficulties in applying X-ray methods of analysis to structures as complex as proteins may at first glance seem insurmountable. Generally speaking, a direct transition from an X-ray photograph to a concrete structure is possible only in especially simple cases. X-ray photographs alone are insufficient for determining the structure. They can serve for this only in combination with other sources of information. In the case of molecules of relatively small dimensions, this information consists in knowing that the elementary cell contains a known limited number of definite atoms and that these atoms are connected with one another by bonds whose length and mutual orientation are more or less well known on the basis of earlier determinations of chemically analogous substances. It is not so difficult then to test a whole series of similar configurations, in order to select from them the one that agrees best with the X-ray data. But with a protein molecule, when it, as in the case of hemoglobin, contains 8000 atoms, this method of “trial and error” exceeds human capabilities.
For a decade and a half there has already existed a method of such treatment of the direct data of X-ray analysis, which gives direct and quite exact indications of many characteristic features of the structure. This is the method of crystallographic Patterson synthesis, or the vector-diagram method. The numbers obtained by us in estimating the brightness of individual reflections
radiographs, serve as the coefficients of the Fourier series
\[ \sum_h \sum_k \sum_l I_{hkl}\cos 2\pi(hx+ky+lz). \tag{4} \]
In this line \(I_{hkl}\) is the square of the amplitude of the reflection \((hkl)\), while \(x,y,z\) are the coordinates of any point in the elementary cell of the structure being determined. If these triple sums are calculated for a significant number of points in the cell, then we obtain what is called a vector diagram. Suppose that at the point \(x_1,y_1,z_1\) there is an atom \(A\), and at the point \(x_2,y_2,z_2\) an atom \(B\). Then on the vector diagram we obtain a peak or a concentration of density at the point with coordinates \(x_1-x_2,\ y_1-y_2,\ z_1-z_2\). It is easy to see that the distance from the origin of the cell to this point is equal, in magnitude and direction, to the line joining the atoms \(A\) and \(B\) in the true structure. In other words, the vector diagram cannot tell us where in the crystal the atoms \(A\) and \(B\) lie, but it does indicate how they are situated in space relative to one another. Furthermore, the height of the corresponding peak is proportional to the product of the masses of \(A\) and \(B\). Fig. 1 illustrates this principle. We cannot go here into the details of constructing a vector diagram, although this is done quite simply if one proceeds from the principles of optical interference. What is most important is to understand the physical essence of these vector diagrams, since they play an extremely large role in modern X-ray structural analysis. Thus, we have obtained certain information about the structure, paying a high price for it, for it is easy to see that as the number of atoms increases the vector diagram rapidly becomes more complicated. In fact, if the structure contains \(n\) atoms, then \(n^2\) vectors can be drawn between them, and all of them are superposed on one another in the vector diagram.
In the analysis of organic molecules of not very large size, this difficulty is overcome by a very simple method. A heavy atom, for example bromine or iodine, is introduced into the molecule under investigation, and then the vectors corresponding to the distances between these heavy atoms stand out sharply on the vector diagram and are always easy to recognize, since they are few in number and the corresponding peaks are very strong. It is then easy to establish how the iodine or bromine atoms must be arranged in the crystal so that the vectors determined by us are obtained between them. Once the heavy atoms have been fixed, determining the positions of the light atoms becomes a much simpler task. We, to a certain extent, vividly color certain characteristic points in the molecule, much as a microscopist stains cell nuclei. But the hemoglobin molecule contains more than 8000 atoms and, consequently, in its vector diagram there must be more than seventy million mutually superposed
links. No sufficiently heavy atom can be set against such a number.
Nevertheless, precisely in the case of protein molecules there is one simplifying circumstance, namely the fact that polypeptide chains exist. If one assumes that these chains are mutually arranged in a definite order, for example like a row of parallel rods, then many vectors within any one chain will be parallel not only to one another, but also to analogous vectors in other parallel chains. On the vector diagram they will jointly appear as a long
Labels in the figure: “True cell”; “Origin”; “Vector diagram”; \(A\), \(B\), \(C\); “vectors between atoms of one chain”; “vectors between atoms of different chains.”
Fig. 1.
rod of high density, which will pass through the origin of the cell and will be parallel to all the indicated vectors in the crystal itself. It is further easy to see that the ends of all vectors between atoms of one chain and atoms of another chain also form a rod parallel to the first and at a distance from it determined by it.
Fig. 1 illustrates what has been said. The vectors between atoms within the same chains are shown by dashed lines from two atoms to all their neighbors, but it is clear that the same vectors may also be drawn from any other pair of atoms of the chain. All these vectors, by their ends, will lie on the vector diagram in the narrow region \(A\). Similar vectors from atoms of one chain to atoms of another chain are drawn with thin lines, and on the vector diagram their ends fall into two regions \(B\). Therefore, if on the vector diagram we see a high ridge extending from its origin in some definite direction, it will be very probable that in the true cell there are also
chains of atoms extending in the same direction. If, moreover, we find high ridges parallel to the ridge passing through the origin, this will be weighty evidence of how the parallel chains are mutually arranged in the crystal.
In studying crystalline hemoglobin from horse blood, Perutz found similar parallel ridges on partial vector diagrams. This led him to construct a complete three-dimensional Patterson synthesis for hemoglobin. The principle of this synthesis, in accordance with what has been said above, is very simple. For each point of the cell with coordinates \(xyz\) we calculate the triple sum (4), using all the measured intensities \(I_{hkl}\) as coefficients, and the resulting sum gives us the Patterson density at the point \(xyz\). The question is one of the laboriousness of the corresponding work. In the case of crystalline hemoglobin, Perutz measured about 28,000 reflections. To obtain a vector diagram with sufficient accuracy, it is necessary to compute triple sums at the points of the cell corresponding to division of the \(a\) axis into 120 parts and of the \(b\) and \(c\) axes into 60 parts. The total number of terms in all the sums (4) is equal to \(28{,}000 \times 120 \times 60 \times 60\), i.e. \(1.21 \times 10^{10}\).
In practice, the analysis is no longer so frightening, since there are a very large number of computational and mechanical improvements that considerably reduce this work. Even so, however, measuring 28,000 reflections and calculating all the triple sums for a single hemoglobin crystal required four years. When this work was planned, there were no guarantees that its results would justify the efforts expended, but, fortunately, these fears did not come true.
Part of the results obtained is shown in Fig. 2, where three cross sections through the three-dimensional vector diagram are given. The elementary cell of hemoglobin is monoclinic, and in the sections shown the \(a\) axis rises from left to right; the \(b\) axis is perpendicular to the plane of the diagram, while the \(c\) axis is vertical in the drawing.
The left-hand diagram gives the section through the origin of coordinates O; this origin itself lies on the left, in the middle of the \(c\) axis. The densities are expressed by means of contour lines (isogyres). In the section we clearly see a high ridge passing through the origin parallel to the \(a\) axis. The peaks on it have intervals of about \(5 \mathring{\mathrm A}\). The same diagram shows several less pronounced ridges above and below the one just indicated. In the middle diagram a section is given at a height of one-tenth of the \(b\) axis, and here we see two more high ridges parallel to those mentioned earlier. At an even higher level of the vector diagram (right-hand section) still another ridge is visible, parallel to the first (the one passing through the origin). In Fig. 3 a projection of these three ridges is given, viewed along the \(a\) axis so that they are visible by their ends. The ridge pass-
Fig. 2.
passing through the origin of coordinates, is here the center of the figure, and we can see how all the other rows are arranged in space around it.
Thus, Perutz’s results indicate that in hemoglobin there are polypeptide chains extending parallel to the \(a\)-axis of the crystal, that they are at a mutual distance of \(10.5\,\text{\AA}\), and are arranged approximately according to the law of close packing of round rods.
Scale \(1\text{ cm}=7\text{\AA}\)
Fig. 3.
Further, a row of peaks in each chain with a mutual distance of about \(5\,\text{\AA}\) indicates that this period is characteristic of the internal structure of the chains, since it is precisely this period that gives rise to the many vectors responsible for the intensity of the corresponding peaks of the vector diagram. From this follows a further important consequence. If the chains are characterized by such a mutual distance, and each chain by such an internal period of repetition, then the specific gravity of the hemoglobin crystal will tell us what mass is associated with each repeating motif. The result of the calculation shows that this mass is very close to the average mass of three amino-acid residues.
Further, if within the chain three residues are stretched to the utmost possible extent so that scheme (2) is exactly repeated, then this will require 10 Å. Consequently, our polypeptide chains must in some way be either folded or twisted. In exactly the same way the bond N—H—O, having a length of 2.65 Å, cannot serve as a bond between different chains; and it has already been indicated how important precisely these bonds are. Obviously, these bonds are located between atoms of the same chains and, one must suppose, it is precisely these bonds that hold the chain in a folded state or twist it.
These concrete indications of the vector diagram confirm certain conclusions that Astbury had made considerably earlier on the basis of the very diffuse rings obtained in X-ray photographs of wool and other forms of keratin, which is a “denatured” fibrous protein. Astbury came to the conclusion that $\alpha$-keratin is a polypeptide chain folded many times, with an internal repeat period of three residues of about 5 Å.
A legitimate question may be raised: how well grounded are our assumptions that it is precisely the vectors within the chains that will be expressed especially strongly, whereas there also exists a very large number of vectors between atoms in the chains and atoms in the side groups, as well as between atoms of one side group and atoms of another? Should it not be expected that the vector diagram represents a by no means ordered set of peaks?
The answer is that, first, as was shown above, the backbone of the protein, i.e. its chains, constitutes exactly one half of all the atoms of the protein. Secondly, if the chains really all have approximately the form of rods and if these rods are parallel to one another, then there is necessarily a significant concentration of atoms around the central axes, when viewed from the end. As an example we shall consider the simplest model of a chain with three residues and a repeat period of 5 Å, and suppose that this chain is constructed according to the law of a helix with three groups $R$ for each turn. The bond lengths given above, as well as the magnitudes of the angles, close such a helix into a cylinder with radius 1.5 Å. On the average a side chain contains about 4 atoms, not counting hydrogen; typical are leucine chains, for which $R = \mathrm{CH_2\cdot CH(CH_3)_2}$. Fig. 4 shows, in a highly idealized form, such a structure when viewed along the chains. For reasons set forth below, the mutual distance between chains is taken to be 9.5 Å horizontally and 14 Å vertically. Circles outline those atoms that belong to the backbones of the chains, 12 in number for each repeat period, since for each turn of the helix there are 3 leucine groups with four atoms in each one.
(the mutual arrangement of these atoms is not shown in the drawing, since the purpose of the diagram is to demonstrate the relative concentration of atoms). The external outlines of the leucine group correspond to the van der Waals radii characteristic of the mutual packing of organic groups not directly connected with one another by a chemical bond. The density inside the circles, in which all the atoms are connected with one another by covalent bonds, is approximately 10 times greater than in the spaces between them. It is natural to suppose that these masses of atoms, with their high concentration, will determine the maxima (peaks) on the Patterson projections, and that this will apply both to the short vectors within each chain and to those vectors which connect atoms of one chain with atoms of another.
Fig. 4.
The mutual arrangement of the chains depicted by us in Fig. 4 has been adopted on the basis of the recent determination of the structure of myoglobin carried out by Kendrew. This structure is considerably simpler than the structure of hemoglobin, since the myoglobin molecule has a molecular weight four times smaller. Patterson projections along the principal axes of the unit cell reveal chains very close to those established in hemoglobin, and with the same small repeat period within the chains, equal to \(5 \text{ Å}\). These chains lie in planes parallel to the face \(b\), and have direction \([201]\). The Patterson projection onto a plane perpendicular to these chains proved unusually simple and expressive. It shows that the arrangement of the chains is very close to that given in Fig. 4, the horizontal distances being \(9.5 \text{ Å}\); the chains are arranged in packets at distances of \(14 \text{ Å}\) along the vertical \(b\). The same projection shows that the chains are not situated exactly one above another in the direction \(b\), as shown in Fig. 4, but are slightly displaced. The fact that the side groups obviously do not fill all the space in Fig. 4 is undoubtedly due to the necessity of allowing room for the water particles associated with each molecule. For details we refer the reader to the original work, which is to appear in the near future, but we must
It should be noted here as well that the analysis of myoglobin that has been carried out compels us to regard all the conclusions presented above with greater confidence and, in particular, gives us a first approximation to the picture of the distribution of chains in this protein.
All the work so far carried out on the X-ray structural analysis of proteins still contains too large an element of conjecture. We must always be prepared, without hesitation, to abandon any proposed model, however attractive it may seem to us, if it is experimentally refuted. We are in many respects like a mountaineer striving to climb a difficult summit, who gladly makes use of any clefts and ledges on which he can secure himself. The results achieved are still very meager, but in the distance a very valuable prize is gleaming. Apparently, there are very fundamental reasons why nature chose polypeptide chains as the carriers of all forms of living matter. And we shall be able to discover this reason only when the secrets of the structure of protein molecules have been solved.