Encoding structural heterogeneity in mmCIF
A structure is a model of many copies of a molecule, and those copies are not identical. A coordinate file records how often each alternate appears, but not which alternates appear together. This page states what the format encodes today, where that becomes underspecified, and an extension that records the missing information.
How the PDB encodes heterogeneity today
Several parts of the PDBx/mmCIF dictionary already record that the copies of a molecule differ from one another, some of them historical. The extension proposed here is one more such mechanism, and its addition is motivated by the shortcomings of the categories listed below. Of these, the two this document is about are _atom_site.label_alt_id and _atom_site.occupancy and their interplay; the others are either outside its scope or de facto obsolete.
- Alternate conformations_atom_site.label_alt_id · _atom_site.occupancy
- Per atom: which alternate it belongs to, and the fraction of copies in which it is present. The everyday mechanism.
- Sequence microheterogeneity_entity_poly_seq.hetero
- Flags a position where the polymer itself is not one sequence — more than one monomer is modelled at one position.
- Grouping of alternates_atom_sites_alt · _atom_sites_alt_ens · _atom_sites_alt_gen
- Names altloc letters and collects them into ensembles. Present in the dictionary; effectively unused in the archive.
- Multiple models_atom_site.pdbx_PDB_model_num
- A whole further copy of the coordinates. NMR ensembles and ensemble refinement use this: each model is one complete state.
- Alternate-specific bonds_struct_conn
- Each bond partner is named by a full atom key including the altloc letter, so a bond can belong to one alternate and not another.
- Displacement, not discrete states_atom_site.B_iso_or_equiv · _atom_site_anisotrop · _pdbx_refine_tls
- How much an atom moves about its position, and how groups of atoms move together — not which discrete alternatives they occupy.
What mmCIF encodes today
The letter and the occupancy
A coordinate file describes an average over many copies of a molecule. (An X-ray structure is an average because diffraction measures every copy in the crystal at once, but what follows is about what the file can write down, not about how the data were collected.)
A few small categories declare what the molecule is, and one large category, _atom_site, holds the coordinates and points back at them through shared keys. Below is a complete, valid file for three residues of crambin, with no heterogeneity in it at all — every atom is fully present, at occupancy 1.00 and with no alternate-location letter.
Heterogeneity enters through two per-atom columns: _atom_site.label_alt_id says which alternate an atom belongs to, and _atom_site.occupancy says in what fraction of the copies it is present. Here one arginine side chain is modelled in two positions, A at 0.67 and B at 0.33, summing to one within the residue. Its backbone carries no letter at all: it is single-conformer and shared by both alternatives.
The mechanism above works because everything it relates sits inside one residue. It stops working as soon as a relationship spans more than one.
The file below is the calcium site of calmodulin, exactly as deposited. Two stretches of the chain run through it, and each is modelled in several alternates. Residues 20–24 have two, lettered B and D. Residues 25–31 have four, lettered A, B, C and D. Each set sums to 1.00, so each is one occupancy group in the ordinary sense. Both stretches coordinate the same calcium ion, so they are physically coupled.
Nothing in the file says so, and the letters do not line up. Select a letter under the viewer: B and D each light up in both stretches, while A and C exist only in the second one. Nothing states whether “B” at residue 20 is the same physical state as “B” at residue 26 — and a reader who assumes that matching letters mean matching states will be wrong. This is the file as it exists today: _atom_site and nothing else.
Both columns describe one thing at a time: how often it appears, on its own. Neither says which alternates appear in the same copy. Take a structure with two nearby sites. The top site holds either X or Y; the bottom site holds either P or Q; each of the four is modelled at occupancy 0.50, present in half the copies of the crystal. Three physically different crystals then produce a byte-for-byte identical file.
| P | Q | occ | |
|---|---|---|---|
| X | 25 | 25 | 50 |
| Y | 25 | 25 | 50 |
| occ | 50 | 50 |
| P | Q | occ | |
|---|---|---|---|
| X | 50 | 0 | 50 |
| Y | 0 | 50 | 50 |
| occ | 50 | 50 |
| P | Q | occ | |
|---|---|---|---|
| X | 0 | 50 | 50 |
| Y | 50 | 0 | 50 |
| occ | 50 | 50 |
Each table counts, out of 100 copies, how often a pair of occupants is found in the same copy — the four interior cells are the joint distribution. The numbers along the edges are the row and column sums of that interior: the marginals. A marginal is precisely what one occupant’s occupancy is — how often it appears at all, added up over whatever the other site happens to be doing — so all three crystals have the same four marginals, and this is everything the deposited file carries about the two sites:
| atoms of | _atom_site.label_alt_id | _atom_site.occupancy |
|---|---|---|
| top site (X) | A | 0.50 |
| top site (Y) | B | 0.50 |
| bottom site (P) | A | 0.50 |
| bottom site (Q) | B | 0.50 |
Four rows of _atom_site.occupancy, one per occupant, all 0.50, and the letters restart at A in each site because a letter is scoped to the residue it sits in — the same overloaded-letter problem the deposited 5E1N file has above. The interiors are completely different, and there is no column in which an interior cell could be written. In the first crystal the two sites are filled independently; in the second, X is only ever found with P; in the third, only ever with Q.
This is the central problem. A structure determination measures an aggregate — a density averaged over every copy in the crystal — and an average is a one-body quantity: it carries each site’s marginal and discards which occupants shared a copy. Recovering the populations that co-occur is what turns an average back into an ensemble, and it is what a heterogeneity state is: one concrete combination of alternates, together with the fraction of copies that are in it. Without them a file lists parts; with them it lists whole molecules and how many of each there are. The extension’s central category is a way of writing one interior cell per row.
▶The ensemble underneath, and where these numbers would come from
What is being averaged is, physically, a Boltzmann-weighted ensemble of conformations: a continuum of states with populations set by their free energies. The joint occupancies here are that ensemble coarse-grained onto whichever handful of alternates the model happens to name — a discrete, deliberately coarse summary of it, and the coarsest one that still says which parts of the molecule move together.
Defining, representing, comparing and validating such ensembles — and the gap between what current experiments deliver and what a ground-truth ensemble would need — is the subject of Wankowicz & Bonomi, From Possibility to Precision in Macromolecular Ensemble Prediction (arXiv:2505.01919). Part IV takes up the same question from this format’s side: a schema that can hold a joint distribution does not by itself produce one.
Where the altlocs and occupancy become insufficient
Seven cases, in increasing order of difficulty. The first is handled completely by what exists today; each one after it asks for something the letter and the occupancy cannot supply. The two sites above are case 5, in the same shape and with real occupants. The last is expressible here too, but at a price worth naming, so it points at Part IV rather than at a case.
- 1A side chain modelled in two rotamers, entirely inside one residue. The letter relates only things that are already together, and the occupancies sum to one there. This one works today. 1EJG_rotamer
- 2Two stretches of one chain that move together, several residues apart — and that use different sets of letters, so there is no letter to match on and nothing that says they are coupled at all. 5E1N_ca_site
- 3A network boundary that falls inside a residue: the amide hydrogen follows the previous residue while the side chain is its own rotamer, both under the same letter. No key built from chain, residue and letter can separate them. 5E1N_gln8_split
- 4A ligand present in 22% of copies, with two poses inside that 22%, so their occupancies sum to 0.22 rather than to 1. And the 78% where it is absent has no atoms at all — nothing for a per-atom mechanism to label. 7HHS_apo_bound
- 5Two adjacent pockets whose fillings depend on each other — the interlude’s two sites, with real occupants. Four occupancies are deposited, and they are the same four whether the pockets are paired one way, paired the other, or independent: three different pieces of chemistry, one file. constructed_two_pocket
- 6Two alternates in different parts of the structure that cannot co-occur — they clash at 2.14 Å — with nothing in their grouping to forbid it, because they are not alternatives of each other. 5E1N_arg74_clash
- 7Two independent fragments and a water ordered only when both are bound. Its occupancy is a product of theirs, and no sum of occupancies equals a product. Expressible, by writing the four combinations out — what it costs is that the two fragments must then share a bundle, so their independence is implied rather than stated. the price of enumerating
The proposal
The proposed categories
The extension is five optional categories, and they do three separate jobs: name the alternate states, say which of them exclude one another, and record which of them occur together, in what proportion. They point into _atom_site through the keys it already has and add no column to it, so a program that does not know them reads exactly the coordinates it reads today. Each is applied to a real case in Part III.
Names a network: a set of atoms that together constitute one alternate state. Each row selects atoms out of _atom_site by chain, an inclusive residue range and an altloc letter, optionally narrowed to a single named atom. One network is usually several rows sharing an alt_group_id, so its membership need not be contiguous — that is what lets one name span two stretches of a chain, or claim one atom out of a residue.
| id | required | ||
| alt_group_id | required | ||
| auth_asym_id | required | ||
| auth_seq_id_start | required | ||
| auth_seq_id_end | required | ||
| label_alt_id | required | ||
| label_atom_id | required |
One row per network, naming the coexistence group it shares with its alternatives — the mutually-exclusive set, which is what a crystallographer calls an occupancy group. At most one member of a group is present in any copy, so their occupancies sum to at most 1; a sum below 1 leaves a fraction of copies with none of them. It says nothing about networks at different sites — that is the joint, and it lives in the next category.
| alt_group_id | required | ||
| coexistence_group_id | required | ||
| state_kind | required |
One row per combination that actually occurs. Its occupancy is the joint occupancy of the whole combination — a network’s own occupancy, the number _atom_site carries, is the sum over every state containing it. bundle_id is the unit of correlation: states sharing one enumerate a joint distribution and sum to 1, and networks in different bundles are independent.
| id | required | ||
| bundle_id | required | ||
| occupancy | required | ||
| provenance | required | ||
| details | required |
One row per network present in a state. A state’s members are spread over several rows because an mmCIF cell holds a single value, and a network absent in a state simply contributes no row — that omission is also how a state names something with no atoms to label, such as the fraction of copies in which a ligand is not there at all.
| state_id | required | ||
| alt_group_id | required |
A sparse NOT-only list: each row says two networks may not co-occur. For a zero in the joint where there is no distribution to write — two networks that are otherwise independent, no populations known, one impossible pairing. Not needed inside a coexistence group, whose members already exclude each other, and not needed inside a bundle, which already forbids by omission everything it does not list. Absent from most files; there is exactly one such row in the prototype annotation of 5E1N.
| id | required | ||
| rule | required | ||
| alt_group_id | required | ||
| alt_group_ids | required |
First, alternatives that share a coexistence_group_id are mutually exclusive: at most one of them is present in any one copy. Second, the always-present single-conformer part of the structure is the implicit network base: occupancy 1, carrying no membership rows because its atoms are simply everything that no other network claims. Only alternates are ever listed. Third, and this is the one that is new: a state lists the networks that are present, and a bundle’s states are exhaustive. A site left empty in some state is spelled by leaving it out, and any combination a bundle does not list has joint occupancy zero. That is what lets a state name a thing with no atoms in it — the 78% of copies in which a ligand is simply not there — and it is why a bundle needs no exclusion rows of its own.
The cases
Explicit state grouping
The deposited file from Part I has the letter B in two stretches of one chain, and nothing in it says whether those two Bs are the same physical state or two unrelated choices that happen to share a letter. A reader has to guess from the occupancies. Two small side tables remove the guess.
_pdbx_alt_groups gives the state a name and says which atoms carry it — by chain, residue range and altloc letter. Because a name can be repeated across rows, one name can span a whole stretch of the chain, which a letter cannot do: a letter is scoped to the residue it sits in. _pdbx_heterogeneity_hierarchy then puts each name in a coexistence group, the set of states that exclude one another — which is what a crystallographer already calls an occupancy group.
Below, the two stretches become six named networks in two such groups. The correlation that was previously only implied by the numbers is now written down, and the letter is relieved of a job it was never able to do. No atom changes; the two tables are additive.
Membership below the residue (specifying alt groups with atom-precision)
A residue-range key is enough almost everywhere, but not everywhere. A network boundary can fall inside a residue, and then the key fails.
In the file below, Gln8 carries two independent choices at once. Its backbone amide hydrogen follows the conformation of the preceding residues — the N–H points back at the previous carbonyl — so it belongs with residues 6–7. Its side chain is its own rotamer, in a separate occupancy group with different occupancies. And both choices use the same letters: the atoms of Gln8 with label_alt_id = A belong to two different networks. A key of (chain, residue range, altloc) cannot separate them; both are “chain A, residue 8, alternate A”.
The optional _pdbx_alt_groups.label_atom_id is the escape hatch: when a membership row names an atom, it claims exactly that atom and nothing else. A membership table with it is precisely as expressive as a per-atom state label, and it reuses fields that are already present on _atom_site.
▶Why not simply add a column to _atom_site?
The same information can be carried by adding one column to the coordinate table — a per-atom token naming the group each atom belongs to. A prototype annotation of 5E1N does exactly that, and it separates this case effortlessly, because every atom names its own group and nothing has to be inferred from a key. The case is not rare: in that file there are 28 (chain, residue, altloc) triples whose atoms fall in two different groups.
_pdbx_alt_groups.label_atom_id reaches all 28 without touching _atom_site, at a cost of 28 extra rows in a side table. That is the reason the proposal takes this route: a new column on the coordinate table is read by every program that parses coordinates, on every file, including the overwhelming majority that carry no heterogeneity annotation at all — whereas a side table is read only by the programs that want it.
Which alternates go together
Everything so far names alternates and says which of them exclude one another. Neither says which alternates are found in the same copy, and that is the quantity the interlude at the end of Part I was about: _atom_site.occupancy is a list of marginals, one per network, and the same list is produced by physically different structures.
_pdbx_het_state writes the missing quantity down directly, one row per combination that occurs. A row’s occupancy is the joint occupancy — the fraction of copies in exactly that combination, not the occupancy of any member — and the networks making it up are listed in _pdbx_het_state_members, one per row, because an mmCIF cell holds a single value.
The bundle: how far a state reaches
A state is a joint assignment, and the obvious worry about joint assignments is that they grow: if a state had to name every alternate in the structure, a protein with forty partial waters would need a table of astronomical size. bundle_id is what stops that. Networks whose occupancies are entangled — where knowing one changes the distribution of another — share a bundle, and that bundle’s states enumerate their joint distribution and nothing else. Networks in different bundles are independent: their occupancies multiply, and they are never written out together.
The second case is the one the marginals cannot decide, and it is the pair of pockets from the interlude, with real occupants. A phenol ligand or a glycol sits in a top pocket at 0.50 each; a second glycol sits in a bottom pocket at 0.30, a third at 0.20, or the pocket is empty. Those four numbers are everything a deposited file carries today, and they are consistent with any number of different pairings. Six states say which pairing it is.
A combination that cannot occur
A bundle is the right tool when you know the populations. We can imagine a case where you know only that one combination is impossible — no proportions, no correlation, just a pairing that cannot happen. That is a single zero in the joint distribution, and building a whole bundle around it would mean inventing the other numbers.
Below, an arginine side chain is modelled in three alternates and a nearby water at partial occupancy. They are not alternatives of each other — one is a rotamer, the other a solvent site — so no coexistence group relates them, and nothing says the water is more likely with one rotamer than another. But in alternate B the guanidinium nitrogen lands 2.14 Å from the water: not a hydrogen bond, a clash. One NOT row records exactly that, and claims nothing else.
The two mechanisms do not overlap, and a file should never use both for the same combination: a bundle already forbids by omission everything it does not list, so a NOT among a bundle’s own members would be redundant at best and contradictory at worst. Correlated sites with populations you know get a bundle; a stray impossible pairing between otherwise independent sites gets a row here.
Open problems
Does this blow up?
Within a bundle, yes, in principle; across a structure, no. The cost of the state table is exponential in the number of mutually correlated switches inside a single bundle — independent sites are separate bundles, and separate bundles add rather than multiply. A structure with several hundred partial waters and alternate side chains that do not talk to each other is several hundred bundles of one or two states apiece: linear in the number of partial sites, not two to the power of anything.
A genuine blow-up therefore needs one bundle of fifteen or twenty mutually entangled switches — and it seems impossible to refine, measure or otherwise determine a fifteen-way joint distribution from one experiment, so I don’t think this is a problem in practice. Something like BinaryCIF would also compress these rather repetitive tables away quite nicely. What remains open is whether some class of structure has a bundle much bigger than anything we have looked at.
Where the joint actually comes from
The schema holds any joint distribution perfectly. The data, in general, do not. The average density is a one-body quantity: at each site it carries that site’s marginal occupancy and nothing more. The joint — which bottom-pocket occupant goes with which top-pocket occupant — is a correlation between sites, a many-body quantity, and standard refinement against Bragg data does not produce it afaik, let alone CryoEM data. Two crystals with the same marginals and completely different correlations give the same average density (this is the point the interlude makes with three tables).