Binary Files
Several of TeraChem's scratch and restart files are written as raw binary rather than text. This page documents their on-disk layout so you can read them from your own tools (for example, loading molecular-orbital coefficients into Python for analysis).
Conventions
Unless a section below says otherwise, TeraChem's binary files follow these rules:
| Property | Value |
|---|---|
| Number format | IEEE-754 64-bit floating point (double, 8 bytes per value) |
| Byte order | Native — little-endian on all supported (x86-64) platforms |
| Header | None — the file is exactly the raw array of values |
| Matrix storage | Row-major (C order) |
Byte order
These files store numbers in the machine's native byte order. On every platform TeraChem currently supports that is little-endian, and the readers below assume little-endian. The files are not portable to big-endian hardware without byte-swapping.
Because the matrix files carry no header, the matrix dimension is recovered from the file size:
Orbital coefficients & AO overlap (c0, ca0, cb0, sao)
These files all share the same layout: an \(n_\text{ao} \times n_\text{ao}\) matrix of
doubles (\(n_\text{ao}^2\) values, \(n_\text{ao}^2 \times 8\) bytes), written row-major
with no header.
| File | Written when | Contents |
|---|---|---|
c0 |
Restricted / closed-shell SCF | MO coefficient matrix |
ca0 |
Unrestricted SCF | α (alpha) MO coefficients |
cb0 |
Unrestricted SCF | β (beta) MO coefficients |
sao |
save_sao yes |
AO overlap matrix \(S\) (symmetric) |
The files live in the scratch directory (scrdir, default ./scr). Variants with
a method suffix (c0.casscf, c0.cisno) and per-frame MD files (c0.0, c0.1,
…, and ca0.<n>/cb0.<n>) use the identical byte layout.
Index convention
On disk the coefficient matrix is laid out in the conventional orientation:
- rows are atomic orbitals (AOs), and
- columns are molecular orbitals (MOs).
The coefficient of AO \(a\) in MO \(p\) is at flat index a*nao + p. Equivalently,
reshaping the file to an (nao, nao) array in row-major order gives C[a, p], and
column p is the AO-coefficient vector of MO p.
For sao the stored matrix is the symmetric AO overlap \(S_{ab}\), so the
row/column (transpose) distinction does not matter.
AO ordering
The AO axis follows TeraChem's internal basis-function ordering (the same ordering used for Molden export). See the basis-set pages for how basis functions are organized; a full per-shell ordering specification is beyond the scope of this page.
Reading the file
import numpy as np
def read_c0(path):
"""Read a TeraChem c0/ca0/cb0/sao file into an (nao, nao) array.
C[a, p] = coefficient of AO a in MO p
(for sao, the symmetric AO overlap S[a, b]).
"""
data = np.fromfile(path, dtype='<f8') # little-endian float64
nao = int(round(len(data) ** 0.5))
assert nao * nao == len(data), "file is not a square number of doubles"
return data.reshape(nao, nao)
C = read_c0('scr/c0')
mo0 = C[:, 0] # AO coefficients of MO 0
#include <stdio.h>
#include <stdlib.h>
#include <math.h>
/* Read a TeraChem c0/ca0/cb0/sao file: nao*nao little-endian doubles,
row-major, no header. C[a*nao + p] = coefficient of AO a in MO p. */
double *read_c0(const char *path, int *nao_out) {
FILE *fp = fopen(path, "rb");
if (!fp) { perror(path); return NULL; }
fseek(fp, 0, SEEK_END);
long bytes = ftell(fp);
rewind(fp);
long count = bytes / (long)sizeof(double); /* = nao*nao */
int nao = (int)lround(sqrt((double)count));
double *C = malloc((size_t)nao * nao * sizeof(double));
fread(C, sizeof(double), (size_t)nao * nao, fp);
fclose(fp);
*nao_out = nao;
return C; /* index as C[a*nao + p] */
}
Worked example
This is a constructed \(2 \times 2\) example (not the output of a real calculation), chosen so the bytes are easy to follow. Rows are AOs, columns are MOs:
With \(n_\text{ao} = 2\), the file is \(2 \times 2 \times 8 = 32\) bytes. Its raw contents are:
00000000: 0000 0000 0000 e03f 6666 6666 6666 e63f .......?ffffff.?
00000010: 0000 0000 0000 e03f 6666 6666 6666 e6bf .......?ffffff..
Reading those bytes back with the Python reader above reproduces the matrix:
The first double (00 00 00 00 00 00 e0 3f, little-endian) is 0.5 = C[0, 0];
the fourth (66 66 66 66 66 66 e6 bf) is -0.7 = C[1, 1], confirming the
row-major, AO-row/MO-column layout.
State density matrices (opdm_AO_<Mult>_<n>)
When the effective-number-of-unpaired-electrons analysis is enabled with
cas_enue yes (see
CASSCF → Effective number of unpaired electrons),
TeraChem writes the AO-basis one-particle reduced density matrix of each
analyzed state to its own file. The analysis is available for CASSCF,
CASCI, and FOMO-CASCI.
| File | Written when | Contents |
|---|---|---|
opdm_AO_<Mult>_<n>.bin |
cas_enue yes (CASSCF / CASCI / FOMO-CASCI) |
One-particle density matrix \(D_{ab}\) of state <n> with multiplicity <Mult> |
<Mult> is the spin-multiplicity label (Singlet, Doublet, Triplet, …) and
<n> is the 1-based state index within that multiplicity — for example
opdm_AO_Singlet_1.bin for the lowest singlet and opdm_AO_Triplet_2.bin for the
second triplet. One file is written per analyzed state, so a state-averaged job
produces several.
The layout is identical to the
orbital-coefficient files: an
\(n_\text{ao} \times n_\text{ao}\) matrix of doubles, row-major, with no header, so
the same reader recovers it (the result is the density matrix \(D_{ab}\) rather than
a coefficient matrix). The files are written to the scratch directory
(scrdir, default ./scr).
Hessian (Hessian.bin)
A vibrational-frequency calculation writes the Cartesian Hessian (matrix of
energy second derivatives) to Hessian.bin in the scratch directory
(scrdir, default ./scr). Unlike the matrix files above, this file has a
header describing the geometry it was computed at, so it does not follow the
"no header" convention.
Hessian.bin is produced whenever a numerical Hessian is built —
run frequencies, run initcond, or a TCPB/IMD Hessian request. TeraChem always
computes the Hessian numerically, by finite differences of analytic gradients;
there is no analytic-Hessian path. The finite-difference stencil and step are
controlled by nderivpoints (3, 5, or 7; default 3) and displacement (in Bohr,
default 0.005).
The on-disk layout, in order, is:
| Field | Type | Count | Notes |
|---|---|---|---|
natoms |
int (4 B) |
1 | Number of atoms N used for the Hessian (QM/MM link atoms excluded) |
nderivpoints |
int (4 B) |
1 | Finite-difference stencil: 3, 5, or 7 |
displacement |
double (8 B) |
1 | Finite-difference step, in Bohr |
| geometry | double (8 B) |
4 × N |
Reference geometry: 4 values per atom, {x, y, z, Z} — Cartesian position in Bohr and the atomic number Z |
| Hessian | double (8 B) |
(3N)² |
The 3N × 3N Cartesian Hessian, row-major |
| dipole derivatives | double (8 B) |
3 × 3N |
\(\partial\mu_{x,y,z}/\partial R\) for each of the 3N coordinates (atomic units) |
Key conventions:
- Units: the Hessian is in Hartree/Bohr² and is not mass-weighted (mass weighting is applied later, in the frequency analysis). The geometry and step are in Bohr.
- Ordering: the Hessian is the full (both triangles) symmetric matrix in
row-major order, indexed as
H[3*i + a][3*j + b]for atoms \(i,j\) and Cartesian components \(a,b \in \{x,y,z\}\). - Dipole-derivative block: normally 3 rows (one per dipole component) of length
3N. With
polarizability yesthis block is enlarged to 12 × 3N — the first 3 rows are the dipole derivatives and the remaining 9 are polarizability derivatives. - Endianness/precision: native little-endian IEEE-754, as for all files on this page.
Resumable companion file Hessian.restart
A numerical Hessian also writes Hessian.restart in the scratch directory —
the same header followed by one record per completed displacement (the
displacement index plus that displacement's gradient). It lets an interrupted
frequency job resume where it left off, and is not removed on completion. The
human-readable results (frequencies and normal modes) are written separately to
Frequencies.dat and the Molden file.
Other binary files
- CI vectors (CASSCF / CI): the configuration-interaction coefficient vectors
are written as flat 1-D arrays of
doubles (one value per determinant), with no header. Detailed documentation is planned. - Internal formats: TeraChem also writes developer-internal binary files — a
tagged checkpoint file (used by the validation scripts) and a composite restart
file (
moinfo) carrying an integer header plus matrices. These are implementation details and are not documented here.