Skip to content

Binary Files

Several of TeraChem's scratch and restart files are written as raw binary rather than text. This page documents their on-disk layout so you can read them from your own tools (for example, loading molecular-orbital coefficients into Python for analysis).

Conventions

Unless a section below says otherwise, TeraChem's binary files follow these rules:

Property Value
Number format IEEE-754 64-bit floating point (double, 8 bytes per value)
Byte order Native — little-endian on all supported (x86-64) platforms
Header None — the file is exactly the raw array of values
Matrix storage Row-major (C order)

Byte order

These files store numbers in the machine's native byte order. On every platform TeraChem currently supports that is little-endian, and the readers below assume little-endian. The files are not portable to big-endian hardware without byte-swapping.

Because the matrix files carry no header, the matrix dimension is recovered from the file size:

\[ n_\text{ao} = \sqrt{\text{file size in bytes} / 8}. \]

Orbital coefficients & AO overlap (c0, ca0, cb0, sao)

These files all share the same layout: an \(n_\text{ao} \times n_\text{ao}\) matrix of doubles (\(n_\text{ao}^2\) values, \(n_\text{ao}^2 \times 8\) bytes), written row-major with no header.

File Written when Contents
c0 Restricted / closed-shell SCF MO coefficient matrix
ca0 Unrestricted SCF α (alpha) MO coefficients
cb0 Unrestricted SCF β (beta) MO coefficients
sao save_sao yes AO overlap matrix \(S\) (symmetric)

The files live in the scratch directory (scrdir, default ./scr). Variants with a method suffix (c0.casscf, c0.cisno) and per-frame MD files (c0.0, c0.1, …, and ca0.<n>/cb0.<n>) use the identical byte layout.

Index convention

On disk the coefficient matrix is laid out in the conventional orientation:

  • rows are atomic orbitals (AOs), and
  • columns are molecular orbitals (MOs).

The coefficient of AO \(a\) in MO \(p\) is at flat index a*nao + p. Equivalently, reshaping the file to an (nao, nao) array in row-major order gives C[a, p], and column p is the AO-coefficient vector of MO p.

For sao the stored matrix is the symmetric AO overlap \(S_{ab}\), so the row/column (transpose) distinction does not matter.

AO ordering

The AO axis follows TeraChem's internal basis-function ordering (the same ordering used for Molden export). See the basis-set pages for how basis functions are organized; a full per-shell ordering specification is beyond the scope of this page.

Reading the file

import numpy as np

def read_c0(path):
    """Read a TeraChem c0/ca0/cb0/sao file into an (nao, nao) array.

    C[a, p] = coefficient of AO a in MO p
    (for sao, the symmetric AO overlap S[a, b]).
    """
    data = np.fromfile(path, dtype='<f8')       # little-endian float64
    nao = int(round(len(data) ** 0.5))
    assert nao * nao == len(data), "file is not a square number of doubles"
    return data.reshape(nao, nao)

C = read_c0('scr/c0')
mo0 = C[:, 0]          # AO coefficients of MO 0
#include <stdio.h>
#include <stdlib.h>
#include <math.h>

/* Read a TeraChem c0/ca0/cb0/sao file: nao*nao little-endian doubles,
   row-major, no header.  C[a*nao + p] = coefficient of AO a in MO p. */
double *read_c0(const char *path, int *nao_out) {
    FILE *fp = fopen(path, "rb");
    if (!fp) { perror(path); return NULL; }
    fseek(fp, 0, SEEK_END);
    long bytes = ftell(fp);
    rewind(fp);
    long count = bytes / (long)sizeof(double);   /* = nao*nao */
    int nao = (int)lround(sqrt((double)count));
    double *C = malloc((size_t)nao * nao * sizeof(double));
    fread(C, sizeof(double), (size_t)nao * nao, fp);
    fclose(fp);
    *nao_out = nao;
    return C;   /* index as C[a*nao + p] */
}

Worked example

This is a constructed \(2 \times 2\) example (not the output of a real calculation), chosen so the bytes are easy to follow. Rows are AOs, columns are MOs:

\[ C = \begin{pmatrix} 0.5 & 0.7 \\ 0.5 & -0.7 \end{pmatrix} \]

With \(n_\text{ao} = 2\), the file is \(2 \times 2 \times 8 = 32\) bytes. Its raw contents are:

00000000: 0000 0000 0000 e03f 6666 6666 6666 e63f  .......?ffffff.?
00000010: 0000 0000 0000 e03f 6666 6666 6666 e6bf  .......?ffffff..

Reading those bytes back with the Python reader above reproduces the matrix:

nao = 2
[[ 0.5  0.7]
 [ 0.5 -0.7]]

The first double (00 00 00 00 00 00 e0 3f, little-endian) is 0.5 = C[0, 0]; the fourth (66 66 66 66 66 66 e6 bf) is -0.7 = C[1, 1], confirming the row-major, AO-row/MO-column layout.

State density matrices (opdm_AO_<Mult>_<n>)

When the effective-number-of-unpaired-electrons analysis is enabled with cas_enue yes (see CASSCF → Effective number of unpaired electrons), TeraChem writes the AO-basis one-particle reduced density matrix of each analyzed state to its own file. The analysis is available for CASSCF, CASCI, and FOMO-CASCI.

File Written when Contents
opdm_AO_<Mult>_<n>.bin cas_enue yes (CASSCF / CASCI / FOMO-CASCI) One-particle density matrix \(D_{ab}\) of state <n> with multiplicity <Mult>

<Mult> is the spin-multiplicity label (Singlet, Doublet, Triplet, …) and <n> is the 1-based state index within that multiplicity — for example opdm_AO_Singlet_1.bin for the lowest singlet and opdm_AO_Triplet_2.bin for the second triplet. One file is written per analyzed state, so a state-averaged job produces several.

The layout is identical to the orbital-coefficient files: an \(n_\text{ao} \times n_\text{ao}\) matrix of doubles, row-major, with no header, so the same reader recovers it (the result is the density matrix \(D_{ab}\) rather than a coefficient matrix). The files are written to the scratch directory (scrdir, default ./scr).

Hessian (Hessian.bin)

A vibrational-frequency calculation writes the Cartesian Hessian (matrix of energy second derivatives) to Hessian.bin in the scratch directory (scrdir, default ./scr). Unlike the matrix files above, this file has a header describing the geometry it was computed at, so it does not follow the "no header" convention.

Hessian.bin is produced whenever a numerical Hessian is built — run frequencies, run initcond, or a TCPB/IMD Hessian request. TeraChem always computes the Hessian numerically, by finite differences of analytic gradients; there is no analytic-Hessian path. The finite-difference stencil and step are controlled by nderivpoints (3, 5, or 7; default 3) and displacement (in Bohr, default 0.005).

The on-disk layout, in order, is:

Field Type Count Notes
natoms int (4 B) 1 Number of atoms N used for the Hessian (QM/MM link atoms excluded)
nderivpoints int (4 B) 1 Finite-difference stencil: 3, 5, or 7
displacement double (8 B) 1 Finite-difference step, in Bohr
geometry double (8 B) 4 × N Reference geometry: 4 values per atom, {x, y, z, Z} — Cartesian position in Bohr and the atomic number Z
Hessian double (8 B) (3N)² The 3N × 3N Cartesian Hessian, row-major
dipole derivatives double (8 B) 3 × 3N \(\partial\mu_{x,y,z}/\partial R\) for each of the 3N coordinates (atomic units)

Key conventions:

  • Units: the Hessian is in Hartree/Bohr² and is not mass-weighted (mass weighting is applied later, in the frequency analysis). The geometry and step are in Bohr.
  • Ordering: the Hessian is the full (both triangles) symmetric matrix in row-major order, indexed as H[3*i + a][3*j + b] for atoms \(i,j\) and Cartesian components \(a,b \in \{x,y,z\}\).
  • Dipole-derivative block: normally 3 rows (one per dipole component) of length 3N. With polarizability yes this block is enlarged to 12 × 3N — the first 3 rows are the dipole derivatives and the remaining 9 are polarizability derivatives.
  • Endianness/precision: native little-endian IEEE-754, as for all files on this page.

Resumable companion file Hessian.restart

A numerical Hessian also writes Hessian.restart in the scratch directory — the same header followed by one record per completed displacement (the displacement index plus that displacement's gradient). It lets an interrupted frequency job resume where it left off, and is not removed on completion. The human-readable results (frequencies and normal modes) are written separately to Frequencies.dat and the Molden file.

Other binary files

  • CI vectors (CASSCF / CI): the configuration-interaction coefficient vectors are written as flat 1-D arrays of doubles (one value per determinant), with no header. Detailed documentation is planned.
  • Internal formats: TeraChem also writes developer-internal binary files — a tagged checkpoint file (used by the validation scripts) and a composite restart file (moinfo) carrying an integer header plus matrices. These are implementation details and are not documented here.