This note records a practical workflow for training an Allegro machine-learning potential from VASP molecular-dynamics data. The example is a 48-atom H/Pb/C/I/N perovskite cell, but the checks are meant to be reusable: make the DFT labels consistent, sample configurations from the target thermodynamic region, train with a conservative Allegro configuration, and validate the potential with both test-set errors and short MD diagnostics.

Background

Allegro is an E(3)-equivariant interatomic potential implemented as an extension of the NequIP framework. The model learns local, symmetry-respecting atomic representations and maps them to energies and forces. Its locality makes it attractive for larger molecular-dynamics simulations, especially when deployed through the pair_allegro LAMMPS interface.

The main lesson from this workflow is not a new hyperparameter trick. Most unstable potentials I encountered were caused by ordinary engineering mismatches: DFT labels generated with inconsistent settings, sampling trajectories that did not represent the intended temperature window, unit mismatches in optional short-range terms, and deploying a final checkpoint instead of the best validation checkpoint.

Principle

A machine-learning potential only learns the distribution it sees. For a finite-temperature MD potential, the useful dataset should therefore combine consistent DFT labels with configurations drawn from the temperature and structural region in which the model will be used. Before training, I check four quantities in the assembled extxyz file: number of frames, species counts, finite energies and forces, and the minimum pair distance distribution.

I also keep the first training pass simple. Optional pair potentials such as ZBL are useful for high-energy close-contact events, but they introduce another unit and mapping surface. For the first stable model, I train Allegro without ZBL, validate the learned potential, and only then decide whether an explicit short-range term is needed.

Installation

Allegro is distributed on PyPI as nequip-allegro. It depends on NequIP, PyTorch, e3nn, and Lightning. On a cluster, I prefer an explicit conda environment and fixed versions, then call the console scripts by absolute path inside Slurm jobs.

conda create -n allegro python=3.10
conda activate allegro

# Install a PyTorch build appropriate for the local CUDA driver first.
pip install torch

# Allegro is published as nequip-allegro on PyPI.
pip install nequip-allegro

nequip-train --help
nequip-compile --help
nequip-package --help

The environment used here contained PyTorch 2.7.1, e3nn 0.5.6, NequIP 0.11.1, and Allegro 0.6.3.

LAMMPS Interface

Allegro models are used in LAMMPS through the pair_nequip_allegro plugin. For Allegro specifically, the LAMMPS pair style is allegro. The model must first be compiled from a checkpoint or packaged model.

# TorchScript model for LAMMPS or other inference targets.
nequip-compile \
  best_epoch_0999.ckpt \
  formal_nve970_lammps.nequip.pth \
  --device cuda \
  --mode torchscript

# AOTInductor export for the Allegro LAMMPS pair style.
nequip-compile \
  best_epoch_0999.ckpt \
  formal_nve970_lammps.nequip.pt2 \
  --device cuda \
  --mode aotinductor \
  --target pair_allegro

A minimal LAMMPS fragment is:

units metal
atom_style atomic
read_data mapbi3.data

pair_style allegro
pair_coeff * * formal_nve970_lammps.nequip.pth H Pb C I N

timestep 0.00025
thermo 100
run 10000

The order after pair_coeff * * maps LAMMPS atom types to model type names. This mapping, together with the LAMMPS unit system, should be checked before interpreting any MD instability as a model failure.

Workflow

The dataset used here follows an NVT -> NVE -> stride sampling protocol.

  1. Run an equilibrated finite-temperature NVT trajectory with the same DFT settings intended for labeling.
  2. Select representative NVT snapshots as starting structures for several short NVE branches.
  3. Run each NVE branch with consistent VASP settings and a conservative timestep.
  4. Discard the initial transient part of each branch and sample every fixed number of frames.
  5. Collect all sampled frames into one ASE-readable extxyz file and audit it before training.

For this run, five NVE branches were started from late NVT snapshots. Each branch used 3000 MD steps; the first 100 steps were discarded, and every 15th frame was collected. The final training dataset contained 970 frames with 48 atoms per frame.

python3 prepare_nve_branches.py

python3 collect_nve_samples.py \
  --job-root nve5_encut600_330K_20260526 \
  --output nve5_encut600_stride15_discard100.extxyz \
  --discard 100 \
  --stride 15

The corresponding published helper scripts are:

Training Configuration

The model was trained with a single GPU. The split was 80/10/10 for train/validation/test. The cutoff radius was 6 Angstrom, the scalar feature width was 128, and the model used two Allegro layers. ZBL was disabled for this first stable potential.

The full configuration is available as formal_nve970.yaml. The essential parts are:

run: [train, test]

cutoff_radius: 6.0
chemical_symbols: [H, Pb, C, I, N]
model_type_names: ${chemical_symbols}

data:
  _target_: nequip.data.datamodule.ASEDataModule
  split_dataset:
    file_path: <DATA_ROOT>/nve5_encut600_stride15_discard100.extxyz
    train: 0.8
    val: 0.1
    test: 0.1
  transforms:
    - _target_: nequip.data.transforms.NeighborListTransform
      r_max: ${cutoff_radius}
    - _target_: nequip.data.transforms.ChemicalSpeciesToAtomTypeMapper
      chemical_symbols: ${chemical_symbols}

trainer:
  _target_: lightning.Trainer
  max_epochs: 1000
  devices: 1
  num_nodes: 1
  callbacks:
    - _target_: lightning.pytorch.callbacks.ModelCheckpoint
      monitor: val0_epoch/weighted_sum
      mode: min
      save_top_k: 5
      save_last: true
      auto_insert_metric_name: false
      filename: best_epoch_{epoch:04d}

num_scalar_features: 128

training_module:
  _target_: nequip.train.EMALightningModule
  loss:
    _target_: nequip.train.EnergyForceLoss
    per_atom_energy: true
    coeffs:
      total_energy: 1.0
      forces: 1.0
  optimizer:
    _target_: torch.optim.Adam
    lr: 0.0005
  model:
    _target_: allegro.model.AllegroModel
    compile_mode: compile
    model_dtype: float32
    r_max: ${cutoff_radius}
    l_max: 2
    num_layers: 2
    num_scalar_features: ${num_scalar_features}
    num_tensor_features: 64

The Slurm driver that trains, packages, compiles, and launches the first MD smoke test is linked here: run_train_then_md.slurm.

Packaging and ASE Validation

After training, I use the best validation checkpoint rather than the final checkpoint by habit. The selected checkpoint was best_epoch_0999.ckpt.

nequip-package build \
  best_epoch_0999.ckpt \
  formal_nve970.nequip.zip

nequip-compile \
  best_epoch_0999.ckpt \
  formal_nve970_ase_cuda.nequip.pth \
  --device cuda \
  --mode torchscript \
  --target ase \
  --data-path nve5_encut600_stride15_discard100.extxyz \
  --num-frames static \
  --num-nodes static

The ASE calculator is loaded through NequIPCalculator.from_compiled_model. The scripts used for the NVT and NVE diagnostics are linked here:

Result

The formal training job completed normally after 1000 epochs. The test-set metrics were:

  • forces_mae = 0.0188637 eV/Angstrom
  • per_atom_energy_mae = 0.000543379 eV/atom
  • total_energy_mae = 0.0260822 eV
  • weighted_sum = 0.00970355

I then ran short dynamics tests. Weakly coupled Langevin NVT runs were stable but hotter than the 330 K target, so I did not treat them as production MD. A conservative NVE check was more diagnostic: 0.25 fs timestep, 2 ps total time, 48 atoms, and initial temperature 330 K.

NVE total energy and temperature curves
NVE diagnostic for the compiled Allegro model. Top: total energy in eV, with the red axis showing drift in meV per atom relative to the initial total energy. Bottom: instantaneous temperature in K; the dashed line is 330 K. Over 2 ps, the total-energy drift is 0.000218 eV, or 0.00455 meV per atom, while the final temperature is 333.3 K. This indicates stable short-time integration for the learned potential under a conservative timestep.

Practical Lessons

  • Check the actual DFT output used for labels, not only the input templates. The training set should not mix label levels unless this is intentional and documented.
  • Sample from the intended thermodynamic region. Structures from a nominal trajectory are not enough if velocities, temperatures, or equilibration windows are inconsistent.
  • Audit the final extxyz file before training: frame count, species count, finite energies and forces, cell/PBC fields, stress fields, and close contacts.
  • Start with a no-ZBL model unless the dataset and deployment units have been checked carefully. If ZBL is added later for VASP/LAMMPS metal units, the unit setting should be consistent with eV and Angstrom.
  • Deploy the best validation checkpoint, not automatically the last checkpoint.
  • Use NVE energy conservation as an early MD diagnostic. Thermostat behavior can obscure whether the learned PES itself is stable.
  • For LAMMPS, check both the pair_coeff atom-type mapping and the unit system before debugging the neural network.

References