Leverage QCArchive data for creating QC ML models. AP-Net2 has been re-implemented in PyTorch with newer versions to come.
Code re-implemented from TensorFlow version located here
To install the package, run the following command:
conda env create -f environment.yml
conda activate qcml
pip install -e .Some tests require pretrained model artifacts from the QCMLForge Hugging Face repository. Enable downloads before running pytest:
export QCMLFORGE_AUTO_DOWNLOAD_PRETRAINED=1
python -m pytest tests/Pretrained-model tests are skipped when this environment variable is not enabled or when their required artifacts cannot be resolved. Tests that do not need pretrained models still run normally.
If you get an OS.Error when running qcml related to torch-geometric, you likely need to install a specific version through the following example:
# If you want the CUDA version
pip uninstall torch-geometric
export TORCH=2.10.0
export CUDA=cu126 # for cuda version 12.6
pip install torch-geometric==2.7.0 -f https://data.pyg.org/whl/torch-${TORCH}+${CUDA}.html
# If you want the CPU version
pip uninstall torch-geometric
export TORCH=2.10.0
pip install torch-geometric==2.7.0 -f https://data.pyg.org/whl/torch-${TORCH}+cpu.htmlA QCArchive+QCMLForge+Cybershuttle workshop demo is available here. This demonstrates a complete workflow of using QCArchive to generate QM data and then training AP-Net models with QCMLForge.
To run the AtomModel inference, run the following command:
import apnet_pt
import qcelemental
mol_mon = qcelemental.models.Molecule.from_data("""0 1
16 -0.8795 -2.0832 -0.5531
7 -0.2959 -1.8177 1.0312
7 0.5447 -0.7201 1.0401
6 0.7089 -0.1380 -0.1269
6 0.0093 -0.7249 -1.1722
1 1.3541 0.7291 -0.1989
1 -0.0341 -0.4523 -2.2196
units angstrom
""")
mols = [mol_mon for _ in range(3)] # example of using multiple molecules
multipoles = apnet_pt.pretrained_models.atom_model_predict(
mols,
compile=False,
batch_size=2,
)
print(multipoles)
# multipoles = [[np.array(q) for q in qs], [[np.array(d) for d in ds], [np.array(qp) for qp in qps]]]To run the APNet2Model inference, run the following command:
import apnet_pt
import qcelemental
mol_dimer = qcelemental.models.Molecule.from_data("""
0 1
O 0.000000 0.000000 0.000000
H 0.758602 0.000000 0.504284
H 0.260455 0.000000 -0.872893
--
0 1
O 3.000000 0.500000 0.000000
H 3.758602 0.500000 0.504284
H 3.260455 0.500000 -0.872893
""")
mols = [mol_dimer for _ in range(3)]
interaction_energies = apnet_pt.pretrained_models.apnet2_model_predict(
mols,
compile=False,
batch_size=2,
)
print(interaction_energies)
# interaction_energies = np.array((N, 5)), where [[E_total, E_elst, E_exch, E_ind, E_disp]...]
# [[-2.67262248 -3.41546969 2.33096213 -0.61418158 -0.97393334]
# [-2.67262368 -3.41547084 2.33096209 -0.61418151 -0.97393341]
# [-2.67261759 -3.41546469 2.33096199 -0.61418152 -0.97393337]]By default this uses the APNet2 ensemble trained by QCMLForge, which reaches
0.204 kcal/mol total MAE on the paper's 150 000-dimer Splinter test split. Pass
weights="ap2_tf_paper" to use the ensemble published with the AP-Net2 paper
instead, converted from TensorFlow and verified against the original
predictions, or weights="qcmlforge_v1" for the pre-AtomMPNN-fix ensemble
that used to be the default:
interaction_energies = apnet_pt.pretrained_models.apnet2_model_predict(
mols,
compile=False,
batch_size=2,
weights="ap2_tf_paper",
)See docs/apnet2-pretrained-weights.md for what each weight set contains and how they score, and docs/apnet2-tensorflow-weights.md for the TensorFlow parity data and the loading caveats.
To train the model, run the following command:
python3 ./train_models.py \
--train_ap2 \
--ap_model_path ./models/example/ap2_example.pt \
--n_epochs 5 Optional experiment tracking is documented in Training with Weights & Biases.
models/ap2_tf_paper/ holds the published TensorFlow ensemble converted to PyTorch
checkpoints, which reproduce the original model's predictions to float32
accumulation noise. See
Running APNet2 with the original TensorFlow weights.
Code re-implemented from TensorFlow version located here
To train the model, run the following command:
python3 ./train_models.py \
--train_am \
--am_model_path ./models/example/am_example.pt \
--n_epochs 5 - Extend AtomMPNN to predict Hirshfeld ratios
- Add classical induction model for AP3
The free-atom polarizabilities come from libmbd. To cite Hirshfeld model, please cite libmbd and the original paper to give appropriate credit for their indirect contributions.