Pipeline for converting the curated reaction dataset into NOMAD archive files for upload to the polymerization OASIS.
copol_prediction/processed_data.csv ← input
│
│ ↓ create_database_json.py
▼
dump/database_json/polymerization_*.json per-measurement JSONs
dump/database_json/monomers/*.json (schema-compliant with
│ PolymerizationReactionInput
│ from FAIRmat-NFDI/nomad-
│ polymerization-reactions)
│
│ ↓ convert_to_archives.py + convert_monomers.py
▼
database/output/polymerization/*.archive.json NOMAD archives
database/output/monomers/*.archive.json (ready to upload)
From the repository root:
rm -rf dump/database_json
rm -rf database/output/polymerization database/output/monomers
mkdir -p database/output/polymerization database/output/monomers
# Install optional dependencies for database scripts
pip install '.[database]'
# 1) processed_data.csv → schema-compliant per-reaction JSONs
python database/create_database_json.py
# 2) JSONs → NOMAD archives (needs the nomad-polymerization CLI, see below)
python database/convert_to_archives.py
# 3) Monomer property files → monomer archives
python database/convert_monomers.py --source copol_prediction/api/molecule_propertiesReads copol_prediction/processed_data.csv and writes one schema-compliant JSON per measurement to dump/database_json/. The output validates against PolymerizationReactionInput in FAIRmat-NFDI/nomad-polymerization-reactions: reactivity ratios are nested under r_values and conf_intervals, the schema-canonical key names (monomer1_smiles, polymerization_method, calculation_method, r-product) are used, and PCA-reduced polymerisation-type / method embeddings + RDKit solvent descriptors are added.
DOIs in the source field are recovered from source_filename where the LLM-populated source column is empty (~half the rows) — every output JSON carries a canonical https://doi.org/<DOI> URL when one exists.
# Default: reads copol_prediction/processed_data.csv
python database/create_database_json.py
# Or point at the paper-frozen snapshot:
python database/create_database_json.py --input copol_prediction/paper_dataset/processed_data.csv
# Or use a legacy per-reaction JSON-directory input layout:
python database/create_database_json.py --input dump/processed_reactions/Output:
dump/database_json/polymerization_<doi>_<idx>.json— one per measurement row in the CSV.dump/database_json/monomers/monomer_*.json— copied fromcopol_prediction/api/molecule_properties/for every monomer that appears in the dataset.
Converts JSON files to NOMAD archive files (.archive.json) using the nomad-polymerization CLI tool.
nomad-polymerization-reactions package installed as part of the [database] extras provides the
nomad-polymerization command, which is used under the hood by this script to convert each JSON file
to an archive.
Usage:
Standard usage (converts all files from dump/database_json):
python database/convert_to_archives.pyArchive files are saved by default in database/output/:
database/output/polymerization/*.archive.json- Polymerization reactionsdatabase/output/monomers/*.archive.json- Monomers
Convert single file:
python database/convert_to_archives.py dump/database_json/polymerization_10.1002_actp.1983.010340208_1.jsonDifferent input directory:
python database/convert_to_archives.py dump/other_json_folderCustom output directory:
python database/convert_to_archives.py --output dump/custom_archivesSave archive files in the same directory as JSON files:
python database/convert_to_archives.py --same-dirModes:
--mode polymerization- For polymerization reactions--mode monomer- For monomers- Without
--mode, the mode is automatically detected from the filename
Builds one NOMAD monomer archive per distinct monomer in the curated reaction
table. Coverage is driven by --reactions-csv, not by the source directory:
every monomer{1,2}_smiles in processed_data.csv is guaranteed an archive,
even if its precomputed quantum-feature file is missing — those become minimal
{smiles, name} stubs (valid per the upstream MonomerInput schema).
# Default: writes 865 archives to database/output/monomers/
python database/convert_monomers.pyThe run ends with a coverage report listing how many archives carry rich
features vs. are stubbed, plus the SMILES of any stub-only entries — rerun
copol_prediction/monomer_feature_calculation.py to backfill descriptors.
Useful flags:
# Different feature-file source (default: copol_prediction/api/molecule_properties)
python database/convert_monomers.py --source copol_prediction/output/molecule_properties
# Custom output directory
python database/convert_monomers.py --output dump/monomer_archives
# Re-archive monomers even if a target archive already exists
python database/convert_monomers.py --no-skip
# One-off: re-archive a single source file under a custom name
python database/convert_monomers.py --fix "C=COCC1CO1.OCCO.json" "2-(oxiran-2-ylmethoxy)ethanol"Display names are taken from the CSV's monomer{1,2}_name columns, with
copolextractor.utils.smiles_to_name() as a fallback and the SMILES as last
resort. The archive files can then be uploaded directly to NOMAD.
For a single bundle suitable for upload, tar+gzip the output directory:
tar czf database/output/monomers.tar.gz -C database/output monomers/The shipped database/output/monomers.tar.gz is regenerated from
monomers/ this way and tracks it 1:1.
Creates the combined multi-panel figure and prints basic dataset statistics (counts + ranges).
python database/analysis/plot_combined_database_figure.pyOutput:
database/analysis/figures/dataset_analysis.pdfdatabase/analysis/figures/dataset_analysis.png
Boiling points input:
database/analysis/solvent_boiling_points_c.json(curated solvent boiling points in °C;nullmeans “exclude”)
Creates stacked-area temporal trend plots for polymerization types and methods.
Uses the local publication_year column (no Crossref calls).
python database/analysis/plot_polymerization_trends.pyOutput:
database/analysis/figures/polymerization_trends.pdfdatabase/analysis/figures/polymerization_trends.png