Skip to content

Psyclone batch driver - #5

Open
andrewcoughtrie wants to merge 4 commits into
mainfrom
psyclone_batch_driver
Open

Psyclone batch driver#5
andrewcoughtrie wants to merge 4 commits into
mainfrom
psyclone_batch_driver

Conversation

@andrewcoughtrie

Copy link
Copy Markdown
Owner

PR Summary

Sci/Tech Reviewer:
Code Reviewer:

Code Quality Checklist

  • I have performed a self-review of my own code
  • My code follows the project's style guidelines
  • Comments have been included that aid understanding and enhance the readability of the code
  • My changes generate no new warnings
  • All automated checks in the CI pipeline have completed successfully

Testing

  • I have tested this change locally, using the LFRic Core rose-stem suite
  • If required (e.g. API changes) I have also run the LFRic Apps test suite using this branch
  • If any tests fail (rose-stem or CI) the reason is understood and acceptable (e.g. kgo changes)
  • I have added tests to cover new functionality as appropriate (e.g. system tests, unit tests, etc.)
  • Any new tests have been assigned an appropriate amount of compute resource and have been allocated to an appropriate testing group (i.e. the developer tests are for jobs which use a small amount of compute resource and complete in a matter of minutes)

trac.log

Security Considerations

  • I have reviewed my changes for potential security issues
  • Sensitive data is properly handled (if applicable)
  • Authentication and authorisation are properly implemented (if applicable)

Performance Impact

  • Performance of the code has been considered and, if applicable, suitable performance measurements have been conducted

AI Assistance and Attribution

  • Some of the content of this change has been produced with the assistance of Generative AI tool name (e.g., Met Office Github Copilot Enterprise, Github Copilot Personal, ChatGPT GPT-4, etc) and I have followed the Simulation Systems AI policy (including attribution labels)

Documentation

  • Where appropriate I have updated documentation related to this change and confirmed that it builds correctly

PSyclone Approval

  • If you have edited any PSyclone-related code (e.g. PSyKAl-lite, Kernel interface, optimisation scripts, LFRic data structure code) then please contact the HPC Optimisation Team

Sci/Tech Review

  • I understand this area of code and the changes being added
  • The proposed changes correspond to the pull request description
  • Documentation is sufficient (do documentation papers need updating)
  • Sufficient testing has been completed

(Please alert the code reviewer via a tag when you have approved the SR)

Code Review

  • All dependencies have been resolved
  • Related Issues have been properly linked and addressed
  • CLA compliance has been confirmed
  • Code quality standards have been met
  • Tests are adequate and have passed
  • Documentation is complete and accurate
  • Security considerations have been addressed
  • Performance impact is acceptable

This branch takes the simpler alternative to the persistent daemon:
transform stale algorithm files in one short-lived batch process per component,
then keep the existing per-file make rules as a safety net.

What changed
------------
* Added `infrastructure/build/psyclone/psyclone_batch.py`.
  - Discovers stale `*.x90` files in `WORKING_DIR`.
  - Matches make's script precedence (`<stem>.py` over `global.py`).
  - Imports PSyclone once in the parent.
  - Forks children to process files in parallel (copy-on-write reuse of the
    imported interpreter and per-file state isolation).
  - Captures child output and reports per-file failures.
  - Never fails the build itself: failed files are left for make's existing
    per-file path to retry and report.
* Reworked `psyclone_psykal.mk` into three explicit phases:
  1. `psyclone-preprocess`
  2. `psyclone-batch`
  3. `psyclone-generate`
  This avoids make's early out-of-date decision and ensures the batch sees all
  preprocessed inputs before generation rules are checked.
* Added worker sizing + cap logic (`PSYCLONE_MAX_WORKERS`, default 8) and
  switched from recursive expansion to simply-expanded variables so `nproc`
  is not forked per recipe.
* Added `$(WORKING_DIR)/kernel` creation before the batch/per-file phases,
  because PSyclone 3.3.1 requires `-okern` target directory to already exist.
* Removed daemon-specific machinery:
  - deleted `psyclone_client.py`
  - deleted `psyclone_server.py`
  - deleted `psyclone_procs.py`
  - removed owner-pid export block from `lfric.mk`
  - replaced server smoke test with batch-focused tests
* Added `infrastructure/build/psyclone/README.md` documenting design,
  variables, safety model and debug knobs.

Tests
-----
* Added `infrastructure/build/psyclone/tests/batch_test.py` covering:
  - fidelity vs direct `psyclone`
  - isolation between files in one batch
  - stale-only rebuild behaviour
  - failure containment
  - skipping hand-written `SOURCE_DIR/psy/*_psy.f90` overrides
* Updated `infrastructure/build/psyclone/tests/Makefile` with `batch` target.

Validation run on this branch
-----------------------------
* `python3 -m py_compile psyclone_batch.py tests/batch_test.py` -> OK
* `python3 infrastructure/build/psyclone/tests/batch_test.py` -> PASS
* `make no-optimisation/invoke` in
  `infrastructure/build/psyclone/tests` with
  `LFRIC_BUILD=<repo>/infrastructure/build CORE_ROOT_DIR=<repo>` -> PASS
  and shows the batch pre-pass doing the work:
  `PSyclone: 2 algorithm file(s), 7.3s to load PSyclone, 0.2s to transform`

Why this shape
--------------
This captures most of the performance win (single import amortised over many
files) with far less operational complexity than a detached daemon: no FIFOs,
no lock/orphan handling, no process ownership tracking, and no long-lived
process state to reset between jobs.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant