Skip to content

Handling exceptions when applications do not report errors but time out in stateful transitions#868

Draft
PawelPlesniak wants to merge 2 commits intoprep-release/fddaq-v5.6.0from
PawelPlesniak/IncompleteStatefulCommandTransition
Draft

Handling exceptions when applications do not report errors but time out in stateful transitions#868
PawelPlesniak wants to merge 2 commits intoprep-release/fddaq-v5.6.0from
PawelPlesniak/IncompleteStatefulCommandTransition

Conversation

@PawelPlesniak
Copy link
Copy Markdown
Collaborator

Description

Fixes issue #803

If a segment does not reach the target state, it is marked as in error. Error recovery with the supervisor will address what happens if an application completes this outside of the designated window. This is defined in #840

Type of change

  • New feature / enhancement
  • Optimization
  • Bug fix
  • Breaking change
  • Documentation

List of required branches from other repositories

N/A

Change log

WHAT HAS CHANGED.

Suggested manual testing checklist

There is a configuration with fake daq applications that is defined to take 70 seconds to complete conf for all of the applications. These applications will fail to complete the transition within the defined timeout, and as such the relevant segments will be placed into an error state.

drunc-unified-shell ssh-standalone config/tests/nestedConfig.data.xml test-config issue803 boot conf 

Developer checklist

Prior to marking this as "Ready for Review"

Tests ran on: WHAT HOSTNAME from release RELEASE_NAME

Unit tests - some tests can't be ran on the CI. This is documented. If this PR checks a feature that can't be tested with CI, this has been marked appropriately.

Integration tests - the daqsystemtest_integtest_bundle requires a lot of resources, and connections to the EHN1 infrastructure. Check the cross referenced list if you can't run these. The developer needs to run at least the .

  • Unit tests (pytest --marker) passed
    • With relevant marker
    • Without marker
  • Integration tests passed
    • Only daqsystemtest_integtest_bundle.sh -k minimal_system_quick_test.py
    • Full daqsystemtest_integtest_bundle.sh
  • Testing skipped as there are no core code changes in this PR, this only relates to documentation/CI workflows

Final checklist prior to marking this as "Ready for Review"

  • Code is clearly commented.
  • New unit tests have been added, or is documented in # ISSUE NUMBER
  • A suitable reviewer has been chosen from this list.

Reviewer checklist

  • This branch has been rebased with develop prior to testing.
  • Suggested manual tests show changes.
  • CI workflows fails documented (if present)
  • Integration tests passed
    • Only concern yourself if failures related to drunc are in the log files
    • If non-drunc failure appears:
      • Validate failure in fresh working area
      • Contact Pawel if unsure

Once the features are validated and both the unit and integration tests pass, the PRs is ready to be merged.

Prior to merging

Choose one of the following an complete all substeps
  • Changes only affect the Run Control, are in a single repository, and do not affect the end user.
    • Changes are documented in docstrings and code comments
    • Wiki has been updated if architectural or endpoint changes
  • Otherwise
    • Workflow changes demonstrated in the Change Log (if necessary)
    • Wiki has been updated (if necessary)
    • #daq-sw-librarians Slack channel notified (see below)

Once completed, the reviewer can merge the PR.

Notification message for a Slack channel

Note - this should be to #dunedaq-integration for general workflow that isn't during a release candidate period, and to #daq-release-prep otherwise.

For an single merge that changes the user workflow

The CCM WG has an isolated PR ready to merge that affects user workflows. The PR is:

_URL_

I will leave time for any comments, otherwise will merge these at the end of the work day _Insert your time zone_.

For co-ordinated merge

The CCM WG has a set of co-ordinated merges ready to merge. The PRs are:

_URL_

_URL_


I will leave time for any comments, otherwise will merge these at the end of the day.

@PawelPlesniak
Copy link
Copy Markdown
Collaborator Author

image Interestingly, there is already some logic for this...

@PawelPlesniak
Copy link
Copy Markdown
Collaborator Author

image In the case where a second application also fails to complete a transition in time, the same error gets thrown. This is likely caused by the nested structure, and the fact that there are multiple layers to this configuration. A robust solution to this problem will take longer to achieve, but I will continue working on it.

@PawelPlesniak PawelPlesniak changed the title Generating an environment for which the issue can be recreated Handling exceptions when applications do not report errors but time out in stateful transitions Mar 31, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants