Conversation
Member
Author
|
not solving the issue in Windows. Investigating further |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
📝 Description
Fixes #14579
Running the same ROM simulation several times with more than one OpenMP thread gave results that differed at round-off level (~1e-15 relative L2, ~1e-12 absolute). The first differing time step changed randomly between runs. Single-thread runs were bit-identical.
Cause: in
Build, the elements loop was declared#pragma omp for schedule(guided, 512) nowait. Because ofnowait, threads that ran out of element chunks started the conditions loop straight away. Condition and element contributions were then added to the sameA/bentries (the boundary nodes) at the same time.AtomicAddmakes each addition safe but doesn't fix the order of the additions, so the last bits of the assembled system changed from run to run. The Galerkin projection Φᵀ A Φ spreads that into every reduced entry, and the nonlinear iterations carry it forward.Fix: remove
nowaitfrom the elements loop, so elements are fully assembled before conditions start. The cost is one implicit barrier perBuild.Before change:

After change:

🆕 Changelog
nowaitfrom the elements loop inBuild([RomApp] Consistency on ConvectionDiffusion ROMs #14579):GlobalROMBuilderAndSolverLeastSquaresPetrovGalerkinROMBuilderAndSolverAnnPromGlobalROMBuilderAndSolverAnnPromLeastSquaresPetrovGalerkinROMBuilderAndSolver