jameslamb's Description of Work
This PR ensures that the check-omp-pragmas check only checks LightGBM's own first-party source code, not:
- code pulled in via submodules
- other untracked code in the repo (e.g. code in Python virtual environments)
For some situations, this makes local runs of pre-commit run --all-files faster and removes a source of false positives.
While touching this, also added a few more C/C++ extension types to the check:
-
.cu(CUDA C++ source files) -
.cuh(CUDA C++ headers) -
.tpp(C++ template files)
Notes for Reviewers
How I found this
I've recently switched from using conda for Python development here to virtualenvs, and frequently do stuff like this:
python -m venv .venv
source .venv/bin/activate
pip install numpy scipy scikit-learn
# etc., etc.
When I started doing that, I found that check-omp-pragmas at time was noticably slower, and occasionally reported findings in files that weren't part of LightGBM.
How I tested this
Manually removed a num_threads() class from the following places:
-
include/LightGBM/bin.hhere - some places in
external_libs/eigen - a copy of
bin.hplaced in.venv/
Saw what I expected... only the first-party LightGBM one was reported and the hook failed.
$ pre-commit run --all-files check-omp-pragmas
check-omp-pragmas........................................................Failed
- hook id: check-omp-pragmas
- exit code: 1
checking that all OpenMP pragmas specify num_threads()
include/LightGBM/bin.h:68: #pragma omp parallel for schedule(static)
Found '#pragma omp parallel' not using explicit num_threads() configuration. Fix those.
For details, see https://www.openmp.org/spec-html/5.0/openmpse14.html#x54-800002.6