Skip to content

[enhancement](scan) Optimize Parquet V2 predicate filtering and fixed-binary decimal decoding - #66396

Open
Gabriel39 wants to merge 2 commits into
apache:masterfrom
Gabriel39:dev/parquet-v2-optimizations-master
Open

[enhancement](scan) Optimize Parquet V2 predicate filtering and fixed-binary decimal decoding#66396
Gabriel39 wants to merge 2 commits into
apache:masterfrom
Gabriel39:dev/parquet-v2-optimizations-master

Conversation

@Gabriel39

@Gabriel39 Gabriel39 commented Aug 3, 2026

Copy link
Copy Markdown
Contributor

What problem does this PR solve?

Related PR: #66360, #66379

Problem Summary:

This PR brings two complementary Parquet V2 optimizations from branch-4.1 to master:

  1. reduce predicate-filtering overhead and correctly propagate late runtime filters through the scanner and reader layers;
  2. avoid unnecessarily decoding same-scale fixed-binary decimals through Int256 when their physical width permits a narrower native type.

Parquet V2 predicate filtering (#66360)

  • Keep identity selection-vector state implicit and compact selected rows in bulk.
  • Retain the selection scratch high-water mark across scanner batches and specialize the first compaction from implicit identity.
  • Refresh late runtime-filter requests at safe row-group boundaries.
  • Re-run footer-statistics pruning and reset adaptive predicate state for unopened row groups after a refresh.
  • Preserve real COUNT(*) carrier values while runtime filters are pending.
  • Initialize refreshed JNI predicates and attribute refresh work to TableReader, FileReader, and Parquet profiles.
  • Preserve Hudi and Paimon child-reader predicate state.
  • Remove query-scoped dictionary-filter cache state.
  • Share immutable VDirectInPredicate pruning materialization across split-local expression clones.
  • Add correctness-checked selection-vector and direct-IN lifecycle microbenchmarks.

Fixed-binary decimal decoding (#66379)

  • Decode same-scale Parquet FIXED_LEN_BYTE_ARRAY decimals with int32_t, int64_t, or Int128, selected from the physical width, instead of always using Int256.
  • Use unaligned full-width big-endian loads for 4-, 8-, and 16-byte values while preserving sign extension for shorter widths.
  • Validate target precision before narrowing and preserve both permissive and strict conversion behavior.
  • Keep rescaling and wider values on the existing Int256 path.
  • Cover positive and negative precision boundaries, shorter signed inputs, strict rollback, permissive null marking, and the complete int32 source domain.

Performance

The results below come from the original branch-4.1 PR benchmarks in #66360 and #66379. This PR applies the same implementation to master; the performance benchmarks were not re-run as part of this forward-port.

Selection-vector processing

The benchmark validates every surviving original row ID after the timed region. It used the same Clang -O3 -DNDEBUG -mavx2 benchmark source for the branch base, the pre-fix PR, and the final implementation, with one pinned CPU, three warmups, eight adjacent A-B-B-A quartets, and at least 0.3 seconds per invocation. Negative percentages are improvements.

Operation Final selectivity Final vs pre-fix PR Final vs branch base
Identity initialization 100% -15.23% -99.12%
Row filter 1% -24.23% -23.76%
Row filter 50% -16.10% -34.50%
Row filter 90% -45.91% -45.95%
Row filter 100% -31.72% -17.32%
Successive filters 1% -33.25% -35.79%
Successive filters 50% -29.80% -35.10%
Successive filters 90% -25.27% -25.16%
Successive filters 100% -24.93% -23.72%

Compared with the branch base, identity initialization improved by 99.12%, row filtering improved by 17.32%-45.95%, and successive filtering improved by 23.72%-35.79%. All final-vs-base paired-ratio CVs were at most 5.85%.

Direct-IN expression lifecycle

FileScannerExpr/direct_in_clone_prepare_open isolates deep-clone, prepare, and open for an already prepared direct-IN runtime filter. Set construction and the original fragment prepare/open are outside the timed region. The shared and forced-rematerialization implementations ran in the same Release binary on one pinned CPU, with 10 repetitions and at least 0.5 seconds per repetition.

IN values Rematerialize median Shared median Speedup
128 207.470 us 1.634 us 126.9x
1,024 1.674 ms 1.642 us 1,019.5x
8,192 13.514 ms 1.672 us 8,082.2x
65,536 108.337 ms 1.650 us 65,663.1x

The shared path remains approximately constant because split-local clones reuse immutable pruning state instead of rebuilding it for every split.

Reader-level regression guardrail

The Parquet reader benchmark covered nullable INT32 predicate scans with a lazy payload for both PLAIN and dictionary encoding. Across 1%, 10%, 50%, and 90% selectivity, CPU-time changes ranged from -1.34% to +1.41% with mixed signs, so it did not detect a material aggregate reader-level regression.

Fixed-binary decimal decoding

The benchmark decoded 65,536 values per iteration through DataTypeDecimalSerDe::read_column_from_parquet, pinned to one CPU, with three warmups followed by 10 repetitions in A-B-B-A order.

Target / physical width Before median CPU After median CPU Speedup CPU reduction
Decimal32 / 4 bytes 1,359,514 ns 88,527 ns 15.36x 93.49%
Decimal64 / 8 bytes 1,634,757 ns 90,609 ns 18.04x 94.46%
Decimal128 / 16 bytes 2,206,272 ns 152,823 ns 14.44x 93.07%

The optimized path was 14.44x-18.04x faster in this benchmark, reducing CPU time by 93.07%-94.46%. The benchmark host was heavily loaded and CPU frequency scaling was enabled, so the exact ratios are noisy; however, the before/after median ranges did not overlap in any A-B-B-A stage.

Validation on master

  • 705 filtered ASAN BE unit tests from 53 suites passed, covering Parquet V2, FileScannerV2, TableReader, Hudi/Paimon/JNI readers, SelectionVector, direct-IN pruning, and decimal SerDe.
  • BE formatting and strict clang-format checks passed.
  • git diff --check passed.

Release note

None

Check List (For Author)

  • Test

    • Regression test
    • Unit Test
    • Manual test (add detailed scripts or steps below)
    • No need to test or manual test. Explain why:
      • This is a refactor/code format and no logic has been changed.
      • Previous test can cover this change.
      • No code files have been changed.
      • Other reason
  • Behavior changed:

    • No.
    • Yes.
  • Does this need documentation?

    • No.
    • Yes.

Check List (For Reviewer who merge this PR)

  • Confirm the release note
  • Confirm test cases
  • Confirm document
  • Add branch pick label

…pache#66360)

Backport the selected Parquet V2 direct-predicate filtering changes from
on this branch.

- keep identity selection-vector state implicit and compact selected
rows in bulk
- retain the selection scratch high-water mark across scanner batches
and specialize first compaction from implicit identity
- refresh late runtime-filter requests at safe row-group boundaries
- re-run footer-statistics pruning and reset adaptive predicate state
for unopened row groups after a refresh
- preserve real COUNT(*) carrier values while runtime filters are
pending
- initialize refreshed JNI predicates and attribute refresh work to
TableReader/FileReader/Parquet profiles
- preserve Hudi/Paimon child-reader predicate state
- remove query-scoped dictionary-filter cache state
- share immutable `VDirectInPredicate` pruning materialization across
split-local expression clones
- add correctness-checked selection and direct-IN lifecycle
microbenchmarks

- `./run-be-ut.sh --run
--filter='FileScannerV2Test.*:*Parquet*:*TableReaderTest.*:Hudi*ReaderTest.*:Paimon*ReaderTest.*:SelectionVectorTest.*:DictionaryFilterCostTest.*'
-j48`
  - 639 tests from 47 test suites passed under ASAN
- targeted late-RF, COUNT(*), dictionary-snapshot, shared-IN-state, and
SelectionVector tests
  - 19 tests from 5 test suites passed under ASAN
- Release benchmark build and smoke run
- expected registrations: 228 decoder, 92 kernel, 25 selection, 167
reader, and 8 expression-lifecycle cases
- all 25 selection and 8 expression-lifecycle cases executed with zero
benchmark errors
- `git diff --check`

The final benchmark source validates every surviving original row ID
after the timed region. Base, pre-fix PR, and final binaries use the
same benchmark source and Clang `-O3 -DNDEBUG -mavx2` on the same host.
Each comparison uses one pinned CPU, three warmups, eight adjacent
A-B-B-A quartets, and at least 0.3 seconds per invocation. The table
reports median paired CPU-time ratios; negative values are improvements.

| Operation | Final selectivity | Final vs pre-fix PR | Final vs branch
base |
|---|---:|---:|---:|
| Identity initialization | 100% | -15.23% | -99.12% |
| Row filter | 1% | -24.23% | -23.76% |
| Row filter | 50% | -16.10% | -34.50% |
| Row filter | 90% | -45.91% | -45.95% |
| Row filter | 100% | -31.72% | -17.32% |
| Successive filters | 1% | -33.25% | -35.79% |
| Successive filters | 50% | -29.80% | -35.10% |
| Successive filters | 90% | -25.27% | -25.16% |
| Successive filters | 100% | -24.93% | -23.72% |

All final-vs-base paired-ratio CVs are at most 5.85%. The previous
16.43%/59.94% dense row-filter regressions and 27.92%-61.05%
successive-filter regressions are no longer present. Retaining `_owned`
avoids repeated value initialization; the implicit-identity
specialization removes the remaining source/coordinate branches from the
first compaction.

`FileScannerExpr/direct_in_clone_prepare_open` isolates deep-clone,
prepare, and open for an already prepared direct-IN runtime filter. Set
construction and the original fragment prepare/open are outside the
timed region. Shared and forced-rematerialization implementations run in
the same Release binary on one pinned CPU, with 10 repetitions and at
least 0.5 seconds per repetition.

| IN values | Rematerialize median | Shared median | Speedup |
|---:|---:|---:|---:|
| 128 | 207.470 us | 1.634 us | 126.9x |
| 1,024 | 1.674 ms | 1.642 us | 1,019.5x |
| 8,192 | 13.514 ms | 1.672 us | 8,082.2x |
| 65,536 | 108.337 ms | 1.650 us | 65,663.1x |

The shared path remains approximately constant because split clones
reuse the immutable, fragment-originated pruning state; the
rematerialization path scales linearly with runtime-filter cardinality.

The earlier reader-level comparison used an identical Parquet-only
Release benchmark binary on the same host, with one pinned CPU, warm
cache, three warmups, A-B-B-A order, 10 repetitions, and at least one
second per repetition. It covers nullable INT32 predicate scans with a
lazy payload for PLAIN and dictionary encoding.

| Encoding | Selectivity | CPU time vs base | Wall time vs base |
|---|---:|---:|---:|
| Dictionary | 1% | +0.44% | +0.46% |
| Dictionary | 10% | +1.13% | +1.22% |
| Dictionary | 50% | +1.41% | +1.51% |
| Dictionary | 90% | +0.41% | +0.53% |
| PLAIN | 1% | -0.65% | -0.66% |
| PLAIN | 10% | +0.32% | +0.38% |
| PLAIN | 50% | -1.21% | -1.18% |
| PLAIN | 90% | -1.34% | -1.29% |

The reader-level point estimates span -1.34% to +1.41% CPU time with
mixed signs, so this benchmark did not detect a material aggregate
regression. It starts at `format::parquet::ParquetReader`; it does not
cover scanner scheduling or end-to-end SQL execution.
## Proposed changes

- Decode same-scale Parquet `FIXED_LEN_BYTE_ARRAY` decimals with an
`int32_t`, `int64_t`, or `Int128` source selected from the physical
width instead of always using `Int256`.
- Use unaligned full-width big-endian loads for 4-, 8-, and 16-byte
values, while preserving sign extension for shorter widths.
- Validate the target decimal precision before narrowing, preserve
permissive/strict conversion behavior, and keep rescaling and wider
values on the existing `Int256` path.
- Add regression coverage for positive and negative precision
boundaries, shorter signed inputs, strict rollback, permissive null
marking, and the complete int32 source domain.

## Validation

- ASAN BE unit tests: 8 tests from `DataTypeSerDeParquetTest` passed.
- Release microbenchmark: 65,536 values per iteration through
`DataTypeDecimalSerDe::read_column_from_parquet`, pinned to one CPU.
Each stage used 3 warmups followed by 10 repetitions in ABBA order.

| Target / physical width | Before median CPU | After median CPU |
Speedup | CPU reduction |
| --- | ---: | ---: | ---: | ---: |
| Decimal32 / 4 bytes | 1,359,514 ns | 88,527 ns | 15.36x | 93.49% |
| Decimal64 / 8 bytes | 1,634,757 ns | 90,609 ns | 18.04x | 94.46% |
| Decimal128 / 16 bytes | 2,206,272 ns | 152,823 ns | 14.44x | 93.07% |

The benchmark host was heavily loaded and CPU frequency scaling was
enabled, so the exact ratios are noisy. However, the before/after median
ranges did not overlap in any ABBA stage. A final optimized-build smoke
run measured median CPU times of 95,077 ns, 94,635 ns, and 145,417 ns
with CPU CVs of 0.41%, 1.76%, and 0.64%, respectively.
@hello-stephen

Copy link
Copy Markdown
Contributor

Thank you for your contribution to Apache Doris.
Don't know what should be done next? See How to process your PR.

Please clearly describe your PR:

  1. What problem was fixed (it's best to include specific error reporting information). How it was fixed.
  2. Which behaviors were modified. What was the previous behavior, what is it now, why was it modified, and what possible impacts might there be.
  3. What features were added. Why was this function added?
  4. Which code was refactored and why was this part of the code refactored?
  5. Which functions were optimized and what is the difference before and after the optimization?

@Gabriel39

Copy link
Copy Markdown
Contributor Author

run buildall

@Gabriel39

Copy link
Copy Markdown
Contributor Author

/review

@Gabriel39
Gabriel39 marked this pull request as ready for review August 3, 2026 14:06
@Gabriel39
Gabriel39 requested a review from yiguolei as a code owner August 3, 2026 14:06
@Gabriel39 Gabriel39 changed the title [enhancement](scan) Forward-port Parquet V2 predicate and decimal optimizations [enhancement](scan) Optimize Parquet V2 predicate filtering and fixed-binary decimal decoding Aug 3, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants