Prevent overflow when tracking long MongoDB keys
Removed a premature integer-to-float conversion in LexicographicKeyRangeTracker and added regression coverage for long byte and string keys.
apache/beam · #40330
Python I/O correctness
Apache Beam can now calculate progress for very long lexicographic keys without overflowing during MongoDB and other range-tracked reads.
Problem
LexicographicKeyRangeTracker converted an arbitrarily large key-distance integer to float before dividing it by the range. Long MongoDB keys could exceed the float conversion limit and fail a real Dataflow read with OverflowError while reporting progress.
Approach
Uses Python's true division directly so it computes the ratio without first converting the full numerator to a float, then exercises the behavior with 400-character byte and string keys.
Impact and scope
- Resolves a reported production Dataflow failure in ReadFromMongoDB and closes GH-26492.
- Protects the shared lexicographic range tracker, so the correction applies beyond a single MongoDB call site.
- Keeps the runtime change to one expression and preserves the existing fractional-progress result for ordinary keys.
- Adds focused regression coverage for both bytes and text, the two key forms handled by the tracker.
Validation
- Passed 28 LexicographicKeyRangeTracker unit tests and 131 MongoDB I/O tests, for 159 focused tests in total.
- All 96 executed hosted checks passed; one CodeQL job was neutral and six change-detection jobs were skipped.
- A Beam contributor approved the change before merge, and the functional commit is authored by Goutam Adwant.