This patch optimizes the performance of prefetching when iterating on
large recordsets. The optimization is *transparent*, i.e., it requires
no code change for it to apply.
The worst-case scenario for prefetching is the following: the ORM has to
fetch a field for a given `record`, `record._prefetch_ids` (its prefetch
set) is *very* large (many thousands), and most records in the prefetch
set are already in cache. This requires the ORM to iterate a lot on the
prefetch set in order to make a batch of records not having the field in
cache.
records = model.browse(ids) # large recordset
for record in records:
record.foo # fetch 'foo' every 1k records
When running such a loop on an empty cache, the overhead of prefetching
(determine a batch) grows as the loop progresses. The time complexity
of this loop is actually O(N²)...
The overhead of prefetching is minimal when `record._prefetch_ids` is
about the size of a prefetching unit, i.e., 1k records. This commit
modifies the iterator method such that every record returned by the
iterator has a prefetch set of maximum 1k records.
We measured the time taken by the loop above on an empty cache, before
and after this commit, on a simple model (res.partner.category) with
100k records. The third measure is a reference one: a `_read` on all
prefetched fields (to fill in the cache) followed by the loop.
Total time Time per 1k records
Before this commit 3.690s 30ms - 45ms
After this commit 1.176s 12ms
Read then loop 1.161s -
The measures are enlightening: the overhead of prefetching was more than
200% of the reference time, and the time to prefetch records grows as
the iteration goes on! This commit reduces the overhead of prefetching
to less than 2% of the reference time, and make it scale gracefully with
data size.
closesodoo/odoo#59239
X-original-commit: d0d54d63dbf49a97e7fea36fe7fb4b860a0b6606
Signed-off-by: Raphael Collet (rco) <rco@openerp.com>