[FIX] core: optimize prefetching on large recordsets

This patch optimizes the performance of prefetching when iterating on
large recordsets.  The optimization is *transparent*, i.e., it requires
no code change for it to apply.

The worst-case scenario for prefetching is the following: the ORM has to
fetch a field for a given `record`, `record._prefetch_ids` (its prefetch
set) is *very* large (many thousands), and most records in the prefetch
set are already in cache.  This requires the ORM to iterate a lot on the
prefetch set in order to make a batch of records not having the field in
cache.

    records = model.browse(ids)     # large recordset
    for record in records:
        record.foo                  # fetch 'foo' every 1k records

When running such a loop on an empty cache, the overhead of prefetching
(determine a batch) grows as the loop progresses.  The time complexity
of this loop is actually O(N²)...

The overhead of prefetching is minimal when `record._prefetch_ids` is
about the size of a prefetching unit, i.e., 1k records.  This commit
modifies the iterator method such that every record returned by the
iterator has a prefetch set of maximum 1k records.

We measured the time taken by the loop above on an empty cache, before
and after this commit, on a simple model (res.partner.category) with
100k records.  The third measure is a reference one: a `_read` on all
prefetched fields (to fill in the cache) followed by the loop.

                            Total time      Time per 1k records
    Before this commit      3.690s          30ms - 45ms
    After this commit       1.176s          12ms
    Read then loop          1.161s          -

The measures are enlightening: the overhead of prefetching was more than
200% of the reference time, and the time to prefetch records grows as
the iteration goes on!  This commit reduces the overhead of prefetching
to less than 2% of the reference time, and make it scale gracefully with
data size.

closes odoo/odoo#59239

X-original-commit: d0d54d63dbf49a97e7fea36fe7fb4b860a0b6606
Signed-off-by: Raphael Collet (rco) <rco@openerp.com>
This commit is contained in:
Raphael Collet
2020-10-06 10:29:38 +00:00
parent c1393a713c
commit 5d6b697ada
+7 -2
View File
@@ -5485,8 +5485,13 @@ Fields:
def __iter__(self):
""" Return an iterator over ``self``. """
for id in self._ids:
yield self._browse(self.env, (id,), self._prefetch_ids)
if len(self._ids) > PREFETCH_MAX and self._prefetch_ids is self._ids:
for ids in self.env.cr.split_for_in_conditions(self._ids):
for id_ in ids:
yield self._browse(self.env, (id_,), ids)
else:
for id in self._ids:
yield self._browse(self.env, (id,), self._prefetch_ids)
def __contains__(self, item):
""" Test whether ``item`` (record or field name) is an element of ``self``.