From 5d6b697adaf4a6a4e1d89b2911313f42c17df9a5 Mon Sep 17 00:00:00 2001 From: Raphael Collet Date: Fri, 25 Sep 2020 14:20:54 +0000 Subject: [PATCH] [FIX] core: optimize prefetching on large recordsets MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit This patch optimizes the performance of prefetching when iterating on large recordsets. The optimization is *transparent*, i.e., it requires no code change for it to apply. The worst-case scenario for prefetching is the following: the ORM has to fetch a field for a given `record`, `record._prefetch_ids` (its prefetch set) is *very* large (many thousands), and most records in the prefetch set are already in cache. This requires the ORM to iterate a lot on the prefetch set in order to make a batch of records not having the field in cache. records = model.browse(ids) # large recordset for record in records: record.foo # fetch 'foo' every 1k records When running such a loop on an empty cache, the overhead of prefetching (determine a batch) grows as the loop progresses. The time complexity of this loop is actually O(N²)... The overhead of prefetching is minimal when `record._prefetch_ids` is about the size of a prefetching unit, i.e., 1k records. This commit modifies the iterator method such that every record returned by the iterator has a prefetch set of maximum 1k records. We measured the time taken by the loop above on an empty cache, before and after this commit, on a simple model (res.partner.category) with 100k records. The third measure is a reference one: a `_read` on all prefetched fields (to fill in the cache) followed by the loop. Total time Time per 1k records Before this commit 3.690s 30ms - 45ms After this commit 1.176s 12ms Read then loop 1.161s - The measures are enlightening: the overhead of prefetching was more than 200% of the reference time, and the time to prefetch records grows as the iteration goes on! This commit reduces the overhead of prefetching to less than 2% of the reference time, and make it scale gracefully with data size. closes odoo/odoo#59239 X-original-commit: d0d54d63dbf49a97e7fea36fe7fb4b860a0b6606 Signed-off-by: Raphael Collet (rco) --- odoo/models.py | 9 +++++++-- 1 file changed, 7 insertions(+), 2 deletions(-) diff --git a/odoo/models.py b/odoo/models.py index 2d57c223e99..6dbb0d7acbf 100644 --- a/odoo/models.py +++ b/odoo/models.py @@ -5485,8 +5485,13 @@ Fields: def __iter__(self): """ Return an iterator over ``self``. """ - for id in self._ids: - yield self._browse(self.env, (id,), self._prefetch_ids) + if len(self._ids) > PREFETCH_MAX and self._prefetch_ids is self._ids: + for ids in self.env.cr.split_for_in_conditions(self._ids): + for id_ in ids: + yield self._browse(self.env, (id_,), ids) + else: + for id in self._ids: + yield self._browse(self.env, (id,), self._prefetch_ids) def __contains__(self, item): """ Test whether ``item`` (record or field name) is an element of ``self``.