Purpose
=======
The work entries generation is not scaling properly
for large employee datasets. A lot of improvements
have been made to reduce the number of query to the
database. However, one of the remaining bottleneck
was the new records insertion, made 1 by 1 on the
database.
Due to the fact that the model doens't have a lot
of columns, and doens't have columns containing
too much data (HTML fields for example), we could
imagine inserting the records by batch, for example
1000 by 1000, as it is done with SELECT.
The results weren't satisfying enough, so we prefered
using the new cron triggers mechanism, to re-trigger
the same cron at the end of its execution in another
transaction, to avoid exceeding the non return point
from which posgreSQL is pedaling in the semoule.
With this commit, it takes less than 30 seconds to
generate the work entries for 1000 employees, instead
of 3 min, and allows to generate the work entries in
a linear execution time instead of a exponential one.
Create method analysis:
=======================
Note, that only the call to the create method is
tracked, not the preprocessing time to retrieve
the work entries values.
When inserting the records 1 by 1:
----------------------------------
records - Elapsed Time (s) - AVG Time per record (s)
10 - 0.0045862197875976 - 0.00045862197875976
432 - 0.1101946830749511 - 0.00025508028489572
192 - 0.0525383949279785 - 0.00027363747358322
522 - 0.1418645381927490 - 0.00027177114596312
892 - 0.2803149223327636 - 0.00031425439723404
8800 - 3.4928441047668457 - 0.00039691410281441
88000 - 188.82338452339172 - 0.00214572027867490
176000 - 1172.4313135147095 - 0.00666154155406084
We observe that from 10.000 new record, the create
method is not scaling anymore before this commit.
For a company with 1000 employees, the mean work
entries by month is 1000*2*21=42.000 work entries,
and the time to create the records is not acceptable.
When inserting the records 1000 by 1000
---------------------------------------
records - Elapsed Time (s) - AVG Time per record (s)
10 - 0.003406763076782 - 0.0003406763076782
432 - 0.075028181076049 - 0.0001736763450834
192 - 0.030718326568603 - 0.0001599912842114
522 - 0.084228038787841 - 0.0001613563961452
892 - 0.145947217941284 - 0.0001636179573332
8800 - 2.204506397247314 - 0.0002505120905962
88000 - 181.9736533164978 - 0.0020678824240511
176000 - 1146.928646564483 - 0.0065166400372981
We observe a sligh improvement, but cleary not enough
and not worth the complexity of inserting the records
in batch on the database.
Note: The following explanation is speculative and could
be validated with a real analysis, but according to the
results we obtained with the cron, this is most likely
to be true.
In fact, this is a limitation of PosgreSQL. In the
implementation that manages the transactions, the number
of inserts becomes greater than certain memory limits,
suddenly it falls on a slower alternative storage (this
is the problem with SELECT and tuples of ids ).
When we inject 100,000 ids into a query, PosgreSQL parses
the query (already there it could have some trouble), then
it stores the ids in a data structure: hash if not too
large, other (on the file system) otherwise. Then it
executes the query and uses the data structure to check
the validity (membership).
Specification
=============
Instead of inserting the record in batch in the same
transaction, it would be more interesting to split the
transaction into several transactions:
Each transaction process N employees who have not yet been
processed, until there are no more. With the new cron
trigger mechanism, it is possible to:
- When the cron runs, it processes N employees
(N to be determined)
- If there is still some, he retriggers itself at the end
of his transaction
Like that, the cron is scheduled 1x / day, but it is
retriggered as many times as necessary each month.
Regarding the number of employees to process, with 100
employees, we can expect:
100 * 2 (morning / evening) * 21 (working days) = 4200
work entries to generate, which is manageable given
the above measures.
closesodoo/odoo#76488
Taskid: 2646056
Signed-off-by: Yannick Tivisse (yti) <yti@odoo.com>
Before this commit, the module dependencies: hr_contract and hr_work_entry were defined in community.
As they could be integrated with some community applications, like hr_attendance or hr_holidays, it is more coherent to move them to community.
Taskid: 2222790