[FIX] base: serialization bug when acquiring cron

Start Odoo with multiple cron threads (e.g. --max-cron-threads=4),
trigger many crons at once, there is a chance one of the cron thread
fails due to a serialization error.

Inside of the `_acquire_one_job` function, the query evaluates many rows
to find one that fit multiple requirements. Two of those requirements
are (1) that the `nextcall` of the row is in the past or (2) that it
exists a cron trigger for that cron with a `call_at` in the past.

In case the `nextcall` of one of those rows is modified or the cron
triggers are removed by another transaction then there can be a
serialisation failure in the current transaction. This serialisation
error is important, it prevents the current cron worker from acquiring a
cron job that has been processed in another cron worker.

The problem is that that postgres doesn't tell which row was modified by
the other transaction (=processed by another worker cron) so it is not
possible to just skip that cron and continue with the others.

Our solution is to limit the WHERE clause of the `_acquire_one_job`
function to a single row. In case there is a serialization failure we
know the cron was processed in another job and we can skip it.

Closes #96584

closes odoo/odoo#96926

X-original-commit: 684750a0c2c83acffb32e555bd3140a2c76d7219
Signed-off-by: Julien Castiaux <juc@odoo.com>
Co-authored-by: Julien Castiaux <juc@odoo.com>
This commit is contained in:
Lin Wenwen
2022-07-28 10:27:58 +02:00
committed by Julien Castiaux
co-authored by Julien Castiaux
parent d3c935ff27
commit c06cee44fe
2 changed files with 42 additions and 8 deletions
+11
View File
@@ -0,0 +1,11 @@
China, 05-28-2022
I hereby agree to the terms of the Odoo Individual Contributor License
Agreement v1.0.
I declare that I am authorized and able to make this agreement and sign this
declaration.
Signed,
Lin Wenwen 514706098@qq.com https://github.com/lin-ww
+31 -8
View File
@@ -99,17 +99,22 @@ class ir_cron(models.Model):
if not jobs:
return
cls._check_modules_state(cron_cr, jobs)
job_ids = tuple([job['id'] for job in jobs])
while True:
job = cls._acquire_one_job(cron_cr, job_ids)
for job_id in (job['id'] for job in jobs):
try:
job = cls._acquire_one_job(cron_cr, (job_id,))
except psycopg2.extensions.TransactionRollbackError:
cron_cr.rollback()
_logger.debug("job %s has been processed by another worker, skip", job_id)
continue
if not job:
break
_logger.debug("job %s acquired", job['id'])
_logger.debug("another worker is processing job %s, skip", job_id)
continue
_logger.debug("job %s acquired", job_id)
# take into account overridings of _process_job() on that database
registry = odoo.registry(db_name)
registry[cls._name]._process_job(db, cron_cr, job)
_logger.debug("job %s updated and released", job['id'])
_logger.debug("job %s updated and released", job_id)
except BadVersion:
_logger.warning('Skipping database %s as its base version is not %s.', db_name, BASE_VERSION)
@@ -191,7 +196,17 @@ class ir_cron(models.Model):
@classmethod
def _acquire_one_job(cls, cr, job_ids):
""" Acquire one job for update from the job_ids tuple. """
"""
Acquire for update one job that is ready from the job_ids tuple.
The jobs that have already been processed in this worker should
be excluded from the tuple.
This function raises a ``psycopg2.errors.SerializationFailure``
when the ``nextcall`` of one of the job_ids is modified in
another transaction. You should rollback the transaction and try
again later.
"""
# We have to make sure ALL jobs are executed ONLY ONCE no matter
# how many cron workers may process them. The exlusion mechanism
@@ -204,7 +219,15 @@ class ir_cron(models.Model):
# the other workers don't select it too.
# (ii) is implemented via the `WHERE` statement, when a job has
# been processed, its nextcall is updated to a date in the
# future and the optionnal trigger is removed.
# future and the optional triggers are removed.
#
# Note about (ii): it is possible that a job becomes available
# again quickly (e.g. high frequency or self-triggering cron).
# This function doesn't prevent from acquiring that job multiple
# times at different moments. This can block a worker on
# executing a same job in loop. To prevent this problem, the
# callee is responsible of providing a `job_ids` tuple without
# the jobs it has executed already.
#
# An `UPDATE` lock type is the strongest row lock, it conflicts
# with ALL other lock types. Among them the `KEY SHARE` row lock