Files
odoo_source/addons/website_event/models/website_visitor.py
T
Romain Derie d348bed1ad [IMP] website, *: use upsert to improve visitor perf
* im_livechat, test_event_full, website_blog, website_crm,
  website_event, website_event_track, website_event_track_quiz,
  webite_livechat, website_sale

There is 6 main changes in this commit:

1. Using raw SQL Upsert instead of the ORM methods. While raw SQL should
generally be avoided, it makes sense for such a low level behavior which
is impacting every flows.
Indeed, tracking visitors is a generic behavior done on all pages and
controllers. It is important to optimize it to reduce processing time
and SQL Queries.
Benchmark of that change alone:
> Rendering a tracked page improves from ~19.5ms to ~17ms (using `ab`
  with 1000 loop) and the requests involved in the tracking process are
  reduced from 8 SQL Queries to 3:
  - 1 request to upsert the visitor
  - 1 request to fetch the visitor data
  - 1 request to add the tracking record

2. Adding in that upsert query the `visitor.track` insert, creating both
records in one go, bringing the query count from 3 to 2.

3. Refactoring of the `parent_id` behavior that was introduced in stable
with [1]. The purpose was to keep track of multiple visitor linked to a
same user to merge the tracking together. Especially useful for tracking
a same visitor on different devices (when logged in).
Only one visitor was kept as active, others would be archived and their
tracks would be set/moved to the main partner.
Removing those duplicate visitor was not possible because those archived
duplicated visitor were holding the devices notification push token.
Since [2], those token were moved to their own table, all related to the
main visitor.
We can then now safely remove those duplicate visitors after merging
their track to the main visitor. Thus, the `parent_id` field is no more
useful. Removing it removes a layer of complexity.
Note that thanks to this part, the `active` field can also be removed.

4. Deeper functionnal change, inspired from Plausible: The access_token
is no more stored in a cookie but is the result of a hashing method
based on <IP Adress, User Agent>.
The reason behind that change is that, in an upcoming refactoring,
sessions won't be stored anymore unless absolutely needed (login, add to
cart..). It will also ship a no cookies policy, trying to get rid of all
cookies.
This change is bringing some functional changes:
- Since the IP is included in the hash to generate the token, it means
  that:
  A. If an anonymous user switch IP (eg from 4G to wifi), it is
     considered as a new visitor.
  B. If 2 anonymous users with the exact same user agent (same browser,
     same browser version, same exact os or phone) are on the same IP,
     those will be considered as the same visitor.
- Since the request host is not included in the hash, it means that
  visiting a DB from 2 differents URLs (domain and/or ip) on the same
  device and same browser will result in a shared visitor.
  It shouldn't imply any issue as this is A. not wrong and B. mostly
  used for tests.
As all this is only related to non logged in user, it shouldn't be a
real issue as anonymous visitors are not supposed to be meant to be
business critical, even if we use them for "a bit more" than simple
analytics data.

5. The access_token is now replaced by the partner_id once the user logs
in, so:
- We don't need to either search on the partner_id field or the
access_token field (depending if the user is logged in or not), we can
only use the access_token row/field to do both.
- On logout, everything works out of the box as the access_token will be
regenerated since there is no partner_id anymore.
- On login, if an access_token matches the user's partner_id, that
visitor is returned.
If there is no such token, a new visitor is created for that partner_id.
In both 2 cases, tracks are moved to that visitor and the anonymous
visitor is removed.
- We can remove the code that was in charge of checking if the
access_token / visitor cookie was wrong (coming from another user eg,
different user login on same device). Indeed, such collision is not
possible anymore as the access_token automatically match the logged in
user.
- We can remove the code that was in charge of checking if the
access_token / visitor cookie was wrong (coming from a logged in user
while the current visitor is not loggedin). Such collision is not
possible anymore as the access_token is (re)generated as an anonymous
token (hash) when not logged in.

6. There is no more check to prevent a track to be created if there was
already a track for that URL in the last 30 minutes.
While this can easily be re-introduced (one CTE on the upsert), it was
adding ~100ms (from ~20 to ~110ms) to the request on a big database as
Odoo where there is ~100 millions tracks and ~100 millions visitors.
It has been validated that it was not a real issue as it is not
fundamentally wrong. If a visitor visited 20 times a product or a
specific page in that short amount of time, you might want to know that
because the user is most likely interested by it.

Changes (1+2), 3, (4+5) and 6 are all independant from each other and
could have existed on their own.

[1]: https://github.com/odoo/odoo/commit/c6b8a44b970a46dcd87a4e2cb1ad52fa340b209f
[2]: https://github.com/odoo/enterprise/pull/16781/commits/f75090fe8b42484e89e933976e8441d2f5eb9415

task-2867045

closes odoo/odoo#87857

Related: odoo/enterprise#28004
Related: odoo/upgrade#3566
Signed-off-by: Romain Derie (rde) <rde@odoo.com>
2022-06-07 16:31:20 +02:00

96 lines
4.6 KiB
Python

# -*- coding: utf-8 -*-
# Part of Odoo. See LICENSE file for full copyright and licensing details.
from odoo import api, fields, models
from odoo.osv import expression
class WebsiteVisitor(models.Model):
_name = 'website.visitor'
_inherit = ['website.visitor']
event_registration_ids = fields.One2many(
'event.registration', 'visitor_id', string='Event Registrations',
groups="event.group_event_registration_desk")
event_registration_count = fields.Integer(
'# Registrations', compute='_compute_event_registration_count',
groups="event.group_event_registration_desk")
event_registered_ids = fields.Many2many(
'event.event', string="Registered Events",
compute="_compute_event_registered_ids", compute_sudo=True,
search="_search_event_registered_ids",
groups="event.group_event_registration_desk")
@api.depends('partner_id, event_registration_ids.name')
def name_get(self):
""" If there is an event registration for an anonymous visitor, use that
registered attendee name as visitor name. """
res_dict = dict(super().name_get())
# sudo is needed for `event_registration_ids`
for visitor in self.sudo().filtered(lambda v: not v.partner_id and v.event_registration_ids):
res_dict[visitor.id] = visitor.event_registration_ids[-1].name
return list(res_dict.items())
@api.depends('event_registration_ids')
def _compute_event_registration_count(self):
if self.ids:
read_group_res = self.env['event.registration']._read_group(
[('visitor_id', 'in', self.ids)],
['visitor_id'], ['visitor_id'])
visitor_mapping = dict(
(item['visitor_id'][0], item['visitor_id_count'])
for item in read_group_res)
else:
visitor_mapping = dict()
for visitor in self:
visitor.event_registration_count = visitor_mapping.get(visitor.id) or 0
@api.depends('event_registration_ids.email', 'event_registration_ids.mobile', 'event_registration_ids.phone')
def _compute_email_phone(self):
super(WebsiteVisitor, self)._compute_email_phone()
for visitor in self.filtered(lambda visitor: not visitor.email or not visitor.mobile):
linked_registrations = visitor.event_registration_ids.sorted(lambda reg: (reg.create_date, reg.id), reverse=False)
if not visitor.email:
visitor.email = next((reg.email for reg in linked_registrations if reg.email), False)
if not visitor.mobile:
visitor.mobile = next((reg.mobile or reg.phone for reg in linked_registrations if reg.mobile or reg.phone), False)
@api.depends('event_registration_ids')
def _compute_event_registered_ids(self):
# include parent's registrations in a visitor o2m field. We don't add
# child one as child should not have registrations (moved to the parent)
for visitor in self:
all_registrations = visitor.event_registration_ids
visitor.event_registered_ids = all_registrations.mapped('event_id')
def _search_event_registered_ids(self, operator, operand):
""" Search visitors with terms on events within their event registrations. E.g. [('event_registered_ids',
'in', [1, 2])] should return visitors having a registration on events 1, 2 as
well as their children for notification purpose. """
if operator == "not in":
raise NotImplementedError("Unsupported 'Not In' operation on visitors registrations")
all_registrations = self.env['event.registration'].sudo().search([
('event_id', operator, operand)
])
if all_registrations:
visitor_ids = all_registrations.with_context(active_test=False).visitor_id.ids
else:
visitor_ids = []
return [('id', 'in', visitor_ids)]
def _inactive_visitors_domain(self):
""" Visitors registered to events are considered always active and should not be deleted. """
domain = super()._inactive_visitors_domain()
return expression.AND([domain, [('event_registration_ids', '=', False)]])
def _merge_visitor(self, target):
""" Override linking process to link registrations to the final visitor. """
self.event_registration_ids.visitor_id = target.id
registration_wo_partner = self.event_registration_ids.filtered(lambda registration: not registration.partner_id)
if registration_wo_partner:
registration_wo_partner.partner_id = target.partner_id
return super()._merge_visitor(target)