Files
odoo_source/addons/website/models/ir_http.py
T
Romain Derie d348bed1ad [IMP] website, *: use upsert to improve visitor perf
* im_livechat, test_event_full, website_blog, website_crm,
  website_event, website_event_track, website_event_track_quiz,
  webite_livechat, website_sale

There is 6 main changes in this commit:

1. Using raw SQL Upsert instead of the ORM methods. While raw SQL should
generally be avoided, it makes sense for such a low level behavior which
is impacting every flows.
Indeed, tracking visitors is a generic behavior done on all pages and
controllers. It is important to optimize it to reduce processing time
and SQL Queries.
Benchmark of that change alone:
> Rendering a tracked page improves from ~19.5ms to ~17ms (using `ab`
  with 1000 loop) and the requests involved in the tracking process are
  reduced from 8 SQL Queries to 3:
  - 1 request to upsert the visitor
  - 1 request to fetch the visitor data
  - 1 request to add the tracking record

2. Adding in that upsert query the `visitor.track` insert, creating both
records in one go, bringing the query count from 3 to 2.

3. Refactoring of the `parent_id` behavior that was introduced in stable
with [1]. The purpose was to keep track of multiple visitor linked to a
same user to merge the tracking together. Especially useful for tracking
a same visitor on different devices (when logged in).
Only one visitor was kept as active, others would be archived and their
tracks would be set/moved to the main partner.
Removing those duplicate visitor was not possible because those archived
duplicated visitor were holding the devices notification push token.
Since [2], those token were moved to their own table, all related to the
main visitor.
We can then now safely remove those duplicate visitors after merging
their track to the main visitor. Thus, the `parent_id` field is no more
useful. Removing it removes a layer of complexity.
Note that thanks to this part, the `active` field can also be removed.

4. Deeper functionnal change, inspired from Plausible: The access_token
is no more stored in a cookie but is the result of a hashing method
based on <IP Adress, User Agent>.
The reason behind that change is that, in an upcoming refactoring,
sessions won't be stored anymore unless absolutely needed (login, add to
cart..). It will also ship a no cookies policy, trying to get rid of all
cookies.
This change is bringing some functional changes:
- Since the IP is included in the hash to generate the token, it means
  that:
  A. If an anonymous user switch IP (eg from 4G to wifi), it is
     considered as a new visitor.
  B. If 2 anonymous users with the exact same user agent (same browser,
     same browser version, same exact os or phone) are on the same IP,
     those will be considered as the same visitor.
- Since the request host is not included in the hash, it means that
  visiting a DB from 2 differents URLs (domain and/or ip) on the same
  device and same browser will result in a shared visitor.
  It shouldn't imply any issue as this is A. not wrong and B. mostly
  used for tests.
As all this is only related to non logged in user, it shouldn't be a
real issue as anonymous visitors are not supposed to be meant to be
business critical, even if we use them for "a bit more" than simple
analytics data.

5. The access_token is now replaced by the partner_id once the user logs
in, so:
- We don't need to either search on the partner_id field or the
access_token field (depending if the user is logged in or not), we can
only use the access_token row/field to do both.
- On logout, everything works out of the box as the access_token will be
regenerated since there is no partner_id anymore.
- On login, if an access_token matches the user's partner_id, that
visitor is returned.
If there is no such token, a new visitor is created for that partner_id.
In both 2 cases, tracks are moved to that visitor and the anonymous
visitor is removed.
- We can remove the code that was in charge of checking if the
access_token / visitor cookie was wrong (coming from another user eg,
different user login on same device). Indeed, such collision is not
possible anymore as the access_token automatically match the logged in
user.
- We can remove the code that was in charge of checking if the
access_token / visitor cookie was wrong (coming from a logged in user
while the current visitor is not loggedin). Such collision is not
possible anymore as the access_token is (re)generated as an anonymous
token (hash) when not logged in.

6. There is no more check to prevent a track to be created if there was
already a track for that URL in the last 30 minutes.
While this can easily be re-introduced (one CTE on the upsert), it was
adding ~100ms (from ~20 to ~110ms) to the request on a big database as
Odoo where there is ~100 millions tracks and ~100 millions visitors.
It has been validated that it was not a real issue as it is not
fundamentally wrong. If a visitor visited 20 times a product or a
specific page in that short amount of time, you might want to know that
because the user is most likely interested by it.

Changes (1+2), 3, (4+5) and 6 are all independant from each other and
could have existed on their own.

[1]: https://github.com/odoo/odoo/commit/c6b8a44b970a46dcd87a4e2cb1ad52fa340b209f
[2]: https://github.com/odoo/enterprise/pull/16781/commits/f75090fe8b42484e89e933976e8441d2f5eb9415

task-2867045

closes odoo/odoo#87857

Related: odoo/enterprise#28004
Related: odoo/upgrade#3566
Signed-off-by: Romain Derie (rde) <rde@odoo.com>
2022-06-07 16:31:20 +02:00

403 lines
16 KiB
Python

# Part of Odoo. See LICENSE file for full copyright and licensing details.
import contextlib
import logging
from lxml import etree
import os
import unittest
import time
import pytz
import werkzeug
import werkzeug.routing
import werkzeug.utils
from functools import partial
import odoo
from odoo import api, models
from odoo import SUPERUSER_ID
from odoo.exceptions import AccessError, MissingError
from odoo.http import request
from odoo.tools.safe_eval import safe_eval
from odoo.osv.expression import FALSE_DOMAIN
from odoo.addons.http_routing.models import ir_http
from odoo.addons.http_routing.models.ir_http import _guess_mimetype
from odoo.addons.portal.controllers.portal import _build_url_w_params
logger = logging.getLogger(__name__)
def sitemap_qs2dom(qs, route, field='name'):
""" Convert a query_string (can contains a path) to a domain"""
dom = []
if qs and qs.lower() not in route:
needles = qs.strip('/').split('/')
# needles will be altered and keep only element which one is not in route
# diff(from=['shop', 'product'], to=['shop', 'product', 'product']) => to=['product']
unittest.util.unorderable_list_difference(route.strip('/').split('/'), needles)
if len(needles) == 1:
dom = [(field, 'ilike', needles[0])]
else:
dom = FALSE_DOMAIN
return dom
def get_request_website():
""" Return the website set on `request` if called in a frontend context
(website=True on route).
This method can typically be used to check if we are in the frontend.
This method is easy to mock during python tests to simulate frontend
context, rather than mocking every method accessing request.website.
Don't import directly the method or it won't be mocked during tests, do:
```
from odoo.addons.website.models import ir_http
my_var = ir_http.get_request_website()
```
"""
return request and getattr(request, 'website', False) or False
class Http(models.AbstractModel):
_inherit = 'ir.http'
@classmethod
def routing_map(cls, key=None):
key = key or (request and request.website_routing)
return super(Http, cls).routing_map(key=key)
@classmethod
def clear_caches(cls):
super()._clear_routing_map()
return super().clear_caches()
@classmethod
def _slug_matching(cls, adapter, endpoint, **kw):
for arg in kw:
if isinstance(kw[arg], models.BaseModel):
kw[arg] = kw[arg].with_context(slug_matching=True)
qs = request.httprequest.query_string.decode('utf-8')
return adapter.build(endpoint, kw) + (qs and '?%s' % qs or '')
@classmethod
def _generate_routing_rules(cls, modules, converters):
website_id = request.website_routing
logger.debug("_generate_routing_rules for website: %s", website_id)
domain = [('redirect_type', 'in', ('308', '404')), '|', ('website_id', '=', False), ('website_id', '=', website_id)]
rewrites = dict([(x.url_from, x) for x in request.env['website.rewrite'].sudo().search(domain)])
cls._rewrite_len[website_id] = len(rewrites)
for url, endpoint in super()._generate_routing_rules(modules, converters):
if url in rewrites:
rewrite = rewrites[url]
url_to = rewrite.url_to
if rewrite.redirect_type == '308':
logger.debug('Add rule %s for %s' % (url_to, website_id))
yield url_to, endpoint # yield new url
if url != url_to:
logger.debug('Redirect from %s to %s for website %s' % (url, url_to, website_id))
_slug_matching = partial(cls._slug_matching, endpoint=endpoint)
endpoint.routing['redirect_to'] = _slug_matching
yield url, endpoint # yield original redirected to new url
elif rewrite.redirect_type == '404':
logger.debug('Return 404 for %s for website %s' % (url, website_id))
continue
else:
yield url, endpoint
@classmethod
def _get_converters(cls):
""" Get the converters list for custom url pattern werkzeug need to
match Rule. This override adds the website ones.
"""
return dict(
super()._get_converters(),
model=ModelConverter,
)
@classmethod
def _auth_method_public(cls):
""" If no user logged, set the public user of current website, or default
public user as request uid.
"""
if not request.session.uid:
website = request.env(user=SUPERUSER_ID)['website'].get_current_website() # sudo
if website:
request.update_env(user=website._get_cached('user_id'))
if not request.uid:
super()._auth_method_public()
@classmethod
def _register_website_track(cls, response):
if request.env['ir.http'].is_a_bot():
return False
if getattr(response, 'status_code', 0) != 200 or request.httprequest.headers.get('X-Disable-Tracking') == '1':
return False
template = False
if hasattr(response, '_cached_page'):
website_page, template = response._cached_page, response._cached_template
elif hasattr(response, 'qcontext'): # classic response
main_object = response.qcontext.get('main_object')
website_page = getattr(main_object, '_name', False) == 'website.page' and main_object
template = response.qcontext.get('response_template')
view = template and request.env['website'].get_template(template)
if view and view.track:
request.env['website.visitor']._handle_webpage_dispatch(website_page)
return False
@classmethod
def _match(cls, path):
if not hasattr(request, 'website_routing'):
website = request.env['website'].get_current_website()
request.website_routing = website.id
return super()._match(path)
@classmethod
def _pre_dispatch(cls, rule, arguments):
super()._pre_dispatch(rule, arguments)
for record in arguments.values():
if isinstance(record, models.BaseModel) and hasattr(record, 'can_access_from_current_website'):
try:
if not record.can_access_from_current_website():
raise werkzeug.exceptions.NotFound()
except AccessError:
# record.website_id might not be readable as
# unpublished `event.event` due to ir.rule, return
# 403 instead of using `sudo()` for perfs as this is
# low level.
raise werkzeug.exceptions.Forbidden()
@classmethod
def _get_web_editor_context(cls):
ctx = super()._get_web_editor_context()
if request.is_frontend_multilang and request.lang == cls._get_default_lang():
ctx['edit_translations'] = False
return ctx
@classmethod
def _frontend_pre_dispatch(cls):
super()._frontend_pre_dispatch()
if not request.context.get('tz'):
with contextlib.suppress(pytz.UnknownTimeZoneError):
tz = request.geoip.get('time_zone', '')
request.update_context(tz=pytz.timezone(tz).zone)
website = request.env['website'].get_current_website()
user = request.env.user
# This is mainly to avoid access errors in website controllers
# where there is no context (eg: /shop), and it's not going to
# propagate to the global context of the tab. If the company of
# the website is not in the allowed companies of the user, set
# the main company of the user.
website_company_id = website._get_cached('company_id')
if user.id == website._get_cached('user_id'):
# avoid a read on res_company_user_rel in case of public user
allowed_company_ids = [website_company_id]
elif website_company_id in user._get_company_ids():
allowed_company_ids = [website_company_id]
else:
allowed_company_ids = user.company_id.ids
request.update_context(
allowed_company_ids=allowed_company_ids,
website_id=website.id,
**cls._get_web_editor_context(),
)
request.website = website.with_context(request.context)
@classmethod
def _dispatch(cls, endpoint):
response = super()._dispatch(endpoint)
cls._register_website_track(response)
return response
@classmethod
def _get_frontend_langs(cls):
# _get_frontend_langs() is used by @http_routing:IrHttp._match
# where is_frontend is not yet set and when no backend endpoint
# matched. We have to assume we are going to match a frontend
# route, hence the default True. Elsewhere, request.is_frontend
# is set.
if getattr(request, 'is_frontend', True):
website_id = request.env.get('website_id', request.website_routing)
res_lang = request.env['res.lang'].with_context(website_id=website_id)
return [code for code, *_ in res_lang.get_available()]
else:
return super()._get_frontend_langs()
@classmethod
def _get_default_lang(cls):
if getattr(request, 'is_frontend', True):
website = request.env['website'].sudo().get_current_website()
return request.env['res.lang'].browse([website._get_cached('default_lang_id')])
return super()._get_default_lang()
@classmethod
def _get_translation_frontend_modules_name(cls):
mods = super()._get_translation_frontend_modules_name()
installed = request.registry._init_modules.union(odoo.conf.server_wide_modules)
return mods + [mod for mod in installed if mod.startswith('website')]
@classmethod
def _serve_page(cls):
req_page = request.httprequest.path
page_domain = [('url', '=', req_page)] + request.website.website_domain()
published_domain = page_domain
# specific page first
page = request.env['website.page'].sudo().search(published_domain, order='website_id asc', limit=1)
# redirect withtout trailing /
if not page and req_page != "/" and req_page.endswith("/"):
# mimick `_postprocess_args()` redirect
path = request.httprequest.path[:-1]
if request.lang != cls._get_default_lang():
path = '/' + request.lang.url_code + path
if request.httprequest.query_string:
path += '?' + request.httprequest.query_string.decode('utf-8')
return request.redirect(path, code=301)
if page and (request.website.is_publisher() or page.is_visible):
_, ext = os.path.splitext(req_page)
response = request.render(page.view_id.id, {
'deletable': True,
'main_object': page,
}, mimetype=_guess_mimetype(ext))
return response
return False
@classmethod
def _serve_redirect(cls):
req_page = request.httprequest.path
domain = [
('redirect_type', 'in', ('301', '302')),
# trailing / could have been removed by server_page
'|', ('url_from', '=', req_page.rstrip('/')), ('url_from', '=', req_page + '/')
]
domain += request.website.website_domain()
return request.env['website.rewrite'].sudo().search(domain, limit=1)
@classmethod
def _serve_fallback(cls):
# serve attachment before
parent = super()._serve_fallback()
if parent: # attachment
return parent
# minimal setup to serve frontend pages
if not request.uid:
cls._auth_method_public()
cls._frontend_pre_dispatch()
cls._handle_debug()
request.params = request.get_http_params()
website_page = cls._serve_page()
if website_page:
website_page.flatten()
cls._register_website_track(website_page)
cls._post_dispatch(website_page)
return website_page
redirect = cls._serve_redirect()
if redirect:
return request.redirect(
_build_url_w_params(redirect.url_to, request.params),
code=redirect.redirect_type,
local=False) # safe because only designers can specify redirects
@classmethod
def _get_exception_code_values(cls, exception):
code, values = super()._get_exception_code_values(exception)
if isinstance(exception, werkzeug.exceptions.NotFound) and request.website.is_publisher():
code = 'page_404'
values['path'] = request.httprequest.path[1:]
if isinstance(exception, werkzeug.exceptions.Forbidden) and \
exception.description == "website_visibility_password_required":
code = 'protected_403'
values['path'] = request.httprequest.path
return (code, values)
@classmethod
def _get_values_500_error(cls, env, values, exception):
View = env["ir.ui.view"]
values = super()._get_values_500_error(env, values, exception)
if 'qweb_exception' in values:
try:
# exception.name might be int, string
exception_template = int(exception.name)
except ValueError:
exception_template = exception.name
view = View._view_obj(exception_template)
if exception.html and exception.html in view.arch:
values['view'] = view
else:
# There might be 2 cases where the exception code can't be found
# in the view, either the error is in a child view or the code
# contains branding (<div t-att-data="request.browse('ok')"/>).
et = view.with_context(inherit_branding=False)._get_combined_arch()
node = et.xpath(exception.path) if exception.path else et
line = node is not None and len(node) > 0 and etree.tostring(node[0], encoding='unicode')
if line:
values['view'] = View._views_get(exception_template).filtered(
lambda v: line in v.arch
)
values['view'] = values['view'] and values['view'][0]
# Needed to show reset template on translated pages (`_prepare_environment` will set it for main lang)
values['editable'] = request.uid and request.website.is_publisher()
return values
@classmethod
def _get_error_html(cls, env, code, values):
if code in ('page_404', 'protected_403'):
return code.split('_')[1], env['ir.ui.view']._render_template('website.%s' % code, values)
return super()._get_error_html(env, code, values)
@api.model
def get_frontend_session_info(self):
session_info = super(Http, self).get_frontend_session_info()
geoip_country_code = request.geoip.get('country_code')
geoip_phone_code = request.env['res.country']._phone_code_for(geoip_country_code) if geoip_country_code else None
session_info.update({
'is_website_user': request.env.user.id == request.website.user_id.id,
'geoip_country_code': geoip_country_code,
'geoip_phone_code': geoip_phone_code,
})
if request.env.user.has_group('website.group_website_publisher'):
session_info.update({
'website_id': request.website.id,
'website_company_id': request.website._get_cached('company_id'),
})
return session_info
class ModelConverter(ir_http.ModelConverter):
def to_url(self, value):
if value.env.context.get('slug_matching'):
return value.env.context.get('_converter_value', str(value.id))
return super().to_url(value)
def generate(self, env, args, dom=None):
Model = env[self.model]
# Allow to current_website_id directly in route domain
args['current_website_id'] = env['website'].get_current_website().id
domain = safe_eval(self.domain, args)
if dom:
domain += dom
for record in Model.search(domain):
# return record so URL will be the real endpoint URL as the record will go through `slug()`
# the same way as endpoint URL is retrieved during dispatch (301 redirect), see `to_url()` from ModelConverter
yield record