Files
odoo_source/addons/website/tools.py
T
Julien Castiaux 06cc322e7e [REF] core: HTTPocalypse (13) http_routing/website
This commit is the 13th commit of a comprehensive refactor of our HTTP
framework. See odoo/odoo#78857 for complete historic, discussions and
rationnals.

Here be dragons.

First and foremost, `http_routing` is a technical module that aim to
provide the minimum viable compatibility code between portal and
website. Its primary job is to take care of the lang inserted in the
path of URLs, e.g. `en` in `/en/my_blog`. Both when routing a request
with a lang in the URL and when rendering templates with multilang
support.

Next to `http_routing` is website, the module used by customers to
create pretty web page accessible online. Website uses a different
routing logic than the backend, webpages are **not** registered in the
routing map of werkzeug but instead delivered by website dedicated
code. It means that **all** request targeting a website page thrown at
the werkzeug router **fail** with a HTTP 404 error. The reality is that
website abuses the fallback mechanism (`_serve_fallback`) to deliver
its pages.

Using `_serve_fallback` as the standard way to deliver pages is broken
by design. It is the least crappy way to deliver content as long as
website page don't have a dedicated path prefix. Using a path prefix
(e.g. `p` in `/p/fr/mon_blog`) it would have been possible to route the
request to a dedicated endpoint using the same router as the backend and
with no change to the HTTP dispatching code. Sadly, the business does
not want such prefix so we have to stick with a broken design.

It is broken because prior to serving a page, website needs to setup A
LOT of stuff on the system. It needs to ensure a proper user is set on
the environment, it needs to setup the GeoIP database, it also needs
to determine the lang the user requested the page and it also needs to
save multiple attributes on the request objet itself (`is_frontend`,
`is_frontend_multilang`, `routing_iteration`, `website`, `lang`,
`rerouting` and `website_routing`). Since `base/ir.http@_match()` will
fail, everything must be set either prior to calling this method or in
`website/ir.http@_serve_fallback`.

---

The original implementation was overriding the `_dispatch` method which
was responsible to call the four `_match()`, `_authenticate()`,
`_postprocess_args` and finally `WebRequest._dispatch()`. The override
was very special, here is an attempt to explain it:

1) try to match an endpoint using the backend router, 404-errors are
   ignored.
2) include the geoip stuff.
3) authenticate using the `auth` @route argument if an endpoint matched
   in (1), otherwise authenticate with the public user.
4) if not endpoint matched or if a frontend endpoint matched in (1):
  a) call `_add_dispatch_parameters` which sets many arguments on the
     `request` object, including the lang found in the request cookies
  b) try to extract a lang from the URL: abort with a redirection when
     the lang is missing or wrong, remove the lang from the request
     path when it is set (updating both `request.lang` and the cookie).
5) return the result of `super()._dispatch()`

Note that `_serve_fallback()` is called during `super()._dispatch()`
when the path still does not route to an endpoint. Thanks to the
`_dispatch` overrides in http_routing and website, it is garanteed that
the system is setup prior to calling `_serve_fallback`.

---

Because it is now `http.py@Request._serve_ir_http()` that is responsible
of calling the four`_match()`, `_authenticate()`, `_pre_dispatch` and
finally `_(http|json)_dispatch()` it is no more possible to override it
to take over the dispatching to perform the http_routing/website magic.
The prior implementation can not work with the new design thus is has
been refactored too.

To render a website page, there must be a user configured on the
environment (not None) and the various special attributes must be set on
the request object. The special `lang` attribute is popped from the
request path but backend endpoints must be delivered in priority.

Using the new design, it has been decided to override the `_match`
method to implement the lang-in-path logic, to override both
`_pre_dispatch` and `_serve_fallback` to call `_add_dispatch_parameters`
which have been renamed `_frontend_pre_dispatch` and to also grant the
public user in the `_serve_fallback` override.

The `_handle_error` override in website has similar needs, the function
is called upon error (4xx/5xx) in order to render a pretty
website-looking page. Because such error can occurs as early as in
`_match()` (page not found and no fallback), when nothing has been setup
yet, website is yet again responsible for setuping everything: request,
orm, frontend.

Many other small improvement are not described here. Hopefully the added
comments in the source code are enought.

PR: odoo#78857
Task: 2571224
2022-02-24 13:30:50 +00:00

157 lines
4.9 KiB
Python

# Part of Odoo. See LICENSE file for full copyright and licensing details.
import contextlib
from lxml import etree
from unittest.mock import Mock, MagicMock, patch
from werkzeug.exceptions import NotFound
from werkzeug.test import EnvironBuilder
import odoo
from odoo.tests.common import HttpCase, HOST
from odoo.tools.misc import DotDict, frozendict
@contextlib.contextmanager
def MockRequest(
env, *, path='/mockrequest/', routing=True, multilang=True,
context=frozendict(), cookies=frozendict(), country_code=None,
website=None, sale_order_id=None, website_sale_current_pl=None,
):
lang_code = context.get('lang', env.context.get('lang', 'en_US'))
env = env(context=dict(context, lang=lang_code))
request = Mock(
# request
httprequest=Mock(
host='localhost',
path=path,
app=odoo.http.root,
environ=dict(
EnvironBuilder(
path=path,
base_url=HttpCase.base_url()
).get_environ(),
REMOTE_ADDR=HOST,
),
cookies=cookies,
referrer='',
),
type='http',
future_response=odoo.http.FutureResponse(),
params={},
redirect=env['ir.http']._redirect,
session=DotDict(
odoo.http.DEFAULT_SESSION,
geoip={'country_code': country_code},
sale_order_id=sale_order_id,
website_sale_current_pl=website_sale_current_pl,
),
db=None,
env=env,
registry=env.registry,
cr=env.cr,
uid=env.uid,
context=env.context,
lang=env['res.lang']._lang_get(lang_code),
website=website,
)
if website:
request.website_routing = website.id
# The following code mocks match() to return a fake rule with a fake
# 'routing' attribute (routing=True) or to raise a NotFound
# exception (routing=False).
#
# router = odoo.http.root.get_db_router()
# rule, args = router.bind(...).match(path)
# # arg routing is True => rule.endpoint.routing == {...}
# # arg routing is False => NotFound exception
router = MagicMock()
match = router.return_value.bind.return_value.match
if routing:
match.return_value[0].routing = {
'type': 'http',
'website': True,
'multilang': multilang
}
else:
match.side_effect = NotFound
with contextlib.ExitStack() as s:
odoo.http._request_stack.push(request)
s.callback(odoo.http._request_stack.pop)
s.enter_context(patch('odoo.http.root.get_db_router', router))
yield request
# Fuzzy matching tools
def distance(s1="", s2="", limit=4):
"""
Limited Levenshtein-ish distance (inspired from Apache text common)
Note: this does not return quick results for simple cases (empty string, equal strings)
those checks should be done outside loops that use this function.
:param s1: first string
:param s2: second string
:param limit: maximum distance to take into account, return -1 if exceeded
:return: number of character changes needed to transform s1 into s2 or -1 if this exceeds the limit
"""
BIG = 100000 # never reached integer
if len(s1) > len(s2):
s1, s2 = s2, s1
l1 = len(s1)
l2 = len(s2)
if l2 - l1 > limit:
return -1
boundary = min(l1, limit) + 1
p = [i if i < boundary else BIG for i in range(0, l1 + 1)]
d = [BIG for _ in range(0, l1 + 1)]
for j in range(1, l2 + 1):
j2 = s2[j -1]
d[0] = j
range_min = max(1, j - limit)
range_max = min(l1, j + limit)
if range_min > 1:
d[range_min -1] = BIG
for i in range(range_min, range_max + 1):
if s1[i - 1] == j2:
d[i] = p[i - 1]
else:
d[i] = 1 + min(d[i - 1], p[i], p[i - 1])
p, d = d, p
return p[l1] if p[l1] <= limit else -1
def similarity_score(s1, s2):
"""
Computes a score that describes how much two strings are matching.
:param s1: first string
:param s2: second string
:return: float score, the higher the more similar
pairs returning non-positive scores should be considered non similar
"""
dist = distance(s1, s2)
if dist == -1:
return -1
set1 = set(s1)
score = len(set1.intersection(s2)) / len(set1)
score -= dist / len(s1)
score -= len(set1.symmetric_difference(s2)) / (len(s1) + len(s2))
return score
def text_from_html(html_fragment):
"""
Returns the plain non-tag text from an html
:param html_fragment: document from which text must be extracted
:return: text extracted from the html
"""
# lxml requires one single root element
tree = etree.fromstring('<p>%s</p>' % html_fragment, etree.XMLParser(recover=True))
return ' '.join(tree.itertext())