[IMP] website: list pages before controllers in sitemap

This shouldn't change anything regarding SEO, but is worth a try.
It will also impact the link suggestion when creating a link in the
editor, but having pages listed first also have sense there, or at least
it won't be worst.

The idea for the SEO part is that if the sitemap order (if too long)
would have some impact as crawlers might have a limited crawling budget
for your website and it will only crawl the first pages it finds.
Are those "first pages" impacted by the sitemap order? It's almost sure
it's not, as Google definitely knows how to crawl on its own, and is
even probably ignoring the sitemap most of the time.

Also, pages:
- Are probably always important content since you created manually a page to
  write something, while (some) controllers might just be content you
  care less about. Pages are probably always important while we can't
  say that for controllers.
- Should be fewer in number than controllers most of the time
- Have a lastmod set, as opposed to controllers
- May be created at any time, meaning a new crawler visit is needed,
  while controllers are almost never added in production, installing a
  new module is something very rare. Exception is about record's
  controllers which are "created" at any time like pages (eg a new
  product controller page)

For all those reasons, this commit reverse the pages vs controllers
order in the sitemap.

This is coming from our prod where some pages are yet not indexed while
they have been published months ago.

closes odoo/odoo#152350

Signed-off-by: Jérémy Kersten <jke@odoo.com>
This commit is contained in:
Romain Derie
2024-02-06 18:49:04 +00:00
parent 22c5940706
commit 4a0087af5a
+25 -23
View File
@@ -1280,6 +1280,31 @@ class Website(models.Model):
of the same.
:rtype: list({name: str, url: str})
"""
# ==== WEBSITE.PAGES ====
# '/' already has a http.route & is in the routing_map so it will already have an entry in the xml
domain = [('url', '!=', '/')]
if not force:
domain += [('website_indexed', '=', True), ('visibility', '=', False)]
# is_visible
domain += [
('website_published', '=', True), ('visibility', '=', False),
'|', ('date_publish', '=', False), ('date_publish', '<=', fields.Datetime.now())
]
if query_string:
domain += [('url', 'like', query_string)]
pages = self._get_website_pages(domain)
for page in pages:
record = {'loc': page['url'], 'id': page['id'], 'name': page['name']}
if page.view_id and page.view_id.priority != 16:
record['priority'] = min(round(page.view_id.priority / 32.0, 1), 1)
if page['write_date']:
record['lastmod'] = page['write_date'].date()
yield record
# ==== CONTROLLERS ====
router = self.env['ir.http'].routing_map()
url_set = set()
@@ -1344,29 +1369,6 @@ class Website(models.Model):
yield page
# '/' already has a http.route & is in the routing_map so it will already have an entry in the xml
domain = [('url', '!=', '/')]
if not force:
domain += [('website_indexed', '=', True), ('visibility', '=', False)]
# is_visible
domain += [
('website_published', '=', True), ('visibility', '=', False),
'|', ('date_publish', '=', False), ('date_publish', '<=', fields.Datetime.now())
]
if query_string:
domain += [('url', 'like', query_string)]
pages = self._get_website_pages(domain)
for page in pages:
record = {'loc': page['url'], 'id': page['id'], 'name': page['name']}
if page.view_id and page.view_id.priority != 16:
record['priority'] = min(round(page.view_id.priority / 32.0, 1), 1)
if page['write_date']:
record['lastmod'] = page['write_date'].date()
yield record
def get_website_page_ids(self):
if not self.env.user.has_group('website.group_website_restricted_editor'):
# Note that `website.pages` have `0,0,0,0` ACL rights by default for