Files
odoo_source/addons/crm/security
David Beguin e20cc70d08 [IMP] crm : add and apply Predictive Lead Scoring
This new Predictive Lead Scoring (aka PLS) replaces the stage based probability mechanism.
It uses Naive Bayes probability computation theorem.

The lead probability is now computated automatically based on multiple criteria (= frequency fields).
Some are mandatory and will always be used, some are optional and can be activated
in the settings.
- Mandatory fields :
    - team_id
    - stage_id
    - tag_ids

- Optional fields :
    - email_state (validity of the email_from field: correct / incorrect or empty if email_from is empty)
    - phone_state (validity of the phone field !! not the mobile one !!: correct / incorrect or empty if phone is empty)
    - country_id
    - state_id
    - source_id (UTM Source)

PLS uses a start date to consider only the lead created after that date to generate the frequency table
and to recompute the lead probability. This date is also configurable in the CRM settings.
Default start date is date.todays() (at module install).

The settings for PLS are NOT company related :
PLS Fields = Many2many to a model that store only the fields allowed for PLS computation.
PLS Start Date = Date
As Date and Many2many fields are not allowed for config_parameters,
Char fields are used to store the configuration
and Date and m2m fields are used to ease the configuration by the user in the config panel.
Date and m2m fields are both computed fields based on their corresponding Char fields.
Char field for Date store a strigified date.
char field for Many2many store a comma separated string which is a list of activated fields.

Each mandatory field is used in a specific way
- Team_id : the frequency table is split for every team_id, plus once for leads that don't have a team_id.
Considering Lead A from Team 1 and Lead B from Team 2.
Even if Lead A and Lead B have the exact same attributes (country, phone, email,stage, ...),
their probability will most likely not be the same.
That's because each teams compute their leads' probability based on their own won/lost leads.

- Stage_id : if no optional fields are activated and if there is no tag on the lead,
this will make PLS work quite the same way as it was working with removed stage based probability,
as only the stage will count in the computation.
But here, instead of fixing manually the probability for each stage,
this stage related probability is automatically computed based on past experience (won and lost leads).

- Tag_ids : each tag is considered separatelly, as if the lead had only one tag.
(no link between tag on same lead is made). To avoid that a tag take to much importance if his subset is too
small, we include the tag frequencies in the frequency table only if at least 50 won or lost leads had this tag.

The frequency table is a table where all won and lost leads are aggregated by fields
(mandatory or optional if activated). For each team_id / field couple, we store the number of
won and lost leads that has that fields values.
For example: Lead A is assigned to team 1 and client comes from France and has been won.
If we consider only this lead A, the frequency table will looks like :
   id   |  variable   | value | won_count | lost_count | team_id
--------+-------------+-------+-----------+------------+---------
    1     country_id     FR         1           0            1

To computed the probability of a lead, we get all the records for the frequency table
that match all fields (per team_id) and we compute a won score and a lost score.
The probability is computed and normalized based on those scores
- P = S(Won) / (S(Won) + S(Lost))

Considering two variables A and B :
- S(Won) ∝ P(A∩B | Won)*P(Won) = P(A|Won) * P(B|Won) * P(Won)
- S(Lost) ∝ P(A∩B | Lost)*P(Lost) = P(A|Lost) * P(B|Lost) * P(Lost)

To overcome the 'zero frequency problem',
we do not start at 0 for the won or lost count for each variable / value combination.
It suggested to start at 1, but to avoid that a small subset take to much weight compared to
a larger one, we start at 0.1.

For example :
If we start at 1 : If team 1 has few records from France,
it can get high probability to win even if every lead where lost.

  variable   | value | won_count | lost_count | team_id | Probability
-------------+-------+-----------+------------+---------+-------------
 country_id     FR         1           7            1         0.125
 country_id     FR        198         7306          2         0.0263

If we start at 0.1 :

  variable   | value | won_count | lost_count | team_id | Probability
-------------+-------+-----------+------------+---------+-------------
 country_id     FR         0.1         6.1          1         0.0161
 country_id     FR        197.1       7305.1        2         0.0263

The more we have records for each couple, the more the computation become precise.

The allow the user to decide himself which probability should be set on the lead,
a manual probabiltiy system has been implemented using 2 stored fields + 1 computed field :
- probability (stored): Probability set on the lead, can be set manually in the lead form.
- automated_probability (stored): Probability computed by PLS.
- is_automated_probability (computed): If both previous fields are equal, the probability is considered
as automatic. and each time the automated_probability is updated, the probability is aligned.
If both are not equal, the probability is considered as manual. The automated_probability is still
updated but the probability is not aligned. The user can reset the probability in automatic mode
by clicking on the estimated probability displayed in edit mode in the lead form.

probability and automated_probability are not computed fields as they are manually computed at write and create.
The computation is triggered if one of the activated frequency fields is modified or set.
To trigger also the computation when modifying the lead values on the lead views (form and kanban),
_onchange methods have been added for each frequency field.

Also, a cron has been added to run every day to rebuild the frequency table and to recompute all active
and pending leads (not won nor lost) probability.

Task ID : 1925439
PR #33589

remove rerun cron on action lost and won
2019-06-11 07:04:33 +00:00
..