Some use case like testing performance or upgrade scripts required a database
with prefilled data, covering basic corner cases. A solution can be to
create data a procedural way.
This commit proposes an API to easily populate a model, usually by giving a
list of possible values for each field or by giving a compute method that will
be based on raw values of other fields.
The basic way to define how to populate a new field is to override `_populate_factories`,
a method that returns a sequence of pairs `(field_name, factory)`.
The definition of a field is a "factory", a function that returns a neverending iterator
combining its value(s) with the values of the iterator given in parameter.
Some factory helpers are given in `tools.populate.py`:
- `iterate(vals, weighs)` ensures that one record is created for each value
by iterating on them, then resumes as `random.choice` on those vals following weights
once the first iteration is finished.
- `cartesian(vals, weights)` makes a cartesian product of its own values with the values
of its input iterator, then resumes as a randomized generator.
- `compute(function)` calls the given function with the current values dict and a random object,
and assigns the current field to the returned value.
- ...
Each iterator yields dictionaries of field values, and the factory should add a
value for the current field(s). The yielded dictionaries also contain a pseudo_field
`"__complete"`, that indicates whether this step is some randomized data
to reach the expected count of records. A falsy value indicates that the iterator
is still covering mandatory cases. This indicates whether a cartesian product is
finished, or an `iterate` has consumed all its values.
The order of the factories is quite important, since some computed fields may need
other fields to be defined, and `cartesian` factories should always be at the beginning
to avoid having too many combination. That is why the factories are given as a list of
pairs instead of a dictionary; this makes it easier to insert elements at any place.
Example:
field A: cartesian([T, F])
field B: cartesian([0, 1])
field C: iterate([a, b, c, d, e])
field D: compute(1-B)
_c is shortcut for __complete
_ is a random value, or result of a random value
```
iter | root | field A | field B | field C | field D | result
0 {_c:F} {... A:T} {... B:0} {...C:a} {...D:1} T,0,a,1 complete:False
{... B:1} {...C:b} {...D:0} T,1,b,0 complete:False
{... A:F} {... B:0} {...C:c} {...D:1} F,0,c,1 complete:False
{... B:1} {...C:d} {...D:0} F,1,d,0 complete:False
1 {_c:T} {... A:_} {... B:_} {...C:e,_c:F} {...D:_} _,_,e,_ complete:False
2 {_c:T} {... A:_} {... B:_} {...C:_} {...D:_} _,_,_,_ complete:True
```
X-original-commit: 4c0182dafa584853ed83a166096f45c33c06a245