Archiving a solve¶
How to keep a solve: the spec, the data it was solved with, and what came
back, so you can read the answer later, ask the question again, or hand both
to someone else. The reference is
load_archive and the entries
beside it.
Archive as you solve¶
That writes spec.yaml, one sources/<key>.parquet per key the file
declares, catalog.parquet saying what each file holds, and answer/ holding
everything the solve produced. Each source is held as the table the solve read,
which sps.tidy returns, with
specsolve_run added: a dimension as its labels and specsolve_position, a
parameter as its dims and value, a relation as its columns:
case/
spec.yaml
sources/cost.parquet
sources/load.parquet
…
sources.parquet (specsolve_run, source, digest) — what each of them is
catalog.parquet (specsolve_run, path, …, column, dim) — what each file holds and over which dimensions
answer/record.parquet how it terminated, what it reached, when, under what name, and on what
answer/metrics.parquet what the build and its solves took
answer/primal/p.parquet one file per variable
answer/dual/power_balance.parquet
The suffix picks the container. Anything without a .zip suffix is a
directory, as above. .zip packs the same members into one file:
sps.solve, model.solve and sps.solve_over take archive=. Nothing
else writes one.
The record says what produced the answer. answer/record.parquet names
the solver, its version and the options it ran with, and the specsolve and
mathspec versions that built the model. So you can install the same
environment again. The options are one JSON string with sorted keys. An
option that changes the answer, such as a time limit, a gap, a seed or a
tolerance, keeps its value. A value that JSON has no number for, such as an
infinite time limit, is the string "inf", "-inf" or "nan". Any other
option keeps its name and has the value <not recorded>. An archive often
goes to shared storage, and a Gurobi licence credential such as WLSSecret is
passed as an option.
result.provenance gives the same five fields without an archive.
To keep the value of an option that is not on the solver's list, name it. Name no credential:
sps.solve('dispatch.yaml', sources, solver_options={'mip_max_nodes': 1000}, record_options=['mip_max_nodes'])
Read it back¶
case = sps.load_archive('case/')
case.result.objective # what it reached
case.result.primal('p') # the values it came back with
sps.solve(case.spec, case.sources) # the same question, asked again
load_archive reads it whole. The sources come back as tables, the
answer's frames are in memory, and nothing has to be kept alive afterwards. A
zip unpacks to a scratch directory that is gone when the call returns:
Pass into= when you want the extracted tree as well, to query it with an
engine that reads parquet. A directory archive is read where it lies and
refuses an into=.
Read one too big to hold¶
scan_archive reads nothing until asked. The sources come back as paths,
and each frame is read off disk at the call that asks for it. A source file
also holds specsolve_run, which a build ignores:
archived = sps.scan_archive('sweep.zip', 'sweep/')
archived.sweep.scan('p') # read at the collect, one name at a time
Scan the archive that does not fit in memory, and the one you will read
little of: a load reads every name, a scan only the ones you ask for. What is
scanned has to outlive what it reads off, so a zip needs an into= you will
keep.
load_archive |
scan_archive |
|
|---|---|---|
| a source | the table the member holds | the path to it |
| the answer | frames in memory | read at the call that asks |
| a sweep's answer | held: sweep.primal('p') |
spilled: sweep.scan('p') |
a zip's into= |
optional, and scratch without one | required, and kept |
Read what a solve cost¶
metrics is a Metrics: how big the model was, how many solves the clocks
cover, and wall-clock seconds in each phase, as one value
(the attributes).
case = sps.load_archive('case/')
case.metrics.rows # how big the model was
case.metrics.solves # how many solves the clocks cover
case.metrics.build_seconds # turning declarations into frames
The row covers the model's whole life, and solves says how long that is.
sps.solve builds the model it solves, so its archive reads solves of 1. A
model solved more than once before it was archived carries the sum:
with sps.build('dispatch.yaml', sources) as model:
model.solve()
model.solve(archive='second/') # solves: 2, and the clocks cover both
A sweep records each slice's own metrics as sweep.metrics instead
(sweeps).
Keep the answer an update produced¶
An updated model answers the merged data, so archive it where it is answered:
with sps.build('dispatch.yaml', sources) as model:
model.update({'p_max': doubled}).solve(archive='case/')
Archive a sweep¶
archive= on solve_over writes the sweep's answer at the paths a single
solve uses, one file per name. A rolling horizon's answer/primal/soc.parquet
is over snapshot, as sweep.primal('soc') returns it. Each file carries
specsolve_run, as a single solve's does, and the readers drop it:
axis = sps.EachWindow('snapshot', steps=24, lookahead=24, into='t')
sps.solve_over('window.yaml', sources, axis, carry={'soc_initial': 'soc'}, archive='roll/')
The archive carries the axis and the carry, so the sweep runs again from the file alone:
archived = sps.load_archive('roll/')
archived.sweep.primal('soc') # (snapshot, value), off answer/primal/soc.parquet
sps.solve_over(archived.spec, archived.sources, archived.axis, carry=archived.carry)
A source the axis cuts is held uncut, the axis column first. A parameter given as one number over a window's local index is held as a table over the axis, because a window of each length reads the number over labels of its own. Each slice cuts from the archive the tables it attached.
Keep the windows where you will read them per window. keep_windows=True
also writes each window's frames, lookahead rows included, under
answer/windows/, with owned.parquet for what each window owns and
catalog.parquet for what each of these files holds. Then per_window=True
reads off the archive:
sps.solve_over('window.yaml', sources, axis, carry={'soc_initial': 'soc'}, archive='roll/', keep_windows=True)
sps.load_archive('roll/').sweep.primal('soc', per_window=True) # (snapshot_start, t, value)
Without them, per_window=True and an expression the file never named are
refused. Solve again from the archive to get them back.
Archive a sweep too large to hold¶
spill_to= writes each slice's frames as the fold goes, so the sweep holds
one slice at a time. archive= packs the whole sweep. Pass both and the
archive reads its answer off the spill, so the sweep is archived without ever
being held:
axis = sps.EachCoordinate('scenario')
sps.solve_over('dispatch.yaml', sources, axis, spill_to='work/', archive='sweep/')
archived = sps.scan_archive('sweep/')
archived.sweep.scan('p') # keyed by scenario, read at the collect
scan_archive leaves the answer on disk, and scan reads it. load_archive
reads it into memory, where it fits, and sweep.primal('p') answers on that
one.
Read a directory of them¶
A directory of archives is a table per glob, and every row of every table
carries specsolve_run, the archive's own name:
The recipes are reading a directory of runs.
What an archive will not take¶
A sweep cut by a hand-built axis. A list of (key, sources) is a set of
sources per slice, and the call is refused before the first slice is solved:
archive= takes a sweep cut by EachCoordinate or EachWindow, which say how one
set of sources was cut and so how the archive can be re-run. A hand-built list
is a set of sources per slice, which are unrelated questions — archive one
solve each.
A directory that already holds something. A directory archive is written
whole, never merged into what is there. A .zip target is replaced.