Blog ยท 6 September 2026

What can LCA database publishing borrow from software releases?

Software projects use continuous integration to run agreed checks whenever something changes. A small experiment with two BAFU releases suggests ideas for making LCA database improvements more visible and repeatable.

Circular LCA database release workflow linking human improvement and review, CI checks, and publication

Publishing an LCA database and releasing software are different jobs. The data requires scientific judgment, documentation, review, and knowledge of how it will be used.

Still, one habit from software development may be worth exploring.

An idea borrowed from software

Software teams often use continuous integration, usually shortened to CI. When someone proposes a change, a shared system builds the software and runs an agreed set of automated checks. The result is available before the change becomes a release.

CI does not decide whether the software is good, and it does not replace reviewers. It makes routine checks repeatable, gives every candidate the same treatment, and keeps a record of what passed.

LCA database publication does not need to copy this model. But it may prompt useful ideas: Which checks could run the same way for every candidate? Which improvements could be measured? Which results should be kept with the release?

Starting from checks that already exist

LCA database producers already perform quality checks. Some live in review procedures, spreadsheets, scripts, or the experience of the people preparing the release. This experiment explores what might become easier if some of those checks, along with a few structural checks, could run consistently for every candidate and compare their result with the previous version.

We explored that question with two successive public releases of the Swiss Federal Administration LCI database, BAFU:2025 and BAFU:2026 v1. Both archives were loaded unchanged into the same version of Volca.

The newer release improved the quality report across several checks. The comparison recognized those improvements while keeping the previous release as a visible baseline.

The experiment concerns release workflow and does not assign a quality score to BAFU. Public data makes the steps reproducible.

Comparing a release with its baseline

A quality report describes the current state of a database. Depending on the rules available, it can inspect reference products, allocation totals, duplicate content, metadata, physical consistency, chemical identifiers, suspicious quantities, or names that may become ambiguous in another format.

Running that report on every release is already useful. Comparing it with the previous release adds another view:

  • signals that were already present;
  • signals that have disappeared as the database improved;
  • situations that appear for the first time and may deserve review.

This comparison is sometimes called a quality ratchet. Existing signals remain visible without forcing every historical case to be resolved at once. Each release can improve the baseline, while an organization chooses which kinds of new signal should pause publication.

The threshold is a policy choice, not a universal definition of quality. It may start with a small set of structural conditions and evolve as the publication process matures.

Comparing meaning rather than identifiers

Database identifiers are not always stable between releases. A producer may regenerate UUIDs even when the underlying process remains the same.

The Volca example therefore matches quality signals using their content:

  • the check and severity;
  • activity and product names;
  • geography;
  • the explanation attached to the signal.

It also handles presentational differences between versions, such as a geography moving from the end of an activity name into a dedicated location field. This prevents a harmless representation change from looking like a new situation.

The matching rules are visible and can be discussed. A database producer may need different rules, especially when names, classifications, or identifiers follow another convention.

The idea is not tied to one database format

Volca runs these checks on its normalized database model after import. The experiment can therefore be adapted to databases delivered as:

  • EcoSpold 1 or EcoSpold 2;
  • ILCD;
  • SimaPro CSV;
  • Brightway Excel.

A producer can keep its existing authoring format. A consultancy could apply the same checks to successive versions of datasets prepared for a client. An institution could test the approach on one publication step without changing the rest of its workflow.

The baseline and candidate should use the same Volca version and reference configuration. If the source format changes between releases, the conversion deserves a separate test so that format differences are not confused with content changes.

Explore it manually before thinking about CI

The idea is easier to understand in the interface than in a CI configuration file.

Open a database in hosted Volca and select its Data quality tab. Volca runs the available checks and presents each one with an explanation and the activities concerned. This is useful as a manual review even if no automated release process exists.

When two versions are loaded side by side, the same tab displays a Compare with selector. Choosing the previous version changes the report into a comparison of what is new, fixed, or unchanged. The detailed comparison can also be downloaded as CSV for review or archiving.

Volca also has a separate Database comparison view. It answers a different question by showing which activities were added, removed, or changed, down to their exchanges. The two views complement each other:

  • Data quality applies explicit checks and compares their signals;
  • Database comparison shows how the database content itself changed.

A producer or consultant can start with either view as a manual step. The same quality reports can also be requested from Python through the Volca API, which makes it possible to automate the comparison in an existing validation script. If the checks prove useful and stable, that Python check can later run automatically in CI. Automation is an optional continuation of the workflow, not a prerequisite for trying it.

A first experiment can remain deliberately modest:

  1. open the current database and inspect the Data quality report;
  2. load a candidate version and compare it with the baseline;
  3. review the improvements and other differences;
  4. decide which checks, if any, would be useful to repeat automatically.

Which checks would be useful next?

The current Volca report covers reusable structural checks such as reference products, allocation totals, duplicates, metadata, physical consistency, chemical identifiers, suspicious quantities, and export-sensitive naming.

Database producers and consultants often apply additional rules that are specific to their methods or clients. They may check mandatory classifications, accepted units, date coverage, naming conventions, expected geography, required documentation fields, or consistency between a process and its exchanges.

A good candidate for automation has a rule that can be stated precisely, enough context to identify the affected process or exchange, and a reproducible example of the expected result.

We would like to learn which checks are still repeated by hand during real database publication work. Some may be too dependent on expert judgment to automate, which is useful to know. Others may fit naturally into an extensible quality report.

Adding a check to the report also makes it available to the comparison. The release gate discovers checks from the report rather than maintaining a separate fixed list.

Keep the boundary clear

Automated structural checks do not assess whether a dataset is representative, whether an emission model is appropriate, or whether a source is reliable. Scientific and methodological review remain essential.

Automation can support the repetitive part of that work. It can apply the same explicit rules to every candidate, recognize improvements, and preserve the result alongside the release.

Different database producers will choose different practices. The software analogy is useful if it encourages practical questions about the next publication: Which existing checks could be repeated automatically? Which evidence should be archived? Which new rule would save reviewers from finding the same kind of issue by hand again?

Try the workflow

The easiest starting point is the hosted interface: explore Volca and open the Data quality tab of a database. A version comparison becomes available when another database is loaded alongside it.

The BAFU releases used for the experiment are publicly documented here:

The same experiment can be run with any two versions of a database that Volca can import. It provides one example of how a publication workflow can make quality checks more repeatable.

See the comparison workflow

Volca database comparison configured for BAFU v1 against BAFU, matched by name, location and product, with the summary option selected.

The database comparison screen before execution. BAFU:2026 v1 is compared with BAFU:2025 by name, location and product. The summary option provides an immediate overview; the full comparison can continue down to the exchanges.