Skip to content

Flow Mapping

Flow mapping is the process of linking elementary flows in your LCA databases to characterization factors in your LCIA methods. Good mapping coverage is essential for accurate LCIA results.

Different databases and methods use different names for the same substance. For example, “Carbon dioxide, fossil” in EcoSpold 2 might be “CO2 (fossil)” in an ILCD method. VoLCA resolves these differences automatically through a matching cascade.

For each characterization factor in a method, VoLCA tries to match it to a database flow in this order; the first hit wins, and the strategy that matched is recorded:

StepStrategyDescription
1UUIDexact elementary-flow UUID match
2Namenormalized-name match
3Synonymmatched via a loaded synonym group (e.g. "CO2" ↔ "Carbon dioxide, fossil")
4CASCAS-number match, kept within the line’s own substance
–Unmatchednone of the above hit, and the line characterizes no flow through the cascade

Steps 2 to 4 find every database flow of that name, synonym or CAS number, and the line then goes to the one that reads its compartment first, by the rule of the next section. Between two that read it alike, the one bearing the line’s own name, so Waste water and Waste water/m3 each stay on the flow written in their unit. A line no flow reads is unmatched, and the coverage counts say so.

A line carrying one factor per location reaches every flow of the same name and medium that reads it, not only the one it was matched to.

How a flow meets a factor’s subcompartment

Section titled “How a flow meets a factor’s subcompartment”

A method writes a factor per substance, medium and subcompartment; a database files each flow under its own. A flow reads a factor at three places, in this order, and nowhere else:

  1. its own subcompartment;
  2. the line its substance writes for the whole medium, with no subcompartment;
  3. the place an if_absent row of the compartment mapping sends it to.

A flow that states no subcompartment is an unspecified emission. What unspecified means in a method depends on its format:

Formatunspecified in a factor line
SimaPro CSV(unspecified) is the line for the whole medium (place 2)
ILCD, JSON-LD, the engine’s CSVone subcompartment like any other (place 1)

So under an ILCD package, carbon dioxide emitted to “urban air close to ground” reads the line written for that subcompartment, and nothing else; the line at “unspecified” is for the flows filed at unspecified. Under a SimaPro method, the same flow reads (unspecified) when the method writes nothing for urban air.

The engine borrows no factor across subcompartments beyond these three places. A method that means an emission to the sea, or a long-term one, to count for nothing writes that line; one that writes nothing leaves the flow at zero.

data/compartments.csv (replaced with [[compartment-mappings]], see Configuration) holds two kinds of row, named in its seventh column, kind:

  • same (the default when the column is empty): one place written another way. SimaPro’s Air / low. pop. is EcoSpold 2’s air / non-urban air or from high stacks. The row rewrites a flow and a factor line alike, and never joins two places one vocabulary keeps apart.

  • if_absent: two different places. This row

    soil,forestry,,soil,non-agricultural,,if_absent

    says that a flow emitted to forest soil reads the factor written for non-agricultural soil, but only under a method whose collection never writes forestry. The EF 3.1 package has no forest soil, so the row holds there; a method with forestry factors of its own keeps them. The collection decides, not the impact category: a category that happens to write no forestry line belongs to a method that does distinguish forests.

The built-in if_absent rows, the last rows of data/compartments.csv, come from the published EcoSpold 2 mapping of the EF 3.1 flow list, except the one that sends the SimaPro lake to surface water.

The cascade above runs once, when a method is loaded: it attaches each of the method’s factor lines to a database flow. Scoring asks the opposite question, one flow at a time: which factor applies to this flow? That lookup has more rungs, because a flow can be reached without any factor line having resolved to it. In order, with the name explain-cf gives each:

StepNameDescription
1flow_idthe method declares a factor for this exact flow
2same_unit_namea factor line declared in this flow’s own unit, for a name that carries a unit suffix
3exact_namename and subcompartment match a factor line
4compartment_defaultthe substance’s line for the whole medium, when it writes none for this subcompartment
5if_absentthe line at the place an if_absent row sends this flow to
6cas_numbera factor for the same substance, found by CAS number, at one of the three places above
7region_base_namethe flow’s name ends in a region the method does not distinguish
8energy_contentan energy resource takes its family’s factor per unit of energy, bridged by the calorific value in its name
9ore_base_elementa graded ore takes the factor of the element its amount measures

Both cascades leave a trail, and one call reads it back:

Terminal window
GET /api/v1/db/{db}/method/{method}/explain-cf/{flow}

The explanation field is a list of sentences the engine writes itself. Show them as they are rather than rewording the codes:

The method sets no factor for this flow’s subcompartment, so its default for the whole compartment, “Carbon dioxide, fossil”, applies.

That line was tied to this flow’s name through a known synonym when the method was loaded.

The factor applied is 1.0 kg CO2 eq per kg.

Alongside them, match names the rung, the method line and the strategy that attached it, stepsTried lists the rungs walked before the one that answered, and outcome is one of three:

  • characterized, a factor applies;
  • conversion_refused, a factor was found but the flow’s unit cannot be converted to the basis the factor is written in, so the flow scores nothing while looking characterized;
  • no_factor, nothing in the method reaches this flow.

In the web interface, the same answer opens under any row of the Contributing Flows table. Each row also carries the short version, the “Found by” column, so a whole table is annotated without asking per row. A blank there means the method’s tables never walked that flow, which a flow arriving from a dependency database has not been; it does not mean the flow is uncharacterized.

Terminal window
# Summary: matched/unmatched counts and rates
volca --config volca.toml --db ecoinvent flow-mapping <METHOD_UUID>
# Detailed: which CFs matched and by what strategy
volca --config volca.toml --db ecoinvent flow-mapping <METHOD_UUID> --matched
# Gaps: CFs with no database match
volca --config volca.toml --db ecoinvent flow-mapping <METHOD_UUID> --unmatched
# Gaps: database flows with no characterization
volca --config volca.toml --db ecoinvent flow-mapping <METHOD_UUID> --uncharacterized

Bridged names: coverage an exact-name tool would miss

Section titled “Bridged names: coverage an exact-name tool would miss”

The cascade above is a strength when you compute in VoLCA, but it hides a portability trap. When a database names a substance differently from the method that characterizes it – Bromomethane versus Methane, bromo-, Halon 1001, same CAS number – VoLCA still scores the flow by matching on the CAS number or a synonym. A tool that matches factors by their exact name, as SimaPro and many downstream consumers do, has no such bridge: it scores that flow as zero, without warning.

The characterization-coverage report lists exactly those flows – the ones a method scores only through a bridge – grouped under the name the method itself uses, so the fix is to rename the database’s flow to that name. It reports one entry per loaded method collection, so two versions of a method can be compared side by side.

Terminal window
# Every flow a method scores only through a name bridge, per loaded collection
GET /api/v1/db/{db}/characterization-coverage

Flow synonym sets teach VoLCA how to translate flow names between systems. The engine’s own set, Default flow synonyms, is built in and active unless a configuration switches it off. A configuration that lists [[flow-synonyms]] entries gets exactly what it lists, so a custom CSV set stands beside the built-in one only when both are named.

Register a custom set in volca.toml (see Configuration):

[[flow-synonyms]]
name = "my-synonyms"
path = "/data/my-synonyms.csv"
active = true
[[flow-synonyms]]
name = "Default flow synonyms" # keeps the built-in set beside yours

The CSV has a name1,name2,direction,cas,note header, the format of the engine’s own data/flows.csv: one synonym pair per row, then three optional columns, a direction (input, output, or empty for both), the CAS number of the substance, and a free-form note. A running server can also accept one at runtime through the web UI or POST /api/v1/flow-synonyms/upload.

When you have multiple databases loaded (e.g. ecoinvent + agribalyse), VoLCA can link flows between them using depends in the config:

[[databases]]
name = "agribalyse"
path = "/data/agribalyse4.csv"
load = true
depends = ["ecoinvent"] # background flows resolved in ecoinvent

After loading, finalize the cross-database links:

Terminal window
# Via API
POST /api/v1/db/agribalyse/finalize

This resolves agribalyse’s background activities against ecoinvent, enabling full inventory computation across databases.

GET /api/v1/db/{db}/method/{methodId}/mapping - coverage stats
GET /api/v1/db/{db}/method/{methodId}/flow-mapping - per-flow mapping detail
GET /api/v1/db/{db}/characterization-coverage - flows scored only through a name bridge
GET /api/v1/db/{db}/method/{methodId}/explain-cf/{flowId} - why one flow scores with its factor
{ "name": "get_flow_mapping", "arguments": { "database": "ecoinvent", "method_id": "..." } }
{ "name": "get_characterization_coverage", "arguments": { "database": "agribalyse" } }
{ "name": "explain_cf", "arguments": { "database": "ecoinvent", "method_id": "...", "flow_id": "..." } }
  • Flow Mapping Audit – detect and close mapping gaps using the post-scoring suggester, compare_impacts, and the PubChem synonym snapshot.