AI Knowledge Hub

What is a bordereaux data dictionary and why does it matter for AI mapping?

Quick answer

A bordereaux data dictionary is a reference document that defines every field in an organisation's target data structure, including its meaning, expected format, valid values and any validation rules. It differs from a canonical data model, which defines the structure itself, by focusing on the precise definition of each field within that structure. A clear, well-maintained data dictionary improves the reliability of both manual and AI-assisted mapping, because it gives clear reference points for what each piece of incoming bordereaux data should become.

What to remember

Key takeaways

  • A data dictionary defines the meaning, format and valid values of every field in a target data structure.
  • It complements, rather than replaces, a canonical data model, which defines the structure itself.
  • Ambiguous or outdated field definitions reduce the reliability of AI-assisted mapping, not just manual mapping.
  • A data dictionary needs a named owner and a change process, since it will drift out of date without active maintenance.

Every organisation that maps bordereaux from multiple coverholders into a single internal structure eventually runs into the same problem: two people, or two systems, interpreting the same field differently.

Is "Gross Premium" before or after brokerage? Does "Location" mean the risk address or the coverholder's registered office? Without a clear answer written down somewhere, these questions get resolved informally, inconsistently, and often differently by different analysts on different days.

A data dictionary exists to prevent that. It is the reference document that says, precisely, what each field in the target structure means and what values are acceptable.

What a data dictionary contains

A bordereaux data dictionary typically defines, for each field in the target structure, its name, a precise business definition, the expected data type and format, any valid value list or range, whether the field is mandatory, and any calculation or derivation rule if the field is not taken directly from source data.

For a field like gross premium, that might mean specifying exactly which components are included, what currency and decimal convention applies, and how the field should be populated if the source bordereau does not provide it directly.

How this differs from a canonical data model

A canonical data model defines the overall structure: what entities exist, such as policy, risk and transaction, and how they relate to each other.

The data dictionary operates at a more granular level. It defines the individual fields within that structure. The two documents work together: the canonical data model provides the skeleton, and the data dictionary provides the precise meaning of each part of it.

An organisation can have a well-designed canonical data model and still suffer from inconsistent mapping if the fields within it are not clearly defined.

How organisations traditionally documented field meaning

Before formal data dictionaries became common, field definitions were often held informally: in the memory of a small number of experienced analysts, scattered across mapping spreadsheets built for individual coverholders, or documented inconsistently in onboarding notes.

This works reasonably well while the same small team maintains the mappings. It becomes a liability as teams change, coverholder volumes grow, or new systems are introduced, since new starters and new tools have no single, authoritative reference to work from.

Why AI mapping quality depends on the data dictionary

AI-assisted mapping tools work by comparing incoming bordereaux fields against the target structure and deciding which target field each source field most likely corresponds to.

That comparison is only as good as the definitions it is comparing against. If the target field "Net Premium" has no clear definition, an AI system has no reliable way to decide whether an ambiguous source column belongs there or under a different field, and neither would a human analyst working from the same incomplete information.

A clear data dictionary gives AI mapping something precise to match against, which tends to produce more consistent, more confident mapping decisions and fewer ambiguous exceptions that require manual resolution. It does not remove the need for human review of mapping decisions, particularly for new or unusual source formats, but it improves the starting point for that review.

Creating and maintaining a usable data dictionary

Building a data dictionary is not a one-off exercise. Definitions need a named owner, typically within a data or operations function working closely with underwriting and finance, since many fields carry both operational and financial meaning that requires agreement across those teams.

Changes to definitions should be version controlled and communicated to anyone relying on the mapping, whether that is a person maintaining templates or a team overseeing an AI-assisted mapping tool. A data dictionary that is created once and never revisited will drift out of date as products, terminology and reporting requirements change, undermining the consistency it was created to provide.

Example

A managing general agent maintains a canonical data model for premium bordereaux, but the definitions behind several fields, such as "gross premium" and "net premium", have never been written down precisely. Different analysts have historically interpreted them slightly differently when building mapping templates.

The MGA creates a data dictionary that defines each field precisely, including how gross and net premium should be calculated and which deductions apply. Once the AI mapping tool is configured against these clearer definitions, it produces more consistent mapping decisions, and exceptions are more often genuine data issues rather than definitional disagreements.

FAQs

  • Is a data dictionary the same as a canonical data model?

    No. The canonical data model defines the structure and relationships of the target data, such as policy, risk and transaction entities. The data dictionary defines the precise meaning and rules for each individual field within that structure. The two work together rather than replacing one another.

  • Who should own a bordereaux data dictionary?

    Ownership typically sits with a data or operations function working closely with underwriting and finance. Many field definitions carry both operational and financial implications, so cross-functional agreement is usually needed to keep the definitions accurate and accepted across the business.

  • Does using AI reduce the need for a data dictionary?

    No. AI reduces the manual effort of building mapping templates, but it still relies on clear target definitions to map data confidently. A poor or missing data dictionary limits how reliable AI-assisted mapping can be, regardless of which tool is used.

What's next?

See it on your own bordereaux template

Send us your target BDX format and we'll show how AI can transform typical market bordereaux into your required structure.

Our latest insurance insights