Skip to main content

How we detect duplicate invoices

How Xelix's two-phase system (rules-based detection + Machine Learning) identifies and risk-rates potential duplicate invoice pairs.

Learn how Xelix flags potential duplicate invoices in your account. We'll walk through the two-phase detection process, from an initial rules-based sweep to a detailed Machine Learning analysis, and how we assign a risk rating to each pair we raise šŸ”


Contents:


1. The two-phase detection process

Xelix's duplicate detection system runs in two main phases: first, we use a set of simple rules to identify potential duplicate invoice pairs. Then, we run a much more detailed Machine Learning analysis on those candidates to remove false positives and assign a risk rating.


2. Phase 1: Rules-based detection

We use a broad rules-based system to initially identify potential duplicate invoice pairs, comparing each invoice's number, date and amount.

We identify a potential duplicate pair if a pair meets all the conditions for any one of the following four cases:

  • All match

All of the following must be true:

  • The two invoices have similar invoice numbers, differing by only two characters or fewer

  • The two invoices have identical dates and amounts

  • Amount match

All of the following must be true:

  • The two invoices have similar invoice numbers

  • The two invoices have completely different dates

  • The two invoices have identical amounts

  • Date match

All of the following must be true:

  • The two invoices have similar invoice numbers

  • The two invoices have identical dates

  • The two invoices' amounts differ in at least one of the following ways:

    • The gross amount of one invoice is equal to the net amount of the other

    • The two gross amounts differ by a factor of 10, 100 or 1000 (e.g. Ā£123.45 / Ā£1,234.50)

    • The two gross amounts differ by less than one major monetary unit (e.g. Ā£123.45 / Ā£122.79)

    • The two gross amounts differ by one digit only (e.g. Ā£123.45 / Ā£133.45)

  • Different number match

All of the following must be true:

  • The two invoices have different invoice numbers

  • The two invoices have identical dates and amounts


3. Phase 2: Machine Learning analysis and risk rating

Every potential duplicate pair that makes it through the rules-based phase then goes through a much more detailed analysis. We compute over 500 datapoints based on the properties of the two invoices, and how they compare to other invoices, both from the same supplier and within your data more broadly.

These datapoints are fed into a number of Machine Learning models, which determine how likely a potential duplicate pair is to be a true duplicate. If a pair is likely to be a true duplicate, we raise it in the platform as a Duplicate Pair, with a unique Duplicate Pair ID.

We also use these datapoints to assign each duplicate pair a risk rating: High, Medium or Low. This represents how confident we are that we've identified a true duplicate:


4. What happens next

Once a duplicate pair is raised, it appears in the Duplicate Invoices tab, where you can review, classify and resolve it. To learn how, check Resolve Duplicate Pairings.

Want to go further? Here's what else you can explore:

Did this answer your question?