Learn how Xelix flags potential duplicate invoices in your account. We'll walk through the two-phase detection process, from an initial rules-based sweep to a detailed Machine Learning analysis, and how we assign a risk rating to each pair we raise š
Contents:
1. The two-phase detection process
Xelix's duplicate detection system runs in two main phases: first, we use a set of simple rules to identify potential duplicate invoice pairs. Then, we run a much more detailed Machine Learning analysis on those candidates to remove false positives and assign a risk rating.
2. Phase 1: Rules-based detection
We use a broad rules-based system to initially identify potential duplicate invoice pairs, comparing each invoice's number, date and amount.
We identify a potential duplicate pair if a pair meets all the conditions for any one of the following four cases:
All match
All of the following must be true:
The two invoices have similar invoice numbers, differing by only two characters or fewer
The two invoices have identical dates and amounts
Amount match
All of the following must be true:
The two invoices have similar invoice numbers
The two invoices have completely different dates
The two invoices have identical amounts
Date match
All of the following must be true:
The two invoices have similar invoice numbers
The two invoices have identical dates
The two invoices' amounts differ in at least one of the following ways:
The gross amount of one invoice is equal to the net amount of the other
The two gross amounts differ by a factor of 10, 100 or 1000 (e.g. £123.45 / £1,234.50)
The two gross amounts differ by less than one major monetary unit (e.g. £123.45 / £122.79)
The two gross amounts differ by one digit only (e.g. £123.45 / £133.45)
Different number match
All of the following must be true:
The two invoices have different invoice numbers
The two invoices have identical dates and amounts
3. Phase 2: Machine Learning analysis and risk rating
Every potential duplicate pair that makes it through the rules-based phase then goes through a much more detailed analysis. We compute over 500 datapoints based on the properties of the two invoices, and how they compare to other invoices, both from the same supplier and within your data more broadly.
These datapoints are fed into a number of Machine Learning models, which determine how likely a potential duplicate pair is to be a true duplicate. If a pair is likely to be a true duplicate, we raise it in the platform as a Duplicate Pair, with a unique Duplicate Pair ID.
We also use these datapoints to assign each duplicate pair a risk rating: High, Medium or Low. This represents how confident we are that we've identified a true duplicate:
4. What happens next
Once a duplicate pair is raised, it appears in the Duplicate Invoices tab, where you can review, classify and resolve it. To learn how, check Resolve Duplicate Pairings.
Want to go further? Here's what else you can explore:
Tailor the system above to your organisation's invoicing data: Duplicate Invoice Configurations
Guarantee certain kinds of duplicates are always raised: Custom Duplicate Groups

