A pay equity dataset normally needs four groups of information: worker identity and sex, job and worker-category data, pay components, and working-time or measurement-period data. Under Directive (EU) 2023/970, employers need enough information to calculate organisation-wide and category-level gender pay gaps, median gaps, complementary or variable pay metrics and quartile distributions. Practical fields commonly include a stable worker ID, sex, job title, job family, grade, category of workers, location, employment status, contracted hours or FTE, gross annual and hourly pay, basic salary, variable compensation, currency and effective dates. Additional fields such as tenure or performance may be useful for investigation, but they should not be confused with fields expressly required by the Directive.

pay equity data requirements

Jurisdiction: European Union

Start With the Outputs You Must Calculate

A useful dataset is designed backward from the required outputs. Article 9 requires the gender pay gap, median gender pay gap, equivalent metrics for complementary or variable components, the proportion of women and men receiving those components, quartile pay bands and category-level pay gaps. Article 10 can require deeper category-level analysis where specified conditions are met. Those outputs tell the employer what information must be traceable at worker level. If a dataset does not distinguish base salary from variable compensation or cannot connect a worker to the relevant category, the missing field may prevent a required calculation or make the result difficult to explain.

Worker Identity and Sex Data Need Stable Definitions

A stable worker identifier is essential because compensation, job and working-time data often come from different systems. The identifier lets analysts join those records without relying on names, which can change or create matching errors. Sex data is also needed for the Directive's female and male pay-gap measures. Employers should define how the field is sourced, which record is authoritative and how missing or inconsistent values are handled. Data protection and national-law requirements should be considered when extracting and using personal data. The analytical dataset should contain only what is necessary for the stated purpose and should be governed with appropriate access controls.

Job and Worker-Category Data Supports Equal-Value Comparisons

Job title alone is rarely enough for reliable pay-equity analysis. Titles can vary across business units even when the work is similar, and identical titles can hide materially different responsibilities. Employers commonly need job family, grade or level, location and a defensible category-of-workers field or the underlying job-evaluation information used to derive it. Article 4 links work of equal value to objective and gender-neutral criteria, while Articles 9 and 10 use categories of workers for detailed analysis. Capturing the job architecture therefore helps the employer move from a headline organisation-wide gap to a comparison that can identify where differences actually sit.

Compensation Data Should Be Split Into Meaningful Components

The analytical dataset should normally distinguish ordinary basic wage or salary from complementary or variable components. That separation is necessary because Article 9 requires specific variable-pay measures and category-level breakdowns. Depending on the organisation, component fields may include bonus, commission, overtime, shift premium, allowance, benefit values and other reward. Not every payroll code needs its own analytical column, but the mapping from source codes to analytical components should be documented. Gross annual and corresponding gross hourly pay are also important because Article 3 defines pay level using those measures. The objective is to preserve enough detail to calculate required metrics and investigate why two workers' total reward differs.

Working-Time Data Prevents Misleading Comparisons

Part-time work, different weekly hours and mid-year employment can distort a simple annual-pay comparison. Employers therefore commonly need contracted hours, FTE, full-time or part-time status and relevant start or end dates. The correct treatment depends on the metric, the Directive and national implementation methodology. Analysts should not simply annualise or hourly-normalise every value without checking the purpose of the calculation. The dataset should preserve the original value and the working-time context so that a transformed measure can be traced back to the source. This is especially important where one metric uses gross annual pay and another uses the corresponding gross hourly pay.

Additional Explanatory Variables Need a Clear Purpose

For deeper investigation, an employer may collect variables such as tenure, relevant experience, performance rating, education, qualification, location, job level or supervisory responsibility. These can help test possible explanations for a pay difference. However, including a variable in a statistical model does not automatically make the factor legally valid or gender neutral. The employer should document why the variable is relevant, how it is measured and whether it is consistently available. Some factors may reflect historical inequalities rather than justify them. Analytical convenience should therefore not be confused with legal justification under equal-pay rules.

Every Field Should Have Data Lineage

A defensible dataset records where each field came from, the date it represents and any transformation applied. Useful metadata can include source system, extraction date, source field name, currency, effective date, data owner and transformation rule. This is especially important when payroll, HRIS, sales compensation and benefits systems are combined. Without lineage, analysts can discover a pay difference but be unable to explain whether it reflects the source record, a conversion rule or a manual correction. A data dictionary and transformation log turn the dataset from a one-off spreadsheet into a repeatable analytical asset that can support future reporting cycles and investigations.

Frequently Asked Questions

Is job title enough for a pay equity analysis?

Usually not. Job titles can be inconsistent. Employers often need job family, level, grade, location and a defensible worker-category framework to support meaningful comparisons.

Do I need tenure and performance data?

They are not universal Article 9 reporting fields, but they can be useful explanatory variables for deeper analysis when they are relevant, consistently measured and legally appropriate.

Should base salary and bonus be stored separately?

Yes, where possible. Article 9 requires separate analysis of ordinary basic wage or salary and complementary or variable components, so preserving that distinction improves reporting and investigation.

Related Guides

Official Sources

Use this as a starting point

Requirements and practices differ by jurisdiction and organisation. Check current local law, official guidance and professional advice for a specific situation.