Data & IA

How to prove your AI model complies: bias metrics and evidence for the regulator

An artificial intelligence model can achieve good performance levels, generate business impact, and be technically ready for production. However, in regulated sectors, these aspects may not be sufficient to meet the requirements of current regulations.

Banks, insurance companies, and other organizations that work with sensitive information need to go a step further and have evidence demonstrating how the AI model behaves with different groups, what information it was built with, and the path the data took to arrive at a result. This ability to reconstruct and trace the results is essential to respond with concrete evidence to any potential request from the regulatory body.

The incorporation of artificial intelligence expands the traditional scope of data governance. It is no longer just about ensuring that the information is reliable, secure and of high quality before feeding it into a model; it is also about governing the model, observing its results and managing the new data and assets it generates.

In that scenario, bias metrics take on a concrete role: they allow statements such as "the model does not discriminate" to be transformed into evidence that can be analyzed. 

In fact, the metrics of bias in AI For compliance, they must form part of a larger structure that includes traceability, documentation, responsibilities, and monitoring mechanisms throughout the entire lifecycle.

In this article we analyze what changes when a model must respond to a regulator, what type of metrics allow us to evaluate fairness and possible biases, how we accompanied a client who needed to demonstrate the behavior of their model to the BCRA, and what information should be prepared before facing an audit.

Why “it is governed” is no longer enough for the BCRA

Data governance didn't begin with artificial intelligence. Regulated organizations have been working for years on quality, security, privacy, access, policies, accountability, and compliance. What changes with AI is that it has broadened the scope of what needs to be governed.

We can see this in a very concrete transformation: before, much of the data journey could end in a report or dashboard used by a person to make a decision, whereas, by using AI, that data can feed a model that generates new results, new content, or new information, which can then re-enter the circuit.

The chain extends, which means having to govern the input information, but also understanding what model processes it, what it produces, under what criteria and what happens to what it generates.

This requirement is especially relevant in the financial system. The current consolidated text on 'Minimum Requirements for the Management and Control of Technology and Information Security Risks', updated by the BCRA Communication “A” 8401, It establishes that the affected entities must Implement effective internal control and technology risk management practices, and demonstrate an understanding of those risks.

It also contains a specific section on software development, acquisition and maintenance, which includes requirements for systems and applications and software life cycle management.

The Communication “A” 8398, For its part, it introduced adjustments to the regime, extended its scope to registered payment service providers included in the PSP Registry of the Central Bank of the Argentine Republic, and reinforced the requirements related to third parties.

This implies that, in an environment where technology is part of regulated processes, To say that a model is governed needs to be able to be translated into controls, documentation, and verifiable evidence.. In other words, an audit needs to demonstrate what actually happened.

Regulated organizations have been working for years on quality, security, privacy, access, policies, accountability, and compliance. What changes with AI is that the scope of governance has expanded.

What bias metrics does a regulator require today?

The first clarification is important: the BCRA communications mentioned in this article do not establish a closed list of metrics. fairness that should be applied to all artificial intelligence models.

However, an organization may need to submit bias metrics in response to a specific regulatory requirement. This leads to a central question: determining what should be measured will depend on the model, the use case, the risks involved, and what needs to be demonstrated.

It is not the same to evaluate a model that recommends products as one that intervenes in risk processes, hiring, evaluation of people or decisions that may generate different consequences for different groups.

As a technical reference, and not as a regulatory checklist of the BCRA, different metrics can be used to evaluate fairness according to the characteristics and risks of each model.

He NIST AI Risk Management Framework It recommends selecting metrics according to the context, purpose, and needs of the evaluation, while the NIST AI Metrology Center It gathers specific metrics to analyze possible differences between groups, such as the following:

  • Demographic Parity It allows you to analyze whether the probability of obtaining a certain result differs between the groups considered. This analysis can help identify behaviors that require further evaluation.
  • Equalized Odds It offers the possibility of comparing the model's performance between groups, taking into account both correct and incorrect results. The goal is to observe whether the model exhibits significant differences in how it classifies the various populations analyzed.

These metrics should not be interpreted in isolation or as an automatic test of compliance. A difference between groups may be a signal that needs to be analyzed considering the data, the purpose of the model, and the context in which it is used.

In a model audit, the goal should not be to accumulate indicators, but to build evidence that explains what risk was being assessed, why certain metrics were selected, what results were obtained, and how they were interpreted.

Bias metrics, in short, acquire regulatory value when they are part of a broader and traceable body of evidence on the behavior of the model.

Case study: How evidence was gathered for a real audit

Let's bring the discussion down to a concrete need, through the recent experience of one of our clients, who needed to demonstrate to the BCRA that a model was not incurring biases related to gender.

Their challenge was not only to assert that the model had been correctly developed or that its overall performance was good, but to obtain metrics and highlight their behavior.

This case illustrates a fundamental difference between developing AI in general and doing so within a regulated organization. A model can generate the expected business impact and still require an additional layer before being deployed in production. the ability to demonstrate that it operates within the parameters that the organization needs to guarantee.

That requirement also changes when the conversation about [the topic] should begin. compliance. If metrics are only sought when an audit arrives, some of the necessary information may be scattered, not preserved, or it may be difficult to reconstruct the exact conditions under which the model was trained or run.

The key is to anticipate that scenario and assess, before a requirement arrives, whether the organization has the necessary information and evidence to explain to the BCRA how a model behaves and support the results it produces.

Incorporating auditability from the beginning of the model's life cycle allows anticipating a potential requirement by defining what information should be kept, what metrics need to be monitored, and what evidence will be needed to reconstruct and explain its behavior when required.

This capability doesn't depend on a single tool, nor can it be concentrated in a single team. The challenge combines technology, processes, and people, and requires governed, high-quality data, model evaluation, and clearly defined responsibilities to transform that information into traceable and available evidence.

The key is to anticipate that scenario and assess, before a requirement arrives, whether the organization has the necessary information and evidence to explain to the BCRA how a model behaves and support the results it produces.

End-to-end traceability: what you need to be able to show

Bias metrics address one part of the requirement. The other part lies in being able to explain where the result we are measuring comes from, and that's where traceability comes in.

The New government needs extend the entire data chain. If an AI system generates a response, document, or any other asset that reaches a customer, the organization needs to be able to understand what happened from the origin to that output.

In practical terms, end-to-end traceability should allow for the reconstruction of issues such as:

  • Source of the data: what were the sources that fed the model and who was responsible for that information.
  • Quality and transformations: what controls were applied, what modifications were made to the information and under what rules.
  • Dataset used: what set and version of the data were used to train, validate or run the model, so that it is possible to accurately reconstruct what information was involved at each stage.
  • Model version: which model produced the result and under what configuration.
  • Metrics used: what performance, quality and bias indicators were analyzed and what were the results.
  • Decisions and approvals: who validated the model, under what criteria and with what documentation.
  • Output generated: what response, data, conversation, document, or other asset emerged from the model and what treatment it subsequently received.
  • Subsequent changes: what modifications were made to the solution and what new validations were performed.

This point is especially relevant because AI incorporates types of assets that were not always present in traditional governance schemes.

For example, a chatbot can generate a document for a customer that did not previously exist in a traditional database, but is now a company-generated asset and may need to be preserved, tracked, and audited. 

The same applies to conversations. It's no longer enough to simply manage tables and structured records. It may be necessary to understand what a person asked, what an AI agent responded, what information it relied on, and what content left the organization as a result of that interaction.

Guide to assembling your evidence folder

The purpose of an evidence folder is to enable an organization to reconstruct, using verifiable information, how a model was developed, validated, deployed, and monitored. To achieve this, it is necessary to gather and organize information that documents the various stages of its lifecycle, the decisions made, and the controls implemented. 

These are some of the main aspects you should consider:

1. Description and purpose of the model. Detail what need it solves, what process it intervenes in, what its scope is, and what decisions it makes or contributes to making.

2. Those responsible. Identify who is responsible for the model, the data and its operation, as well as the participation of the business, technology, risk, security and compliance areas.

3. Data and lineage. Document the sources of information, their origin and path, the quality rules applied, the transformations made and the versions used.

4. Model documentation. Record the model version, its methodology, the variables used, the training process, and the criteria defined for its validation and approval.

5. Performance Metrics. Retain the indicators used to evaluate the model's performance and demonstrate whether it achieves the expected results for the use case.

6. Bias and fairness metrics. Document which groups and potential biases were assessed, what metrics were used, why they were selected, and what the results were.

7. Validation evidence. Gather the tests and controls performed, their results, the observations identified, and the necessary approvals before going into production.

8. Outputs and assets generated. Record, where appropriate, the responses, conversations, documents, or other content produced by the model that need to be traced or audited.

9. Monitoring. Define which aspects of the model are monitored once in production, how frequently, using which indicators, and how potential deviations are managed.

10. Change history. Maintain a record of updates, new versions and modifications made, along with the validations performed after each change.

The logic is simple: if a regulator asks a question, the answer shouldn't depend on retrospectively reconstructing information distributed among teams, tools, and vendors. There should be a chain of evidence.

AI incorporates types of assets that were not always present in traditional governance schemes.

Governing AI involves expanding the boundaries of data governance

Artificial intelligence does not replace the need for data governance., makes it wider.

The models need clean, secure, traceable, and high-quality information. But once they begin operating, they also generate new results, content, and data that re-enter the ecosystem. This cycle necessitates extending governance.

The challenge is especially relevant for organizations that are already bringing models into production, because the greater the impact of AI on customers, processes, or sensitive decisions, the greater the need to know what is happening and to be able to demonstrate it.

AI adds a new layer, and it's no longer enough to know that the data is governed. It's necessary to be able to demonstrate how that data fed into a model, what that model produced, how its behavior was evaluated, and what evidence exists to justify it.

This is the point where bias metrics cease to be merely a technical discussion and become a tool for compliance, algorithmic transparency, and risk management. For regulated organizations, preparing them before an audit can make a considerable difference.

Prepare your evidence folder before the next requirement

If your organization needs to know how prepared it is to respond to an audit, you can also Schedule a private working session with the Data & AI team at IT Patagonia to identify gaps and evolution priorities

en_US