Author: Timothy Powell, CPA, CHCP | August 19, 2026
Artificial intelligence companies developing healthcare revenue-cycle tools need large volumes of real claims data. Two of the most valuable sources are the 837 electronic claim and the 835 electronic remittance advice.
The 837 reports what the provider billed. The 835 reports how the payer adjudicated that claim, including payments, denials, adjustments, deductibles, and coinsurance. When properly matched, these transactions can help train AI to identify underpayments, predict denials, estimate collectability, and recognize payer behavior.
Where Can AI Companies Obtain the Data?
The most practical sources are organizations already authorized to receive and maintain the transactions:
- Hospitals and physician groups;
- Health plans;
- Healthcare clearinghouses;
- Revenue-cycle management companies; and
- Companies purchasing or financing healthcare accounts receivable.
An AI company can contract with one of these organizations to develop or operate a defined application. Because the work may involve creating, receiving, maintaining, or transmitting protected health information (PHI), the AI company will frequently be a business associate.
The parties must execute a business associate agreement (BAA) specifying the permitted uses of the data, required safeguards, approved subcontractors, breach-reporting responsibilities, and what happens to the information when the engagement ends.
A BAA is Not Permission to Build any Model
Signing a BAA does not give an AI company unlimited authority to use a hospital’s claims.
If an AI vendor receives PHI to identify denials for Hospital A, it cannot automatically add those claims to a general model sold to Hospitals B through Z. The use must be authorized by the agreement and permitted by the HIPAA Privacy Rule.
Contracts should address model training directly, including the following:
- Whether client data may be used for training;
- Whether information from different clients may be combined;
- Who owns the trained model and derived data;
- Whether the vendor may retain information after termination; and
- How the vendor will prevent PHI from appearing in model outputs.
De-identification Creates Another Pathway
AI companies may use properly de-identified 835 and 837 data without treating it as PHI. HIPAA recognizes two de-identification methods: Safe Harbor and Expert Determination.
Safe Harbor requires the removal of specified identifiers. Expert Determination permits a qualified expert to determine that the risk of identifying an individual is very small.
Removing patient names is not sufficient. Claims may contain medical-record numbers, claim-control numbers, subscriber identifiers, service dates, addresses, free-text fields, rare diagnoses, and unusual combinations of procedures.
Safe Harbor may also eliminate dates needed to train timely-filing or payment-delay models. Expert Determination may preserve more useful relationships while still reducing re-identification risk.
Limited and Synthetic Data
A limited data set is not fully de-identified. It remains PHI, requires a data-use agreement, and may be used only for specified purposes such as research, public health, or healthcare operations.
Synthetic 835 and 837 files provide another option. They are useful for teaching transaction structure and testing software, although they may not accurately reproduce real payer behavior.
The Correct Division of Responsibility
AI companies can obtain claims data through carefully defined provider, payer, clearinghouse, or revenue-cycle relationships. They can also license properly de-identified data or generate synthetic transactions.
The essential rule is simple: standardized data is not public data. Permission to process an 835 or 837 for one customer is not necessarily permission to use it to train a commercial AI product.
Healthcare organizations must control the data, legal agreements must control its use, and AI companies must design their models around those limitations.
This article was originally published on RACmonitor.