API CoA email parser — from the supplier PDF to the batch record without retyping.

The certificate of analysis (CoA) for the pharma API arrives as a PDF by email and almost always someone retypes it into the electronic batch record (EBR), with the transcription risk that implies. iLEAN Connect captures the email at second zero, extracts the specifications from the PDF as an anchored task and proposes the posting to the EBR. QA signs — Connect does not.

← See all iLEAN Connect capabilities

Pharma QA lead reviewing an API certificate of analysis on screen — iLEAN Connect has extracted the PDF from the email and posted it to the electronic batch record, waiting for signature
The problem

The critical data lives in a PDF that every supplier formats its own way.

Every API batch entering a pharma plant brings its CoA: identity, purity, assay, impurity profile, moisture, residual solvents. It is the document that says whether the batch can go into production and under which specification. The most critical piece of the whole upstream process. And yet:

  1. It arrives by email. Not by EDI, not by API, not through a supplier portal. By email, as an attachment.
  2. The PDF is formatted by each supplier its own way. An exported spreadsheet table, running text, a scan with a stamp, the same value in three languages. There is no standard across suppliers.
  3. Someone retypes it into the EBR. A warehouse technician, a quality control analyst — somebody reads it and writes it into another system. Every time. Every batch.

The risk is not the QA signature — QA does its job. The risk is transcription: the value read wrong because the decimal comma sat somewhere else, the unit mixed up (ppm vs %), the result pasted into the previous batch's row. And the real friction is QA time spent typing instead of reviewing.

How it fits the IRIS system

iLEAN Connect — the filler between the supplier mailbox and the EBR.

The CoA is not plant data — it is data that arrives from outside, and it hardly ever reaches the decision maker in time. Connecting the supplier email to the EBR was always possible, but joining heterogeneous formats cost a fortune, so almost nobody did it. That is exactly what AI knocked to the floor. Connect is the piece that fills that gap — without ripping out the ERP, without ripping out the EBR.

Connect reads the CoA. The agent recontextualises it into the EBR schema. QA reviews it side by side with the PDF. QA signs — never the other way round.

How Connect works applied to the pharma API CoA:

  • Listening outward. Connect watches the API supplier mailbox at second zero. Nobody has to forward anything or call a meeting: the CoA enters the system the very second the supplier sends it.
  • Anchored extraction. It identifies the attachment as a CoA (by sender, subject or content), applies the right extractor for that supplier and deposits the result into a normalised schema: parameter, value, unit, specification, result, batch, analysis date. The AI does not generate free data; it recontextualises what is already in the PDF.
  • Multilingual without touching the original document. If the CoA comes in German, Hindi or Chinese, the extraction deposits the value in the EBR language — but the original PDF is stored intact as evidence, with its hash and its timestamp.
  • A proposal to QA, not a direct commit. The record enters the EBR in a «pending QA review» state, with the PDF attached, the extraction model used and a visible diff: what the PDF says and what is about to be posted. QA reviews, adjusts if needed and signs. The signature stays Annex 11 / 21 CFR Part 11 ready if your EBR already is.

Connect carries the data, the agent structures it, the person signs. That is the rule. This is exactly what the IRIS system calls the Connect piece: the plant's ear turned outward.

See the full IRIS architecture →

Before and after

Retyping the CoA vs. an anchored parser with QA signature.

AspectManual retyping into the EBRWith iLEAN Connect (CoA parser → EBR)
Time from CoA to EBRMinutes per batch × number of suppliersSeconds — the proposed data is ready for QA
Transcription riskDecimal comma, unit, previous batch's rowExtraction anchored to the PDF, diff visible to QA
Supplier formatEach one its own — the operator adaptsOne extractor per supplier, relearned without code
CoA languageMental translation by the operatorMultilingual extraction into the EBR language
Traceability for the auditorPDF in a folder, EBR in a system, no linkPDF + hash + model + diff anchored to the batch
GMP electronic signatureQA signs after retypingQA signs after review — the retyping disappears
Impact estimate

Impact estimate for your plant — to be validated with your numbers.

The block below is an estimate to be validated with the actual data of your plant. We put it on the table so the committee has an order of magnitude; we refine it during the diagnostic.

  • Pharma API plant with 10–20 active suppliers, CoA as PDF by email, EBR/MES already deployed and GMP validated, several thousand batches a year.
  • Connect pilot with 2–3 typical suppliers (the Pareto of the CoAs): first value expectable within a few weeks — the extraction already feeds the EBR in a «pending QA» state.
  • Indicative payback between 4 and 9 months, depending on QA hours spent retyping and the cost of investigating transcription errors found in audit.
  • Hard lever: ≥ 30% reduction of QA time spent typing — turned into QA time spent reviewing, which is where QA adds value. A single deviation avoided from a bad posting pays for the pilot.

And the quality director's reasonable doubt

«What if the AI hallucinates while reading the CoA and posts an incorrect value?» — hallucination is a problem of free generation, not of anchored tasks. In tasks where the AI merely recontextualises a value from a PDF into a normalised schema, the best models pushed the error below 1.5% [1]. And even so, what is critical is never signed automatically: Connect proposes the posting, QA signs. The three safety rings are there precisely for this.

[1] OpenAI paper "Why Language Models Hallucinate", 2025 — on the reliability of AI in anchored tasks.

Frequently asked

What people ask about the pharma API CoA parser

What is a CoA and why is it so hard to post it into the EBR?

The CoA (certificate of analysis) is the document the API supplier sends with every batch, carrying the results of the critical tests (identity, purity, assay, impurities, moisture). In pharma, every API batch entering the plant must be reflected in the electronic batch record (EBR) with its specifications — and the CoA is the source. The problem is that every supplier has its own PDF format (a table, running text, a different language), so almost always someone retypes or pastes it by hand, with the transcription risk that implies.

How does Connect read a CoA PDF whose format changes from supplier to supplier?

Connect captures the API supplier email at second zero, identifies the attachment as a CoA (by sender, subject or by the content itself) and applies the right extractor for that supplier. The extraction is anchored: the AI does NOT generate free data, it recontextualises data already present in the PDF into a normalised schema (parameter, value, unit, specification, result, batch). If the format changes, it is relearned without touching code. That is exactly where modern AI made affordable what was unfeasible 10 years ago.

Does this work for Annex 11 / GMP and electronic signature?

The parser on its own does not sign — and that is precisely the design. Connect posts the extracted data into the EBR/MES in a «pending QA review» state, with the original PDF attached, the document hash, the capture timestamp and the extraction model used. Quality Assurance reviews the proposal side by side with the PDF and signs electronically. The signature stays Annex 11 / 21 CFR Part 11 ready if your EBR already is — Connect does not break what you have, it feeds it faster.

What if the CoA arrives in German, Hindi or Chinese?

It is translated at extraction level, not at document level. Today's AI is exceptionally good at what the book calls anchored tasks — reading a value in one language and depositing it into the EBR schema in the system language, preserving unit and specification. The original PDF is stored intact as evidence and the extracted value keeps a note of the model used. If the CoA comes digitally signed by the supplier (CFDI, eIDAS, etc.), Connect verifies it and records it in the batch dossier.

How long does it take to glue Connect between the supplier mailbox and our EBR?

Bringing a first CoA live — one supplier, one API, one specific EBR/MES — usually delivers first value within a few weeks. The integration weighs more on the EBR side (rarely on the email: email is standard) than on the extraction side. After the first supplier, adding others is marginal cost — the extractor is trained with two or three samples of the new format and it is done. Send us your plant data and we return an order of magnitude within 48h.

Let's talk

Tell us your case and within 48h we will send you the estimated ROI of this AI project for your pharma API plant.

We work on the real data of your plant, not on ours. Diagnostic with no commitment.

Request estimated ROI within 48h See iLEAN Connect