z
zffgsr

Jose Rocha

@zffgsr

PDF and invoice data extraction into Excel

Portogallo
Inglese, Portoghese
Alcune informazioni sono riportate in lingua inglese.
Chi sono
Systems engineering graduate at ISEP in Porto. In my internship at an IT consultancy I built document pipelines: tools generating Word and Excel files from raw project data, down to the XML when templates fought back. Document extraction is that problem in reverse, and it's what I offer here. PDFs, scans and invoices in; clean Excel, CSV or JSON back, with a note saying which pages I couldn't read rather than guesses filling the gaps. Excel certified (FCUP), Cambridge C2 English. I work in English and Portuguese. I don't do handwriting recognition - accuracy isn't good enough to charge for.... Continua a leggere

Competenze

z
zffgsr
Jose Rocha
offline • 
Tempo di risposta medio: 1 ora

Consulta i miei servizi

Inserimento dati
I will extract data from PDF and scanned documents into excel, CSV or json
Inserimento dati
I will extract invoice and receipt data into excel with ai ocr

Portfolio

Esperienza lavorativa

OPTIMIZER

IT Trainee

OPTIMIZER • Full time

Mar 2026 - Jul 20264 mos

Built document automation tooling for an IT consultancy's project management practice - the machinery that turns structured data into finished documents. What I did: - Wrote generation pipelines producing Word, Excel and PowerPoint files from raw project data, working at the OOXML level when templates would not cooperate: table cell properties, section breaks, image relationships, character escaping. - Built an Excel generator that emits live formulas rather than pre-computed values, so the output stays a working spreadsheet instead of a static report. - Mapped and documented 19 end-to-end business processes, then turned the repeatable ones into tools usable by staff with no technical background. - Enforced one hard rule across every tool: never generate a document from assumed data. If an input was missing, the tool asked instead of inventing. Why this matters for the work I offer here: Document extraction is the same problem in reverse. Building generators taught me exactly how documents fall apart - where tables break across pages, how merged cells and inconsistent headers happen in the first place, why a date column arrives as text in three different formats. That is the knowledge I use pulling data back out of PDFs, scans and invoices. It also taught me where automated output goes wrong quietly, which is the expensive kind. That is why every delivery comes with a note listing what I could not read with confidence, instead of plausible guesses filling the gaps.