AI Developer
RemotePhilippines
PhilippinesRemoteFull Time
Full Time
Job Summary
Design and ship end-to-end document extraction pipelines that ingest PDFs and images, converting raw inputs into clean, queryable JSON or relational data. Build FastAPI services with async endpoints for large batches, handling timeouts and retries while integrating with S3, Pub/Sub, and downstream systems. Implement OCR strategies ranging from native text layers to layout-aware parsing and image-based recognition, applying prompt engineering with OpenAI and Anthropic APIs to extract structured data from complex layouts like tables and multi-column text.
Required Qualifications
- Hands-on experience building OCR or document extraction pipelines in production
- Strong Python skills — clean, maintainable code with proper error handling
- Practical experience with FastAPI (routing, dependency injection, async, middleware)
- Prompt engineering experience with OpenAI or Anthropic APIs — not just calling the API, but designing reliable extraction chains
- Familiarity with PDF internals: text layers, bounding boxes, embedded fonts, page structure
- Experience with vision-language models (GPT-4V, Claude 3 vision) for image-heavy documents
- Comfortable with AI-driven development (fully Developer-in-the-loop)
- Experience with Cloud OCR: AWS Textract, Google Document AI, or Azure Form Recognizer
- Experience with LangChain, LlamaIndex, or similar orchestration frameworks
- Vector search / RAG pipelines for document Q&A
- Docker, basic CI/CD, and cloud deployment (AWS / GCP / Azure)
- Experience with agentic workflows (tool use, multi-step LLM chains)
- Write, test, and iterate prompts for OpenAI (GPT-4o, GPT-4 Turbo) and Anthropic (Claude) models
- Apply prompt engineering techniques: chain-of-thought, few-shot, structured output forcing, tool use/function calling
- Build extraction agents that combine OCR output with LLM reasoning for ambiguous or complex documents
- Evaluate and benchmark prompt strategies; document what works and why
Desired Qualifications
- Native PDF text extraction | pdfplumber, PyMuPDF, pdfminer — fast, accurate when the text layer exists; must detect and fall back when it doesn't
- Layout-aware parsing | Preserve reading order across columns, tables, and mixed content blocks
- Image-based OCR | Tesseract, EasyOCR, or cloud OCR (AWS Textract, Google Document AI, Azure Form Recognizer) for scanned inputs
- Table extraction | Structured output from tabular data — row/column alignment, merged cells, nested tables
- Output formats | Flat text, structured JSON, markdown — output type driven by downstream use case
- You care about output quality — you're not happy until the extraction is clean and reliable
- You test your prompts like you test your code — systematically, with real data
- You know when to use an LLM and when a regex is the better tool
- You can communicate tradeoffs clearly to non-technical stakeholders
- You're comfortable in a fast-moving, remote-first environment
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.