Grounded “Chat-with-your-docs” RAG app
Retrieval-augmented Q&A over a real document set with streaming answers + inline citations. Building in the open with a public 20–50-pair eval set and a published quality score.
AI Application Engineer · LLM-powered products · Remote (UTC+5)
Full-stack engineer wiring LLM APIs, RAG pipelines and AI agents into production Next.js apps on AWS. A real-time C++/Rust flight-simulation background carries straight into streaming, latency budgets and inference cost. I’m building a RAG + agent portfolio in the open — actively shipping, not “coming soon”.
About
I’m a full-stack engineer who treats LLMs as a production concern, not a demo. Over 4+ years I’ve shipped MERN and Next.js apps on AWS, and from March 2024 I led full-stack delivery at Apex IT Solutions where AI-assisted workflows and LLM API integration entered real client apps. My current role engineering real-time flight-simulation software in C++/Rust gives me a systems lens that maps directly onto AI-app concerns — token streaming, latency budgets, backpressure and inference cost. Right now I’m building my RAG and agent portfolio in the open: I believe LLM features should ship with a small eval set, guardrails (I track the OWASP LLM Top 10), observability and a real cost-per-query number — not stay in a notebook. My stack is Next.js/TypeScript + Node + Postgres/pgvector on AWS. I’m based in Pakistan (UTC+5) with 4–6 hours of daily Gulf and Malaysia overlap, and I work remote or via EOR.
Focus
OpenAI/Anthropic wired into Next.js/Node with streaming responses and structured output.
Chunking, embeddings, pgvector/Chroma retrieval and grounded answers with citations.
Tool-using / function-calling agents, multi-step workflows, MCP-style tool servers.
Small eval sets, quality measurement, input/output guardrails (OWASP LLM Top 10 aware).
Model routing (cheap vs frontier), caching, token budgeting, $/1k-request reporting.
Fast AI-feature prototypes built to graduate into production code.
The AI Lab
In progress — not “coming soon”. Each item ships with a live URL, public GitHub and one real metric (eval score / cost-per-query) when v1 lands. Follow the build on GitHub.
Grounded Q&A with streaming + citations. Will publish a 20–50-pair eval set and a quality score.
Tool-using assistant with a hand-written MCP-style tool server, guardrails and logged runs.
Routes simple vs complex queries to cheap vs frontier models and reports $ per 1k requests.
Work
Every card is labelled by its true state — shipped, in progress, or planned. No invented metrics.
Retrieval-augmented Q&A over a real document set with streaming answers + inline citations. Building in the open with a public 20–50-pair eval set and a published quality score.
A tool-using assistant backed by a hand-written MCP-style tool server, with input/output guardrails and logged runs. Scoped on the public roadmap.
Routes simple vs complex queries to cheap vs frontier models and reports $ per 1k requests — making the cost story concrete.
Production full-stack build (via Apex) — representative of the app surface where I integrate LLM features (support assistants, document search). No AI metrics claimed for this build.
Shipped commerce build (via Apex) — the natural home for LLM product search, recommendations and support agents in future work.
The edge
As a flight-simulation software engineer I work in performance-critical, real-time C++ where every millisecond and frame budget matters. That maps directly onto AI-app concerns — and it’s why my LLM features stay fast and cheap.
p95 latency budgets & token streaming UX
Backpressure, batching & throughput
Observability & cost-per-query discipline
Process
Decide what “good” means before writing the feature.
20–50 cases first — evals, not vibes.
Prompt / RAG / agent, scored against the evals.
Safety, logging and a cost-per-query number.
Watch quality and $/query in production.
Credentials
FAST-NUCES (National University of Computer and Emerging Sciences), Islamabad
Panaverse
English (professional working) · Urdu (native)
FAQ
Yes — remote from Pakistan (UTC+5) with 4–6 hours of daily overlap with Gulf and Malaysian hours; contract and EOR friendly.
I wire LLM APIs, RAG pipelines and AI agents into production Next.js applications on AWS — application-layer AI engineering rather than model training.
A real-time C++/Rust flight-simulation background carries straight into streaming, latency budgets and inference-cost control — the parts that make AI features feel fast and stay affordable.
I’m building a RAG and agent portfolio in the open — actively shipping, with in-progress work labelled honestly rather than hidden behind “coming soon”.
Contact
Remote / EOR-friendly · 4–6h overlap with GST & MYT · Islamabad–Rawalpindi, Pakistan (UTC+5).