U S M A N   A H S A N
Usman Ahsan · AI Application Engineer

AI Application Engineer · LLM-powered products · Remote (UTC+5)

I ship LLM features into real products — not notebooks.

Full-stack engineer wiring LLM APIs, RAG pipelines and AI agents into production Next.js apps on AWS. A real-time C++/Rust flight-simulation background carries straight into streaming, latency budgets and inference cost. I’m building a RAG + agent portfolio in the open — actively shipping, not “coming soon”.

Usman Ahsan Pervez — AI Application Engineer (LLM-powered products)
Islamabad–Rawalpindi, Pakistan· UTC+5· Remote / EOR-friendly · 4–6h overlap with GST & MYT· English (professional working)
OpenAIAnthropicRAGpgvectorLangChainMCPNext.jsTypeScriptAWSEvalsOpenAIAnthropicRAGpgvectorLangChainMCPNext.jsTypeScriptAWSEvals

About

What I actually do with LLMs

Usman Ahsan Pervez

I’m a full-stack engineer who treats LLMs as a production concern, not a demo. Over 4+ years I’ve shipped MERN and Next.js apps on AWS, and from March 2024 I led full-stack delivery at Apex IT Solutions where AI-assisted workflows and LLM API integration entered real client apps. My current role engineering real-time flight-simulation software in C++/Rust gives me a systems lens that maps directly onto AI-app concerns — token streaming, latency budgets, backpressure and inference cost. Right now I’m building my RAG and agent portfolio in the open: I believe LLM features should ship with a small eval set, guardrails (I track the OWASP LLM Top 10), observability and a real cost-per-query number — not stay in a notebook. My stack is Next.js/TypeScript + Node + Postgres/pgvector on AWS. I’m based in Pakistan (UTC+5) with 4–6 hours of daily Gulf and Malaysia overlap, and I work remote or via EOR.

Focus

Capabilities

LLM API integration

OpenAI/Anthropic wired into Next.js/Node with streaming responses and structured output.

RAG pipelines

Chunking, embeddings, pgvector/Chroma retrieval and grounded answers with citations.

AI agents & tools

Tool-using / function-calling agents, multi-step workflows, MCP-style tool servers.

Evals & guardrails

Small eval sets, quality measurement, input/output guardrails (OWASP LLM Top 10 aware).

Cost & latency

Model routing (cheap vs frontier), caching, token budgeting, $/1k-request reporting.

Prototyping

Fast AI-feature prototypes built to graduate into production code.

The AI Lab

Building in the open

In progress — not “coming soon”. Each item ships with a live URL, public GitHub and one real metric (eval score / cost-per-query) when v1 lands. Follow the build on GitHub.

In progress

Chat-with-your-docs (RAG)

Next.js · pgvector · OpenAI/Anthropic

Grounded Q&A with streaming + citations. Will publish a 20–50-pair eval set and a quality score.

Planned

Multi-tool agent + MCP server

TypeScript · MCP · guardrails

Tool-using assistant with a hand-written MCP-style tool server, guardrails and logged runs.

Planned

Cost-aware model router

TypeScript · caching · routing

Routes simple vs complex queries to cheap vs frontier models and reports $ per 1k requests.

Work

Selected work

Every card is labelled by its true state — shipped, in progress, or planned. No invented metrics.

In progress

Grounded “Chat-with-your-docs” RAG app

Next.js · TypeScript · pgvector · OpenAI/Anthropic · AWS

Retrieval-augmented Q&A over a real document set with streaming answers + inline citations. Building in the open with a public 20–50-pair eval set and a published quality score.

Planned

Multi-tool AI agent + MCP tool server

TypeScript · Node · MCP · function calling · guardrails

A tool-using assistant backed by a hand-written MCP-style tool server, with input/output guardrails and logged runs. Scoped on the public roadmap.

Planned

Cost-aware model router

TypeScript · Node · OpenAI/Anthropic · caching

Routes simple vs complex queries to cheap vs frontier models and reports $ per 1k requests — making the cost story concrete.

Shipped

Delivrex — Medical Logistics Platform

Full-stack web · responsive UI · backend

Production full-stack build (via Apex) — representative of the app surface where I integrate LLM features (support assistants, document search). No AI metrics claimed for this build.

Shipped

Homesware — UK E-commerce

E-commerce · catalog · checkout

Shipped commerce build (via Apex) — the natural home for LLM product search, recommendations and support agents in future work.

The edge

The real-time systems edge

As a flight-simulation software engineer I work in performance-critical, real-time C++ where every millisecond and frame budget matters. That maps directly onto AI-app concerns — and it’s why my LLM features stay fast and cheap.

In flight sim

Real-time loops

In production

p95 latency budgets & token streaming UX

In flight sim

Deterministic timing

In production

Backpressure, batching & throughput

In flight sim

Telemetry-first

In production

Observability & cost-per-query discipline

Process

How I ship an LLM feature

Define task + success criteria

Decide what “good” means before writing the feature.

Build a small eval set

20–50 cases first — evals, not vibes.

Prototype & measure

Prompt / RAG / agent, scored against the evals.

Guardrails + observability + cost

Safety, logging and a cost-per-query number.

Ship behind a flag & iterate

Watch quality and $/query in production.

Credentials

Education & certifications

2017 – 2021

BS, Electrical & Electronics Engineering

FAST-NUCES (National University of Computer and Emerging Sciences), Islamabad

2022 – 2023

Web 3.0 & Metaverse Developer (Certification)

Panaverse

English (professional working) · Urdu (native)

FAQ

Questions, answered

Can I hire a remote AI application engineer?

Yes — remote from Pakistan (UTC+5) with 4–6 hours of daily overlap with Gulf and Malaysian hours; contract and EOR friendly.

What kind of AI work do you do?

I wire LLM APIs, RAG pipelines and AI agents into production Next.js applications on AWS — application-layer AI engineering rather than model training.

How does your background help with AI products?

A real-time C++/Rust flight-simulation background carries straight into streaming, latency budgets and inference-cost control — the parts that make AI features feel fast and stay affordable.

Is your AI portfolio live or coming soon?

I’m building a RAG and agent portfolio in the open — actively shipping, with in-progress work labelled honestly rather than hidden behind “coming soon”.

Contact

Have an AI build in mind?

Remote / EOR-friendly · 4–6h overlap with GST & MYT · Islamabad–Rawalpindi, Pakistan (UTC+5).