This job is no longer available
This job expired on 19/08/2026. It no longer accepts applications.
AI Engineer – AI‑First Backend & LLM Specialist
VentureDive · Lahore
Job description
About the role
We are seeking an AI Engineer who embraces an AI‑first mindset and has a solid backend development background. You will design, build, and operate production‑grade AI systems where intelligence is baked into the core architecture rather than added as an afterthought.
Key responsibilities
- Develop scalable backend services with Python frameworks such as FastAPI, Flask, or Django to support AI‑driven applications.
- Implement, fine‑tune, and deploy large language models (e.g., GPT‑4, Claude, Llama 3, Mistral) using Hugging Face, PyTorch, or TensorFlow.
- Design advanced Retrieval‑Augmented Generation pipelines and integrate vector databases like Pinecone, Weaviate, Milvus, PGVector, or Chroma.
- Build AI agents and multi‑agent workflows with LangChain, LlamaIndex, Google ADK, BAML, Agno, or CrewAI.
- Own the full MLOps lifecycle: experiment tracking, CI/CD deployment, monitoring for hallucinations, and latency/cost optimization.
- Architect AI‑native systems that prioritize semantic search, embeddings, and context‑aware data flows.
- Deploy and optimize open‑source LLMs on‑premise using serving engines such as vLLM, SGLang, or Ollama.
Required profile
- Minimum 4 years of professional Python experience, including async programming, type hinting, and high‑performance backend patterns.
- Deep understanding of transformer architectures, attention mechanisms, and tokenization.
- Hands‑on experience with vector embeddings and similarity search optimization.
- Proficiency in building AI data pipelines, data cleaning, ingestion, and prompt engineering.
- Strong DevOps/MLOps background with Docker, Kubernetes, and cloud platforms (AWS, GCP, Azure). Experience with GPU orchestration and model quantization (GGUF, AWQ) is a plus.
Required skills
- Python
- FastAPI, Flask, Django
- Transformers, LLMs (GPT‑4, Claude, Llama 3, Mistral)
- Hugging Face, PyTorch, TensorFlow
- RAG pipelines
- Pinecone, Weaviate, Milvus, PGVector, Chroma
- LangChain, LlamaIndex, Google ADK, BAML, Agno, CrewAI
- vLLM, SGLang, Ollama
- Docker, Kubernetes
- AWS, GCP, Azure
- GPU orchestration, model quantization
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in Pakistan.
A question about this job?
Ask it here: you will get the full job summary by e-mail, right away.
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
VentureDive
Lahore