info@telentra.online +1 343 655 9893
LLM Engineer

LLM Engineer

New York, NY, USA Contract Full Time
RemoteRemoteData Science

Job Description

Our client, an international AI development company based in New York, is currently seeking a "LLM Engineer (RAG, VectorStores, and Protocol Engineering)" to lead strategic product development efforts in a fast-paced and collaborative environment.

This role will focus on implementing scalable vector store integrations, building retrieval pipelines, and enabling advanced communication protocols between intelligent agents.

Key Responsibilities

RAG & VectorStore Systems:

Build and maintain end-to-end RAG pipelines for context-augmented generation

Integrate and optimize vector databases (FAISS, Pinecone, Weaviate, Milvus)

Support the engineering backbone of agentic and RAG-based AI systems

Protocol Engineering:

Implement Model Context Protocol (MCP) to maintain stateful LLM interactions

Build Agent-to-Agent (A2A) communication layers for multi-agent orchestration

Enable persistent memory and context sharing across model calls

Platform Enablement:

Collaborate with cross-functional teams to productionize models and workflows

Ensure seamless data flow and model integration across services

Qualifications & Skills

Deep experience with vector databases and RAG architecture

Familiarity with MCP (Model Context Protocol) and A2A (Agent-to-Agent) design patterns

Solid background in Python, cloud-based ML pipelines, and containerization tools

Experience in operationalizing LLM-based systems in production

Detail-oriented with a strong engineering mindset

Effective communicator with technical and non-technical stakeholders

Self-driven and adaptable in a fast-paced R&D environment

Very strong English communication skills, both written and verbal (essential for global collaboration)

Experience with Contract Analysis and automated document understanding

Nice to Have

Experience with invoice parsing / invoice understanding, including extracting structured data from financial documents

Experience in NLP-to-SQL, natural language querying, or generating structured database queries from unstructured text

Familiarity with legal-tech, document intelligence, or enterprise knowledge extraction systems

Experience with observability, vector hygiene, and evaluation frameworks for RAG/agentic systems

Hands-on work with high-throughput, low-latency AI pipelines

What you'll do

  • Design, train, and evaluate large language models and agentic pipelines
  • Build production-grade inference and orchestration systems
  • Collaborate with product and engineering teams to ship AI features end-to-end
  • Run experiments, benchmark models, and iterate on prompt and fine-tuning strategies
  • Own model quality, latency, and cost across the stack
  • Mentor engineers and contribute to internal ML platform improvements

What we're looking for

  • Strong background in machine learning, NLP, or deep learning
  • Hands-on experience with PyTorch, Hugging Face, or similar frameworks
  • Experience deploying models to production
  • Solid Python engineering skills
  • Familiarity with RAG, vector databases, and agent frameworks
  • Strong communication skills in English

Nice to have

  • Published research or open-source contributions
  • Experience with distributed training or GPU optimization
Location: New York, NY, USA

Apply with confidence

Every Telentra role is vetted, and every applicant is treated confidentially.

1,200+
Placements delivered

Successful hires across USA, UK, South Africa and Türkiye.

92%
Client retention

Long-standing partnerships with employers who hire with us again and again.

90-Day
Replacement guarantee

If a hire doesn't work out in the first 90 days, we replace at no extra cost.

100%
Confidential search

Discreet, NDA-backed head hunting for sensitive and executive briefs.

GDPR
& EEO compliant

Data-privacy compliant processes and fair, bias-aware shortlisting.

Vetted
Candidate screening

Right-to-work, reference and credential checks on every shortlist.