info@telentra.online +1 343 655 9893
LLM Engineer

LLM Engineer

New York, NY, USA Permanent Full Time
RemoteRemoteData Science

Job Description

We are looking for an "LLM Engineer" for our client, an international AI development company based in New York, who will work alongside data scientists, backend engineers, and product stakeholders to design and deploy production-grade Generative AI systems.

Key Responsibilities Build and maintain production-ready backend services supporting Generative AI applications. Design and implement Retrieval-Augmented Generation (RAG) pipelines end-to-end, including ingestion, embedding, retrieval strategy, and evaluation. Orchestrate and optimize machine learning and LLM-based workflows for reliability, latency, and cost efficiency. Develop scalable APIs and service layers enabling integration of AI capabilities into enterprise platforms. Partner with cross-functional teams to deploy AI solutions into existing architectures and operational environments.

Required Qualification 5+ years of backend engineering experience using Python in production environments. Strong hands-on experience building applications powered by Large Language Models. Solid understanding of retrieval architectures, vector databases, embeddings, and prompt orchestration strategies. Experience working with SQL databases (PostgreSQL, MySQL) in performance-sensitive systems. Practical machine learning experience, including feature engineering, evaluation workflows, and model lifecycle considerations. Comfort working in Linux-based environments and using Bash for automation.

Nice to Have Experience building APIs with FastAPI, Flask, or Django. Familiarity with asynchronous Python and microservices-based architectures. Experience deploying containerized workloads using Docker and Kubernetes. Exposure to MLOps practices, including model versioning, monitoring, and reproducibility. Hands-on experience with AWS services such as EC2, S3, Lambda, and SageMaker. Experience designing CI/CD pipelines for ML or AI-enabled systems.

About Our Client Our client is a people-focused organization dedicated to developing impactful products and services that create meaningful value for customers and communities. They foster a collaborative, respectful, and inclusive work environment where employees are encouraged to take ownership, contribute ideas, and grow professionally. The company supports flexibility and work–life balance while maintaining strong performance and accountability standards.

Benefits & Wellbeing Compensation for this role is determined based on competitive market data and may vary depending on geographic location, experience, skills, and qualifications. Specific details will be discussed during the interview process. The company offers a comprehensive benefits package designed to support employees’ well-being, financial security, and professional development. Benefits may include medical, dental, and vision coverage, retirement plan contributions, paid time off, flexible working arrangements, and opportunities for career growth, in accordance with company policies and applicable local regulations. It is an equal opportunity employer and considers all qualified applicants without regard to legally protected characteristics. Applicants must have the legal right to work in the country of employment.

What you'll do

  • Design, train, and evaluate large language models and agentic pipelines
  • Build production-grade inference and orchestration systems
  • Collaborate with product and engineering teams to ship AI features end-to-end
  • Run experiments, benchmark models, and iterate on prompt and fine-tuning strategies
  • Own model quality, latency, and cost across the stack
  • Mentor engineers and contribute to internal ML platform improvements

What we're looking for

  • Strong background in machine learning, NLP, or deep learning
  • Hands-on experience with PyTorch, Hugging Face, or similar frameworks
  • Experience deploying models to production
  • Solid Python engineering skills
  • Familiarity with RAG, vector databases, and agent frameworks
  • Strong communication skills in English

Nice to have

  • Published research or open-source contributions
  • Experience with distributed training or GPU optimization
Location: New York, NY, USA

Apply with confidence

Every Telentra role is vetted, and every applicant is treated confidentially.

1,200+
Placements delivered

Successful hires across USA, UK, South Africa and Türkiye.

92%
Client retention

Long-standing partnerships with employers who hire with us again and again.

90-Day
Replacement guarantee

If a hire doesn't work out in the first 90 days, we replace at no extra cost.

100%
Confidential search

Discreet, NDA-backed head hunting for sensitive and executive briefs.

GDPR
& EEO compliant

Data-privacy compliant processes and fair, bias-aware shortlisting.

Vetted
Candidate screening

Right-to-work, reference and credential checks on every shortlist.