Generative AI Engineer

Afficiency

New York City, UShybrid$100k-$120k/yrPosted Jul 13, 2026
Posting intelligenceActively listed

Skills

elasticsearchkubernetesdatabricksregressionsnowflaketerraformlangchainworkabledockerpythonflaskazuregooglecloudawsllmml

About the role

Company Description

Afficiency is a rapidly growing Insurtech startup whose mission is to provide life insurance to everyone on the platforms they already trust. Located in NYC, we design life insurance products that can be purchased entirely digitally and can be easily embedded into distribution platforms or with agents who sell insurance. The company is experiencing rapid growth and is well-funded. We are looking for new team members to join us on our journey to shake up the life insurance industry. We need individuals who bring passion, curiosity, and a desire for excellence.

Job Description

As a Generative AI Engineer at Afficiency, you will be responsible for designing, developing and deploying Generative AI solutions that enhance our core product platforms and client implementations. You will work closely with engineering, data science, and infrastructure teams to build scalable AI-driven applications using Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), model fine-tuning, and reinforcement learning approaches.

This role is ideal for someone who is based in the NYC Metro Area, passionate about building real-world GenAI applications and bringing them into production, while continuously improving performance, reliability, and user outcomes.

Qualifications

Responsibilities

Deliver GenAI solutions end-to-end

Own technical design and implementation of GenAI applications from discovery through production handoff.

Build APIs/services that integrate with enterprise systems and analytics platforms.

Implement enterprise-grade RAG

Build retrieval systems with hybrid search, filtering, re-ranking, query rewriting, and context optimization.

Implement permission-aware retrieval aligned to entitlements and data access policies.

Establish evaluation and quality controls.

Define metrics for retrieval quality and answer grounding (faithfulness, citation accuracy, coverage).

Create golden datasets, regression tests, and automated evaluation harnesses.

Operationalize GenAI (LLMOps)

Instrument observability (latency, cost, token usage, error rates) and implement safe rollout patterns.

Implement caching, rate limiting, fallbacks, and incident-ready operational practices.

Partner across teams to land solutions

Collaborate with business owners to translate requirements into workable designs.

Work with Security/Compliance to embed guardrails, auditability, and privacy controls.

Provide clear documentation and implementation of playbooks to enable internal teams' post-engagement.

Must Have

Education: Master's degree or equivalent experience required

3+ years in software engineering, data engineering, ML engineering, or applied AI, including recent GenAI delivery in production.

Demonstrated expertise in RAG system design and optimization, including:

chunking + metadata enrichment, hybrid search, re-ranking, retrieval evaluation

grounding/citations and hallucination mitigation patterns

Strong Python and backend engineering skills (FastAPI/Flask), plus strong SQL.

Experience working in regulated or security-conscious environments, with knowledge of:

access controls/entitlements, data privacy, logging/audit trails, secure SDLC practices

Proven ability to work effectively as an IC consultant:

communicate architecture decisions clearly

influence cross-functional stakeholders without direct authority produce high-quality documentation and handoff materials

Nice to Have

Fine-tuning experience (SFT, LoRA/QLoRA) and familiarity with preference optimization concepts (DPO/RLHF)

Vector/hybrid search platforms: Elasticsearch/OpenSearch vector, FAISS, Pinecone, Weaviate, Milvus

LLMOps tooling: MLflow/W&B, OpenTelemetry, prompt registries, evaluation frameworks

Cloud + platform: AWS/Azure/GCP, Docker/Kubernetes, Terraform

Tools & Technologies

LLM frameworks: LangChain, LlamaIndex, Semantic Kernel (optional)

Vector/hybrid search: Open to different skillsets

Data: (Snowflake/Databricks/warehouse), event pipelines, document stores

Observability: logging/tracing/metrics, dashboards, alerting

Additional Information

What We Offer

Competitive salary with equity options

Robust health, dental, and vision benefits for employee and dependents

401k matching contributions

Generous PTO policy

Provided work-from-home equipment

Compensation

This LLM Engineer role pays $100k-$120k/yr. Within typical range for llm engineer roles in United States.

Questions about this role

Click "Apply with AI Applyd" above. We auto-fill the application from your resume and answer screening questions in seconds. No copy and paste, no juggling tabs.

Compensation for LLM Engineer roles in United States varies widely by seniority, employer size, and remote vs onsite arrangement. Check the salary range on this listing when published, or browse our LLM Engineer hub for United States medians across recent openings.

Most applications complete in under 90 seconds. You can track the status in your dashboard and watch the screenshot proof land the moment the application submits.

AI Applyd supports Greenhouse, Lever, Ashby, Workday, iCIMS, SmartRecruiters, Personio, Teamtailor and other major ATS platforms. If we can submit through the platform, we do.

Want AI Applyd to auto-apply to roles like this?

We tailor your resume per posting, fill the forms, and track replies for you.