Staff AI Engineer

Ho Chi Minh City, Vietnam • Full-time
AI Job Summary
  • 8+ years software engineering; 2+ years building and operating LLM-based systems in production.
  • Build multi-step agentic systems and choose topologies (e.g., MCP, deterministic workflows, orchestrator–worker, A2A).
  • Create eval sets, scoring rubrics (incl. LLM-as-judge), and regression measurement to guide architecture decisions.

Role Type

Permanent • Full-time • Mid-level Senior

Description

WE NEED YOU ON OUR TEAM!

We’re hiring a Staff Engineer to be the senior technical authority for AI across the organization. This is not a people-management role – it’s a consulting and architecture role with organization-wide reach: you design AI solutions, set the standards other engineers build to, drive evaluation discipline, and shape how AI enters our products from the design phase onward. You report directly to the CTO and multiply the impact of every squad working with AI. 

What You’ll Do

  • Advise and design AI solutions across the full spectrum – from single LLM calls to complex multi-agent infrastructure – recommending implementation approaches that account for fallbacks, cost/latency budgets, and rate-limit constraints (throttling, TPM/RPM).
  • Architect agentic systems with the right structure for the problem: MCP server design, deterministic workflows, single agents with tool access, orchestrator–worker or A2A multi-agent topologies.
  • Own the evaluation practice that drives technical decisions: eval sets, scoring rubrics, and regression measurement for non-deterministic outputs – and use that evidence to settle when to prompt, when to RAG, when to fine-tune, and when to build an agent, so architecture choices are backed by measurement rather than intuition.
  • Partner with Product from the design phase: consult on where AI genuinely adds value in a feature, what’s feasible at what cost, and what the failure modes will look like – so AI capability shapes the product spec instead of being retrofitted after.
  • Forecast the direction of AI in the next 2-3 years, and start to take actions today so Perform.AI remains technologically relevant in that future.
  • Raise the technical level of engineers across squads: review their designs and code before implementation, define shared standards they follow (prompt conventions, eval-driven development), design and deliver internal AI training programs, and make architecture decisions the whole team builds on.

WHAT YOU SHOULD BRING ALONG!

Core Qualifications

  • 8+ years of software engineering experience, with 2+ years building and operating LLM-based systems in production.
  • Hands-on experience designing and running LLM-based systems in production – beyond API calls: handling fallbacks, cost/latency trade-offs, rate limits (TPM/RPM), and structured output reliability at scale.
  • Experience building multi-step agentic systems and choosing the right topology for the problem – using frameworks like LangChain/LangGraph or equivalents (LlamaIndex, Pydantic AI, Claude Agents SDK), with working knowledge of MCP, context/state management, and agent failure modes.
  • Demonstrated evaluation discipline for non-deterministic systems: building eval sets, defining scoring rubrics (including LLM-as-judge approaches), measuring regression – using tools like LangSmith, Langfuse, or equivalents – and using eval evidence to drive architecture decisions (prompt vs. RAG vs. fine-tune vs. agent).
  • Strong backend fundamentals with the depth to drive the tech stack, not just work within it: expert-level Python, distributed systems, high-volume data pipelines, and the judgment to select and standardize frameworks, libraries, and infrastructure for AI workloads across the team.

Nice to Have

  • Experience in logistics, e-commerce, or other high-volume data domains.
  • Contributions to open source, technical writing, or conference talks.
  • Experience fine-tuning models (managed fine-tuning APIs or open-weight approaches like LoRA/PEFT) and judging when fine-tuning beats prompting or RAG on cost and quality.
  • Experience across the LLM deployment spectrum: working with managed cloud LLM services (AWS Bedrock, Azure OpenAI, Vertex AI, or equivalent) – quotas, regional availability, model version management – familiarity with deploying open-weight models (e.g., Llama, Qwen from Hugging Face) on your own infrastructure with an understanding of the cost/control trade-offs between the two.

WHAT YOU WILL RECEIVE IN RETURN!

We at Perform.AI are dedicated to being a platform for growth for all our team members, regardless of function and location.

  • The opportunity to work in a fast-growing, super exciting, and innovative business building the AI Commerce Operating System for global e-commerce brands. You will be the needle of success in the growth of a global product that will become a key platform behind successful e-commerce brands worldwide.
  • Fully sponsored Claude Code usage – As part of our commitment to staying at the bleeding edge of AI-driven engineering, we provide sponsored access to Claude Code, enabling you to build faster, smarter, and more efficiently with state-of-the-art AI support.
  • Submerge in the bleeding-edge of Agentic Development – We don’t simply supplement an existing process with AI tools, we rework every aspect of the software development life cycle, from requirement gathering, implementation, to testing. We believe the productivity multiplication of AI doesn’t just come from individual applications but also from an entire system that embraces it into every fabric.
  • An environment where everybody never stops growing and focuses on succeeding – we continuously work with you on your strengths and weaknesses across many important dimensions and look at ways for you to address them and further your development.
  • Join a passionate, global team at Perform.AI where you can continuously develop your skills and actively drive our mission and international success.

OUR BENEFITS!

Perform.AI is committed to providing a comprehensive and competitive benefits package that supports the wellbeing, growth, and performance of our employees.

  • Competitive compensation package
  • 13th month salary and ESOP
  • Health insurance and annual health check-up
  • Hybrid working, flexible hours, and unlimited leave
  • Learning support fund for professional development
  • Access to AI tools and innovation initiatives
  • Company-provided laptop & monitor
  • Complimentary lunch and daily snacks
  • Employee care programs and engagement activities

Company Overview

Perform.AI is the AI Commerce Operating System for ambitious brands. In agentic commerce, performing is what gets a brand recommended – so Perform.AI runs the delivery promise, the post-order experience and the return as one interconnected system that acts on what needs fixing, with Logistics carrying the daily work of the delivery promise underneath. Founded in 2016 as Parcel Perform, the company became Perform.AI in September 2026.