1 day ago
Bucharest, Romania +3 moreStaff+
Responsibilities
- Run instrumented spikes and benchmarks for emerging LLMs, agentic systems, vector databases, orchestration frameworks, copilots, and assistants.
- Design and maintain evaluation datasets, prompts, scenarios, telemetry, regression tracking, and benchmarking tools.
- Build focused prototypes to assess architecture, integration patterns, operational constraints, security boundaries, and failure modes.
- Produce decision memos and readiness guidance covering adoption recommendations, guardrails, limitations, operational considerations, and integration requirements.
- Coordinate hand-offs of validated technologies to production engineering teams and support transitions when evaluation context is needed.
- Work with Security, Legal, Compliance, and AI teams on risk assessments, mitigations, governance recommendations, reusable standards, and checklists.
- Mentor engineers, participate in technical hiring, and share evaluation and experimentation knowledge through write-ups, talks, and meetups.
Requirements
- 8+ years of software engineering experience, including significant senior- or Staff-level scope.
- Deep hands-on experience building and evaluating systems based on LLMs and modern AI tooling.
- Experience building agentic or multi-step AI systems involving tool use, orchestration, state, retrieval, or external integrations.
- Strong software engineering fundamentals and the ability to rapidly build high-quality experimental systems.
- Strong knowledge of cloud infrastructure, preferably AWS, and the ability to run experimental workloads securely and cost-consciously.
- Experience with observability, telemetry, testing, benchmarking, distributed components, asynchronous workflows, reliability, and scalability.
- Track record of designing experiments and benchmarks that compare AI models and tools under real-world constraints.
- Experience constructing evaluation datasets, including task selection, labeling, holdout discipline, and maintenance as models improve.
- Working knowledge of LLM-as-judge methods, human evaluation, inter-annotator agreement, statistical significance, regression tracking, telemetry, and versioning.
- Ability to turn ambiguous ideas into scoped evaluation plans, make trade-off decisions, write concise decision memos, and explain technical results to non-specialists.
- Experience collaborating with platform, product, and operations teams and influencing stakeholders without formal authority.
- Curiosity, disciplined measurement, risk awareness, self-direction, and a preference for reusable tools, templates, and playbooks.
Benefits
- Flexible schedules and a combination of self-managed focus time with in-person collaboration depending on role and location.
- Fintech industry environment with opportunities to build and innovate.
- Internal referral bonus program.
- Work from anywhere while traveling for up to 3 months each year.
Tech Stack
Categories
About dLocal
dLocal builds a cross-border payments platform for global merchants selling into emerging markets, offering pay-ins, payouts, marketplace tools, and a fraud/chargeback suite. It supports 900+ local payment methods across 60+ countries to accept wallets, bank transfers, and cash vouchers while settling funds internationally; revenue comes from transaction fees and related services. Founded in 2016 and headquartered in Montevideo, Uruguay, dLocal is a public company listed on NASDAQ (ticker: DLO).
