Writing.io Jobs

Find the best remote jobs. Answer a few questions and we'll deploy a powerful assistant to help you search, create alerts, and more.

1 What roles are you open to?

2 Experience level

3 Work style

Did you know? If memory is enabled, Writing.io can remember your job search preferences and help you to improve your resume, craft customized outreach and more.

Engineer Senior Build/Release/CI Engineer

Designs and maintains CI/CD pipelines and build infrastructure for browser development using AWS and DevOps practices.

Senior Remote Posted about 23 hours ago Himalayas
What this role involves
Senior DevOps Engineer - browser CI/CD + AWSRemote - Anywhere (American timezones preferred)About BraveBrave is on a mission to protect the human right to privacy online.
Read the full description
Engineer Senior Machine Learning Engineer at Talent Inc.

Senior ML Engineer owns end-to-end machine learning products from problem definition through production, using AI-native development tools to ship autonomous systems and canonical data infrastructure.

Senior Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

THE COMPANY

Careerminds is a leader in career transition and coaching solutions, helping organizations support employees through change while enabling workforce growth and development. Our product portfolio includes market-leading Career Transition and Coaching Services as well as Progression, our application for Career Frameworks and progression planning.

THE ROLE

We’re growing our machine learning team. We’re looking for Machine Learning Engineers who own products end to end — from the problem, to production, to the metric that proves it worked.

This role exists because of how we build. A small product strategy team sets direction and priorities; engineers own the work end to end — discovery, design, build, ship, and the result. You’ll have the autonomy of a founder inside your domain and the accountability that comes with it.

That accountability includes the unglamorous half of ML. You own the experiment that doesn’t pan out and the call to kill it, not just the launch. We’d rather you run four honest experiments and ship the one that works than ship four things that all look fine on a dashboard.

AI-native development isn’t an aspiration here — it’s the baseline. Our engineers ship with Claude Code and Claude Design as their default tools, and the leverage that creates is why one engineer can own a product end to end.

We want people already working this way who want to push the ceiling higher, not people who are curious about AI. In the interview we’ll ask you to show us the trail: repos, PRs, or shipped work you built this way.

This is a 100% remote/work-from-home role.

THE KEY RESPONSIBILITIES

Depending on area of focus:

Canonical data and entity resolution

  • Canonical datasets for titles, companies, skills, and industries — the layer every application depends on. Content-addressed IDs, faceted taxonomies, alias graphs accumulated across tens of millions of rows.
  • Rules-based resolution pipelines with LLM escalation, where the accumulated alias graph is the durable asset and escalation volume should fall over time.
  • Nightly agent loops that adjudicate ambiguous entities, propose structural changes, and get gated by invariant checks and blast-radius limits before anything commits.
  • Job ingestion at scale: multi-source feeds, deduplication, freshness, and the indexing economics underneath.

Retrieval, ranking, and matching

  • Job matching v2: two-tower retrieval with cross-encoder reranking, trained on outcome labels rather than clicks. Hard-negative mining, propensity weighting, impression-time logging.
  • Mobility embeddings learned from observed career sequences — the similarity a text encoder can’t recover, where Claims Adjuster and Underwriting Assistant are substitutable despite sharing no vocabulary.
  • Pivot feasibility: given where someone is, what moves are realistic, what’s missing, and which intermediate roles actually worked for peers.

Applied LLMs and agents

  • Fine-tuning where it earns its cost — against outcome labels, not for tasks a well-prompted frontier model already handles.
  • Agentic systems in production with human approval gates: agents that analyze, propose changes as reviewable artifacts, and execute only after a human signs off. We have this pattern running against tens of millions of customer touchpoints a year and want to push it much further.
  • Continuous skills inference from work artifacts rather than static documents — a problem several of our enterprise customers are currently solving for themselves, badly.
  • New product surfaces where the right answer genuinely requires an LLM, and the discipline to notice when it doesn’t.

Across all of it

  • Evaluation infrastructure you’d defend in a design review: time-forward splits, calibration, offline-to-online agreement, and honest handling of feedback-loop degeneration and survivorship bias.
  • Building inside real constraints: GDPR, EU AI Act high-risk classification for employment AI, and client data commitments are design inputs here, not someone else’s problem.

THE MUST-HAVES

  • 5+ years shipping ML systems into production — and you can name the system, the metric before and after, and how you knew the model caused the change.
  • Depth in both classical ML and deep learning (PyTorch or TensorFlow) applied to live products, not notebooks and Kaggle sets.
  • Working fluency with LLMs in production — retrieval, evals, prompt and context engineering, and the judgment to recognize when an LLM is the wrong tool.
  • You already ship with agentic coding tools — Claude Code, Claude Design, or close equivalents — and can point to work you built with them.
  • Software engineering fundamentals strong enough to own your own deploys — Python, Git, cloud (we run AWS), containers, and the patience for genuinely messy, human-authored, self-reported data.

THE NICE-TO-HAVES

  • Entity resolution, record linkage, or taxonomy design at scale
  • Ranking, recommendation, or two-tower retrieval systems
  • Sequence models on longitudinal or event-stream data
  • Embedding and vector retrieval systems in production
  • Experiment design, causal inference, or off-policy evaluation
  • Warehouse-native ML (dbt, Snowflake, or similar)
  • Labor market, HR tech, or people-data domain experience
  • Open-source contributions or publications

At Careerminds, we believe that diversity in thought and cultural background leads to better teams and stronger companies. We seek talented, qualified employees, regardless of race, color, sex/gender (including pregnancy, gender identity, and gender expression), national origin, religion, sexual orientation, disability, age, marital status, citizen status, veteran status, or any other protected classification under country or local law. Careerminds is proud to be an Equal Employment Opportunity Employer.

Come join our team. Together, we’ll help others tell their career stories and land their dream jobs.

Read the full description
Engineer Site Reliability Engineer, Tech Lead at Loadsmart

Site Reliability Engineer Tech Lead builds and maintains critical infrastructure systems, ensures application reliability and SLAs, and leads reliability projects across engineering squads.

Lead Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

ARE YOU INTERESTED IN JOINING AN INNOVATIVE LOGISTICS TECHNOLOGY COMPANY?

Loadsmart is a growth-stage technology company valued at over $1 billion (a true Tech Unicorn)!

We are a collection of industry veterans and user-centered engineers using innovative technology to fearlessly reinvent the future of freight by helping shippers, brokers, warehouses and carriers to move more with less.

With headquarters in Chicago and a globally distributed remote team, Loadsmart continues to attract top talent committed to driving meaningful change. We seek professionals who embody our core values: curiosity, clarity, results, commitment, and teamwork.

In the SRE, Tech Lead role you will build and maintain the company’s internal platform, driving operational excellence and empowering the entire engineering team. You should have experience in analyzing, proposing, and implementing safer systems and processes. Collaborating closely with engineering squads across platform engineering, you will ensure our applications are both safe and reliable. As a hands-on leader, you will stay actively involved in technical work while collaborating closely with internal stakeholders and engineering squads across the organization to ensure our applications are both safe and reliable.

DEPARTMENT: Engineering

LOCATION: Anywhere in Brazil - Remote

WHAT YOU GET TO DO

  • Collaborate with and support our creative, tight-knit development team.
  • Design, deploy, and operate Loadsmart’s critical systems while balancing reliability, cost, and agility.
  • Play a key role in driving reliability projects with engineering teams.
  • Utilize your intuitive problem-solving skills and contagious positive attitude to tackle challenging and exciting issues, inspiring those around you.
  • Collect metrics and understand their business impact, encouraging the team to do the same.
  • Perform troubleshooting and root-cause analysis of system operation issues.
  • Be accountable for the platform’s Service Level Agreements and Objectives.
  • Provide infrastructure support during off-hours as needed
  • Take ownership of software infrastructure projects
  • Seek, give, and receive constructive feedback through code and specification reviews.
  • Familiarity with AI agents and agentic workflows, applying AI across the SDLC (AI-assisted coding), LLMs, MCP servers/gateways, and how emerging AI tooling can improve reliability and operations is a plus.

REQUIRED QUALIFICATIONS:

  • 1-3 years leading Reliability Work across multiple engineering squads
  • Over 5 years of experience in Cloud Computing, SRE/DevOps
  • Proven experience collaborating with internal stakeholders across multiple engineering squads
  • Strong project management skills with a demonstrated ability to delegate and mentor team members
  • Proficient in English communication (both written and spoken) to collaborate in an international team with native and non-native English speakers
  • Detail-oriented with high initiative and self-motivation
  • Strong understanding of software engineering principles and how systems work under the hood
  • In-depth knowledge of modern networking and operating systems
  • Proficiency in AWS, cloud environments, containers, Kubernetes, Docker, and DevOps engineering, including managing tests and CI/CD pipelines
  • Familiarity with automation tools and provisioners like Terraform, Ansible, or Chef
  • Solid troubleshooting and system engineering experience in UNIX/Linux production environments
  • Experience with monitoring, alerting, and incident management
  • Proficiency in automating tasks with scripting languages like Python, Bash, etc
  • Experience or exposure to PostgreSQL and DBA responsibilities is a plus
  • Fluent in English (both written and spoken); comfortable interacting with native English speakers daily.

WORKING AT LOADSMART:

• Competitive base salaries - we believe in rewarding top talent

• Extremely competitive Equity package - become a shareholder in our company!

• Loadie Time Off - PTO and sick days without a limit

At Loadsmart, we believe our biggest asset is our people. We are proud to be an equal opportunity employer, hiring and developing individuals from diverse backgrounds and experiences to add to our collaborative culture. Loadsmart treats all candidates and employees with respect and does not discriminate in our recruiting, hiring, and promoting processes, including on the basis of race, color, religion, sex, age, sexual orientation, gender identity and/or expression, national origin, veteran status, or disability.

It is the policy of Loadsmart that all offers of employment made shall be contingent upon successful completion of electronic background check(s). These checks will be job-related, consistent with business necessity and conducted by our vendor, pursuant to all applicable laws, rules, policies and procedures of our candidates’ specific locale.

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

Read the full description
Engineer Technical Lead - GPU Infrastructure at Tether.io

Lead GPU infrastructure engineering for a Kubernetes-based compute platform, managing containerized GPU resources and inference endpoints at scale.

Lead Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Description

Join Tether and Shape the Future of Digital Finance

At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.

Innovate with Tether

Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.

But that’s just the beginning:

Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.

Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.

Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.

Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.

Why Join Us?

Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.

If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.

Are you ready to be part of the future?

About the job

Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.

The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.

This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.

Responsibilities

Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.

Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.

Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.

Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.

Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.

Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.

Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.

Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.

Hiring. Complete the platform team and set the technical bar for the engineers who join it.

Requirements

Must have

  • Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.

  • Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.

  • GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.

  • High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.

  • Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.

  • Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.

  • HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.

  • Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.

  • Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.

  • A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.

  • Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.

  • Excellent written and spoken English. Most partner and leadership work happens in writing.

  • Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.

Desirable

  • Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).

  • Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.

  • VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).

  • Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.

  • Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.

  • Peer-to-peer or distributed-systems background.

  • Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.

Important information for candidates

Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:

  • Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/

  • Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.

  • Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.

  • Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io

  • We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.

When in doubt, feel free to reach out through our official website.

Read the full description
Engineer Technical Lead - GPU Infrastructure at Tether.io

Technical Lead oversees GPU infrastructure and Kubernetes platform development for Tether Data's managed inference and compute services.

Lead Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Description

Join Tether and Shape the Future of Digital Finance

At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.

Innovate with Tether

Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.

But that’s just the beginning:

Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.

Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.

Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.

Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.

Why Join Us?

Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.

If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.

Are you ready to be part of the future?

About the job

Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.

The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.

This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.

Responsibilities

Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.

Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.

Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.

Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.

Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.

Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.

Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.

Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.

Hiring. Complete the platform team and set the technical bar for the engineers who join it.

Requirements

Must have

  • Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.

  • Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.

  • GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.

  • High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.

  • Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.

  • Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.

  • HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.

  • Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.

  • Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.

  • A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.

  • Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.

  • Excellent written and spoken English. Most partner and leadership work happens in writing.

  • Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.

Desirable

  • Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).

  • Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.

  • VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).

  • Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.

  • Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.

  • Peer-to-peer or distributed-systems background.

  • Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.

Important information for candidates

Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:

  • Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/

  • Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.

  • Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.

  • Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io

  • We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.

When in doubt, feel free to reach out through our official website.

Read the full description
Engineer Technical Lead - GPU Infrastructure at Tether.io

Technical Lead manages GPU infrastructure and Kubernetes-based compute platform, overseeing architecture, performance, and team delivery for a managed inference service.

Lead Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Description

Join Tether and Shape the Future of Digital Finance

At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.

Innovate with Tether

Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.

But that’s just the beginning:

Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.

Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.

Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.

Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.

Why Join Us?

Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.

If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.

Are you ready to be part of the future?

About the job

Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.

The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.

This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.

Responsibilities

Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.

Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.

Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.

Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.

Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.

Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.

Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.

Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.

Hiring. Complete the platform team and set the technical bar for the engineers who join it.

Requirements

Must have

  • Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.

  • Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.

  • GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.

  • High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.

  • Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.

  • Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.

  • HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.

  • Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.

  • Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.

  • A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.

  • Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.

  • Excellent written and spoken English. Most partner and leadership work happens in writing.

  • Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.

Desirable

  • Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).

  • Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.

  • VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).

  • Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.

  • Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.

  • Peer-to-peer or distributed-systems background.

  • Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.

Important information for candidates

Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:

  • Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/

  • Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.

  • Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.

  • Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io

  • We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.

When in doubt, feel free to reach out through our official website.

Read the full description
Engineer Mobile Software Engineer (iOS) - Olo App at Olo

Develops iOS features for a restaurant SaaS platform, writes clean code, participates in code reviews, and collaborates with product teams on the mobile app roadmap.

Mid Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Olo is a leading SaaS platform accelerating digital transformation in the restaurant industry, by helping customers deliver more personalised and profitable guest experiences. As a result, our digital ordering, payment, and guest engagement solutions enable brands to do more with less and make every guest feel like a regular.

While our roots are in NYC, we’re intentionally investing in Belfast and Northern Ireland as a key hub, with an established leadership presence, a local team, and community for the long term. This role is fully remote, offering you flexibility to work from anywhere within NI.

Your new role

With the brand-new Olo app launching later this year, you’ll play an integral part in shaping next-gen features and bringing our post-MVP vision to life while also gathering feedback from the fast-growing user base and fixing bugs they report. You’ll help us plan for the feature roadmap of the app and provide input in the direction of both the iOS and Android codebases. You’ll have the support of a dedicated Mobile Engineering Manager for the Olo app and work alongside four highly experienced mobile engineers, three of whom are based in Northern Ireland.

How you’ll make an impact

Here’s how your day-to-day might look but, like most mission driven tech companies, no two days are exactly the same!

  • Demonstrate a solid understanding of the team’s domain and technology stack, contributing to discussions and development decisions with growing independence.

  • Handle small-to-medium features independently and begin taking ownership of moderately complex tasks with some guidance.

  • Write clean, maintainable code and actively participate in peer code reviews, providing constructive feedback and adhering to coding standards.

  • Collaborate closely with Product to refine requirements, helping to shape solutions that meet business needs effectively.

  • Focus on delivering high-quality software solutions within established timelines, emphasising best practices in software development.

  • Engage in troubleshooting and debugging efforts, showing an ability to resolve common and moderately complex issues with minimal support.

  • Assist in the deployment and monitoring of services, learning how to manage and troubleshoot issues in production environments.

  • Contribute to building and maintaining reliable distributed systems, implementing resilience mechanisms as appropriate.

  • Participate actively in team ceremonies and demonstrate initiative by taking ownership of tasks and helping to unblock others when possible.

  • Engage in continuous learning and self-improvement by exploring new technologies and best practices relevant to the team’s work.

  • Use Claude Code as part of your daily workflow, and grow your skills through hands-on AI training designed to help you become highly effective with modern AI coding agents and IDEs.

  • Demonstrate ownership of the team’s delivery pipeline, ensuring that code quality, testing standards, and deployment practices are continuously optimised.

  • Active participation in on-call duties is required, with specific responsibilities determined by your assigned team and area of expertise.

What will set you up for success

  • Bachelor’s Degree in Computer Science, Software Engineering, or a related field, or equivalent practical experience.

  • 2+ years of experience in software engineering, specifically focused on native mobile developement

  • Proficiency in native iOS mobile development using Swift and being capable of independently implementing moderately complex features and algorithms.

  • Experience using version control tools (e.g., GitHub) and participating in continuous integration/continuous delivery (CI/CD) pipelines (e.g., GitHub Actions).

  • Comfort writing and maintaining unit tests

  • Strong problem-solving and collaboration skills

  • Experience communicating technical details to team members, product managers, and stakeholders to deliver solutions that align with business objectives.

About Olo

Olo is a leading restaurant technology provider with ordering, payment, and guest engagement solutions that help brands increase orders, streamline operations, and improve the guest experience. Each day, Olo processes millions of orders on its open SaaS platform, gathering the right data from each touchpoint into a single source—so restaurants can better understand and better serve every guest on every channel, every time. Over 800 restaurant brands trust Olo and its network of more than 400 integration partners to innovate on behalf of the restaurant community, accelerating technology’s positive impact and creating a world where every restaurant guest feels like a regular. Learn more at olo.com

Applicant Privacy Notice (United Kingdom)

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

Read the full description
Engineer Senior Staff Engineer, Generative AI at Nagarro

Designs and builds scalable generative AI and agentic solutions end-to-end, implementing RAG patterns, prompt engineering, and LLM fine-tuning with production-grade Python code on cloud platforms.

Senior Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Company Description

👋🏼We’re Nagarro.

We are a Digital Product Engineering company that is scaling in a big way! We build products, services, and experiences that inspire, excite, and delight. We work at scale — across all devices and digital mediums, and our people exist everywhere in the world (18000+ experts across 40 countries, to be exact). Our work culture is dynamic and non-hierarchical. We’re looking for great new colleagues. That’s where you come in!

Job Description

REQUIREMENTS:

  • Total experience 8+ years.
  • Deep understanding of LLMs (e.g., GPTs, Llama, Claude, Gemini, Qwen, Mistral, BERT-family models) and their architectures (Transformers)
  • Should have expert-level prompt engineering skills and proven experience implementing RAG patterns
  • High proficiency in Python and standard AI/ML libraries (e.g., LangChain, LlamaIndex, LangGraph, LangSmith, Hugging Face Transformers, Scikit-learn, PyTorch/TensorFlow).
  • Experience implementing RAG architectures and prompt engineering.
  • Strong experience with fine-tuning and distillation techniques and evaluation.
  • Strong experience using managed AI/ML services on the target cloud platform (e.g., Azure Machine Learning Studio, AI Foundry).
  • Strong understanding of vector databases (e.g., Weaviate, Neo4j)
  • understanding of GenAI evaluation metrics (e.g., BLEU, ROUGE, perplexity, semantic similarity, human evaluation).
  • Architect and implement scalable GenAI and Agentic AI solutions end-to-end.
  • Should be able to write high-quality, production-ready Python code with strong testing and maintainability practices.
  • Should be able to productionize AI systems on Azure or AWS, ensuring enterprise-grade reliability and performance.
  • Should be able to build and expose APIs using FastAPI, integrating with databases through an ORM.
  • Should be able to scale GenAI solutions to support enterprise workloads.
  • Collaborate across product and engineering teams to convert business needs into AI-driven solutions.
  • Strong ability to both architect and code GenAI/Agentic AI solutions.
  • Proven production experience with GenAI deployments on Azure or AWS.
  • Strong experience in scaling AI solutions in live environments.
  • Very strong Python programming skills with a track record of clean, efficient, and maintainable code.
  • Should have successfully delivered at least one production GenAI/Agentic AI solution.
  • Must have proficiency with FastAPI and at least one ORM (e.g., SQLAlchemy, Tortoise ORM).
  • Should have familiarity with Model Context Protocol (MCP).
  • Should have contributions to open-source GenAI projects.
  • Good to have experience with React (or some other JS frameworks) for building user-facing interfaces and front-end integrations
  • Excellent communication skills and the ability to collaborate effectively with cross-functional teams.

RESPONSIBILITIES:

  • Understanding the client’s business use cases and technical requirements and be able to convert them into technical design which elegantly meets the requirements.
  • Mapping decisions with requirements and be able to translate the same to developers.
  • Identifying different solutions and being able to narrow down the best option that meets the clients’ requirements.
  • Defining guidelines and benchmarks for NFR considerations during project implementation.
  • Writing and reviewing design document explaining overall architecture, framework, and high-level design of the application for the developers.
  • Reviewing architecture and design on various aspects like extensibility, scalability, security, design patterns, user experience, NFRs, etc., and ensure that all relevant best practices are followed.
  • Developing and designing the overall solution for defined functional and non-functional requirements; and defining technologies, patterns, and frameworks to materialize it.
  • Understanding and relating technology integration scenarios and applying these learnings in projects.
  • Resolving issues that are raised during code/review, through exhaustive systematic analysis of the root cause, and being able to justify the decision taken.
  • Carrying out POCs to make sure that suggested design/technologies meet the requirements.

Qualifications

Bachelor’s or master’s degree in computer science, Information Technology, or a related field.

Read the full description
Engineer Senior Staff Engineer, Generative AI at Nagarro

Designs and deploys production-grade generative AI and agentic solutions using LLMs, RAG patterns, and cloud platforms at enterprise scale.

Senior Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Company Description

👋🏼We’re Nagarro.

We are a Digital Product Engineering company that is scaling in a big way! We build products, services, and experiences that inspire, excite, and delight. We work at scale — across all devices and digital mediums, and our people exist everywhere in the world (18000+ experts across 40 countries, to be exact). Our work culture is dynamic and non-hierarchical. We’re looking for great new colleagues. That’s where you come in!

Job Description

REQUIREMENTS:

  • Total experience 8+ years.
  • Deep understanding of LLMs (e.g., GPTs, Llama, Claude, Gemini, Qwen, Mistral, BERT-family models) and their architectures (Transformers)
  • Should have expert-level prompt engineering skills and proven experience implementing RAG patterns
  • High proficiency in Python and standard AI/ML libraries (e.g., LangChain, LlamaIndex, LangGraph, LangSmith, Hugging Face Transformers, Scikit-learn, PyTorch/TensorFlow).
  • Experience implementing RAG architectures and prompt engineering.
  • Strong experience with fine-tuning and distillation techniques and evaluation.
  • Strong experience using managed AI/ML services on the target cloud platform (e.g., Azure Machine Learning Studio, AI Foundry).
  • Strong understanding of vector databases (e.g., Weaviate, Neo4j)
  • understanding of GenAI evaluation metrics (e.g., BLEU, ROUGE, perplexity, semantic similarity, human evaluation).
  • Architect and implement scalable GenAI and Agentic AI solutions end-to-end.
  • Should be able to write high-quality, production-ready Python code with strong testing and maintainability practices.
  • Should be able to productionize AI systems on Azure or AWS, ensuring enterprise-grade reliability and performance.
  • Should be able to build and expose APIs using FastAPI, integrating with databases through an ORM.
  • Should be able to scale GenAI solutions to support enterprise workloads.
  • Collaborate across product and engineering teams to convert business needs into AI-driven solutions.
  • Strong ability to both architect and code GenAI/Agentic AI solutions.
  • Proven production experience with GenAI deployments on Azure or AWS.
  • Strong experience in scaling AI solutions in live environments.
  • Very strong Python programming skills with a track record of clean, efficient, and maintainable code.
  • Should have successfully delivered at least one production GenAI/Agentic AI solution.
  • Must have proficiency with FastAPI and at least one ORM (e.g., SQLAlchemy, Tortoise ORM).
  • Should have familiarity with Model Context Protocol (MCP).
  • Should have contributions to open-source GenAI projects.
  • Good to have experience with React (or some other JS frameworks) for building user-facing interfaces and front-end integrations
  • Excellent communication skills and the ability to collaborate effectively with cross-functional teams.

RESPONSIBILITIES:

  • Understanding the client’s business use cases and technical requirements and be able to convert them into technical design which elegantly meets the requirements.
  • Mapping decisions with requirements and be able to translate the same to developers.
  • Identifying different solutions and being able to narrow down the best option that meets the clients’ requirements.
  • Defining guidelines and benchmarks for NFR considerations during project implementation.
  • Writing and reviewing design document explaining overall architecture, framework, and high-level design of the application for the developers.
  • Reviewing architecture and design on various aspects like extensibility, scalability, security, design patterns, user experience, NFRs, etc., and ensure that all relevant best practices are followed.
  • Developing and designing the overall solution for defined functional and non-functional requirements; and defining technologies, patterns, and frameworks to materialize it.
  • Understanding and relating technology integration scenarios and applying these learnings in projects.
  • Resolving issues that are raised during code/review, through exhaustive systematic analysis of the root cause, and being able to justify the decision taken.
  • Carrying out POCs to make sure that suggested design/technologies meet the requirements.

Qualifications

Bachelor’s or master’s degree in computer science, Information Technology, or a related field.

Read the full description
Engineer Staff Engineer, Generative AI at Nagarro

Staff engineer designs and deploys production-grade generative AI systems, implementing RAG patterns, fine-tuning LLMs, and building scalable agentic solutions on cloud platforms.

Senior Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Company Description

👋🏼We’re Nagarro.

We are a Digital Product Engineering company that is scaling in a big way! We build products, services, and experiences that inspire, excite, and delight. We work at scale — across all devices and digital mediums, and our people exist everywhere in the world (18000+ experts across 40 countries, to be exact). Our work culture is dynamic and non-hierarchical. We’re looking for great new colleagues. That’s where you come in!

Job Description

REQUIREMENTS:

  • Total experience: 6+ years.
  • Deep understanding of LLMs (e.g., GPTs, Llama, Claude, Gemini, Qwen, Mistral, BERT-family models) and their architectures (Transformers)
  • Should have expert-level prompt engineering skills and proven experience implementing RAG patterns
  • High proficiency in Python and standard AI/ML libraries (e.g., LangChain, LlamaIndex, LangGraph, LangSmith, Hugging Face Transformers, Scikit-learn, PyTorch/TensorFlow).
  • Experience implementing RAG architectures and prompt engineering.
  • Strong experience with fine-tuning and distillation techniques and evaluation.
  • Strong experience using managed AI/ML services on the target cloud platform (e.g., Azure Machine Learning Studio, AI Foundry).
  • Strong understanding of vector databases (e.g., Weaviate, Neo4j)
  • understanding of GenAI evaluation metrics (e.g., BLEU, ROUGE, perplexity, semantic similarity, human evaluation).
  • Architect and implement scalable GenAI and Agentic AI solutions end-to-end.
  • Should be able to write high-quality, production-ready Python code with strong testing and maintainability practices.
  • Should be able to productionize AI systems on Azure or AWS, ensuring enterprise-grade reliability and performance.
  • Should be able to build and expose APIs using FastAPI, integrating with databases through an ORM.
  • Should be able to scale GenAI solutions to support enterprise workloads.
  • Collaborate across product and engineering teams to convert business needs into AI-driven solutions.
  • Strong ability to both architect and code GenAI/Agentic AI solutions.
  • Proven production experience with GenAI deployments on Azure or AWS.
  • Strong experience in scaling AI solutions in live environments.
  • Very strong Python programming skills with a track record of clean, efficient, and maintainable code.
  • Should have successfully delivered at least one production GenAI/Agentic AI solution.
  • Must have proficiency with FastAPI and at least one ORM (e.g., SQLAlchemy, Tortoise ORM).
  • Should have familiarity with Model Context Protocol (MCP).
  • Should have contributions to open-source GenAI projects.
  • Good to have experience with React (or some other JS frameworks) for building user-facing interfaces and front-end integrations
  • Excellent communication skills and the ability to collaborate effectively with cross-functional teams.

RESPONSIBILITIES:

  • Understanding the client’s business use cases and technical requirements and be able to convert them into technical design which elegantly meets the requirements.
  • Mapping decisions with requirements and be able to translate the same to developers.
  • Identifying different solutions and being able to narrow down the best option that meets the clients’ requirements.
  • Defining guidelines and benchmarks for NFR considerations during project implementation.
  • Writing and reviewing design document explaining overall architecture, framework, and high-level design of the application for the developers.
  • Reviewing architecture and design on various aspects like extensibility, scalability, security, design patterns, user experience, NFRs, etc., and ensure that all relevant best practices are followed.
  • Developing and designing the overall solution for defined functional and non-functional requirements; and defining technologies, patterns, and frameworks to materialize it.
  • Understanding and relating technology integration scenarios and applying these learnings in projects.
  • Resolving issues that are raised during code/review, through exhaustive systematic analysis of the root cause, and being able to justify the decision taken.
  • Carrying out POCs to make sure that suggested design/technologies meet the requirements.

Qualifications

Bachelor’s or master’s degree in computer science, Information Technology, or a related field.

Read the full description
Engineer Senior Staff Engineer at Nagarro

Designs and implements scalable generative AI and agentic AI solutions end-to-end using LLMs, RAG patterns, and cloud platforms like Azure/AWS.

Senior Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Company Description

👋🏼We’re Nagarro.

We are a Digital Product Engineering company that is scaling in a big way! We build products, services, and experiences that inspire, excite, and delight. We work at scale — across all devices and digital mediums, and our people exist everywhere in the world (18000+ experts across 40 countries, to be exact). Our work culture is dynamic and non-hierarchical. We’re looking for great new colleagues. That’s where you come in!

Job Description

REQUIREMENTS:

  • Total experience 8+ years.
  • Deep understanding of LLMs (e.g., GPTs, Llama, Claude, Gemini, Qwen, Mistral, BERT-family models) and their architectures (Transformers)
  • Should have expert-level prompt engineering skills and proven experience implementing RAG patterns
  • High proficiency in Python and standard AI/ML libraries (e.g., LangChain, LlamaIndex, LangGraph, LangSmith, Hugging Face Transformers, Scikit-learn, PyTorch/TensorFlow).
  • Experience implementing RAG architectures and prompt engineering.
  • Strong experience with fine-tuning and distillation techniques and evaluation.
  • Strong experience using managed AI/ML services on the target cloud platform (e.g., Azure Machine Learning Studio, AI Foundry).
  • Strong understanding of vector databases (e.g., Weaviate, Neo4j)
  • understanding of GenAI evaluation metrics (e.g., BLEU, ROUGE, perplexity, semantic similarity, human evaluation).
  • Architect and implement scalable GenAI and Agentic AI solutions end-to-end.
  • Should be able to write high-quality, production-ready Python code with strong testing and maintainability practices.
  • Should be able to productionize AI systems on Azure or AWS, ensuring enterprise-grade reliability and performance.
  • Should be able to build and expose APIs using FastAPI, integrating with databases through an ORM.
  • Should be able to scale GenAI solutions to support enterprise workloads.
  • Collaborate across product and engineering teams to convert business needs into AI-driven solutions.
  • Strong ability to both architect and code GenAI/Agentic AI solutions.
  • Proven production experience with GenAI deployments on Azure or AWS.
  • Strong experience in scaling AI solutions in live environments.
  • Very strong Python programming skills with a track record of clean, efficient, and maintainable code.
  • Should have successfully delivered at least one production GenAI/Agentic AI solution.
  • Must have proficiency with FastAPI and at least one ORM (e.g., SQLAlchemy, Tortoise ORM).
  • Should have familiarity with Model Context Protocol (MCP).
  • Should have contributions to open-source GenAI projects.
  • Good to have experience with React (or some other JS frameworks) for building user-facing interfaces and front-end integrations
  • Excellent communication skills and the ability to collaborate effectively with cross-functional teams.

RESPONSIBILITIES:

  • Understanding the client’s business use cases and technical requirements and be able to convert them into technical design which elegantly meets the requirements.
  • Mapping decisions with requirements and be able to translate the same to developers.
  • Identifying different solutions and being able to narrow down the best option that meets the clients’ requirements.
  • Defining guidelines and benchmarks for NFR considerations during project implementation.
  • Writing and reviewing design document explaining overall architecture, framework, and high-level design of the application for the developers.
  • Reviewing architecture and design on various aspects like extensibility, scalability, security, design patterns, user experience, NFRs, etc., and ensure that all relevant best practices are followed.
  • Developing and designing the overall solution for defined functional and non-functional requirements; and defining technologies, patterns, and frameworks to materialize it.
  • Understanding and relating technology integration scenarios and applying these learnings in projects.
  • Resolving issues that are raised during code/review, through exhaustive systematic analysis of the root cause, and being able to justify the decision taken.
  • Carrying out POCs to make sure that suggested design/technologies meet the requirements.

Qualifications

Bachelor’s or master’s degree in computer science, Information Technology, or a related field.

Read the full description
Engineer Senior Staff Engineer, Generative AI+NLP at Nagarro

Senior staff engineer designs, builds, and deploys production-grade generative AI and NLP systems using LLMs, RAG patterns, and cloud platforms at scale.

Senior Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Company Description

👋🏼We’re Nagarro.

We are a Digital Product Engineering company that is scaling in a big way! We build products, services, and experiences that inspire, excite, and delight. We work at scale — across all devices and digital mediums, and our people exist everywhere in the world (18000+ experts across 40 countries, to be exact). Our work culture is dynamic and non-hierarchical. We’re looking for great new colleagues. That’s where you come in!

Job Description

REQUIREMENTS:

  • Total experience 8+ years.
  • Deep understanding of LLMs (e.g., GPTs, Llama, Claude, Gemini, Qwen, Mistral, BERT-family models) and their architectures (Transformers)
  • Should have expert-level prompt engineering skills and proven experience implementing RAG patterns
  • High proficiency in Python and standard AI/ML libraries (e.g., LangChain, LlamaIndex, LangGraph, LangSmith, Hugging Face Transformers, Scikit-learn, PyTorch/TensorFlow).
  • Experience implementing RAG architectures and prompt engineering.
  • Strong experience with fine-tuning and distillation techniques and evaluation.
  • Strong experience using managed AI/ML services on the target cloud platform (e.g., Azure Machine Learning Studio, AI Foundry).
  • Strong understanding of vector databases (e.g., Weaviate, Neo4j).
  • Should be proficient in building complex Deep Learning and GENAI pipelines.
  • Should have experience with computer vision/Deep Learning/NLP and GENAI domain tech stacks.
  • understanding of GenAI evaluation metrics (e.g., BLEU, ROUGE, perplexity, semantic similarity, human evaluation).
  • Architect and implement scalable GenAI and Agentic AI solutions end-to-end.
  • Should be able to write high-quality, production-ready Python code with strong testing and maintainability practices.
  • Should be able to productionize AI systems on Azure or AWS, ensuring enterprise-grade reliability and performance.
  • Should be able to build and expose APIs using FastAPI, integrating with databases through an ORM.
  • Should be able to scale GenAI solutions to support enterprise workloads.
  • Collaborate across product and engineering teams to convert business needs into AI-driven solutions.
  • Strong ability to both architect and code GenAI/Agentic AI solutions.
  • Proven production experience with GenAI deployments on Azure.
  • Strong experience in scaling AI solutions in live environments.
  • Very strong Python programming skills with a track record of clean, efficient, and maintainable code.
  • Should have successfully delivered at least one production GenAI/Agentic AI solution.
  • Must have proficiency with FastAPI and at least one ORM (e.g., SQLAlchemy, Tortoise ORM).
  • Should have familiarity with Model Context Protocol (MCP).
  • Should have contributions to open-source GenAI projects.
  • Good to have experience with React (or some other JS frameworks) for building user-facing interfaces and front-end integrations
  • Excellent communication skills and the ability to collaborate effectively with cross-functional teams.

RESPONSIBILITIES:

  • Understanding the client’s business use cases and technical requirements and be able to convert them into technical design which elegantly meets the requirements.
  • Mapping decisions with requirements and be able to translate the same to developers.
  • Identifying different solutions and being able to narrow down the best option that meets the clients’ requirements.
  • Defining guidelines and benchmarks for NFR considerations during project implementation.
  • Writing and reviewing design document explaining overall architecture, framework, and high-level design of the application for the developers.
  • Reviewing architecture and design on various aspects like extensibility, scalability, security, design patterns, user experience, NFRs, etc., and ensure that all relevant best practices are followed.
  • Developing and designing the overall solution for defined functional and non-functional requirements; and defining technologies, patterns, and frameworks to materialize it.
  • Understanding and relating technology integration scenarios and applying these learnings in projects.
  • Resolving issues that are raised during code/review, through exhaustive systematic analysis of the root cause, and being able to justify the decision taken.
  • Carrying out POCs to make sure that suggested design/technologies meet the requirements.

Qualifications

Bachelor’s or master’s degree in computer science, Information Technology, or a related field.

Read the full description
Engineer Staff Engineer, Generative AI at Nagarro

Designs and builds production-ready generative AI and agentic systems end-to-end, including LLM implementations, RAG architectures, prompt engineering, and cloud deployment on Azure/AWS.

Senior Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Company Description

👋🏼We’re Nagarro.

We are a Digital Product Engineering company that is scaling in a big way! We build products, services, and experiences that inspire, excite, and delight. We work at scale — across all devices and digital mediums, and our people exist everywhere in the world (18000+ experts across 40 countries, to be exact). Our work culture is dynamic and non-hierarchical. We’re looking for great new colleagues. That’s where you come in!

Job Description

REQUIREMENTS:

  • Total experience: 6+ years.
  • Deep understanding of LLMs (e.g., GPTs, Llama, Claude, Gemini, Qwen, Mistral, BERT-family models) and their architectures (Transformers)
  • Should have expert-level prompt engineering skills and proven experience implementing RAG patterns
  • High proficiency in Python and standard AI/ML libraries (e.g., LangChain, LlamaIndex, LangGraph, LangSmith, Hugging Face Transformers, Scikit-learn, PyTorch/TensorFlow).
  • Experience implementing RAG architectures and prompt engineering.
  • Strong experience with fine-tuning and distillation techniques and evaluation.
  • Strong experience using managed AI/ML services on the target cloud platform (e.g., Azure Machine Learning Studio, AI Foundry).
  • Strong understanding of vector databases (e.g., Weaviate, Neo4j)
  • understanding of GenAI evaluation metrics (e.g., BLEU, ROUGE, perplexity, semantic similarity, human evaluation).
  • Architect and implement scalable GenAI and Agentic AI solutions end-to-end.
  • Should be able to write high-quality, production-ready Python code with strong testing and maintainability practices.
  • Should be able to productionize AI systems on Azure or AWS, ensuring enterprise-grade reliability and performance.
  • Should be able to build and expose APIs using FastAPI, integrating with databases through an ORM.
  • Should be able to scale GenAI solutions to support enterprise workloads.
  • Collaborate across product and engineering teams to convert business needs into AI-driven solutions.
  • Strong ability to both architect and code GenAI/Agentic AI solutions.
  • Proven production experience with GenAI deployments on Azure or AWS.
  • Strong experience in scaling AI solutions in live environments.
  • Very strong Python programming skills with a track record of clean, efficient, and maintainable code.
  • Should have successfully delivered at least one production GenAI/Agentic AI solution.
  • Must have proficiency with FastAPI and at least one ORM (e.g., SQLAlchemy, Tortoise ORM).
  • Should have familiarity with Model Context Protocol (MCP).
  • Should have contributions to open-source GenAI projects.
  • Good to have experience with React (or some other JS frameworks) for building user-facing interfaces and front-end integrations
  • Excellent communication skills and the ability to collaborate effectively with cross-functional teams.

RESPONSIBILITIES:

  • Understanding the client’s business use cases and technical requirements and be able to convert them into technical design which elegantly meets the requirements.
  • Mapping decisions with requirements and be able to translate the same to developers.
  • Identifying different solutions and being able to narrow down the best option that meets the clients’ requirements.
  • Defining guidelines and benchmarks for NFR considerations during project implementation.
  • Writing and reviewing design document explaining overall architecture, framework, and high-level design of the application for the developers.
  • Reviewing architecture and design on various aspects like extensibility, scalability, security, design patterns, user experience, NFRs, etc., and ensure that all relevant best practices are followed.
  • Developing and designing the overall solution for defined functional and non-functional requirements; and defining technologies, patterns, and frameworks to materialize it.
  • Understanding and relating technology integration scenarios and applying these learnings in projects.
  • Resolving issues that are raised during code/review, through exhaustive systematic analysis of the root cause, and being able to justify the decision taken.
  • Carrying out POCs to make sure that suggested design/technologies meet the requirements.

Qualifications

Bachelor’s or master’s degree in computer science, Information Technology, or a related field.

Read the full description
Engineer Software Architect at Fingerprint

Lead cross-cutting architecture decisions across engineering teams, designing scalable systems and setting technical standards for a fraud detection platform.

Lead Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Fingerprint empowers enterprises to detect and stop online fraud with the world’s most accurate device intelligence.  We lead our industry with bleeding-edge identification capabilities and work on turning new ideas and discoveries in the fraud detection space into reality. Our customers range from innovative startups to leading enterprise companies, including Plaid, Dropbox, and Booking.com.

Fingerprint is a globally dispersed, 100% remote company. We were named on on the 2026 Forbes Best Startup Employers list and ranked #803 on the 2026 Inc. 5000 list of America’s fastest-growing private companies.

We have raised $77M and are backed by Craft Ventures (Tesla, Facebook, Airbnb ), Nexus Venture Partners ( Postman, Apollo.io, MinIO, Druva) and Uncorrelated Ventures ( Redis, Rollbar,  Gradle).

About the role

You will be Fingerprint’s first dedicated Architect. You will lead cross-cutting architecture — how our systems fit together and where they need to go — working alongside the Staff and Lead engineers who already shape them, without direct reports. The mandate has three parts. First, go deep: build the end-to-end picture nobody currently has time to hold, identify where the platform will strain as we scale, and set the direction to address it. Second, raise the bar: strengthen and extend our design review, API, service, and reliability practices so eight engineering groups can move fast without stepping on each other. Third, make it durable: give cross-cutting architecture work consistent cadence, a durable decision record, and follow-through, so strong individual judgment compounds into platform-level outcomes.

You report directly to the VP of Engineering. That placement is deliberate — you need neutrality across groups so the standards you set with teams are adopted everywhere.

What you’ll do

Own the end-to-end architecture

  • Build and maintain the end-to-end architecture picture — data flows, service boundaries, ownership seams — and keep it current as the org changes.
  • Identify systemic scaling, reliability, and cost risks across the pipeline (ingestion, identification, signals, delivery) and drive architectural changes to address them before they become incidents.
  • Set direction for API design and service patterns across REST, SDK, and MCP surfaces so customers see one coherent platform.
  • Shape the roadmap for foundational initiatives (e.g., cell architecture, multi-region readiness, rate limiting, feature flag infrastructure) and lead the hardest designs yourself.

Lead cross-cutting architecture

  • Serve as full-time technical lead across the engineering org: set the agenda, drive decisions to closure, own the decision record, and follow through on adoption across teams.
  • Partner with Engineering leadership and the Staff/Lead engineers across teams to prioritize cross-cutting work and connect it to team roadmaps.
  • Mentor Staff-level engineers on system design, technical writing, and influence — raise the bar for what “Staff” means at Fingerprint.

Strengthen and extend best practices

  • Build on the design review practices we already have (RFCs, cross-team design reviews) and make them more consistent, lighter-weight, and more useful, so teams choose to use them.
  • Codify and extend standards for observability, quality, security-by-design, and API evolution/versioning, working with Cloud Platform, Security, and product teams.
  • Review and sign off on critical designs; teach through review rather than gatekeeping.

Set us up for scale

  • Translate business trajectory (Tier 1 enterprise customers, new product launches, AI/MCP workflows) into a multi-quarter technical strategy and sequence the work.
  • Advise Engineering leadership and executive team on build/buy, platform investment, and technical risk in plain language.
  • Lead AI adoption in engineering practice: set norms for AI-assisted design, coding, and review across teams, and shape our architecture and documentation so AI agents can work in our systems as effectively as engineers do.
  • Stay hands-on: prototype, write reference implementations and documentation, and dig into production when the problem demands it.

What we’re looking for

  • 10+ years of software engineering experience, including 3+ years operating as an Architect for a 100+ Engineering org — you have owned architecture for a platform, not just a service.
  • Track record of leading through influence: you have led senior engineers you didn’t manage, run design review or architecture forums that people actually used, and driven adoption of standards across an org.
  • Deep expertise in distributed systems and high-throughput, low-latency backend architecture (Go or similar systems language; Kubernetes/AWS; event-driven and data-intensive systems). Fluency across adjacent layers — client SDKs, data pipelines, ML serving — is a strong plus.
  • Experience designing public APIs and SDKs at scale, including versioning, backward compatibility, and multi-surface consistency.
  • Demonstrated ability to anticipate systemic risk and act on it — you can point to failures you prevented, not just ones you fixed.
  • Exceptional written communication. You make complex decisions legible to engineers and executives alike, and you default to async, documented decision-making.
  • AI-native by default. You use AI coding and reasoning tools as a normal part of how you design, prototype, review, and write — and you have opinions, from experience, about where they accelerate engineering work and where they don’t yet.
  • Architect for an AI-assisted org. You think about how codebases, documentation, service boundaries, and APIs should be shaped so that both humans and AI agents can work in them safely — legible structure, strong contracts, automated verification.
  • Pragmatism over purity. You balance long-term architecture with delivery pressure and know when “good enough” is the right call.
  • Comfortable being the first: you have stepped into a dedicated role where the work was previously shared across a group, earned the trust of the people already doing it, and made them more effective rather than displacing them.

Nice to have

  • Experience in fraud detection, identity, device intelligence, or other adversarial domains.
  • Multi-region / cell-based architecture experience.
  • Experience with MCP, LLM tool integrations, or agent-facing API design.
  • Familiarity with ClickHouse, Kafka, or similar high-volume data infrastructure.

Compensation & Transparency

For US-based employees, the cash compensation range for this role is $177,000 – $240,000. We set standard ranges for all US roles based on function, level, and geographic location, benchmarked against similar stage growth companies. To comply with local legislation and provide greater transparency, we share salary ranges on all job postings. However, these ranges are specific to the hiring location and may differ within or outside the US. Offers vary depending on, but not limited to, relevant experience, education, certifications/licenses, skills, training, and market conditions.

Due to regulatory and security reasons, there’s a small number of countries where we cannot have Fingerprint teammates based. Additionally, because Fingerprint is an all-remote company and people can join our workforce from almost any country, we do not sponsor visas. Fingerprint teammates need to be authorized to work from their home location.

We are dedicated to creating an inclusive work environment for everyone. We embrace and celebrate the unique experiences, perspectives and cultural backgrounds that each employee brings to our workplace. Fingerprint strives to foster an environment where our employees feel respected, valued and empowered, and our team members are at the forefront in helping us promote and sustain an inclusive workplace. We highly encourage people from underrepresented groups in tech to apply.

If you are applying as a resident of California, please read our CCPA notice here.

If you are applying as a resident of the EU, please read our GDPR notice here.

  • We have noticed a rise in recruiting impersonations across the industry, where scammers attempt to access candidates’ personal and financial information through fake interviews and offers. All Fingerprint recruiting email communications will always come from the @fingerprint.com domain. Any outreach claiming to be from Fingerprint via other sources should be ignored.*

Due to regulatory and security reasons, there’s a small number of countries where we cannot have Fingerprint teammates based. Additionally, because Fingerprint is an all-remote company and people can join our workforce from almost any country, we do not sponsor visas. Fingerprint teammates need to be authorized to work from their home location.

We are dedicated to creating an inclusive work environment for everyone. We embrace and celebrate the unique experiences, perspectives and cultural backgrounds that each employee brings to our workplace. Fingerprint strives to foster an environment where our employees feel respected, valued and empowered, and our team members are at the forefront in helping us promote and sustain an inclusive workplace. We highly encourage people from underrepresented groups in tech to apply.

If you are applying as a resident of California, please read our CCPA notice here.

If you are applying as a resident of the EU, please read our GDPR notice here.

**We have noticed a rise in recruiting impersonations across the industry, where scammers attempt to access candidates’ personal and financial information through fake interviews and offers. All Fingerprint recruiting email communications will always come from the @fingerprint.com domain. Any outreach claiming to be from Fingerprint via other sources should be ignored.

Read the full description
Engineer Frontend Web Application Developer - Remote at KoboToolbox

Develops and maintains frontend code for large-scale web applications, writing robust TypeScript/React features, reviewing code, and collaborating on architecture decisions.

Mid Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Location: Remote

Availability: 35-40 hours per week

Working hours: US East business hours

Reporting to: Lead developer

KoboToolbox has an immediate opening for a Frontend Web Application Developer to fill a full-time position of approximately 35-40 hours per week, for a commitment of at least 1 year. As a member of our team, you will share in the challenge and excitement of writing code used by over 32,000 organizations around the world. These organizations create data-driven change through the collection and analysis of more than 20 million surveys per month.

Only candidates who already have experience working on large web applications will be considered. Beyond technical acumen, we are seeking a team member who demonstrates curiosity, initiative, and a cooperative approach to problem solving and decision-making.

If you’re passionate about leveraging technology to make a positive impact, we want to hear from you!

Responsibilities

  • Searching and reading an extensive, long-lived code base to understand existing behavior and conventions.
  • Recognizing when existing conventions may no longer be the best approach and suggesting improvements aligned with contemporary best practices.
  • Writing robust, concise, and reusable code with accompanying tests and documentation.
  • Reviewing other developers’ code and providing constructive feedback.
  • Incorporating feedback from code reviews to improve your skills and align with team standards.
  • Using AI-assisted development tools effectively while maintaining responsibility for code quality and correctness.
  • Efficiently investigating bug reports from the support team, creating well-scoped fixes or actionable plans for resolution.
  • Scoping, prioritizing, estimating, and organizing work into manageably sized tasks.
  • Attending regular video-conference check-ins with other members of the technical team.
  • Communicating with the public in conjunction with our support staff or directly through forums, issue trackers, etc.
  • Contributing to discussions about design and architecture collaboratively with the larger team.
  • Performing other related duties as directed by the lead developer.

Required Qualifications

  • Professional experience writing, deploying, and maintaining client-side code for real-world, API-driven single-page applications.
  • Proficiency with TypeScript, React, and related technologies, including styling, state management, and efficient data exchange over HTTP.
  • Unflinching ability to work with legacy technologies such as Backbone and CoffeeScript.
  • Recent experience giving and receiving code reviews.
  • Ability to independently set up and troubleshoot a local development environment, knowing when to ask for help on genuine roadblocks.
  • Interest in data collection (surveying), particularly in humanitarian emergencies and other challenging contexts, and a desire to improve our platform for our users.
  • Proficiency with spoken and written English.
  • Fluency with Git.
  • Overlap with working hours in the Eastern time zone.
  • Average availability of at least 30 hours per week, preferably 35 hours or more.

Preferred Qualifications

Experience with the following is preferred but not required to apply:

  • Using Docker and Docker Compose in a development environment.
  • Programming in Python, ideally with Django (and particularly Django REST Framework).
  • Building user-facing LLM integrations.
  • Surveying with XLSForm, ODK XForm, and OpenRosa.
  • Integrating with Stripe for payment processing.

An added note on qualifications… We know that people from underrepresented groups including women, people of color, LGBTQ+ individuals, people with disabilities, and others are less likely to apply for a role if they don’t meet every qualification listed. We’re more interested in your potential and growth mindset than a perfect checklist match. If this role excites you and you meet many, but not all, of the qualifications, we’d still love to hear from you.

Equal Opportunity Employer

Kobo is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other legally protected characteristic.

  • Genuine Impact: Contribute directly to projects that impact millions of people around the world globally, working alongside the largest international humanitarian organizations as well as thousands of national and small community based partners in 200 countries.
  • Meaningful Work Environment: Join a team that believes work should be meaningful as well as fun, tackling global challenges through innovative data collection and management tools with a proven impact for lasting change.
  • Diverse Team: Be part of an amazing, progressive, and globally diverse team that values diversity, equity, and inclusion across all spectrums.
  • Flexible Work Culture: Enjoy mutual flexibility supported by a culture that prioritizes work-life balance.
  • Professional Development: Benefit from generous professional development options.

Full Time Employee Benefits (U.S. candidates only):

  • Health & Wellness: 5 medical insurance options, dental, and vision (up to 80% premium covered), plus life insurance and Long Term Disability.
  • Financial Security: 401(k) retirement plan with 100% match up to 2%.
  • Work-Life Balance: 20 days paid time off, 12 floating holidays, unlimited sick days, and paid parental leave.

*Please note, if based in the US or Canada this is a full time employee position. In other locations full time employment or contractor positions are options depending on individual circumstances.

Read the full description
Engineer Software Quality Engineer I (UK Remote) at Turnitin

Designs, writes, and executes automated and manual tests while collaborating with development teams to ensure application quality and drive continuous improvement.

Junior Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Company Description

When you join Turnitin, you’ll be welcomed into a company that is a recognized innovator in global education. For over 25 years, Turnitin has partnered with educators and institutions to develop learning integrity solutions that recognize the enduring value of critical thinking in a rapidly changing world. Over 16,000 academic institutions, publishers, and corporations use our services in more than 185 countries around the world: Turnitin Feedback Studio, Clarity, Originality, Gradescope, ExamSoft, Similarity, and iThenticate. Protecting the value of an authentic education is at the heart of who we are.

Experience a remote-first culture that empowers you to work with purpose and accountability in a way that best suits you, supported by a comprehensive package that prioritizes your overall well-being. Our diverse community of colleagues are all unified by a shared desire to make a difference in education.

Turnitin is a global organization with team members in over 35 countries including the United States, Mexico, United Kingdom, Australia, Japan, India, and the Philippines.

Job Description

The Software Quality Engineer I (SQE1) is an integral part of the Quality Engineering department, dedicated to assuring application quality, adherence to requirements, and an exceptional user experience by testing company applications. The main purpose of this role is to design, write, execute, and maintain automated and manual tests, as well as contribute to the overall test automation framework and continuous testing (CT) pipelines.

This role will primarily interact with product development teams. The SQE1 works cross-functionally with Software Developers, Product Managers, and fellow Quality Engineers within an Agile scrum environment to provide rapid, in-process testing results, identify risks, and drive continuous improvement.

Responsibilities:

  • Test Strategy & Planning: Partner with the Scrum team to define, develop, and implement quality assurance best practices and procedures, including test strategies, plans, cases, and other related assets.
  • Test Execution: Design, write, execute, and maintain both automated and manual test cases, ensuring high quality across functional, regression, integration, and scale testing.
  • Development Collaboration: Work closely with development teams to support code releases with functional test validations on multiple platforms and environments.
  • Defect Management: Investigate, record, triage, and track defects to verification and closure.
  • Risk Mitigation: Identify potential quality risks and dependencies early in the Agile cycle; collaborate on mitigation and eradication plans.
  • Modern QA & AI Capabilities: Proactively research, evaluate, and integrate modern QA tools and AI-driven testing capabilities (such as AI-assisted test generation, predictive analysis, and smart automation tools) to improve efficiency and testing coverage.
  • Professional Development (Goal): Commit to continuous professional and technical growth, developing deeper expertise in framework architecture, advanced CI/CD pipelines, and broad-scope test strategy design.

Qualifications

  • Education & Experience: Bachelor’s degree in Computer Science, Engineering, or equivalent practical experience, combined with 4+ years of experience in software quality engineering.

  • Automation & Scripting: Proven hands-on experience in UI Web Automation and API Automation using established frameworks. (Java, Python, Jenkins, docker, Maven, github).

  • QA Methodologies: Strong, high-level understanding of the Software Testing Life Cycle (STLC) and QA concepts, methodologies, and test management.

  • CI/CD/CT: competency with continuous integration and deployment tools (such as Jenkins, Git, etc.).

  • Agile Practices: Practical experience working within Agile frameworks (Scrum & Kanban), estimating work, and managing test data.

  • Modern QA & AI Capabilities: Proactively research, evaluate, and integrate modern QA tools and AI-driven testing capabilities (such as AI-assisted test generation, predictive analysis, and smart automation tools) to improve efficiency and testing coverage.

  • Nice to Have Skills:

    • UI Mobile Automation & Functional/Manual testing.
    • Performance testing experience (e.g., using JMeter).
    • Familiarity with AWS services (S3, Lambda, EC2, CloudWatch) or security/accessibility testing.

Tii Elements:

  • Accountability: Takes ownership of the quality of deliverables and works autonomously within standard guidelines to resolve issues.
  • Collaboration & Influencing: Builds productive working relationships with developers and product teams to resolve quality issues and collaborate on best practices.
  • Quality Focus: Strongly committed to delivering exceptional user experiences and robust software quality.
  • Relationship Building: Cultivates stable, productive working relationships across cross-functional teams.
  • Curiosity: Actively stays up-to-date with new testing tools, test strategies, and emerging technologies (including AI tools).

Additional Information

Total Rewards @ Turnitin

At Turnitin, we believe Total Rewards go far beyond pay. While salary, bonus, or commission are important, they’re only part of the value you receive in exchange for your work.

Beyond compensation, you’ll experience the intrinsic rewards of unleashing your potential and making a positive impact on global education. You’ll also thrive in a culture free of politics, surrounded by humble, inclusive, and collaborative teammates.

In addition, our extrinsic rewards include generous time off and health and wellness programs that provide choice, flexibility, and a safety net for life’s challenges. You’ll also enjoy a remote-first culture that empowers you to work with purpose and accountability in the way that suits you best, all supported by a comprehensive package that prioritizes your overall well-being.

Our Mission is to ensure the integrity of global education and meaningfully improve learning outcomes.

Our Values underpin everything we do.

  • Customer Centric: Our mission is focused on improving learning outcomes; we do this by putting educators and learners at the center of everything we do.
  • Passion for Learning: We are committed to our own learning and growth internally. And we support education and learning around the globe.
  • Integrity: Integrity is the heartbeat of Turnitin—it is the core of our products, the way we treat each other, and how we work with our customers and vendors.
  • Action & Ownership: We have a bias for action. We act like owners. We are willing to change even when it’s hard.
  • One Team: We strive to break down silos, collaborate effectively, and celebrate each others’ successes.
  • Global Mindset: We consider different perspectives and celebrate diversity. We are one team. The work we do has an impact on the world.

Global Benefits

  • Remote First Culture
  • Health Care Coverage*
  • Education Reimbursement*
  • Competitive Paid Time Off
  • Self-Care Days
  • National Holidays*
  • 2 Founder Days + Juneteenth Observed
  • Paid Volunteer Time*
  • Charitable contribution match*
  • Monthly Wellness or Home Office Reimbursement/*
  • Access to Modern Health (mental health platform)
  • Parental Leave*
  • Retirement Plan with match/contribution*

\* varies by country

Seeing Beyond the Job Ad

At Turnitin, we recognize it’s unrealistic for candidates to fulfill 100% of the criteria in a job ad.  We encourage you to apply if you meet the majority of the requirements because we know that skills evolve over time. If you’re willing to learn and unleash your potential alongside us, join our team!

#LI-AP1

Read the full description
Engineer Software Quality Engineer I (UK Remote) at Turnitin

Design, write, and execute automated and manual tests while collaborating with development teams to ensure application quality and drive continuous improvement in an Agile environment.

Junior Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Company Description

When you join Turnitin, you’ll be welcomed into a company that is a recognized innovator in global education. For over 25 years, Turnitin has partnered with educators and institutions to develop learning integrity solutions that recognize the enduring value of critical thinking in a rapidly changing world. Over 16,000 academic institutions, publishers, and corporations use our services in more than 185 countries around the world: Turnitin Feedback Studio, Clarity, Originality, Gradescope, ExamSoft, Similarity, and iThenticate. Protecting the value of an authentic education is at the heart of who we are.

Experience a remote-first culture that empowers you to work with purpose and accountability in a way that best suits you, supported by a comprehensive package that prioritizes your overall well-being. Our diverse community of colleagues are all unified by a shared desire to make a difference in education.

Turnitin is a global organization with team members in over 35 countries including the United States, Mexico, United Kingdom, Australia, Japan, India, and the Philippines.

Job Description

The Software Quality Engineer I (SQE1) is an integral part of the Quality Engineering department, dedicated to assuring application quality, adherence to requirements, and an exceptional user experience by testing company applications. The main purpose of this role is to design, write, execute, and maintain automated and manual tests, as well as contribute to the overall test automation framework and continuous testing (CT) pipelines.

This role will primarily interact with product development teams. The SQE1 works cross-functionally with Software Developers, Product Managers, and fellow Quality Engineers within an Agile scrum environment to provide rapid, in-process testing results, identify risks, and drive continuous improvement.

Responsibilities:

  • Test Strategy & Planning: Partner with the Scrum team to define, develop, and implement quality assurance best practices and procedures, including test strategies, plans, cases, and other related assets.
  • Test Execution: Design, write, execute, and maintain both automated and manual test cases, ensuring high quality across functional, regression, integration, and scale testing.
  • Development Collaboration: Work closely with development teams to support code releases with functional test validations on multiple platforms and environments.
  • Defect Management: Investigate, record, triage, and track defects to verification and closure.
  • Risk Mitigation: Identify potential quality risks and dependencies early in the Agile cycle; collaborate on mitigation and eradication plans.
  • Modern QA & AI Capabilities: Proactively research, evaluate, and integrate modern QA tools and AI-driven testing capabilities (such as AI-assisted test generation, predictive analysis, and smart automation tools) to improve efficiency and testing coverage.
  • Professional Development (Goal): Commit to continuous professional and technical growth, developing deeper expertise in framework architecture, advanced CI/CD pipelines, and broad-scope test strategy design.

Qualifications

  • Education & Experience: Bachelor’s degree in Computer Science, Engineering, or equivalent practical experience, combined with 4+ years of experience in software quality engineering.

  • Automation & Scripting: Proven hands-on experience in UI Web Automation and API Automation using established frameworks. (Java, Python, Jenkins, docker, Maven, github).

  • QA Methodologies: Strong, high-level understanding of the Software Testing Life Cycle (STLC) and QA concepts, methodologies, and test management.

  • CI/CD/CT: competency with continuous integration and deployment tools (such as Jenkins, Git, etc.).

  • Agile Practices: Practical experience working within Agile frameworks (Scrum & Kanban), estimating work, and managing test data.

  • Modern QA & AI Capabilities: Proactively research, evaluate, and integrate modern QA tools and AI-driven testing capabilities (such as AI-assisted test generation, predictive analysis, and smart automation tools) to improve efficiency and testing coverage.

  • Nice to Have Skills:

    • UI Mobile Automation & Functional/Manual testing.
    • Performance testing experience (e.g., using JMeter).
    • Familiarity with AWS services (S3, Lambda, EC2, CloudWatch) or security/accessibility testing.

Tii Elements:

  • Accountability: Takes ownership of the quality of deliverables and works autonomously within standard guidelines to resolve issues.
  • Collaboration & Influencing: Builds productive working relationships with developers and product teams to resolve quality issues and collaborate on best practices.
  • Quality Focus: Strongly committed to delivering exceptional user experiences and robust software quality.
  • Relationship Building: Cultivates stable, productive working relationships across cross-functional teams.
  • Curiosity: Actively stays up-to-date with new testing tools, test strategies, and emerging technologies (including AI tools).

Additional Information

Total Rewards @ Turnitin

At Turnitin, we believe Total Rewards go far beyond pay. While salary, bonus, or commission are important, they’re only part of the value you receive in exchange for your work.

Beyond compensation, you’ll experience the intrinsic rewards of unleashing your potential and making a positive impact on global education. You’ll also thrive in a culture free of politics, surrounded by humble, inclusive, and collaborative teammates.

In addition, our extrinsic rewards include generous time off and health and wellness programs that provide choice, flexibility, and a safety net for life’s challenges. You’ll also enjoy a remote-first culture that empowers you to work with purpose and accountability in the way that suits you best, all supported by a comprehensive package that prioritizes your overall well-being.

Our Mission is to ensure the integrity of global education and meaningfully improve learning outcomes.

Our Values underpin everything we do.

  • Customer Centric: Our mission is focused on improving learning outcomes; we do this by putting educators and learners at the center of everything we do.
  • Passion for Learning: We are committed to our own learning and growth internally. And we support education and learning around the globe.
  • Integrity: Integrity is the heartbeat of Turnitin—it is the core of our products, the way we treat each other, and how we work with our customers and vendors.
  • Action & Ownership: We have a bias for action. We act like owners. We are willing to change even when it’s hard.
  • One Team: We strive to break down silos, collaborate effectively, and celebrate each others’ successes.
  • Global Mindset: We consider different perspectives and celebrate diversity. We are one team. The work we do has an impact on the world.

Global Benefits

  • Remote First Culture
  • Health Care Coverage*
  • Education Reimbursement*
  • Competitive Paid Time Off
  • Self-Care Days
  • National Holidays*
  • 2 Founder Days + Juneteenth Observed
  • Paid Volunteer Time*
  • Charitable contribution match*
  • Monthly Wellness or Home Office Reimbursement/*
  • Access to Modern Health (mental health platform)
  • Parental Leave*
  • Retirement Plan with match/contribution*

\* varies by country

Seeing Beyond the Job Ad

At Turnitin, we recognize it’s unrealistic for candidates to fulfill 100% of the criteria in a job ad.  We encourage you to apply if you meet the majority of the requirements because we know that skills evolve over time. If you’re willing to learn and unleash your potential alongside us, join our team!

#LI-AP1

Read the full description
Engineer Software Quality Engineer I (UK Remote) at Turnitin

Designs, writes, and executes automated and manual tests while collaborating with development teams to ensure application quality and drive continuous improvement.

Junior Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Company Description

When you join Turnitin, you’ll be welcomed into a company that is a recognized innovator in global education. For over 25 years, Turnitin has partnered with educators and institutions to develop learning integrity solutions that recognize the enduring value of critical thinking in a rapidly changing world. Over 16,000 academic institutions, publishers, and corporations use our services in more than 185 countries around the world: Turnitin Feedback Studio, Clarity, Originality, Gradescope, ExamSoft, Similarity, and iThenticate. Protecting the value of an authentic education is at the heart of who we are.

Experience a remote-first culture that empowers you to work with purpose and accountability in a way that best suits you, supported by a comprehensive package that prioritizes your overall well-being. Our diverse community of colleagues are all unified by a shared desire to make a difference in education.

Turnitin is a global organization with team members in over 35 countries including the United States, Mexico, United Kingdom, Australia, Japan, India, and the Philippines.

Job Description

The Software Quality Engineer I (SQE1) is an integral part of the Quality Engineering department, dedicated to assuring application quality, adherence to requirements, and an exceptional user experience by testing company applications. The main purpose of this role is to design, write, execute, and maintain automated and manual tests, as well as contribute to the overall test automation framework and continuous testing (CT) pipelines.

This role will primarily interact with product development teams. The SQE1 works cross-functionally with Software Developers, Product Managers, and fellow Quality Engineers within an Agile scrum environment to provide rapid, in-process testing results, identify risks, and drive continuous improvement.

Responsibilities:

  • Test Strategy & Planning: Partner with the Scrum team to define, develop, and implement quality assurance best practices and procedures, including test strategies, plans, cases, and other related assets.
  • Test Execution: Design, write, execute, and maintain both automated and manual test cases, ensuring high quality across functional, regression, integration, and scale testing.
  • Development Collaboration: Work closely with development teams to support code releases with functional test validations on multiple platforms and environments.
  • Defect Management: Investigate, record, triage, and track defects to verification and closure.
  • Risk Mitigation: Identify potential quality risks and dependencies early in the Agile cycle; collaborate on mitigation and eradication plans.
  • Modern QA & AI Capabilities: Proactively research, evaluate, and integrate modern QA tools and AI-driven testing capabilities (such as AI-assisted test generation, predictive analysis, and smart automation tools) to improve efficiency and testing coverage.
  • Professional Development (Goal): Commit to continuous professional and technical growth, developing deeper expertise in framework architecture, advanced CI/CD pipelines, and broad-scope test strategy design.

Qualifications

  • Education & Experience: Bachelor’s degree in Computer Science, Engineering, or equivalent practical experience, combined with 4+ years of experience in software quality engineering.

  • Automation & Scripting: Proven hands-on experience in UI Web Automation and API Automation using established frameworks. (Java, Python, Jenkins, docker, Maven, github).

  • QA Methodologies: Strong, high-level understanding of the Software Testing Life Cycle (STLC) and QA concepts, methodologies, and test management.

  • CI/CD/CT: competency with continuous integration and deployment tools (such as Jenkins, Git, etc.).

  • Agile Practices: Practical experience working within Agile frameworks (Scrum & Kanban), estimating work, and managing test data.

  • Modern QA & AI Capabilities: Proactively research, evaluate, and integrate modern QA tools and AI-driven testing capabilities (such as AI-assisted test generation, predictive analysis, and smart automation tools) to improve efficiency and testing coverage.

  • Nice to Have Skills:

    • UI Mobile Automation & Functional/Manual testing.
    • Performance testing experience (e.g., using JMeter).
    • Familiarity with AWS services (S3, Lambda, EC2, CloudWatch) or security/accessibility testing.

Tii Elements:

  • Accountability: Takes ownership of the quality of deliverables and works autonomously within standard guidelines to resolve issues.
  • Collaboration & Influencing: Builds productive working relationships with developers and product teams to resolve quality issues and collaborate on best practices.
  • Quality Focus: Strongly committed to delivering exceptional user experiences and robust software quality.
  • Relationship Building: Cultivates stable, productive working relationships across cross-functional teams.
  • Curiosity: Actively stays up-to-date with new testing tools, test strategies, and emerging technologies (including AI tools).

Additional Information

Total Rewards @ Turnitin

At Turnitin, we believe Total Rewards go far beyond pay. While salary, bonus, or commission are important, they’re only part of the value you receive in exchange for your work.

Beyond compensation, you’ll experience the intrinsic rewards of unleashing your potential and making a positive impact on global education. You’ll also thrive in a culture free of politics, surrounded by humble, inclusive, and collaborative teammates.

In addition, our extrinsic rewards include generous time off and health and wellness programs that provide choice, flexibility, and a safety net for life’s challenges. You’ll also enjoy a remote-first culture that empowers you to work with purpose and accountability in the way that suits you best, all supported by a comprehensive package that prioritizes your overall well-being.

Our Mission is to ensure the integrity of global education and meaningfully improve learning outcomes.

Our Values underpin everything we do.

  • Customer Centric: Our mission is focused on improving learning outcomes; we do this by putting educators and learners at the center of everything we do.
  • Passion for Learning: We are committed to our own learning and growth internally. And we support education and learning around the globe.
  • Integrity: Integrity is the heartbeat of Turnitin—it is the core of our products, the way we treat each other, and how we work with our customers and vendors.
  • Action & Ownership: We have a bias for action. We act like owners. We are willing to change even when it’s hard.
  • One Team: We strive to break down silos, collaborate effectively, and celebrate each others’ successes.
  • Global Mindset: We consider different perspectives and celebrate diversity. We are one team. The work we do has an impact on the world.

Global Benefits

  • Remote First Culture
  • Health Care Coverage*
  • Education Reimbursement*
  • Competitive Paid Time Off
  • Self-Care Days
  • National Holidays*
  • 2 Founder Days + Juneteenth Observed
  • Paid Volunteer Time*
  • Charitable contribution match*
  • Monthly Wellness or Home Office Reimbursement/*
  • Access to Modern Health (mental health platform)
  • Parental Leave*
  • Retirement Plan with match/contribution*

\* varies by country

Seeing Beyond the Job Ad

At Turnitin, we recognize it’s unrealistic for candidates to fulfill 100% of the criteria in a job ad.  We encourage you to apply if you meet the majority of the requirements because we know that skills evolve over time. If you’re willing to learn and unleash your potential alongside us, join our team!

#LI-AP1

Read the full description
Engineer Technical Lead - GPU Infrastructure at Tether.io

Technical Lead oversees GPU infrastructure and Kubernetes-based compute platform development, managing architecture and team for Tether Data's managed inference services.

Lead Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Description

Join Tether and Shape the Future of Digital Finance

At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.

Innovate with Tether

Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.

But that’s just the beginning:

Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.

Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.

Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.

Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.

Why Join Us?

Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.

If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.

Are you ready to be part of the future?

About the job

Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.

The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.

This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.

Responsibilities

Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.

Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.

Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.

Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.

Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.

Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.

Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.

Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.

Hiring. Complete the platform team and set the technical bar for the engineers who join it.

Requirements

Must have

  • Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.

  • Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.

  • GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.

  • High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.

  • Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.

  • Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.

  • HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.

  • Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.

  • Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.

  • A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.

  • Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.

  • Excellent written and spoken English. Most partner and leadership work happens in writing.

  • Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.

Desirable

  • Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).

  • Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.

  • VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).

  • Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.

  • Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.

  • Peer-to-peer or distributed-systems background.

  • Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.

Important information for candidates

Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:

  • Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/

  • Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.

  • Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.

  • Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io

  • We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.

When in doubt, feel free to reach out through our official website.

Read the full description
Engineer Technical Lead - GPU Infrastructure at Tether.io

Technical Lead designs and manages GPU infrastructure and Kubernetes-based compute platform for AI inference and managed services.

Lead Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Description

Join Tether and Shape the Future of Digital Finance

At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.

Innovate with Tether

Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.

But that’s just the beginning:

Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.

Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.

Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.

Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.

Why Join Us?

Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.

If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.

Are you ready to be part of the future?

About the job

Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.

The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.

This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.

Responsibilities

Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.

Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.

Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.

Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.

Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.

Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.

Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.

Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.

Hiring. Complete the platform team and set the technical bar for the engineers who join it.

Requirements

Must have

  • Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.

  • Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.

  • GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.

  • High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.

  • Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.

  • Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.

  • HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.

  • Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.

  • Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.

  • A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.

  • Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.

  • Excellent written and spoken English. Most partner and leadership work happens in writing.

  • Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.

Desirable

  • Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).

  • Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.

  • VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).

  • Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.

  • Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.

  • Peer-to-peer or distributed-systems background.

  • Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.

Important information for candidates

Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:

  • Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/

  • Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.

  • Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.

  • Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io

  • We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.

When in doubt, feel free to reach out through our official website.

Read the full description