Writing.io Jobs

Find the best remote jobs. Answer a few questions and we'll deploy a powerful assistant to help you search, create alerts, and more.

1 What roles are you open to?

2 Experience level

3 Work style

Did you know? If memory is enabled, Writing.io can remember your job search preferences and help you to improve your resume, craft customized outreach and more.

Engineer Poland – Director of Software Engineering (Experiences)

Leads software engineering teams and strategy for travel experiences platform, overseeing technical architecture and engineering excellence.

Lead Posted about 7 hours ago Jobicy AI
What this role involves
About Tripadvisor The Tripadvisor Group connects people to experiences worth sharing, and aims to be the world’s most trusted source for travel and experiences. We leverage our brands, technology, and...
Read the full description
Engineer LawnStarter: Principal Quality Engineer

Principal Quality Engineer owns quality standards, leads the QA squad, drives testing strategy and tooling evolution, and scales quality practices across the organization.

Lead Posted about 18 hours ago We Work Remotely — Programming
What this role involves

Headquarters: Mexico
URL: http://lawnstarter.com

About LawnStarter

LawnStarter is the nation's leading on-demand marketplace for lawn care and related services, with over $150M in annual bookings. We're expanding beyond lawn care to become the one-stop shop for all home services.

About Engineering Quality at LawnStarter

Our QA team is growing, and we're investing real money in better test coverage. We already have shared quality standards and a QA squad that organizes itself horizontally across teams. What we don't have is someone dedicated to owning that vision: every QA is heads-down on their own team, so nobody has the room to drive the standards forward, unify how the squad works, or push the state of the art. That's the gap this role fills. It's not a nice-to-have we're adding because things are going well. It's the structure and leadership we need to make what we've already built actually compound.

The Role

You'll be LawnStarter's first Principal Quality Engineer, reporting directly to the Head of Engineering. We already have quality standards, working CI/CD gates, and AI tooling in the mix. Nobody owns the vision behind all of it, drives it forward, or leads the QA squad that keeps it running. You will.

This isn't a QA management job today. You won't spend your day approving tickets or chasing bug counts. You'll take ownership of the standards and tooling that already exist, work collaboratively with the QA squad to improve them, and keep pushing the state of quality engineering forward as the org grows. As the QA squad grows, this role could take on direct people management of QAs, so hiring, coaching, and performance management experience is a real plus even though it's not the day-one job.

You're not starting from zero, and you're not inheriting a mess either. That means the job isn't "invent a strategy nobody has thought about." It's "listen to the QAs who live in this every day, understand where the pain is, and lead them toward what's next." We need someone who reads where LawnStarter is today and builds from there, not someone who shows up with a favorite tool stack and a coverage number they picked before they met the team.

What makes this role different:

  • You own and evolve the quality strategy together with the QA team, instead of building one in a vacuum or just executing someone else's.
  • You have direct sponsorship from the Head of Engineering and the room to keep growing the AI-native tooling already in the mix.
  • Your success is measured by how much better the whole engineering org's quality gets under your leadership, not by a headcount number.

What You'll Own

  • Strategy and vision: The company-wide quality engineering roadmap, plus a quality accountability framework that spells out what engineers, EMs, and PMs each own at every stage of building software.
  • Test strategy across the stack: Unit, component, contract, API, end-to-end, mobile, and non-functional testing (performance, load, accessibility), agreed with each team as their quality contract.
  • AI-native quality tooling: AI agents and skills, including Claude Code style tooling, that any engineer can run on demand for test generation, code review, and quality guidance. You'll keep hunting for the next AI capability worth putting in front of the team.
  • CI/CD and test infrastructure: Required checks, quality gates, E2E and visual regression pipelines, performance budgets, and the mobile test automation strategy across emulators and real device clouds. That includes deciding what actually needs to block a deploy versus what can run async, and keeping the pipeline itself fast.
  • Production quality signal: Tying monitoring into incident response and building dashboards teams actually check, not ones that just exist.
  • Technical mentorship of the QA discipline: QA engineers keep reporting to their own Engineering Managers today. You won't manage their day-to-day or their careers right now. You'll own the technical and craft side of QA org-wide, mentor QA engineers across every squad, and send your input to their EMs to feed performance reviews. That could change as the squad grows into a direct reporting line to this role.

Problems to Solve

A solid quality bar with no one steering it.

We already have shared standards, working quality gates, and AI tooling in the mix. What's missing is someone whose job is to look at all of it, decide what needs to improve next, and actually drive that improvement instead of it being everyone's part-time responsibility.

A QA squad with no dedicated leadership.

Our QAs already organize horizontally across teams, but every one of them is focused on their own team's work day to day. Nobody has the bandwidth to unify how they work, spread what's working on one team to the others, or represent their pain points where decisions get made. You'll give that squad the structure and leadership it's missing.

Standards that need continuous R&D, not a rewrite.

The foundation works. The risk is standing still while the org grows. You'll keep pushing the state of the art, testing new approaches and tools, and deciding what's worth rolling out broadly versus what stays an experiment.

Quality work that's still someone's side project.

Right now, improving quality practice competes with each QA's day-to-day team commitments. You'll make it someone's actual job to carry that forward, and get engineers, EMs, and PMs treating it as planned work instead of something squeezed in.

What Success Looks Like (Year 1)

  • A clear quality vision, owned and communicated: Every product team can point to what "good" means for their testing and why, with you as the person who set that direction and keeps it current.
  • Measurable improvement on existing gates: Flakiness, pipeline cost, and what runs sync versus async on core repositories are visibly better than when you started, not just maintained.
  • At least one new AI-native capability shipped: You've pushed the existing AI tooling forward with something new, like an added test-generation flow or audit, that engineers actually reach for weekly.
  • A QA squad that feels led: The QAs across teams can point to concrete ways their work, tools, or standards improved because someone was finally driving it.

Requirements

Who You Are

AI-native, specifically for quality work. You already use AI to generate tests, not just to autocomplete code. You've set up agentic tooling like Claude Code or MCP-based workflows so engineers can run on-demand test generation or code review against their own repos, and you keep experimenting with what AI can take off a team's plate next. This is unlikely to be a good fit if AI tools mostly make you nervous, or your use of them stops at your own personal productivity.

A strategic thinker, not a tool zealot. Give you a team's current maturity level, and you'll design the testing approach that fits where they are, not the one you used at your last job. You resist the urge to mandate a specific framework or a specific coverage number before you understand the context. Someone who shows up insisting "we need to hit 90% coverage" or "everyone must use Playwright" without first learning how the team works will struggle here.

Deeply hands-on across the whole test pyramid. You've built and maintained component, contract, API, E2E, and performance test suites yourself, not just reviewed slides about them. You're the person who steps in when the engineering team is blocked by the pipeline and can't ship. If your test automation experience stops at reading dashboards other people built, this role will feel out of reach fast.

A patient teacher who still ships. Coaching engineers on testing practice and running enablement sessions is part of the job. So is personally building the CI pipelines and the tooling. You don't write the doc and hope someone else implements it. This is unlikely to be a good fit for someone who only wants to advise, or who avoids hands-on infrastructure work.

Comfortable owning influence without owning headcount, at least for now. You'll mentor QA engineers across every squad on technical craft and send notes to their EMs for performance reviews. You won't be their manager on day one, and you won't set their career path. This works well if you're motivated by discipline-wide technical influence. It won't if you need direct reports from the start to feel engaged. If you've hired, coached, and managed performance for a QA or engineering team before, that's a genuine plus: this role could grow into managing the QA squad directly as it scales.

This Role Is NOT

  • A traditional QA management job, at least not on day one: You won't manage QA engineers day to day or own their career growth right now. Their EMs do that. You own the technical and strategic direction of the QA discipline instead, with the door open to direct management as the squad grows.
  • A "throw out what's there and start over" job: We have real standards and tooling worth building on. You can't parachute in with a pre-packaged process or a favorite tool stack and mandate it regardless of fit.
  • A ticket-approval or gatekeeping role: You're not the last line of defense that blocks releases at the end. You're embedding quality earlier, so there's less to block.
  • A headcount-growth play: Your success isn't measured by how big a QA team you build. It's measured by how much better the whole engineering org's quality gets.

Benefits

Compensation & Benefits

  • Base salary: $80,000-$100,000 USD annually.
  • Fully remote: This is deep-focus analytical work with US-facing partners. We hire the best analyst regardless of city and trust you to manage your environment and overlap hours.
  • Flexible PTO: Measured on outcomes, not hours logged.

LawnStarter provides equal employment opportunities (EEO) to all employees and applicants for employment without regard to race, color, religion, sex, national origin, age, disability, or genetics. We comply with applicable state and local laws governing nondiscrimination in employment.

To apply: https://weworkremotely.com/remote-jobs/lawnstarter-principal-quality-engineer-1

Read the full description
Engineer Site Reliability Engineer, Tech Lead at Loadsmart

Site Reliability Engineer Tech Lead builds and maintains critical infrastructure systems, ensures application reliability and SLAs, and leads reliability projects across engineering squads.

Lead Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

ARE YOU INTERESTED IN JOINING AN INNOVATIVE LOGISTICS TECHNOLOGY COMPANY?

Loadsmart is a growth-stage technology company valued at over $1 billion (a true Tech Unicorn)!

We are a collection of industry veterans and user-centered engineers using innovative technology to fearlessly reinvent the future of freight by helping shippers, brokers, warehouses and carriers to move more with less.

With headquarters in Chicago and a globally distributed remote team, Loadsmart continues to attract top talent committed to driving meaningful change. We seek professionals who embody our core values: curiosity, clarity, results, commitment, and teamwork.

In the SRE, Tech Lead role you will build and maintain the company’s internal platform, driving operational excellence and empowering the entire engineering team. You should have experience in analyzing, proposing, and implementing safer systems and processes. Collaborating closely with engineering squads across platform engineering, you will ensure our applications are both safe and reliable. As a hands-on leader, you will stay actively involved in technical work while collaborating closely with internal stakeholders and engineering squads across the organization to ensure our applications are both safe and reliable.

DEPARTMENT: Engineering

LOCATION: Anywhere in Brazil - Remote

WHAT YOU GET TO DO

  • Collaborate with and support our creative, tight-knit development team.
  • Design, deploy, and operate Loadsmart’s critical systems while balancing reliability, cost, and agility.
  • Play a key role in driving reliability projects with engineering teams.
  • Utilize your intuitive problem-solving skills and contagious positive attitude to tackle challenging and exciting issues, inspiring those around you.
  • Collect metrics and understand their business impact, encouraging the team to do the same.
  • Perform troubleshooting and root-cause analysis of system operation issues.
  • Be accountable for the platform’s Service Level Agreements and Objectives.
  • Provide infrastructure support during off-hours as needed
  • Take ownership of software infrastructure projects
  • Seek, give, and receive constructive feedback through code and specification reviews.
  • Familiarity with AI agents and agentic workflows, applying AI across the SDLC (AI-assisted coding), LLMs, MCP servers/gateways, and how emerging AI tooling can improve reliability and operations is a plus.

REQUIRED QUALIFICATIONS:

  • 1-3 years leading Reliability Work across multiple engineering squads
  • Over 5 years of experience in Cloud Computing, SRE/DevOps
  • Proven experience collaborating with internal stakeholders across multiple engineering squads
  • Strong project management skills with a demonstrated ability to delegate and mentor team members
  • Proficient in English communication (both written and spoken) to collaborate in an international team with native and non-native English speakers
  • Detail-oriented with high initiative and self-motivation
  • Strong understanding of software engineering principles and how systems work under the hood
  • In-depth knowledge of modern networking and operating systems
  • Proficiency in AWS, cloud environments, containers, Kubernetes, Docker, and DevOps engineering, including managing tests and CI/CD pipelines
  • Familiarity with automation tools and provisioners like Terraform, Ansible, or Chef
  • Solid troubleshooting and system engineering experience in UNIX/Linux production environments
  • Experience with monitoring, alerting, and incident management
  • Proficiency in automating tasks with scripting languages like Python, Bash, etc
  • Experience or exposure to PostgreSQL and DBA responsibilities is a plus
  • Fluent in English (both written and spoken); comfortable interacting with native English speakers daily.

WORKING AT LOADSMART:

• Competitive base salaries - we believe in rewarding top talent

• Extremely competitive Equity package - become a shareholder in our company!

• Loadie Time Off - PTO and sick days without a limit

At Loadsmart, we believe our biggest asset is our people. We are proud to be an equal opportunity employer, hiring and developing individuals from diverse backgrounds and experiences to add to our collaborative culture. Loadsmart treats all candidates and employees with respect and does not discriminate in our recruiting, hiring, and promoting processes, including on the basis of race, color, religion, sex, age, sexual orientation, gender identity and/or expression, national origin, veteran status, or disability.

It is the policy of Loadsmart that all offers of employment made shall be contingent upon successful completion of electronic background check(s). These checks will be job-related, consistent with business necessity and conducted by our vendor, pursuant to all applicable laws, rules, policies and procedures of our candidates’ specific locale.

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

Read the full description
Engineer Technical Lead - GPU Infrastructure at Tether.io

Lead GPU infrastructure engineering for a Kubernetes-based compute platform, managing containerized GPU resources and inference endpoints at scale.

Lead Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Description

Join Tether and Shape the Future of Digital Finance

At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.

Innovate with Tether

Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.

But that’s just the beginning:

Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.

Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.

Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.

Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.

Why Join Us?

Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.

If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.

Are you ready to be part of the future?

About the job

Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.

The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.

This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.

Responsibilities

Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.

Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.

Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.

Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.

Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.

Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.

Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.

Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.

Hiring. Complete the platform team and set the technical bar for the engineers who join it.

Requirements

Must have

  • Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.

  • Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.

  • GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.

  • High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.

  • Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.

  • Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.

  • HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.

  • Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.

  • Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.

  • A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.

  • Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.

  • Excellent written and spoken English. Most partner and leadership work happens in writing.

  • Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.

Desirable

  • Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).

  • Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.

  • VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).

  • Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.

  • Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.

  • Peer-to-peer or distributed-systems background.

  • Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.

Important information for candidates

Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:

  • Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/

  • Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.

  • Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.

  • Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io

  • We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.

When in doubt, feel free to reach out through our official website.

Read the full description
Engineer Technical Lead - GPU Infrastructure at Tether.io

Technical Lead oversees GPU infrastructure and Kubernetes platform development for Tether Data's managed inference and compute services.

Lead Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Description

Join Tether and Shape the Future of Digital Finance

At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.

Innovate with Tether

Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.

But that’s just the beginning:

Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.

Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.

Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.

Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.

Why Join Us?

Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.

If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.

Are you ready to be part of the future?

About the job

Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.

The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.

This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.

Responsibilities

Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.

Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.

Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.

Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.

Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.

Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.

Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.

Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.

Hiring. Complete the platform team and set the technical bar for the engineers who join it.

Requirements

Must have

  • Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.

  • Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.

  • GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.

  • High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.

  • Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.

  • Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.

  • HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.

  • Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.

  • Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.

  • A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.

  • Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.

  • Excellent written and spoken English. Most partner and leadership work happens in writing.

  • Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.

Desirable

  • Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).

  • Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.

  • VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).

  • Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.

  • Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.

  • Peer-to-peer or distributed-systems background.

  • Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.

Important information for candidates

Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:

  • Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/

  • Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.

  • Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.

  • Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io

  • We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.

When in doubt, feel free to reach out through our official website.

Read the full description
Engineer Technical Lead - GPU Infrastructure at Tether.io

Technical Lead manages GPU infrastructure and Kubernetes-based compute platform, overseeing architecture, performance, and team delivery for a managed inference service.

Lead Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Description

Join Tether and Shape the Future of Digital Finance

At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.

Innovate with Tether

Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.

But that’s just the beginning:

Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.

Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.

Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.

Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.

Why Join Us?

Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.

If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.

Are you ready to be part of the future?

About the job

Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.

The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.

This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.

Responsibilities

Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.

Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.

Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.

Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.

Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.

Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.

Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.

Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.

Hiring. Complete the platform team and set the technical bar for the engineers who join it.

Requirements

Must have

  • Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.

  • Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.

  • GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.

  • High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.

  • Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.

  • Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.

  • HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.

  • Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.

  • Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.

  • A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.

  • Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.

  • Excellent written and spoken English. Most partner and leadership work happens in writing.

  • Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.

Desirable

  • Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).

  • Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.

  • VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).

  • Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.

  • Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.

  • Peer-to-peer or distributed-systems background.

  • Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.

Important information for candidates

Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:

  • Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/

  • Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.

  • Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.

  • Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io

  • We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.

When in doubt, feel free to reach out through our official website.

Read the full description
Engineer Software Architect at Fingerprint

Lead cross-cutting architecture decisions across engineering teams, designing scalable systems and setting technical standards for a fraud detection platform.

Lead Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Fingerprint empowers enterprises to detect and stop online fraud with the world’s most accurate device intelligence.  We lead our industry with bleeding-edge identification capabilities and work on turning new ideas and discoveries in the fraud detection space into reality. Our customers range from innovative startups to leading enterprise companies, including Plaid, Dropbox, and Booking.com.

Fingerprint is a globally dispersed, 100% remote company. We were named on on the 2026 Forbes Best Startup Employers list and ranked #803 on the 2026 Inc. 5000 list of America’s fastest-growing private companies.

We have raised $77M and are backed by Craft Ventures (Tesla, Facebook, Airbnb ), Nexus Venture Partners ( Postman, Apollo.io, MinIO, Druva) and Uncorrelated Ventures ( Redis, Rollbar,  Gradle).

About the role

You will be Fingerprint’s first dedicated Architect. You will lead cross-cutting architecture — how our systems fit together and where they need to go — working alongside the Staff and Lead engineers who already shape them, without direct reports. The mandate has three parts. First, go deep: build the end-to-end picture nobody currently has time to hold, identify where the platform will strain as we scale, and set the direction to address it. Second, raise the bar: strengthen and extend our design review, API, service, and reliability practices so eight engineering groups can move fast without stepping on each other. Third, make it durable: give cross-cutting architecture work consistent cadence, a durable decision record, and follow-through, so strong individual judgment compounds into platform-level outcomes.

You report directly to the VP of Engineering. That placement is deliberate — you need neutrality across groups so the standards you set with teams are adopted everywhere.

What you’ll do

Own the end-to-end architecture

  • Build and maintain the end-to-end architecture picture — data flows, service boundaries, ownership seams — and keep it current as the org changes.
  • Identify systemic scaling, reliability, and cost risks across the pipeline (ingestion, identification, signals, delivery) and drive architectural changes to address them before they become incidents.
  • Set direction for API design and service patterns across REST, SDK, and MCP surfaces so customers see one coherent platform.
  • Shape the roadmap for foundational initiatives (e.g., cell architecture, multi-region readiness, rate limiting, feature flag infrastructure) and lead the hardest designs yourself.

Lead cross-cutting architecture

  • Serve as full-time technical lead across the engineering org: set the agenda, drive decisions to closure, own the decision record, and follow through on adoption across teams.
  • Partner with Engineering leadership and the Staff/Lead engineers across teams to prioritize cross-cutting work and connect it to team roadmaps.
  • Mentor Staff-level engineers on system design, technical writing, and influence — raise the bar for what “Staff” means at Fingerprint.

Strengthen and extend best practices

  • Build on the design review practices we already have (RFCs, cross-team design reviews) and make them more consistent, lighter-weight, and more useful, so teams choose to use them.
  • Codify and extend standards for observability, quality, security-by-design, and API evolution/versioning, working with Cloud Platform, Security, and product teams.
  • Review and sign off on critical designs; teach through review rather than gatekeeping.

Set us up for scale

  • Translate business trajectory (Tier 1 enterprise customers, new product launches, AI/MCP workflows) into a multi-quarter technical strategy and sequence the work.
  • Advise Engineering leadership and executive team on build/buy, platform investment, and technical risk in plain language.
  • Lead AI adoption in engineering practice: set norms for AI-assisted design, coding, and review across teams, and shape our architecture and documentation so AI agents can work in our systems as effectively as engineers do.
  • Stay hands-on: prototype, write reference implementations and documentation, and dig into production when the problem demands it.

What we’re looking for

  • 10+ years of software engineering experience, including 3+ years operating as an Architect for a 100+ Engineering org — you have owned architecture for a platform, not just a service.
  • Track record of leading through influence: you have led senior engineers you didn’t manage, run design review or architecture forums that people actually used, and driven adoption of standards across an org.
  • Deep expertise in distributed systems and high-throughput, low-latency backend architecture (Go or similar systems language; Kubernetes/AWS; event-driven and data-intensive systems). Fluency across adjacent layers — client SDKs, data pipelines, ML serving — is a strong plus.
  • Experience designing public APIs and SDKs at scale, including versioning, backward compatibility, and multi-surface consistency.
  • Demonstrated ability to anticipate systemic risk and act on it — you can point to failures you prevented, not just ones you fixed.
  • Exceptional written communication. You make complex decisions legible to engineers and executives alike, and you default to async, documented decision-making.
  • AI-native by default. You use AI coding and reasoning tools as a normal part of how you design, prototype, review, and write — and you have opinions, from experience, about where they accelerate engineering work and where they don’t yet.
  • Architect for an AI-assisted org. You think about how codebases, documentation, service boundaries, and APIs should be shaped so that both humans and AI agents can work in them safely — legible structure, strong contracts, automated verification.
  • Pragmatism over purity. You balance long-term architecture with delivery pressure and know when “good enough” is the right call.
  • Comfortable being the first: you have stepped into a dedicated role where the work was previously shared across a group, earned the trust of the people already doing it, and made them more effective rather than displacing them.

Nice to have

  • Experience in fraud detection, identity, device intelligence, or other adversarial domains.
  • Multi-region / cell-based architecture experience.
  • Experience with MCP, LLM tool integrations, or agent-facing API design.
  • Familiarity with ClickHouse, Kafka, or similar high-volume data infrastructure.

Compensation & Transparency

For US-based employees, the cash compensation range for this role is $177,000 – $240,000. We set standard ranges for all US roles based on function, level, and geographic location, benchmarked against similar stage growth companies. To comply with local legislation and provide greater transparency, we share salary ranges on all job postings. However, these ranges are specific to the hiring location and may differ within or outside the US. Offers vary depending on, but not limited to, relevant experience, education, certifications/licenses, skills, training, and market conditions.

Due to regulatory and security reasons, there’s a small number of countries where we cannot have Fingerprint teammates based. Additionally, because Fingerprint is an all-remote company and people can join our workforce from almost any country, we do not sponsor visas. Fingerprint teammates need to be authorized to work from their home location.

We are dedicated to creating an inclusive work environment for everyone. We embrace and celebrate the unique experiences, perspectives and cultural backgrounds that each employee brings to our workplace. Fingerprint strives to foster an environment where our employees feel respected, valued and empowered, and our team members are at the forefront in helping us promote and sustain an inclusive workplace. We highly encourage people from underrepresented groups in tech to apply.

If you are applying as a resident of California, please read our CCPA notice here.

If you are applying as a resident of the EU, please read our GDPR notice here.

  • We have noticed a rise in recruiting impersonations across the industry, where scammers attempt to access candidates’ personal and financial information through fake interviews and offers. All Fingerprint recruiting email communications will always come from the @fingerprint.com domain. Any outreach claiming to be from Fingerprint via other sources should be ignored.*

Due to regulatory and security reasons, there’s a small number of countries where we cannot have Fingerprint teammates based. Additionally, because Fingerprint is an all-remote company and people can join our workforce from almost any country, we do not sponsor visas. Fingerprint teammates need to be authorized to work from their home location.

We are dedicated to creating an inclusive work environment for everyone. We embrace and celebrate the unique experiences, perspectives and cultural backgrounds that each employee brings to our workplace. Fingerprint strives to foster an environment where our employees feel respected, valued and empowered, and our team members are at the forefront in helping us promote and sustain an inclusive workplace. We highly encourage people from underrepresented groups in tech to apply.

If you are applying as a resident of California, please read our CCPA notice here.

If you are applying as a resident of the EU, please read our GDPR notice here.

**We have noticed a rise in recruiting impersonations across the industry, where scammers attempt to access candidates’ personal and financial information through fake interviews and offers. All Fingerprint recruiting email communications will always come from the @fingerprint.com domain. Any outreach claiming to be from Fingerprint via other sources should be ignored.

Read the full description
Engineer Technical Lead - GPU Infrastructure at Tether.io

Technical Lead oversees GPU infrastructure and Kubernetes-based compute platform development, managing architecture and team for Tether Data's managed inference services.

Lead Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Description

Join Tether and Shape the Future of Digital Finance

At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.

Innovate with Tether

Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.

But that’s just the beginning:

Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.

Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.

Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.

Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.

Why Join Us?

Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.

If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.

Are you ready to be part of the future?

About the job

Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.

The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.

This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.

Responsibilities

Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.

Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.

Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.

Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.

Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.

Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.

Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.

Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.

Hiring. Complete the platform team and set the technical bar for the engineers who join it.

Requirements

Must have

  • Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.

  • Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.

  • GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.

  • High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.

  • Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.

  • Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.

  • HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.

  • Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.

  • Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.

  • A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.

  • Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.

  • Excellent written and spoken English. Most partner and leadership work happens in writing.

  • Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.

Desirable

  • Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).

  • Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.

  • VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).

  • Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.

  • Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.

  • Peer-to-peer or distributed-systems background.

  • Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.

Important information for candidates

Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:

  • Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/

  • Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.

  • Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.

  • Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io

  • We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.

When in doubt, feel free to reach out through our official website.

Read the full description
Engineer Technical Lead - GPU Infrastructure at Tether.io

Technical Lead designs and manages GPU infrastructure and Kubernetes-based compute platform for AI inference and managed services.

Lead Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Description

Join Tether and Shape the Future of Digital Finance

At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.

Innovate with Tether

Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.

But that’s just the beginning:

Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.

Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.

Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.

Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.

Why Join Us?

Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.

If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.

Are you ready to be part of the future?

About the job

Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.

The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.

This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.

Responsibilities

Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.

Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.

Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.

Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.

Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.

Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.

Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.

Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.

Hiring. Complete the platform team and set the technical bar for the engineers who join it.

Requirements

Must have

  • Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.

  • Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.

  • GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.

  • High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.

  • Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.

  • Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.

  • HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.

  • Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.

  • Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.

  • A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.

  • Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.

  • Excellent written and spoken English. Most partner and leadership work happens in writing.

  • Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.

Desirable

  • Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).

  • Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.

  • VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).

  • Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.

  • Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.

  • Peer-to-peer or distributed-systems background.

  • Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.

Important information for candidates

Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:

  • Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/

  • Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.

  • Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.

  • Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io

  • We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.

When in doubt, feel free to reach out through our official website.

Read the full description
Engineer Technical Lead - GPU Infrastructure at Tether.io

Technical Lead oversees GPU infrastructure and Kubernetes-based platform development, managing compute resources and managed inference endpoints for Tether Data's AI infrastructure.

Lead Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Description

Join Tether and Shape the Future of Digital Finance

At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.

Innovate with Tether

Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.

But that’s just the beginning:

Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.

Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.

Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.

Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.

Why Join Us?

Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.

If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.

Are you ready to be part of the future?

About the job

Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.

The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.

This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.

Responsibilities

Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.

Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.

Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.

Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.

Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.

Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.

Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.

Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.

Hiring. Complete the platform team and set the technical bar for the engineers who join it.

Requirements

Must have

  • Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.

  • Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.

  • GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.

  • High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.

  • Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.

  • Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.

  • HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.

  • Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.

  • Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.

  • A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.

  • Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.

  • Excellent written and spoken English. Most partner and leadership work happens in writing.

  • Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.

Desirable

  • Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).

  • Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.

  • VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).

  • Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.

  • Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.

  • Peer-to-peer or distributed-systems background.

  • Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.

Important information for candidates

Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:

  • Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/

  • Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.

  • Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.

  • Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io

  • We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.

When in doubt, feel free to reach out through our official website.

Read the full description
Engineer Technical Lead (Python, AWS, Terraform Experience) at Netcompany

Technical Lead drives team technical delivery, designs system architecture, and writes code across Python, AWS, Terraform, and multiple other tech stacks.

Lead Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Company Description

Are you ready to join the forefront of technology innovation with Netcompany?

As one of the fastest growing technology companies, we are disrupting the marketplace and revolutionizing the way businesses operate. Our vision is to be the leading digital challenger in Europe whilst evolving the next generation of IT consulting.

Operating across both public and private sectors, we offer a comprehensive range of services from application development and seamless cloud migration to program delivery and service operations, our offerings are designed to meet diverse business needs.

Job Description

This is an exciting opportunity within technology consultancy, offering fast-track career development and the chance to explore cutting-edge technologies. As a Tech Lead, you will play a crucial role in driving your team technically to deliver high-quality code, while also supporting the full life cycle of software development, including hands-on coding.

On the specific programme, we manage multiple services, including both run-and-maintain activities and transformative work by designing and developing new features. You will work with a wide range of technologies including Python, C#, Java, TypeScript, React, Terraform, and AWS/Azure. We always use the most suitable technology for each project, giving you the opportunity to learn and work with different languages.

Can you see yourself defining and ensuring best practice and quality assurance across multiple services? Are you looking to be involved in all parts of the process, from design and development, to ensuring that we deliver a high-quality end product to our clients?

Qualifications

Essential:

  • Experience leading technical teams, with a proven ability to define technical strategy and architecture for multiple services.

  • Full-stack or backend technical background with experience in Python and TypeScript (or similar). Cloud infrastructure experience with AWS or Azure and IaC with Terraform.

  • Experience delivering UK Government digital services, with a strong understanding of the GDS/NHS Service Standard and assessment process.

  • Strong collaborative leadership skills, working effectively across multidisciplinary teams including Product, UCD and Delivery.

  • Ability to balance and prioritise technical, user and business needs when making decisions.

  • Champions user-centred ways of working, ensuring research and design are appropriately prioritised throughout the development lifecycle.

  • Excellent communication skills and professional attitude, including presentation to non-technical audiences.

  • Proven track record of driving quality assurance, code reviews, and process optimisation across multiple services.

  • Experience leading projects from start to finish, managing technical and non-technical stakeholders, and translating customer requirements into technical designs.

  • Full life cycle delivery experience, from analysis and design through to implementation and QA.

  • Mentoring and coaching developers, fostering a collaborative and high-performing team culture.

  • Highly ambitious, wanting to excel in your career.

Desirable:

  • Experience delivering digital services within multidisciplinary product teams.

  • Knowledge of accessibility and inclusive design, including WCAG standards.

  • Experience coaching and developing engineers and building collaborative team cultures.

  • Experience delivering services in complex organisations or programmes with multiple stakeholders and competing priorities.

Candidates MUST be willing to travel to client site anywhere in the UK when needed and MUST have the right to work in the UK.

Hybrid working model available.

Additional Information

Benefits include

  • Hybrid working model with some flexible working
  • 25 days’ holiday
  • Private Medical Health care via Vitality
  • Pension contribution, Life Assurance
  • Professional certifications supported as part of learning and development.
  • A range of retail discounts to enhance your lifestyle, encompassing restaurants, supermarkets, travel, leisure activities and health and well-being services.
  • Access to our Employee Resource Groups, our groups represent diverse backgrounds and provide a platform for colleagues to connect, learn, and support one another.

Company information

At Netcompany, we pride ourselves on our entrepreneurial spirit and our capacity for doing things differently. Our culture is built on fostering low bureaucracy, emphasizing high agility and promoting flexibility, enabling everyone to contribute their best.

Our journey began in the UK with the acquisition of Hunter Macdonald in 2017. As one of Northern Europe’s most accomplished IT companies, we have expanded our headcount globally to 7400+ employees and have offices in UK, Denmark, Norway, Poland, Holland and Vietnam.

We are a Disability Confident Employer and are committed to creating an inclusive and diverse environment that celebrates every individual. Our recruitment processes are based on individual merit. If you require any reasonable adjustments or additional support during the interview process, please email us at [email protected] for assistance.

#LI-RS1

Read the full description
Engineer Technical Lead (Python, AWS, Terraform Experience) at Netcompany

Technical lead directs multiple engineering teams on cloud-based services, writes Python/TypeScript code, and ensures architecture quality across AWS/Azure infrastructure and Terraform deployments.

Lead Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Company Description

Are you ready to join the forefront of technology innovation with Netcompany?

As one of the fastest growing technology companies, we are disrupting the marketplace and revolutionizing the way businesses operate. Our vision is to be the leading digital challenger in Europe whilst evolving the next generation of IT consulting.

Operating across both public and private sectors, we offer a comprehensive range of services from application development and seamless cloud migration to program delivery and service operations, our offerings are designed to meet diverse business needs.

Job Description

This is an exciting opportunity within technology consultancy, offering fast-track career development and the chance to explore cutting-edge technologies. As a Tech Lead, you will play a crucial role in driving your team technically to deliver high-quality code, while also supporting the full life cycle of software development, including hands-on coding.

On the specific programme, we manage multiple services, including both run-and-maintain activities and transformative work by designing and developing new features. You will work with a wide range of technologies including Python, C#, Java, TypeScript, React, Terraform, and AWS/Azure. We always use the most suitable technology for each project, giving you the opportunity to learn and work with different languages.

Can you see yourself defining and ensuring best practice and quality assurance across multiple services? Are you looking to be involved in all parts of the process, from design and development, to ensuring that we deliver a high-quality end product to our clients?

Qualifications

Essential:

  • Experience leading technical teams, with a proven ability to define technical strategy and architecture for multiple services.

  • Full-stack or backend technical background with experience in Python and TypeScript (or similar). Cloud infrastructure experience with AWS or Azure and IaC with Terraform.

  • Experience delivering UK Government digital services, with a strong understanding of the GDS/NHS Service Standard and assessment process.

  • Strong collaborative leadership skills, working effectively across multidisciplinary teams including Product, UCD and Delivery.

  • Ability to balance and prioritise technical, user and business needs when making decisions.

  • Champions user-centred ways of working, ensuring research and design are appropriately prioritised throughout the development lifecycle.

  • Excellent communication skills and professional attitude, including presentation to non-technical audiences.

  • Proven track record of driving quality assurance, code reviews, and process optimisation across multiple services.

  • Experience leading projects from start to finish, managing technical and non-technical stakeholders, and translating customer requirements into technical designs.

  • Full life cycle delivery experience, from analysis and design through to implementation and QA.

  • Mentoring and coaching developers, fostering a collaborative and high-performing team culture.

  • Highly ambitious, wanting to excel in your career.

Desirable:

  • Experience delivering digital services within multidisciplinary product teams.

  • Knowledge of accessibility and inclusive design, including WCAG standards.

  • Experience coaching and developing engineers and building collaborative team cultures.

  • Experience delivering services in complex organisations or programmes with multiple stakeholders and competing priorities.

Candidates MUST be willing to travel to client site anywhere in the UK when needed and MUST have the right to work in the UK.

Hybrid working model available.

Additional Information

Benefits include

  • Hybrid working model with some flexible working
  • 25 days’ holiday
  • Private Medical Health care via Vitality
  • Pension contribution, Life Assurance
  • Professional certifications supported as part of learning and development.
  • A range of retail discounts to enhance your lifestyle, encompassing restaurants, supermarkets, travel, leisure activities and health and well-being services.
  • Access to our Employee Resource Groups, our groups represent diverse backgrounds and provide a platform for colleagues to connect, learn, and support one another.

Company information

At Netcompany, we pride ourselves on our entrepreneurial spirit and our capacity for doing things differently. Our culture is built on fostering low bureaucracy, emphasizing high agility and promoting flexibility, enabling everyone to contribute their best.

Our journey began in the UK with the acquisition of Hunter Macdonald in 2017. As one of Northern Europe’s most accomplished IT companies, we have expanded our headcount globally to 7400+ employees and have offices in UK, Denmark, Norway, Poland, Holland and Vietnam.

We are a Disability Confident Employer and are committed to creating an inclusive and diverse environment that celebrates every individual. Our recruitment processes are based on individual merit. If you require any reasonable adjustments or additional support during the interview process, please email us at [email protected] for assistance.

#LI-RS1

Read the full description
Engineer Site Reliability Engineer, Tech Lead at Loadsmart

Site Reliability Engineer Tech Lead designs and operates critical infrastructure systems, drives reliability projects across engineering teams, and ensures platform performance and SLAs.

Lead Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

ARE YOU INTERESTED IN JOINING AN INNOVATIVE LOGISTICS TECHNOLOGY COMPANY?

Loadsmart is a growth-stage technology company valued at over $1 billion (a true Tech Unicorn)!

We are a collection of industry veterans and user-centered engineers using innovative technology to fearlessly reinvent the future of freight by helping shippers, brokers, warehouses and carriers to move more with less.

With headquarters in Chicago and a globally distributed remote team, Loadsmart continues to attract top talent committed to driving meaningful change. We seek professionals who embody our core values: curiosity, clarity, results, commitment, and teamwork.

In the SRE, Tech Lead role you will build and maintain the company’s internal platform, driving operational excellence and empowering the entire engineering team. You should have experience in analyzing, proposing, and implementing safer systems and processes. Collaborating closely with engineering squads across platform engineering, you will ensure our applications are both safe and reliable. As a hands-on leader, you will stay actively involved in technical work while collaborating closely with internal stakeholders and engineering squads across the organization to ensure our applications are both safe and reliable.

DEPARTMENT: Engineering

LOCATION: Anywhere in Brazil - Remote

WHAT YOU GET TO DO

  • Collaborate with and support our creative, tight-knit development team.
  • Design, deploy, and operate Loadsmart’s critical systems while balancing reliability, cost, and agility.
  • Play a key role in driving reliability projects with engineering teams.
  • Utilize your intuitive problem-solving skills and contagious positive attitude to tackle challenging and exciting issues, inspiring those around you.
  • Collect metrics and understand their business impact, encouraging the team to do the same.
  • Perform troubleshooting and root-cause analysis of system operation issues.
  • Be accountable for the platform’s Service Level Agreements and Objectives.
  • Provide infrastructure support during off-hours as needed
  • Take ownership of software infrastructure projects
  • Seek, give, and receive constructive feedback through code and specification reviews.
  • Familiarity with AI agents and agentic workflows, applying AI across the SDLC (AI-assisted coding), LLMs, MCP servers/gateways, and how emerging AI tooling can improve reliability and operations is a plus.

REQUIRED QUALIFICATIONS:

  • 1-3 years leading Reliability Work across multiple engineering squads
  • Over 5 years of experience in Cloud Computing, SRE/DevOps
  • Proven experience collaborating with internal stakeholders across multiple engineering squads
  • Strong project management skills with a demonstrated ability to delegate and mentor team members
  • Proficient in English communication (both written and spoken) to collaborate in an international team with native and non-native English speakers
  • Detail-oriented with high initiative and self-motivation
  • Strong understanding of software engineering principles and how systems work under the hood
  • In-depth knowledge of modern networking and operating systems
  • Proficiency in AWS, cloud environments, containers, Kubernetes, Docker, and DevOps engineering, including managing tests and CI/CD pipelines
  • Familiarity with automation tools and provisioners like Terraform, Ansible, or Chef
  • Solid troubleshooting and system engineering experience in UNIX/Linux production environments
  • Experience with monitoring, alerting, and incident management
  • Proficiency in automating tasks with scripting languages like Python, Bash, etc
  • Experience or exposure to PostgreSQL and DBA responsibilities is a plus
  • Fluent in English (both written and spoken); comfortable interacting with native English speakers daily.

WORKING AT LOADSMART:

• Competitive base salaries - we believe in rewarding top talent

• Extremely competitive Equity package - become a shareholder in our company!

• Loadie Time Off - PTO and sick days without a limit

At Loadsmart, we believe our biggest asset is our people. We are proud to be an equal opportunity employer, hiring and developing individuals from diverse backgrounds and experiences to add to our collaborative culture. Loadsmart treats all candidates and employees with respect and does not discriminate in our recruiting, hiring, and promoting processes, including on the basis of race, color, religion, sex, age, sexual orientation, gender identity and/or expression, national origin, veteran status, or disability.

It is the policy of Loadsmart that all offers of employment made shall be contingent upon successful completion of electronic background check(s). These checks will be job-related, consistent with business necessity and conducted by our vendor, pursuant to all applicable laws, rules, policies and procedures of our candidates’ specific locale.

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

Read the full description
Engineer Technical Lead - GPU Infrastructure at Tether.io

Technical Lead oversees GPU infrastructure and Kubernetes-based compute platform development, managing containerization, inference endpoints, and platform observability for Tether Data's AI services.

Lead Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Description

Join Tether and Shape the Future of Digital Finance

At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.

Innovate with Tether

Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.

But that’s just the beginning:

Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.

Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.

Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.

Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.

Why Join Us?

Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.

If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.

Are you ready to be part of the future?

About the job

Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.

The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.

This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.

Responsibilities

Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.

Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.

Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.

Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.

Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.

Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.

Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.

Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.

Hiring. Complete the platform team and set the technical bar for the engineers who join it.

Requirements

Must have

  • Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.

  • Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.

  • GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.

  • High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.

  • Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.

  • Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.

  • HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.

  • Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.

  • Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.

  • A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.

  • Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.

  • Excellent written and spoken English. Most partner and leadership work happens in writing.

  • Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.

Desirable

  • Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).

  • Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.

  • VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).

  • Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.

  • Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.

  • Peer-to-peer or distributed-systems background.

  • Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.

Important information for candidates

Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:

  • Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/

  • Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.

  • Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.

  • Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io

  • We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.

When in doubt, feel free to reach out through our official website.

Read the full description
Engineer Technical Lead - GPU Infrastructure at Tether.io

Technical Lead oversees GPU infrastructure and Kubernetes-based compute platform development, managing distributed systems and platform observability for AI inference services.

Lead Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Description

Join Tether and Shape the Future of Digital Finance

At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.

Innovate with Tether

Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.

But that’s just the beginning:

Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.

Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.

Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.

Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.

Why Join Us?

Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.

If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.

Are you ready to be part of the future?

About the job

Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.

The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.

This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.

Responsibilities

Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.

Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.

Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.

Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.

Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.

Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.

Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.

Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.

Hiring. Complete the platform team and set the technical bar for the engineers who join it.

Requirements

Must have

  • Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.

  • Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.

  • GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.

  • High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.

  • Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.

  • Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.

  • HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.

  • Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.

  • Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.

  • A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.

  • Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.

  • Excellent written and spoken English. Most partner and leadership work happens in writing.

  • Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.

Desirable

  • Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).

  • Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.

  • VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).

  • Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.

  • Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.

  • Peer-to-peer or distributed-systems background.

  • Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.

Important information for candidates

Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:

  • Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/

  • Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.

  • Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.

  • Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io

  • We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.

When in doubt, feel free to reach out through our official website.

Read the full description
Engineer Technical Lead - GPU Infrastructure at Tether.io

Technical Lead manages GPU infrastructure and Kubernetes platform for distributed compute and inference services at Tether Data.

Lead Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Description

Join Tether and Shape the Future of Digital Finance

At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.

Innovate with Tether

Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.

But that’s just the beginning:

Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.

Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.

Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.

Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.

Why Join Us?

Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.

If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.

Are you ready to be part of the future?

About the job

Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.

The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.

This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.

Responsibilities

Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.

Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.

Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.

Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.

Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.

Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.

Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.

Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.

Hiring. Complete the platform team and set the technical bar for the engineers who join it.

Requirements

Must have

  • Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.

  • Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.

  • GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.

  • High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.

  • Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.

  • Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.

  • HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.

  • Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.

  • Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.

  • A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.

  • Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.

  • Excellent written and spoken English. Most partner and leadership work happens in writing.

  • Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.

Desirable

  • Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).

  • Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.

  • VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).

  • Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.

  • Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.

  • Peer-to-peer or distributed-systems background.

  • Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.

Important information for candidates

Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:

  • Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/

  • Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.

  • Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.

  • Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io

  • We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.

When in doubt, feel free to reach out through our official website.

Read the full description
Engineer Staff Software Engineer (Platform Engineering)

Leads platform engineering technical direction and strategy as part of the engineering leadership team.

Lead Posted about 24 hours ago Himalayas
What this role involves
You will be part of the Engineering leadership team at ServiceTitan responsible for the technical direction of our product.
Read the full description
Engineer Staff Software Engineer, AI-Native Systems

Designs and builds backend infrastructure and systems for an AI-native virtual care platform as a technical lead.

Lead Remote Posted 1 day ago Jobicy AI
What this role involves
Staff Software Engineer, AI-Native Systems (Tech Lead) Location: Remote (US) We are Virtual care only works if the infrastructure behind it does. Wheel builds that infrastructure — the systems that...
Read the full description
Engineer Director of Engineering, Leverage

Directs engineering strategy, team, and technical execution for fleet management software platform.

Lead Posted 1 day ago Jobicy AI
What this role involves
A little about us…Fleetio is a modern software platform that helps thousands of organizations worldwide manage their fleet operations. Transportation technology is a hot market, and we’re leading the charge...
Read the full description
Engineer Senior DevOps Lead

Leads DevOps infrastructure, CI/CD pipelines, and cloud operations for an enterprise AI platform.

Lead Remote Posted 1 day ago Jobicy AI
What this role involves
Location: Berlin, Germany Remote Status: Fully Remote (#LI-Remote) LivePerson (NASDAQ: LPSN) is a leader in trusted enterprise conversational AI and digital transformation. The world’s leading brands use our award-winning Conversational...
Read the full description