Part-Time Contract | Remote | Approximately 20 Hours per Week
About the Role
A mission-critical SaaS platform used by approximately 800 flooring businesses to manage leads, estimates, product selection, inventory, scheduling, installations, invoicing, and payments is being modernized from a legacy .NET/ASP, SQL Server, stored-procedure-heavy environment to React, Node.js, TypeScript, native PostgreSQL, REST APIs, and AWS. Much of the initial conversion has been completed using AI-assisted development. The immediate priority is to clean up, test, validate, and safely release the converted system while existing customers remain active.
We are looking for a hands-on Senior DevOps/Platform Engineer who can strengthen the delivery platform, observability, security, and production safeguards required for frequent, low-risk releases.
Responsibilities
- Assess the current AWS infrastructure, CI/CD pipelines, environments, deployment process, and production risks.
- Improve automated build, test, security-scanning, deployment, and rollback workflows.
- Support small, frequent releases through feature flags, staged rollouts, health checks, and automated validation.
- Establish safe coexistence between legacy and modern services during the migration.
- Build and maintain infrastructure as code for repeatable, auditable environment provisioning.
- Strengthen monitoring, centralized logging, distributed tracing, alerting, dashboards, and incident-response procedures.
- Define service-level indicators for availability, latency, error rates, database health, queues, and critical business workflows.
- Help diagnose performance and stability issues involving AWS, Node.js services, APIs, PostgreSQL, Babelfish, and legacy .NET components.
- Support the transition from Babelfish and SQL Server-oriented behavior to native PostgreSQL.
- Review capacity, connection-pool, timeout, retry, queue, and load-shedding configurations.
- Improve resilience for asynchronous and event-driven workloads using services such as Lambda, SQS, EventBridge, API Gateway, Redis, and dead-letter queues.
- Implement secure secrets management, access controls, audit logging, vulnerability management, backup, and recovery procedures.
- Verify that rollback remains safe after application or database writes, including backward-compatible schema changes.
- Review AI-generated infrastructure and deployment changes for correctness, security, maintainability, and operational risk.
- Document deployment standards, incident procedures, environment configuration, and operational ownership.
- Partner closely with engineering and product teams, with meaningful overlap during US working hours.
Required Qualifications
- At least 7 years of professional DevOps, SRE, platform engineering, or cloud infrastructure experience.
- Strong hands-on experience operating production workloads in AWS.
- Proven ownership of CI/CD architecture, automated deployments, monitoring, incident response, and rollback.
- Experience supporting production systems built with Node.js, React, PostgreSQL, and .NET.
- Strong understanding of:
- Infrastructure as code
- Networking, DNS, load balancing, and API gateways
- Containers and serverless workloads
- Logs, metrics, traces, alerts, and production dashboards
- Secrets management and least-privilege access
- Backups, disaster recovery, and production readiness
- Database migrations and backward-compatible schema changes
- Connection pooling, capacity planning, retries, idempotency, and failure recovery
- Experience introducing staged deployments, feature flags, canary releases, or blue-green deployments.
- Ability to troubleshoot across application, infrastructure, network, and database layers.
- Experience working on incremental modernization where legacy and modern systems must run in parallel.
- Strong written and spoken English with the ability to communicate risk and recommendations clearly.
- Ability to work independently and deliver senior-level ownership within a part-time schedule.
Preferred Qualifications
- Experience with AWS Lambda, SQS, EventBridge, API Gateway, Redis, CloudWatch, and container-based AWS services.
- Experience migrating SQL Server workloads to PostgreSQL, particularly through Babelfish.
- Experience with Terraform, CloudFormation, or similar infrastructure-as-code tooling.
- Experience operating workflow-heavy SaaS involving inventory, scheduling, invoicing, payments, or other production-critical transactions.
- Familiarity with OpenTelemetry, distributed tracing, and structured logging.
- Experience reviewing infrastructure or deployment code generated through Claude Code, Codex, Cursor, or similar AI tools.
- Experience improving engineering platforms for teams releasing daily or near-daily.
What Success Looks Like
During the initial engagement, this person will:
- Produce a clear assessment of the current infrastructure, deployment risks, and release blockers.
- Establish reliable CI/CD and environment-management standards.
- Improve monitoring across the core quote-to-invoice workflow.
- Make deployments smaller, repeatable, observable, and safely reversible.
- Strengthen production readiness for the AI-converted React, Node.js, and PostgreSQL platform.
- Reduce instability associated with Babelfish, database connections, asynchronous workloads, and legacy-modern coexistence.
- Create practical operational documentation the engineering team can maintain after the engagement.
Engagement Details
- Remote, part-time contract
- Approximately 20 hours per week
- Regular overlap with the US-based team required
- Initial focus on release readiness and platform stabilization, followed by modernization, scalability, and support for AI-enabled product capabilities