Senior Software Engineer · Tech Lead · Tel Aviv

I make backend systems fast, cheap and observable, then build the AI on top.

Nine years shipping software, the last several owning highly-available AWS services at 300+ QPS. I find the bottleneck nobody wants to touch, fix it for good, and leave the team with the tooling to keep it fixed. Today I'm freelancing on agents, on-prem LLMs and full-stack products, and open to the right full-time team.

Available now Tel Aviv, Israel Backend · Distributed systems · AI platforms
300+
QPS on distributed AWS services I engineered and ran
10×
latency improvement from end-to-end lifecycle work
−40%
compute and storage cost
70% → stable
websocket error rate after a 2.5-year problem
6 days
of developer time saved every two weeks

Selected work

Five problems from my time at Joyned, each with the situation, the call I made, and what changed.

Cost & performance
10× fasterand 60%+ less usage on the most expensive services

The AWS bill doubled while traffic stayed flat

Situation
Six months of doubling cost with no growth in scale. Browser analytics events sat in a queue for ~1.5s, and the first request that decides where the platform opens took 2+ seconds for every visitor, including large partners. The DevOps vendor was stalled.
What I did
Audited bills and usage, found over-provisioned DynamoDB and inefficient access patterns, and removed an expensive Lambda by integrating API Gateway directly with SQS. Re-architected the frontend event pipeline. When the vendor stalled, I submitted the infra PRs myself and redesigned the deploy flow for zero downtime.
Result
Client events dropped from 1.5s to about 200ms, effectively invisible to users. Cost and latency both fell sharply, the platform loaded faster, and juniors got a documented playbook for cost/performance trade-offs.
AWSDynamoDBLambdaAPI GatewaySQSCloudWatch
Reliability
70% → stableerror rate on real-time infrastructure

Fixing a websocket layer nobody trusted for 2.5 years

Situation
Connections broke on bad tokens and failed reconnects, with error rates reaching 70%. A one-session-per-user limit also blocked notifications from any session a user wasn't currently in, which stalled other features.
What I did
Quantified the failures across services with CloudWatch and dashboards to prove the problem was systemic. Then I decoupled session management from socket control, built automatic connection lifecycle management for frontend engineers, added exponential backoff, and refactored several projects onto it. I pushed for the redesign over patching, and brought the data to the VP R&D.
Result
Live sessions and invites became reliable, developer trust and feature velocity went up, and the team had a foundation for future real-time features, with monitoring hooks, docs and training so it stays healthy.
WebSocketsArchitectureRoot-cause analysisCloudWatch
Observability
Reactive → proactiveone logging standard, front to back

From "customers told us" to "we saw it first"

Situation
Monitoring only alerted on hard failures. Everything else meant digging through an analytics tool built for management dashboards. Customers found bugs before we did.
What I did
Pitched it to the VP R&D and senior engineers as an investment in reliability and velocity, not a bug fix. Built a shared backend package and a logging pipeline with configurable levels into Datadog and CloudWatch, wrote the standards for what to log and how, and built dashboards for both developers and management. Built the infrastructure first so others could unblock a production-only bug.
Result
Bugs and anomalies caught before customers noticed. The standard spread across frontend, backend and multiple repos, engineers adopted it in their own features, and leadership got visibility into feature and system health.
DatadogCloudWatchStandardsMentoring
AI platform
Agents for everyonean orchestration layer devs and non-R&D teams build on

Nobody owned AI agents, so I built the layer

Situation
The company needed a way to build and experiment with AI agents, and no one owned it.
What I did
Shipped the orchestration framework early and grew it on demand, opening it to engineers and to people outside R&D. Alongside it, an AI CI/CD pipeline. Fixed a deployment-size problem with lazy loading, cutting LangChain imports to zero before each execution.
Result
Faster engineering velocity and smoother product integration for internal and external clients. The same instincts (observability, guardrails, developer experience) now drive my freelance agent and on-prem LLM work.
AgentsLangChainBedrockCI/CD
Process & judgment
6 dev-dayssaved every two weeks, no automation needed

Fixing the process before paying to automate it

Situation
A recurring feature-delivery workflow lost nearly a sprint every cycle. Management's answer was a costly automated pipeline.
What I did
Audited requirements, pain points and blockers across dev and workflow teams, then proposed a streamlined step-by-step process with a reversible rollout plan. Introduced tooling where it helped and set clear ownership.
Result
Six developer-days saved every two weeks, better morale, and a clean base to automate later, on the steps that still deserve it.
Process designStakeholdersPragmatism

Experience

Startups to scale-ups, gaming to travel-tech, always close to the code.

Toolbox

What I reach for.

Languages

PythonTypeScriptJavaScriptC#Java

Cloud & data

AWSDynamoDBLambdaSQSAuroraRedshiftBedrockAzurePostgreSQL

Platform

Node.jsDockerZeroMQOpenCVDatadogCloudWatch

How I work

Tech leadershipSystem designMentoringDesign & code reviewSpec-driven AI development

Let's build something that holds up under load.

I'm open to senior backend, tech-lead and AI platform roles in Tel Aviv or remote, and to freelance work. The fastest way to reach me is email.