juliakafarska

joined 3 months ago
 

Originally published at https://blog.light-cloud.com/industry/cloud-pricing-is-a-business-model on the Light Cloud Blog.

By James, Light Cloud

Summary: Egress fees survived for over a decade and died within nine weeks of EU regulation taking hold. Prices that behave like that aren't cost recovery. Cloud pricing complexity is a business model, and it's working.

Cloud Pricing Complexity Is a Business Model, Not an Accident

On July 23, 2021, Cloudflare's CEO co-signed a blog post with a spreadsheet in it. AWS's Egregious Egress estimated that AWS was charging US and Canadian customers roughly 80 times its own bandwidth cost, with markups in some regions approaching 8,000%. AWS never published the rebuttal that would have settled it, the one with its actual costs in a table. It didn't need to. Egress pricing wasn't an error anyone had to defend; it was a mechanism doing its job.

So here's the thesis, stated flat: cloud pricing complexity is a business model. The bill nobody can predict and the bandwidth charge that only applies in one direction aren't accidents of a complicated product. They're structural incentives, built by smart people, working as intended. Treating them as accidents is how you lose to them.

The one-way valve

Data transfer into a hyperscaler is free. Transfer out is metered. A pipe that's free inbound and expensive outbound isn't a pipe, it's a valve, and the valve has a compounding property: every gigabyte you store raises the cost of ever leaving. Egress was never priced as bandwidth. It's an exit tariff that grows with your own success.

Watch what happened when the tariff met regulation. Within nine weeks in early 2024, Google (January 11), AWS (March 5), and Microsoft (March 13) all announced free data transfer out for customers leaving their platforms entirely, each framed as generosity rather than compliance. The trigger was the EU Data Act, which caps switching charges at direct cost today and bans them outright on January 12, 2027. Costs didn't change that quarter. Law did. A price that survives well over a decade and dies only when regulators kill it was strategy all along.

And read the fine print: only the final exit got free. The meter on ordinary operational egress, the kind you pay every month while staying, runs exactly as before.

A discount you need a trading desk for

On-demand is the sticker price almost nobody pays at scale. The real prices sit behind reserved instances and savings plans: one-year and three-year commitments, partially or fully prepaid, with discounts deep enough that refusing them looks irresponsible in any budget review. That's a forward contract on compute. AWS even operates a Reserved Instance Marketplace where customers resell commitments they mis-sized, and the moment your discount program needs a secondary market, it has stopped being a discount program.

An entire profession formed around reading the bill. The FinOps Foundation runs certifications and an annual conference for the discipline of understanding what you're paying for. Sit with that. Your office electricity bill doesn't support a career track.

The commitment machine's subtlest effect is political: once three years of spend are prepaid, your CFO becomes the cloud's advocate inside your own company, because every workload that might move somewhere better is now a workload that strands committed dollars. Lock-in enforced by your own finance department costs the vendor nothing. The discount is the moat.

Complexity as a sorting machine

Every dimension a bill gains makes it harder to predict and easier to segment. Per-request pricing here, per-GB-month there, a separate rate for traffic that crosses an availability zone boundary, and an hourly charge for the NAT gateway somebody enabled years ago that everyone since has been afraid to touch. Classic price discrimination requires the seller to learn each buyer's willingness to pay. A sufficiently complex price sheet does it automatically: buyers with FinOps staff excavate the discounts, and everyone else pays sticker.

The $4,847 bill that started this company, for an app serving 10,000 requests a day, wasn't a malfunction. It was the machine sorting. Light Cloud exists because that bill landed in the tier reserved for people who don't employ someone to fight it. The pricing pages are all public, and still nobody can compute their own invoice from them; both facts are stable, and the second one is the point.

The honest defense

The strongest counterargument deserves stating properly: cloud pricing is complex because the underlying costs are complex. Hardware differs by region, power costs swing with local grids, and metering usage at fine granularity is exactly the thing that lets a hobby project cost pennies a month instead of a license fee. Usage-based pricing is also fairer than the per-core licensing world it replaced, where you paid the same whether the server idled or burned. All of that is true, and I'll concede more: nobody designed this in one villainous meeting. It accreted, one reasonable-looking SKU at a time.

But complex costs don't force complex prices. Utilities face brutally variable input costs and still sell flat tariffs, because absorbing pricing risk is part of the product a utility sells, and hyperscaler margins leave plenty of room to hedge internally. The tell is where the complexity clusters: thick around the places where switching decisions happen, like egress and long-term commitments, and thin where they don't. Accidental complexity would be evenly distributed. This isn't.

This is why Light Cloud prices the way it does. Usage-based, scaling to zero when nothing runs, with a pricing page meant to be computable by one person with a napkin; if you can't predict your bill from it, we treat that as our bug rather than your homework.

A two-person company keeping pricing simple proves little, and I know it; simplicity is cheap when you're small. The commitment is about direction. Complexity added to a pricing page becomes revenue somebody defends later, so the discipline is refusing the first SKU we can't explain in one sentence.

Next time a cloud bill surprises you, don't ask what you did wrong. Ask which mechanism found you: the valve, the commitment, or the sorter. Then ask the sharper question: which one is your architecture feeding right now?

Related reading


This article first appeared on the Light Cloud Blog: https://blog.light-cloud.com/industry/cloud-pricing-is-a-business-model.

 

Originally published at https://blog.light-cloud.com/multi-cloud/multi-cloud-overrated-portability-isnt on the Light Cloud Blog.

By James, Light Cloud

Summary: Three outages inside a month did not prove you need to run production on two clouds. They proved you need a credible, tested ability to leave one. Those are different products with wildly different prices.

Multi-Cloud Is Overrated. Portability Isn't.

The internet broke three times inside a month. On October 20, 2025, a latent race condition in DynamoDB's DNS automation took down AWS us-east-1 and dragged a large slice of the consumer internet with it. Nine days later, an errant configuration change broke Azure Front Door, taking Microsoft 365 and the Azure portal along. Then on November 18, a database permissions change at Cloudflare made a Bot Management feature file double in size, and a long list of dependents went dark for hours.

Each time, the same advice arrived within the hour: this is why you need multi-cloud.

We sell software that models three clouds, so believe me when I say we'd love that advice to be right. It mostly isn't. Active-active multi-cloud is complexity theater for the majority of companies that attempt it, and what last autumn's outage cluster actually argues for is portability: a credible, tested ability to leave. Those are different products. They carry wildly different prices.

What active-active actually costs

Running production on two clouds at once means engineering for the intersection of their feature sets. You give up the managed services that made either cloud attractive, because DynamoDB doesn't run on Azure, and build against the lowest common denominator instead. Then everything doubles. Two IAM models with different permission semantics, two networking stacks, two sets of quotas and failure modes, and an observability layer that has to stitch it all into one picture. Your on-call now debugs two providers instead of one.

There's a people bill too. Every engineer you hire needs fluency in two permission models and two failure vocabularies, or you split the team into provider silos and reinvent the coordination problems you bought a cloud to escape. Hiring gets slower. The pager gets worse.

Data is where the theater collapses. Compute is fairly portable. State has mass. Keeping your primary database live in two clouds means continuous cross-cloud replication, and the EU Data Act's ban on switching charges, fully in force on January 12, 2027, won't help you there, because it covers leaving a provider while the meter on ordinary operational egress keeps running every month you stay. And a standby you've never failed over to isn't a disaster plan. It's a hope with a budget line.

October 29 held a quieter lesson: the Azure portal itself was impaired during the Azure outage. If the failover runbook lives in the cloud that's down, the multi-cloud strategy has a single point of failure with a wiki URL.

Portability is an option you hold

In finance terms, portability is an option: you pay a small ongoing premium for the right, never the obligation, to move. You rarely exercise it. Its value shows up anyway. At contract renewal, a vendor who knows you can leave prices differently than one who knows you can't. During incidents, a restore you've tested elsewhere turns catastrophe into degradation. And in front of regulators it's now a required artifact: DORA, applying to EU financial entities since January 17, 2025, requires documented exit strategies for critical ICT providers, and auditors have started asking to see them.

Most companies can't produce a real one. A credible exit plan is an inventory of what you actually run, a mapping of each piece to a provider-neutral equivalent, a data restore executed at least once somewhere else, and an honest time estimate. If assembling that takes your team a quarter, you don't hold the option. You hold a slideshow about the option.

Holding the option cheaply is an architecture decision, and mostly a boring one: containers over proprietary runtimes where the managed premium isn't earning its keep, and plain PostgreSQL over a proprietary database API unless you've priced the divorce. Add a written map of what runs where, kept current, because inventory is the part that rots fastest. None of this requires a second cloud. It requires saying no to a few conveniences whose real price is the exit.

The case for the second cloud

Here's the steelman, fairly: selective redundancy saved real companies last autumn. A status page hosted on a different provider, DNS with a second resolver, a read-only mode served from another CDN: cheap, and effective on October 20. Some businesses justify full active-active, trading venues among them, where minutes of downtime cost more than years of duplicated infrastructure. And plenty of enterprises are multi-cloud whether they chose it or not, because acquisitions arrive carrying their own stacks.

All conceded. Notice, though, what the wins have in common: thin, stateless layers, chosen deliberately, tested regularly, and priced honestly, which is portability practiced at the edges rather than a second copy of production. Duplicate the cheap layers where failure is loud. Hold the option on everything else.

This distinction is what ICE is built around. Its deterministic graph models your infrastructure across AWS, GCP, and Azure as one structure, which makes the exit-plan artifact, what runs where and what maps to what, a byproduct of normal operation instead of a quarterly archaeology project. Portability as a property you hold rather than a second production you fund.

I'll admit regulation is a tailwind we didn't earn: DORA made the deliverable mandatory while the tooling to produce it barely exists. Small companies get few gifts from Brussels. We plan to use this one.

The next request for your exit plan might come from a regulator, an insurer, or your board after the next headline outage. Could you hand over a tested document by Friday, or would you be scheduling a meeting to plan the plan?

Related reading


This article first appeared on the Light Cloud Blog: https://blog.light-cloud.com/multi-cloud/multi-cloud-overrated-portability-isnt.

 

Originally published at https://blog.light-cloud.com/industry/everyone-builds-the-same-platform-badly on the Light Cloud Blog.

By James, Light Cloud

Summary: Gartner said 80% of large engineering orgs would have platform teams by 2026. They do. And nearly all of them are assembling the same portal from the same parts, with no customers who can say no.

Every Company Is Building the Same Internal Platform, Badly

Gartner predicted that by 2026, 80% of large software engineering organizations would run platform engineering teams, up from 45% in 2022. It's 2026, and from where I sit the prediction landed. Ask around: there's a platform team in your company, and it's building a service catalog, golden-path templates, pipeline glue, a secrets story, per-branch preview environments, and a cost dashboard. So is the platform team at the company across the street. Same spec, same parts, no shared code.

My claim is blunt: nearly every company is building the same internal platform, most are building it badly, and the root cause is a category error. The platform gets treated as a headcount line when it's actually a product, and products built without customers who can say no come out bad.

The same portal, hand-rolled

Spotify open-sourced Backstage in 2020 and donated it to the CNCF, where it has since been adopted by more than 3,400 companies. That number is the tell. Thousands of organizations looked at their internal tooling problem and picked the identical answer, which confirms the problem is common rather than special. But Backstage is a framework, a very good one, and a framework is homework. You staff engineers to assemble your portal from it: writing catalog-info.yaml files for every service, carrying plugins through API churn, wiring your auth, and chasing a catalog that drifts stale the week after launch. The industry's response to "everyone builds the same thing" was a kit for building the same thing.

The category even has its own conference circuit. PlatformCon fills its schedule with talks on golden paths and developer experience, and you could swap the company names between most of the slide decks without anyone noticing. An industry event where everyone describes building the same software is a strange kind of proof that it should be software you can buy.

A product with no market pressure

An internal platform has captive customers. Nobody can churn. Without churn there's no signal that onboarding is broken, and with adoption arriving by mandate, nothing ever tests whether a single engineer would choose the thing voluntarily. The roadmap gets set by whichever team escalates loudest. That's a politics engine sitting where a market should be.

The metrics compound it. Platform teams get measured on adoption, so the quarterly goal becomes migrating more teams onto the platform, which rewards mandates and onboarding pushes rather than anything a user would describe as better. A vendor lives on retention instead: the product has to keep being chosen, month after month, by people holding a cancel button. Remove the cancel button and you've removed the feedback.

When the platform disappoints, engineers route around it with a Makefile and a cron job running kubectl apply, and the platform team's answer is usually a policy forbidding the workaround, which is a move no commercial vendor could survive making. Then there's the cost accounting nobody performs. A five-engineer platform team is the most expensive software subscription in the company, except it can't be cancelled at renewal, its price rises with every salary review, and it serves exactly one customer, so the economics never improve with scale. Internal software gets none of the scrutiny we'd apply to a vendor invoice a tenth its size.

When building it is right

Spotify had real scale pain; Backstage began life solving service discovery for an organization drowning in microservices, and at that size a dedicated platform is obviously justified. Netflix has talked for years about its Paved Road, and for them the platform genuinely is strategy. Regulated industries need integrations no vendor ships. And a small platform team that mostly glues managed services together, rather than rebuilding them, tends to pay for itself. Timing matters as well: a platform extracted from a product you've already scaled encodes real lessons, while a platform built ahead of scale encodes guesses.

I'll concede one more thing: buying has its own failure mode, the vendor platform so generic it fits nobody and so sticky nobody can admit the purchase failed. Skepticism toward platforms-as-products is earned. The test is differentiation. Write your platform team's current quarter on a whiteboard, then write what you'd guess any other company's platform team is doing this quarter, and if the two lists match, differentiated headcount is being spent on undifferentiated software that vendors and open-source projects are already competing to commoditize. Build what makes your company strange. Buy what makes it the same as everyone else.

Light Cloud is our bet on the product version of this argument. The things platform teams keep hand-rolling are what we sell as boring commodities: deploys from GitHub with a preview environment on every branch, and databases without a database subteam, priced to scale to zero when idle. I'm aware this post argues our own book. Discount for that, and the pattern still stands.

The argument predates the company, though. Platform work expands to fill the headcount allocated to it, because there's always one more integration, and no customer exists to say the next one isn't worth paying for.

So run the audit. List what your platform team shipped last quarter, then search GitHub and a vendor directory for each item, and count how many entries nobody else has built. The count won't be high, and the honest version of the exercise also counts the maintenance hours, because a portal is never shipped, only kept alive. What would those engineers build if the platform were already built?

Related reading


This article first appeared on the Light Cloud Blog: https://blog.light-cloud.com/industry/everyone-builds-the-same-platform-badly.

 

Originally published at https://blog.light-cloud.com/dev-tools/anatomy-of-a-3-cent-preview-environment on the Light Cloud Blog.

By James, Light Cloud

Summary: A full, disposable copy of your app for every branch sounds expensive until you see the meter. Here's the arithmetic that gets a preview environment to about three cents, and the engineering that keeps it there.

Anatomy of a 3-Cent Preview Environment

Around three cents. That's the ballpark lifetime cost of the thing that most changed how code review feels on Light Cloud: a full, disposable copy of the application, on its own URL, for every branch. Not a screenshot bot, and not a staging server you book in a spreadsheet. A real environment, created on push, destroyed on merge.

Here's the claim: preview environments stopped being a luxury the moment scale-to-zero billing became real, and the remaining cost isn't infrastructure. It's engineering discipline, and most of that discipline is teardown. This post shows the arithmetic and the machinery, including the parts that bite, because "cheap previews" sounds like marketing until you can check the meter yourself.

Where "around three cents" comes from

A preview on our platform is a container that scales to zero. It bills nothing while idle, and idle is most of its life. The representative preview below assumes fifteen minutes of cumulative active traffic across its entire existence, one vCPU, 512 MiB of memory, and public Cloud Run list prices in a standard region as of August 2026.

Line item List rate This preview uses Cost
CPU $0.000024 / vCPU-second 900 vCPU-seconds $0.0216
Memory $0.0000025 / GiB-second 450 GiB-seconds $0.0011
Requests $0.40 / million ~10,000 requests $0.0040
Total ~$0.027

Call it three cents. The free tier (the first 180,000 vCPU-seconds each month cost nothing) would shrink it further, and a build minute plus image storage adds a little back; the order of magnitude doesn't move. Fifteen active minutes is generous, by the way. In my experience a preview gets opened about twice: once by the author confirming it deployed, once by the reviewer forming an opinion.

That observation carries the whole economic argument. An environment billed by the second costs what it's actually used, and a preview is barely used, so a preview is barely billed. A staging server priced by the month inverts this, charging you for every hour nobody is looking at it, which is nearly all of them. Same compute. Different pricing physics.

The table also hides a floor worth naming: with request-based billing the instance is only allocated while serving traffic, so a preview's eleventh day of existence costs exactly what its first did, which is zero if nobody visits. Longevity is free. Only attention is billed.

The lifecycle, and where it bites

open PR         -> build image -> deploy revision -> post URL on the PR
push new commit -> rebuild     -> replace revision (same URL)
no traffic      -> scale to zero
merge or close  -> destroy revision, DNS record, preview database

The first three lines are the demo. The fourth line is the product. Provisioning is a mostly solved problem with good primitives underneath it; teardown is where preview systems rot, because teardown's failure mode is silent. A missed webhook, a PR closed mid-deploy, a force-pushed branch rename: each one orphans an environment that sits invisible until someone finds it months later, still holding a database. So we treat reconciliation as the core loop, continuously comparing what should exist (open PRs) against what does exist (running previews), instead of trusting that every event arrived exactly once. Event-driven teardown is a rumor. Reconciliation is a fact.

Routing has a trap of its own. Every preview needs a working URL moments after push, and issuing a fresh TLS certificate per preview walks straight into Let's Encrypt rate limits on a busy repository, so previews live under a wildcard certificate and a wildcard DNS record provisioned once. One early decision deletes an entire class of flaky waiting.

Isolation is the second bite. A preview executes branch code, and branch code hasn't passed review yet, so a preview that receives production secrets is a phishing kit you built for yourself. Preview environments get their own scoped credentials and their own data, never production's, and that rule costs us real convenience, because "test it against real data" is the most requested thing previews can't safely be.

Databases are the third bite, and the honest one. Stateless containers scale to zero gracefully; PostgreSQL does not, and the trade-off triangle is real: a schema-only database is cheap and safe but empty, seeded fixtures are useful but drift from reality, and a copy of production data is realistic and a compliance incident waiting for a fork. Our default sits at the cheap, safe end of that triangle, and one rule is absolute: production data never crosses into a preview.

"That's not staging"

The strongest objection deserves its space: a preview with a seeded database won't catch the bug that only appears at production data shapes, won't carry a load test, and mocks the third-party integrations that break in the interesting ways. All true. A team that replaces capacity planning with preview environments will meet reality on a bad day.

But that objection compares instruments doing different jobs. Staging answers "does this survive production conditions"; a preview answers "is this the change we think it is", and it answers inside the review conversation, while opinions are still cheap to change. The expensive failure previews prevent isn't an outage. It's the LGTM that nobody actually looked at.

Around three cents buys a second opinion that loads in a browser. Next time someone says disposable environments are too expensive for your team, ask to see their arithmetic next to this table. Then ask what the last staging-slot conflict cost in engineer-hours, and compare columns.

Related reading


This article first appeared on the Light Cloud Blog: https://blog.light-cloud.com/dev-tools/anatomy-of-a-3-cent-preview-environment.

 

Originally published at https://blog.light-cloud.com/devops/graph-not-text-file on the Light Cloud Blog.

By James, Light Cloud

Summary: Facebook figured this out in 2015: config errors were a major source of outages, and their answer was to compile configuration from source and model its dependencies. A decade later, most of us still review the characters instead of the structure.

Why Our Source of Truth Is a Graph, Not a Text File

That's Facebook's engineering team at SOSP 2015, describing a system absorbing thousands of live configuration changes a day. Their fix wasn't stricter YAML review. They compiled configs from high-level source code, expressed configuration dependencies "similar to the include statement in a C++ program", and validated invariants against the result before anything touched a server. At the sharpest end of the problem, a decade ago, the conclusion was already in: stop treating configuration as text you diff. Treat it as a structure you compute.

A year ago Julia wrote a breakup letter to YAML on this blog. That post was the feelings. This one is the argument, and the argument is why ICE's source of truth is a deterministic graph rather than a directory of text files.

Questions a text file can't answer

The test of a source of truth is whether it can answer the questions you ask during an incident. Line one of the table is the question I got burned by. The rest follow.

The question you actually ask What the text file knows
What breaks if I delete this subnet? Which lines contain the subnet's name
What order do these changes need to apply in? Nothing; order is computed at plan time, then discarded
Has reality drifted from what's declared? Nothing; text can't observe the cloud
Is this rename safe? It reads as a delete plus a create; good luck

Terraform itself concedes the point, quietly. Run terraform graph and it prints the DAG (directed acyclic graph) of resources and dependencies it builds internally before every single plan. The graph exists on every run. It's derived from your text, used to order operations, and thrown away, while review happens on the characters. The most load-bearing artifact in the whole pipeline is the one no reviewer ever sees.

The YAML ecosystem keeps rediscovering the gap. Helm templates YAML with text substitution, Kustomize patches YAML with more YAML, and both exist because the format can't express the relationships everyone actually needs, so the industry bolts string machinery onto a tree structure and calls the result configuration management. Tools that exist to work around the source of truth are testimony about the source of truth.

Determinism is the actual feature

A graph as the source of truth changes three things, and none of them are cosmetic. Hold on: two things, stated properly.

First, identity. In a graph, a resource is a node with an identity that survives renaming, so a rename is a rename. In text, the same edit surfaces as destruction plus creation, and every Terraform operator eventually learns terraform state mv the way you learn most things in this field, which is at night. Second, reproducibility. A deterministic model means the same graph produces the same actions, every time, with drift detected by comparing the graph's expectations against observed reality node by node, rather than by diffing two commits of a file that was never looking at the cloud in the first place.

The scale details are worth sitting with: they report a median config size of 1KB with large ones reaching MBs or GBs, hundreds of thousands of configs, and trillions of configuration checks daily. Nobody reviews that by reading characters. Structure was the only way through.

Text won for a reason

Steelman time, and it's a strong one. Text is the only format every engineer, editor, and tool on earth can open. Git gave it merge machinery, blame, and history for free. It's greppable at 3 a.m. It locks you into no vendor. And the failure mode of the alternative is real: an opaque model nobody can inspect is worse than ugly YAML, because at least ugly YAML can be read in a pager on a bad night. GitOps built genuine operational rigor on all of this.

Every bit of that is conceded, and it shapes the design rather than defeating it. The graph serializes to versionable, diffable text, so the audit trail survives; what changes is what the diff says. A text diff reports "+14 -9 lines". A graph diff reports which nodes changed and what depends on them, which is the difference between describing an edit and describing its blast radius. You keep git. You stop asking git to be a model of your infrastructure, because it never was one.

This is the bet ICE makes concrete: one deterministic graph modeling your infrastructure across AWS, GCP, and Azure, with the graph as the thing you operate on and text as one of its projections. Drift stops being a quarterly surprise and becomes a comparison the tool runs continuously.

I don't claim the graph model is finished territory; serialization formats, review UX, and escape hatches for the weird 5% are open problems we work on in the open. What I'll defend flatly is the direction. Structure first, text as output.

Facebook needed this at hundreds of thousands of configs. You'll feel it at fifty, the first time a plan output surprises you, because the tool held a graph that knew the answer and discarded it before showing you a diff of characters. Why is the throwaway the part you review?

Related reading


This article first appeared on the Light Cloud Blog: https://blog.light-cloud.com/devops/graph-not-text-file.

 

Originally published at https://blog.light-cloud.com/dev-tools/why-a-desktop-app-in-2026 on the Light Cloud Blog.

By James, Light Cloud

Summary: The decade that moved email, docs, design, and project management into the browser left the tools developers live in untouched. Every major development environment is a desktop app. ICE is one on purpose.

Why We're Building a Desktop App in 2026

Ask developers where they spend the working day. The 2025 Stack Overflow survey answers with a list, and the list has a property nobody remarks on:

Development environment Share of developers
VS Code 75.9%
Visual Studio 29%
Notepad++ 27.4%
IntelliJ IDEA 27.1%
Vim 24.3%

Every entry is a desktop app. The decade that moved email, docs, design, and project management into the browser left the tools developers live in untouched, and the company that owns the world's largest web platform ships its flagship editor as a local program.

We're building ICE, our Integrated Cloud Environment, as a desktop app, and in 2026 that reads as contrarian. My claim is that it's the opposite: for instruments, the tools a professional plays for hours a day, desktop never lost, and infrastructure tooling is an instrument that got misfiled as a website.

Documents versus instruments

The browser's wins share a shape: the thing being worked on is a shared artifact, and the URL is the artifact. Docs, tickets, wikis. Distribution beats latency for those, because the collaboration is the product.

Instruments have the opposite shape. An editor, a terminal, a debugger: you inhabit them, they own your keyboard, and they touch local resources constantly. A browser tab can't fully own the keyboard; Cmd-W is always one reflex away from closing your session, and the tab boundary keeps your tool at arm's length from the filesystem, the OS keychain, and your running processes. Latency compounds too. A person who triggers an interaction thousands of times a day feels every added millisecond as friction, which is a large part of why 75.9% of the industry works in a local editor while writing software for the cloud.

There's a quieter dependency too: instruments get customized. Keybindings, themes, dotfiles carried between jobs like family recipes. A tool you shape to your hands is a tool you keep, and the browser makes that shaping shallow, because in a tab the state belongs to the site rather than to you.

An operating session is a place, not a page

Cloud consoles are forms. Stateless, per-request, amnesiac: every visit starts from a dashboard that has forgotten your context, and the context is the job. Working on infrastructure means holding a model across an afternoon, and the tab-shaped version of that model evaporates on every reload. Anyone who has rebuilt a seven-tab investigation after one accidental window close knows the tax.

There's a darker version of this argument, and October 2025 supplied it. During Azure's October 29 outage, the Azure portal itself was impaired, meaning the tool for managing the blast radius was inside the blast radius. The local-first movement wrote the principle down years before that outage made it vivid:

Your model of your infrastructure, the map you need most during an outage, shouldn't be hosted inside the thing that's on fire. ICE's graph lives on your machine. The clouds can be down and the map still opens.

Credentials follow the same logic. A browser-based tool that talks to AWS, GCP, and Azure on your behalf generally means your cloud credentials transit somebody's backend, and that somebody becomes part of your attack surface. The design goal of a desktop app is blunter: keys stay in the OS keychain, API calls go from your machine to your clouds, and we never hold what we can't lose.

Figma is the counterargument

The steelman has a name. Figma beat entrenched native design tools from inside a browser tab, and the reasons were real: nothing to install, a URL is the file, multiplayer by default. Distribution through a link is a genuinely superior adoption model, and no update ever ships late to a browser. I'll concede the blurry boundary too. Plenty of "desktop apps" are Electron, a bundled browser in a native coat, so the war looks over if you squint. And building desktop means carrying the update, signing, and packaging tax that the web abolished; we pay that tax and it's not small.

But look at what Figma's users collaborate on: the canvas itself, an artifact safe to share by URL. Infrastructure's object is a live system wearing credentials, where "anyone with the link" is the beginning of an incident report, and where the collaboration layer already exists and is called git. And Electron's popularity argues my side, quietly: when teams could ship a tab, they still chose a desktop shell to get the keyboard, the tray, the keychain, and the filesystem. The UI technology surrendered. The deployment target didn't.

Defending this choice publicly is the point of making it. ICE is in development, the desktop decision is made, and if it's wrong we'll find out in the most instructive way available. My bet: the people who spend all day operating infrastructure will want what people who spend all day writing code already have, a local instrument with the full model inside it.

Here's a test you can run on yourself. The last time you were paged, how long did you spend finding the right console tab and re-authenticating before you could even look at the problem? Now compare that with how long your editor takes to open from the dock. Muscle memory already voted.

Related reading


This article first appeared on the Light Cloud Blog: https://blog.light-cloud.com/dev-tools/why-a-desktop-app-in-2026.

 

Originally published at https://blog.light-cloud.com/startup/security-as-a-two-person-company on the Light Cloud Blog.

By James, Light Cloud

Summary: Can a two-person company be trusted with your production workloads? That question deserves specific answers instead of vibes. Here's our threat model, what's done, and what honestly isn't.

Security When You're a Two-Person Infrastructure Company

In March 2024, a database engineer noticed his SSH logins were taking about half a second instead of a tenth of one, and he pulled on that thread until it unraveled into the xz-utils backdoor: a multi-year operation in which an attacker patiently earned maintainer trust in a tiny open source project that nearly everything links against. The most sophisticated supply chain attack in memory wasn't aimed at a big company's firewall. It was aimed at one exhausted volunteer.

I bring this up because it reframes the question people politely avoid asking us. The unspoken objection to an infrastructure vendor our size isn't a specific vulnerability. It's a feeling: surely two people can't do security. My claim is that the feeling deserves to be replaced with specific questions and specific answers, because size cuts both ways, and the xz incident is what the failure of a giant, distributed trust model looks like.

What small actually changes

Start with the honest downsides. There's no security team, because there's no team to carve one from. Code review has exactly one reviewer. If both of us are asleep, nobody's awake. A serious compliance questionnaire takes us days we don't have, and a SOC 2 audit costs real money we'd rather spend on engineering. Anyone who tells you smallness is secretly a security advantage across the board is selling something.

But the advantages are real too, and they're structural rather than heroic. Two people is a tiny social attack surface: nobody's going to phish an HR department we don't have, or social-engineer a support tier that doesn't exist. There's no forgotten test cluster from a team that disbanded in 2023. Every credential that exists is known to both of us, every service that runs was started by one of us, and the entire system fits in two heads, which is a property most CISOs would trade a tool budget for.

The xz lesson lands here. That attack worked because the system was too large for anyone to hold: thousands of dependencies, each a trust decision nobody remembers making. Our defense is refusing that shape where we can. Fewer dependencies, pinned and reviewed when they change. Boring, widely-watched components over clever ones. And we buy our lowest layers, tenant isolation included, from a hyperscaler's managed primitives rather than rolling our own, because pretending two people should hand-build multi-tenancy is exactly the kind of confidence you should run from.

The questionnaire, answered in public

The table below is the short version of what due diligence usually asks us, answered the way we answer privately. The honest column is the point.

What you should ask Our answer Status
How are tenants isolated? Workloads run in isolated containers on managed cloud primitives; we don't share a process across customers In production
Who can access production? Two named people, MFA everywhere, least-privilege credentials In production
What happens if one of you disappears? Shared credential custody and written runbooks In production, reviewed rarely
Where are secrets kept? In a managed secret store, scoped per environment; never in the repo In production
Are you SOC 2 certified? Not yet; it's on the roadmap, and we say so instead of implying otherwise Not done
Do you run a bug bounty? No; we take reports by email and answer fast Not done

Two of those rows say "not done". Leaving them visible costs us deals, and hiding them would eventually cost us customers, and of those two prices only one compounds.

One answer belongs outside the table because it's a promise rather than a control. If we're ever breached, you'll hear it from us first, in plain language, with a timeline of what happened and what we changed, and before any lawyer smooths the edges off. Companies a thousand times our size routinely fail that bar. It's the one security capability where being small carries no handicap at all.

The steelman: big vendors really do have things we don't

A serious counterargument deserves its space. A large vendor has a 24/7 security operations center, red teams, dedicated incident response, and auditors on retainer. Those aren't theater. At 3 a.m. during an active intrusion, headcount is a genuine capability, and a SOC 2 report, whatever its limits, at least proves someone examined the controls. If your risk model requires that machinery today, we're the wrong vendor today, and I'd rather say so than argue you out of a reasonable requirement.

What I'd push back on is the inference from big to safe. Blast radius scales with the vendor: when a large platform is breached, the incident arrives with their entire customer list attached. Trust in a vendor of any size ultimately rests on the same two things: whether the isolation between you and other tenants is real, and whether you can leave quickly if your trust turns out to be misplaced. We build for both, and the second one, portability, is the security control almost no questionnaire asks about.

Security for a company like ours isn't a claim to be believed. It's a posture to be inspected, and this post is part of keeping it inspectable. If you'd grill us harder than that table does, send the questions; the honest answers are the cheapest security investment we make. What's the "not done" row your current vendor hasn't shown you?

Related reading


This article first appeared on the Light Cloud Blog: https://blog.light-cloud.com/startup/security-as-a-two-person-company.

 

Originally published at https://blog.light-cloud.com/cloud/the-honest-math-of-repatriation on the Light Cloud Blog.

By James, Light Cloud

Summary: The loudest cloud exit in the industry ended quietly in the summer of 2025 when 37signals' last petabyte left S3. Their published numbers are real. So are the reasons most companies copying them would lose money.

The Honest Math of Cloud Repatriation

The loudest cloud-exit story in the industry ended with almost no noise. In the summer of 2025, when a four-year storage contract expired, 37signals moved its final petabytes off S3 and completed the exit David Heinemeier Hansson had been announcing, itemizing, and gloating about since 2022. AWS even waived about $250,000 in egress fees on the way out, which under the post-Data-Act rules is what goodbye looks like now.

What makes 37signals worth a post isn't that they left the cloud. Companies drift on and off cloud constantly. It's that they published receipts at every step, which makes them the one repatriation story you can actually do arithmetic on, and the arithmetic deserves more honesty than either fan club gives it. My read: their math is real, their savings are real, and most companies who cite them are reading someone else's spreadsheet as if it were their own.

The receipts

Collected from their published posts and the reporting around them, the numbers that matter:

Item Published figure
Cloud spend at peak, 2022 Roughly $3.2M per year
Replacement servers, 2023 About $700K of Dell hardware, recouped within the year
Annual savings by late 2024 Almost $2M per year
S3 exit, 2025 ~10 PB on S3 at ~$1.5M/yr replaced by 18 PB of Pure Storage running under $200K/yr
Projected five-year total Over $10M saved

Take the numbers at face value; nobody has seriously disputed them, and their transparency shames an industry that discusses infrastructure costs the way Victorians discussed ankles. Doubling storage capacity while cutting its annual cost by a factor of seven is not an accounting trick. It's what buying hardware looks like when the hardware market has spent a decade getting absurdly good while cloud storage prices mostly didn't follow it down.

Why it worked for them

Every line of that table rests on properties of 37signals that the table doesn't show. Their load is stable and predictable: mature products, steady subscriber bases, no hypergrowth, no viral spikes, which means capacity planning is a spreadsheet rather than a gamble, and owned hardware is a mortgage on a house they know they'll keep living in. Renting makes sense when you don't know where you'll live next year. They know.

They also brought an ops team that most companies their size don't have, and a stack built for leaving. They run boring, portable components, they wrote their own deployment tooling (Kamal) expressly to make cloud and metal interchangeable, and they'd already sworn off the proprietary managed services that make exits into rewrites. In the vocabulary of our lock-in taxonomy: they'd only ever locked the billing lock, so leaving was a math problem instead of an engineering one. The repatriation didn't create their portability. Their portability is what made the repatriation cheap enough to be worth blogging about.

Why it's a meme everywhere else

Now run the same table for a typical company citing it. Spiky or growing load turns the mortgage back into a gamble: own for peak and idle most of it, or own for average and fall over at peak; elasticity is precisely the product the cloud is good at. The ops team you'd need is payroll the savings must fund before a dollar counts, and two senior infrastructure engineers cost more than a lot of startups' entire cloud bill. If your stack leans on managed databases, queues, and identity, you're not repatriating; you're rebuilding, and the rebuild is the cost that never makes it into the envy math. And a small team taking on datacenter contracts, hardware refresh cycles, and 3 a.m. disk failures is spending its scarcest resource, attention, on the least differentiating work available.

We're the walking counterexample, and it's worth being concrete: a two-person infrastructure company that deliberately builds on hyperscaler primitives, because pretending we should rack servers would be theater. Repatriation math at our scale doesn't just fail to break even. It doesn't reach the starting line.

The middle path is the actual lesson

Here's the steelman for the cloud side, stated fairly: elasticity, managed services, and velocity are worth real premiums for most companies most of the time, 37signals is a special case that generalizes poorly, and DHH's evangelism sometimes elides how special. Every word defensible. And yet the story still carries a general lesson, because the interesting thing 37signals did wasn't leaving. It was being able to leave: knowing their per-workload costs precisely, keeping their stack portable, and treating "where should this run" as a periodically re-asked question instead of an identity.

That's the version that generalizes. The math of rent-versus-own shifts with hardware prices, cloud pricing, your load shape, and your team, which means the right answer has an expiration date, and the companies in trouble aren't the ones on cloud or on metal; they're the ones who can no longer do the math. Per-workload margins nobody tracks, exit plans nobody tests, architectures that made the question unaskable years ago. Portable workloads can move when the math says move, in either direction; everything else stays where it was put, at whatever price appears.

So skip the argument about whether DHH is right. Answer the question his receipts pose: if the math said "move" for one of your workloads next year, could you? And if you don't know the math, that's your answer.

Related reading


This article first appeared on the Light Cloud Blog: https://blog.light-cloud.com/cloud/the-honest-math-of-repatriation.

 

Originally published at https://blog.light-cloud.com/industry/neoclouds-and-the-great-unbundling on the Light Cloud Blog.

By James, Light Cloud

Summary: CoreWeave went from crypto-mining afterthought to a public company with a $99 billion backlog by selling exactly one thing well. The specific companies may wobble. The procurement habit they taught buyers won't.

Neoclouds and the Great Unbundling

As of March 31, 2026, CoreWeave reported a contracted revenue backlog of $99.4 billion. Sit with the number for a second. A company that was mining Ethereum a few years ago, that went public in March 2025 at $40 a share to considerable skepticism, carries committed future revenue approaching the GDP of a small country, with quarters like Q2 2025's $1.21 billion, up 207% year over year, and customers like Nvidia placing $6.3 billion orders.

For fifteen years, the cloud industry's foundational assumption was that the bundle always wins: nobody beats AWS because AWS has everything, and everything is what enterprises buy. CoreWeave, Lambda, Nebius, and the rest of the GPU-first neoclouds just falsified that at nine figures a quarter. My claim is that the falsification matters more than the companies. Whatever happens to any particular neocloud, they've taught a generation of buyers to shop outside the bundle, and that habit doesn't reverse.

How the bundle cracked

The opening was supply: AI demand outran hyperscaler GPU capacity, and a buyer who can't get H100 allocations doesn't care how many other services the vendor offers. But the neoclouds kept the customers supply alone can't explain, because specialist economics turn out to be real. A cloud that sells one thing doesn't carry the tax of two hundred others: no army of half-maintained services, no bundle pricing designed for cross-subsidy, datacenters engineered for exactly one workload shape. For training runs, that focus shows up as price, availability, and performance the generalists struggled to match, and sophisticated buyers noticed.

CoreWeave isn't alone in the cohort, which is part of the proof. Nebius arrived by the strangest route available, carved out of Yandex and relisted on Nasdaq as a European-rooted AI infrastructure company. Lambda grew from selling deep-learning workstations into a GPU cloud with its own gravity. Different origins, different balance sheets, one shared bet: that a cloud can be narrow and win. The strongest confirmation came from the incumbents themselves, who quietly became neocloud customers, Microsoft most famously among them. When the bundle buys from the unbundlers, the argument about whether specialists can compete is over; what remains is negotiating the price.

The noticing is the historic part. AI forced the most conservative procurement departments on earth to unbundle one workload, evaluate it on its own merits, and sign with a vendor whose logo their board had never seen. That's a psychological dam breaking. Once a company has split its GPU spend from its general compute, "we buy everything from one cloud" stops being a law of nature and becomes what it always was: one option, with a price.

What unbundles next

Storage went first, quietly, years ago; the Backblazes and Wasabis proved a single-service cloud can undercut the bundle when the service is a commodity. Databases are mid-unbundling now, with specialist Postgres vendors peeling the most valuable managed service out of the bundle one developer at a time. Inference looks next: unlike training, it's latency-sensitive and spiky, which invites both edge specialists and scale-to-zero economics, and my bet is it splits from training procurement within a couple of years.

And then there's the unbundling we have obvious skin in: ordinary small workloads. While hyperscaler attention and capex chase AI factories, the developer buying a container, a database, and a preview environment is nobody's priority, and specialist platforms serve that buyer better than a 240-service console does. The neoclouds proved the top of the market can be unbundled. The bottom is softer.

The tooling is catching up to the habit, too. Multi-vendor procurement only works if comparing and moving stays cheap, which is why every unbundling wave drags a portability wave behind it, and why the lock-in taxonomy matters more in an unbundled world, not less. A buyer juggling four specialist vendors needs the map of what runs where far more than a buyer with one throat to choke ever did.

The steelman: shortage artifact

The case against reading too much into neoclouds is respectable. They may be a GPU shortage wearing a business model: when supply normalizes, hyperscalers reclaim the workloads, and gravity, egress, data locality, enterprise agreements, reasserts the bundle. CoreWeave specifically is a debt-heavy bet with concentrated customers, and a backlog is a promise, with counterparties, not cash. The dot-com era minted specialist infrastructure companies too, and the survivors' list is short. Maybe this cohort is scaffolding: essential during the boom, absorbed or gone after it.

Concede the company-level risk entirely; some of these firms will have ugly years, and consolidation is likely. The ratchet argument survives anyway, because it doesn't depend on who survives. Buyers who unbundled once keep the muscle: the procurement templates exist, the multi-vendor tooling exists, the board slide that says "we evaluate per workload" exists. Habits, unlike companies, don't need to refinance. The bundle can win any given workload back, and it will; what it can't recover is the presumption that it wins by default. Ask the telecom industry how presumptions age: the phone bundle lost its own spell decades ago, the incumbents are all still here, and nobody has bought the bundle out of reflex since.

Which leaves the question where it always lands on this blog: on your side of the table. The neocloud era's real gift to every buyer is permission, permission to ask, workload by workload, "who is actually best at this?", and the only prerequisite for using it is being able to move. When did your team last ask that question about anything other than GPUs?

Related reading


This article first appeared on the Light Cloud Blog: https://blog.light-cloud.com/industry/neoclouds-and-the-great-unbundling.

 

Originally published at https://blog.light-cloud.com/industry/power-is-the-new-region on the Light Cloud Blog.

By James, Light Cloud

Summary: Three Mile Island is being restarted for a single customer. Hyperscalers are buying reactors because the constraint on cloud capacity stopped being racks or chips and became megawatts, and that breaks a decade of assumptions about regions.

Power Is the New Region

The site of America's most famous nuclear accident is being switched back on for one customer. Under a 20-year power purchase agreement signed with Microsoft, Constellation is restarting Three Mile Island's Unit 1, rebranded the Crane Clean Energy Center, with 835 megawatts targeted for 2028 and every one of them earmarked for AI datacenters. A retired reactor, resurrected, for a software company. Whatever else 2024 is remembered for in this industry, that deal is the moment the constraint changed in public.

Here's the thesis: for fifteen years, cloud capacity planning assumed the scarce inputs were racks, then chips. Both eras are over. The binding constraint on cloud buildout is now electricity, the industry's site selection has started following energy instead of users, and that quietly breaks assumptions about regions, latency, and pricing that everyone's architecture diagrams still encode.

The shopping spree

Microsoft's reactor wasn't an eccentric one-off; it opened a genre. The deals since read like a utility's M&A desk got hold of big tech's checkbook.

Buyer Deal Scale and timeline
Microsoft Constellation PPA, Three Mile Island restart 835 MW, 20 years, targeted 2028
Google First corporate SMR purchase agreement, with Kairos Power ~500 MW across 6-7 small modular reactors, first unit around 2030
Amazon Led a $500M round in SMR developer X-energy; bought a $650M campus next to the Susquehanna nuclear plant SMR fleet ambitions plus nuclear-adjacent land, this decade

Read the table as a confession. Companies whose competence is software are becoming counterparties to reactor restarts and first-of-a-kind SMR deployments, timelines measured in half-decades, because they've concluded the grid won't sell them what they need on any faster schedule. Nobody signs a 20-year PPA for a temporary problem.

The queue is the moat

The mechanics behind the confession are unglamorous. Getting a new datacenter connected to the grid means joining an interconnection queue, and those queues now run years in the good cases, which turns energy procurement into the longest-lead item in the entire capacity supply chain, longer than chips, longer than construction. Money can compress most shortages. It cannot much compress permitting, transmission builds, or turbine order books, which is why capital has started chasing anything that bypasses the queue: retired reactors, on-site generation, campuses bought specifically because they sit next to existing plants.

The scale of the collision is public arithmetic. The hyperscalers plan around $700 billion of combined capital expenditure in 2026, overwhelmingly for AI infrastructure, and infrastructure at that scale is measured in gigawatts. Microsoft has been unusually candid about where that collides with reality: an $80 billion Azure backlog attributed to power constraints, with purchased GPUs sitting idle because there's no electricity to install them under. Sit with that image. The most valuable chips on earth, in warehouses, waiting for a substation.

Which makes power contracts the new moat. Two years ago the competitive question between clouds was who had the best silicon roadmap; the harder question now is who locked in generation, and queue positions, interconnect agreements, and PPAs signed in 2024 are assets rivals cannot replicate at any price on the same timeline. It's the unbundling era's resource layer: the neoclouds proved compute could be bought outside the bundle, and now everyone discovers that what actually gates compute is a commodity older than computing.

What it quietly breaks

Regions used to be a demand-side concept: put capacity where users, data residency, and enterprise customers are, and price it roughly uniformly. Energy-first site selection inverts that. New capacity lands where megawatts are available, Brandenburg or Pennsylvania or wherever a reactor has spare output, which is not necessarily where anyone's users are, and the decade-old assumption that your provider will simply have capacity near your market when you need it stops being safe. Latency budgets meet geology.

Pricing assumptions crack next. Electricity costs now diverge sharply by location and contract vintage, and uniform-ish regional pricing papers over an input cost that stopped being uniform; my bet is the paper doesn't hold, and region-differentiated compute pricing, or scarcity surcharges wearing another name, arrive within a few years. Timeline assumptions were always the most fragile: demand is compounding now, while the table above delivers its megawatts in 2028 and 2030. The gap between those dates is the era we're in, and it's the era in which capacity allocations, waitlists, and quota negotiations became a normal part of buying cloud, a sentence that would have sounded absurd in 2020.

The steelman: constraints attract solutions

The case for calm is respectable. Efficiency is improving fast, inference is being squeezed onto cheaper silicon, and the industry has a long record of demand forecasts embarrassing themselves; some of those SMR deals are best understood as long-dated hedges and press releases, with first-of-a-kind reactors carrying first-of-a-kind risk, and the grid does eventually build out. If AI demand plateaus, today's power panic will look like the fiber glut of 2001, and contrarians buying distressed capacity will feast.

Concede all of it as possible, and note what the calm case requires: believing simultaneously that the companies spending $700 billion are wrong about demand, and that the constraint their own executives call binding will dissolve before it reshapes the market. Even on optimistic timelines, the operative decade runs on scarce power, and market structure formed during scarcity, the contracts, the queue positions, the siting, outlives the scarcity that formed it. That's the actual lesson of every infrastructure cycle, fiber included: the glut ended, the ownership map it created didn't.

There's a small demand-side moral we can't resist, since our whole product exists on the other end of this telescope: when the industry's binding constraint is electricity, workloads that scale to zero stop being a pricing gimmick and start being a grid courtesy. Idle compute burning watts is now everyone's problem. Yours too: do you know which region your next deployment lands in, and do you know what's powering it?

Related reading


This article first appeared on the Light Cloud Blog: https://blog.light-cloud.com/industry/power-is-the-new-region.

 

Originally published at https://blog.light-cloud.com/tutorials/microservices-on-light-cloud-part-1-deploy on the Light Cloud Blog.

By Julia, Light Cloud

Summary: Deploy a React frontend, a Node.js API, a Python FastAPI API and PostgreSQL from one GitHub monorepo on Light Cloud, step by step, with every screen shown.

Microservices Part 1: Deploy a Frontend, Two APIs and a Database

This series

  1. Microservices Part 1: Deploy a Frontend, Two APIs and a Database (this post)
  2. Microservices Part 2: Connect Your Services Safely
  3. Microservices Part 3: Autoscale Under Load
  4. Microservices Part 4: Cold Starts vs Always On

To deploy microservices on Light Cloud, put each service in its own folder of one GitHub repository, create one Light Cloud app per folder by setting its Root directory, and connect the services with environment variables that hold each other's addresses. Light Cloud detects the framework in every folder, builds it, and gives each service its own HTTPS address. No Kubernetes, no Dockerfiles and no YAML.

This is Part 1 of a six-part series. By the end of it you have a small online shop running as three services and a database. Later parts add autoscaling, cold-start tuning, branch environments and request tracing to the same app.

What you will build

Bean There, a tiny coffee shop made of four parts:

  • web: a React (Vite) shop front, served from the edge as a static site.
  • catalog-api: a Node.js (Express) service that owns products and stock.
  • orders-api: a Python (FastAPI) service that owns orders. To place an order it asks catalog-api to reserve stock first.
  • bean-there-db: one PostgreSQL database. Each API owns its own table.

The finished project in the Light Cloud console: catalog-api, orders-api and web deployed, and bean-there-db ready

Live demo: main-web-examples.light-cloud.io. Source code: github.com/light-cloud-com/tutorial-microservices, tag part-1.

The repository has one folder per service:

tutorial-microservices/
  web/           React (Vite)       static site
  catalog-api/   Node.js (Express)  container
  orders-api/    Python (FastAPI)   container
  README.md

Before you start

  • A GitHub account.
  • A Light Cloud account on the Starter plan or higher. The free Hobby plan allows one service and no databases; this project needs two API services and a database.
  • The Light Cloud GitHub App connected to your account. If it is not, the console asks you to connect it the first time you choose Deploy from GitHub.
  • Optional, to test from a terminal: macOS and Linux have curl and OpenSSL built in. On Windows, use PowerShell 7 (winget install Microsoft.PowerShell); the Windows tabs use curl.exe, which ships with Windows 10 and 11. Pick your system on any command below and the page remembers it.

Step 1: Fork the repository

Light Cloud deploys from a repository you own, so start by making your own copy.

  1. Open github.com/light-cloud-com/tutorial-microservices.
  2. Click Fork, keep the name tutorial-microservices, and click Create fork.

tutorial-microservices now appears under your own GitHub account.

Step 2: Create a folder for the project

A folder keeps the four resources of this project together in the console sidebar.

  1. In the Light Cloud console, click New... at the top of the sidebar.
  2. Click New Folder.

The New menu in the Light Cloud sidebar with New Folder at the bottom

  1. Type bean-there as the Name and click Create folder.

The Create a folder dialog with the name bean-there and the Create folder button highlighted

You should see bean-there in the sidebar.

Step 3: Create the PostgreSQL database

Both APIs store their data in one database, so create it first.

  1. Hover over bean-there in the sidebar and click the + button next to it.
  2. Click New Database.

The bean-there folder menu with New Database highlighted

  1. Under Engine, click PostgreSQL.

The Create page with the Database tile selected and the PostgreSQL engine highlighted

  1. Change the Name to bean-there-db.
  2. Open Size and choose Dev, the shared instance. It is ready in seconds and is enough for this tutorial.

The Size dropdown open with Dev, Shared instance highlighted, and Starter and Pro below it

  1. Leave Storage at 1 GB and pick the Region closest to you. Click Create database.

The database form filled in with the name bean-there-db and the Create database button highlighted

You should see the database page with a green Ready badge. In my run it took about six seconds.

The bean-there-db overview page with the Ready badge highlighted

Step 4: Copy the connection string

The connection string is the address and password your services use to reach the database.

  1. Open the Credentials tab.
  2. Under Connection String, click Copy. Paste it somewhere private for the next steps; you will use it twice.

The Connection String card with the password hidden and the Copy button highlighted

The copied string looks like this. Your user name, password and database name are in it; here they are hidden behind stars:

postgresql://u_******:********@bean-there-db-yourworkspace.db.light-cloud.io:5432/db******

The database only accepts encrypted connections, and this string does not say so yet. You will add a short ending to it for each service in the next steps.

Step 5: Deploy catalog-api

catalog-api is the first service because orders-api needs its address.

  1. Click + next to bean-there again and choose Deploy from GitHub.

The bean-there folder menu with Deploy from GitHub highlighted

  1. In Search your repositories..., type tutorial-microservices and click your fork.

The repository search showing tutorial-microservices highlighted

Light Cloud now looks at the root of the repository and shows "We're not sure what this is". That is expected: the root holds three apps, not one. You tell it which folder to use.

The warning We're not sure what this is, with the folder button next to Root directory highlighted

  1. Click the folder button next to Root directory and choose catalog-api.

The folder picker listing catalog-api, orders-api and web, with catalog-api highlighted

The banner turns green: Express / Backend, based on package.json. Light Cloud found Express in the dependencies, so it runs this folder as a server.

  1. The Name field still says tutorial-microservices. Change it to catalog-api. The name becomes part of the service's address.

The detection banner Express Backend, with Name set to catalog-api and Root directory set to catalog-api

  1. Click Advanced - build settings, environment variables, domain, scaling. Port is already 8080; leave the rest as it is.
  2. Under Environment variables, click Add variable twice and fill in:
KEY value
DATABASE_URL your connection string, followed by ?sslmode=require&uselibpqcompat=true
INTERNAL_SECRET a long random string; keep it, orders-api needs the same one

Put together, the two values look like this (stars hide your password and names):

DATABASE_URL     postgresql://u_******:********@bean-there-db-yourworkspace.db.light-cloud.io:5432/db******?sslmode=require&uselibpqcompat=true
INTERNAL_SECRET  cf91d1e26971527e2ec7362e9452f5abb4670e767ab1a0af

To make a random secret, run the command below. You should see one line of 48 random letters and digits; yours will be different:

$ openssl rand -hex 24
cf91d1e26971527e2ec7362e9452f5abb4670e767ab1a0af
# Run this in Git Bash, which comes with Git for Windows.
$ openssl rand -hex 24
cf91d1e26971527e2ec7362e9452f5abb4670e767ab1a0af

The Environment variables section with DATABASE_URL and INTERNAL_SECRET added and their values hidden

The ending ?sslmode=require&uselibpqcompat=true tells the Node.js pg driver to use an encrypted connection the way the PostgreSQL command-line tools do. Without it, catalog-api cannot connect to the database.

  1. Click Deploy.

The catalog-api environments page with Production showing Deploying

You should see the Production environment move from Deploying to Deployed. In my run that took about 75 seconds.

Step 6: Check that catalog-api works

A quick request confirms the service is up and can read the database.

Your service's address follows the pattern https://main/-<app name>-<workspace>.light-cloud.io. You can also copy it from the URL card on the environment's Overview tab.

The catalog-api Overview tab with the Deployed badge and the URL card highlighted

Ask it for its products. You should see the four products catalog-api created on its first start, on one line:

$ curl https://main-catalog-api-yourworkspace.light-cloud.io/products
[{"id":1,"name":"Espresso beans, 1 kg","price_cents":2400,"stock":40},{"id":2,"name":"Pour-over kettle","price_cents":5900,"stock":12},{"id":3,"name":"Ceramic mug","price_cents":1500,"stock":100},{"id":4,"name":"Paper filters, 100 pack","price_cents":600,"stock":250}]
PS> curl.exe https://main-catalog-api-yourworkspace.light-cloud.io/products
[{"id":1,"name":"Espresso beans, 1 kg","price_cents":2400,"stock":40},{"id":2,"name":"Pour-over kettle","price_cents":5900,"stock":12},{"id":3,"name":"Ceramic mug","price_cents":1500,"stock":100},{"id":4,"name":"Paper filters, 100 pack","price_cents":600,"stock":250}]

Now check that the internal route refuses callers without the secret. You should see the status line HTTP/2 401 and the error body (other headers left out here):

$ curl -i -X POST https://main-catalog-api-yourworkspace.light-cloud.io/internal/products/1/reserve \
  -H "content-type: application/json" -d '{"quantity":1}'
HTTP/2 401
content-type: application/json; charset=utf-8
x-powered-by: Express

{"error":"Unauthorized"}
PS> curl.exe -i -X POST https://main-catalog-api-yourworkspace.light-cloud.io/internal/products/1/reserve `
  -H "content-type: application/json" -d '{"quantity":1}'
HTTP/2 401
content-type: application/json; charset=utf-8
x-powered-by: Express

{"error":"Unauthorized"}

Every Light Cloud address is public, so this check is what keeps strangers from changing your stock. This is the part that does it:

// Only other services may call /internal routes.
function requireInternalSecret(req, res, next) {
  if (!INTERNAL_SECRET || req.get("x-internal-secret") !== INTERNAL_SECRET) {
    return res.status(401).json({ error: "Unauthorized" });
  }
  next();
}

Step 7: Deploy orders-api

orders-api follows the same steps, with its own folder and one extra variable.

  1. Click + next to bean-there, choose Deploy from GitHub, and pick your fork again.
  2. Set Root directory to orders-api and change Name to orders-api.

The banner reads FastAPI / Backend, based on requirements.txt. Light Cloud runs FastAPI with uvicorn on port 8000; you do not write a start command.

The detection banner FastAPI Backend for the orders-api folder

  1. Open Advanced and add three variables:
KEY value
DATABASE_URL your connection string, followed by ?sslmode=require
INTERNAL_SECRET the same secret you gave catalog-api
CATALOG_API_URL catalog-api's address, for example https://main-catalog-api-yourworkspace.light-cloud.io/

For example:

DATABASE_URL     postgresql://u_******:********@bean-there-db-yourworkspace.db.light-cloud.io:5432/db******?sslmode=require
INTERNAL_SECRET  cf91d1e26971527e2ec7362e9452f5abb4670e767ab1a0af
CATALOG_API_URL  https://main-catalog-api-yourworkspace.light-cloud.io/

The Environment variables section of orders-api with DATABASE_URL, INTERNAL_SECRET and CATALOG_API_URL

Python's psycopg driver only needs ?sslmode=require. The longer ending in Step 5 is specific to Node.js.

  1. Click Deploy.

This is how orders-api uses those variables when a customer buys something:

async with httpx.AsyncClient(timeout=10) as client:
    reply = await client.post(
        f"{CATALOG_API_URL}/internal/products/{order.product_id}/reserve",
        json={"quantity": order.quantity},
        headers={"x-internal-secret": INTERNAL_SECRET, "x-request-id": request_id},
    )

When it is deployed, check it. You should see:

$ curl https://main-orders-api-yourworkspace.light-cloud.io/health
{"status":"ok","service":"orders-api"}
PS> curl.exe https://main-orders-api-yourworkspace.light-cloud.io/health
{"status":"ok","service":"orders-api"}

Step 8: Deploy the web frontend

The shop front is a static site, so it needs both API addresses at build time.

  1. Click + next to bean-there, choose Deploy from GitHub, and pick your fork.
  2. Set Root directory to web and change Name to web.

The banner reads React / Frontend and the card says Served from the edge: the built files go to a content delivery network, not to a server.

The detection banner React Frontend with Served from the edge highlighted

  1. Open Advanced. Light Cloud has already filled in npm ci, npm run build and the output folder dist. Add two variables:
KEY value
VITE_CATALOG_API_URL catalog-api's address
VITE_ORDERS_API_URL orders-api's address

The web app's Environment variables with VITE_CATALOG_API_URL and VITE_ORDERS_API_URL

Vite copies variables that start with VITE_ into the JavaScript bundle during the build. That is why they must be set before the first deploy.

  1. Click Deploy.

Step 9: Allow the frontend to call the APIs

Open the web address, for example https://main-web-yourworkspace.light-cloud.io/. The page loads, but the product list stays empty.

The Bean There page loaded with no products and an error message

The browser blocked the calls because the APIs only allow http://localhost:5173/, the address used during local development. This rule is called CORS (cross-origin resource sharing): a browser only lets a page on one address read answers from another address if that address says yes. Both APIs read the allowed address from WEB_ORIGIN:

app.use(cors({ origin: WEB_ORIGIN }));

Set it on catalog-api:

  1. In the sidebar, open catalog-api, then Production, then the Settings tab.
  2. In Environment Variables, click Edit, then Add.
  3. Enter WEB_ORIGIN as the key and your web address as the value, without a slash at the end.
  4. Click Save.

The Environment Variables editor with WEB_ORIGIN added and the Save button highlighted

You should see "Environment variables saved". Saving redeploys the service by itself; there is no separate deploy button to press.

Repeat the same four steps for orders-api.

Step 10: Place an order

Reload the shop about a minute after saving. The products appear, each with a Buy button.

Click Buy on any product. You should see "Order #1 placed", the stock go down by one, and the order appear under Latest orders.

The live Bean There shop with the message Order #1 placed: Pour-over kettle

That one click went through all four parts: the browser called orders-api, orders-api asked catalog-api to reserve stock with the shared secret, catalog-api updated the products table, and orders-api saved the order in the orders table.

Troubleshooting

Access to fetch has been blocked by CORS policy

The full message starts with Access to fetch at 'https://main-catalog-api/-...' from origin 'https://main-web/-...' has been blocked by CORS policy. The API's WEB_ORIGIN is missing or does not match the web address exactly. Check for https://, and make sure there is no slash at the end. Saving the variable redeploys the service; wait about a minute and reload.

catalog-api fails to start with a certificate or SSL error

The DATABASE_URL on catalog-api is missing the ending ?sslmode=require&uselibpqcompat=true. Add it in Settings, Environment Variables, and save.

Placing an order says "Catalog service unavailable"

orders-api could not reserve stock. Either CATALOG_API_URL points to the wrong address, or INTERNAL_SECRET is not the same on both services. Check both, then save.

The product list is empty and there is no CORS error

The web app was built before VITE_CATALOG_API_URL and VITE_ORDERS_API_URL were set. Add them in the web app's Settings; saving rebuilds it with the new values.

FAQ

Do I need Kubernetes or Docker to run microservices on Light Cloud?

No. Light Cloud detects each service from its files (package.json, requirements.txt) and builds the container for you. A Dockerfile is optional.

Can my services talk to each other over a private network?

Not today. Services call each other over their public HTTPS addresses, so internal routes need their own check. This tutorial uses a shared secret header.

Why do both APIs use the same database?

To keep the tutorial cheap and simple. Each service owns its own table and never reads the other's, so you can split them into two databases later without changing the code.

Does it cost money when nobody is using the shop?

The APIs start with Min instances set to 0, so they scale to zero when idle. The plan price includes usage worth that price each month.

Can I write the services in other languages?

Yes. Each folder is detected on its own, so a Go, Java or .NET service can sit next to these two in the same repository.


This article first appeared on the Light Cloud Blog: https://blog.light-cloud.com/tutorials/microservices-on-light-cloud-part-1-deploy.

 

Originally published at https://blog.light-cloud.com/tutorials/microservices-on-light-cloud-part-2-connect-services on the Light Cloud Blog.

By Julia, Light Cloud

Summary: Allow several browser origins with CORS, rotate the secret your services share without dropping a request, and add a timeout between services, step by step on Light Cloud.

Microservices Part 2: Connect Your Services Safely

This series

  1. Microservices Part 1: Deploy a Frontend, Two APIs and a Database
  2. Microservices Part 2: Connect Your Services Safely (this post)
  3. Microservices Part 3: Autoscale Under Load
  4. Microservices Part 4: Cold Starts vs Always On

To connect microservices safely on Light Cloud, keep every service address and secret in environment variables, allow each browser origin explicitly with CORS, rotate shared secrets in three saves (receiver accepts both, caller switches, receiver drops the old one), and put a timeout on every call between services. Each save redeploys only the service you changed, so none of this needs downtime.

This is Part 2 of the series. It continues from Part 1, where you deployed Bean There: a React shop front, a Node.js catalog-api, a Python orders-api and PostgreSQL. Here you make the connections between them production-ready.

What you will build

Three changes to the same running shop:

  • Several allowed origins. The APIs answer both the live site and http://localhost:5173/, so you can run the frontend on your laptop against the deployed APIs.
  • Secret rotation without downtime. You replace the secret orders-api uses to call catalog-api while orders keep working.
  • A timeout between services. If catalog-api does not answer within 5 seconds, orders-api says so clearly instead of hanging.

Live demo: main-web-examples.light-cloud.io. Source code: github.com/light-cloud-com/tutorial-microservices, tag part-2.

Before you start

  • Bean There deployed from Part 1: web, catalog-api, orders-api and bean-there-db, all running.
  • Your fork of tutorial-microservices.
  • A terminal: macOS and Linux have curl and OpenSSL built in. On Windows, use PowerShell 7 (winget install Microsoft.PowerShell); the Windows tabs use curl.exe, which ships with Windows 10 and 11.
  • Git, to pull the Part 2 code into your fork.

Step 1: Get the Part 2 code

The Part 2 changes live in catalog-api and orders-api. Pull them into your fork.

  1. Open your fork on GitHub.
  2. Click Sync fork, then Update branch.

Or from a clone of your fork. You should see the pull fast-forward to the Part 2 commit, touching only the two APIs and the README, and the push end with the same commit range (output trimmed):

$ git remote add upstream https://github.com/light-cloud-com/tutorial-microservices.git
$ git pull upstream main
Updating 8023f1a..ddcbae0
Fast-forward
 README.md             | 36 +++++++++++++++---------------------
 catalog-api/server.js | 35 ++++++++++++++++++++++++++++-------
 orders-api/main.py    | 25 +++++++++++++++++--------
 3 files changed, 60 insertions(+), 36 deletions(-)
$ git push origin main
   8023f1a..ddcbae0  main -> main
PS> git remote add upstream https://github.com/light-cloud-com/tutorial-microservices.git
PS> git pull upstream main
Updating 8023f1a..ddcbae0
Fast-forward
 README.md             | 36 +++++++++++++++---------------------
 catalog-api/server.js | 35 ++++++++++++++++++++++++++++-------
 orders-api/main.py    | 25 +++++++++++++++++--------
 3 files changed, 60 insertions(+), 36 deletions(-)
PS> git push origin main
   8023f1a..ddcbae0  main -> main

The push changes files in catalog-api/ and orders-api/ only. Light Cloud redeploys just those two. The web app stays on the commit it already runs, because nothing in web/ changed.

The web app's Production environment still on the Part 1 commit 8023f1a after the push

catalog-api's Production environment deploying the Part 2 commit

You should see catalog-api and orders-api deploying the new commit, and web unchanged. This is the monorepo rule from Part 1 at work: each app has a Root directory, and a push only redeploys the apps whose folder it touched.

Step 2: Allow more than one origin

A browser only lets a page read an API's answer if the API names that page's origin in its CORS header. Part 1 allowed one origin. Now WEB_ORIGIN can hold a comma-separated list. This is the new code in catalog-api:

// One or more browser origins allowed to call this API, comma-separated:
// WEB_ORIGIN=https://main-web-myteam.light-cloud.io,http://localhost:5173/
const WEB_ORIGINS = (process.env.WEB_ORIGIN || "http://localhost:5173/")
  .split(",")
  .map((origin) => origin.trim())
  .filter(Boolean);

app.use(cors({ origin: WEB_ORIGINS }));

orders-api does the same in Python:

WEB_ORIGINS = [o.strip() for o in os.environ.get("WEB_ORIGIN", "http://localhost:5173/").split(",") if o.strip()]
app.add_middleware(CORSMiddleware, allow_origins=WEB_ORIGINS, allow_methods=["*"], allow_headers=["*"])

Add your laptop's address to both APIs:

  1. Open catalog-api, then Production, then the Settings tab.
  2. In Environment Variables, click Edit.
  3. Change WEB_ORIGIN to your web address and http://localhost:5173/, separated by a comma and no spaces:
https://main-web-yourworkspace.light-cloud.io,http://localhost:5173/

The WEB_ORIGIN variable with two origins and the Save button highlighted

  1. Click Save. Repeat for orders-api.

Saving redeploys the service. After about a minute, ask the API as if you were the local frontend. You should see the origin echoed back:

$ curl -I -H "Origin: http://localhost:5173/" https://main-catalog-api-yourworkspace.light-cloud.io/products
HTTP/2 200
date: Fri, 25 Sep 2026 19:04:59 GMT
content-type: application/json; charset=utf-8
content-length: 371
cf-ray: a40c4c243e9fe75e-WAW
cf-cache-status: DYNAMIC
access-control-allow-origin: http://localhost:5173/
PS> curl.exe -I -H "Origin: http://localhost:5173/" https://main-catalog-api-yourworkspace.light-cloud.io/products
HTTP/2 200
date: Fri, 25 Sep 2026 19:04:59 GMT
content-type: application/json; charset=utf-8
content-length: 371
cf-ray: a40c4c243e9fe75e-WAW
cf-cache-status: DYNAMIC
access-control-allow-origin: http://localhost:5173/

-I shows only the headers, and -H pretends to be the local frontend. I cut the output after access-control-allow-origin, the line that matters. An origin that is not in the list gets no Access-Control-Allow-Origin header at all, so the browser blocks it. That is the behaviour you want: a list, never *, for an API that changes data.

Step 3: Rotate the shared secret without downtime

catalog-api refuses calls to /internal routes unless they carry the shared secret. If you change the secret on both services at the same moment, there is a window where one has the new value and the other the old one, and orders fail. The fix is to let catalog-api accept two secrets for a short time.

This is how catalog-api checks the header now:

// The current secret, plus the previous one while a rotation is in progress.
const INTERNAL_SECRET = process.env.INTERNAL_SECRET;
const INTERNAL_SECRET_PREVIOUS = process.env.INTERNAL_SECRET_PREVIOUS;

// Compares in constant time, so response timing does not leak the secret.
function sameSecret(expected, given) {
  if (!expected || !given) return false;
  const a = Buffer.from(expected);
  const b = Buffer.from(given);
  return a.length === b.length && timingSafeEqual(a, b);
}

function requireInternalSecret(req, res, next) {
  const given = req.get("x-internal-secret");
  if (sameSecret(INTERNAL_SECRET, given)) return next();
  if (sameSecret(INTERNAL_SECRET_PREVIOUS, given)) {
    log(req, "internal call used the previous secret");
    return next();
  }
  log(req, "rejected internal call");
  return res.status(401).json({ error: "Unauthorized" });
}

Make a new secret first. You should see one line of 48 random letters and digits; yours will be different. Keep it somewhere private for the next steps:

$ openssl rand -hex 24
cf91d1e26971527e2ec7362e9452f5abb4670e767ab1a0af
# Run this in Git Bash, which comes with Git for Windows.
$ openssl rand -hex 24
cf91d1e26971527e2ec7362e9452f5abb4670e767ab1a0af

A. catalog-api accepts both secrets

  1. Open catalog-api, Production, Settings, and click Edit under Environment Variables.
  2. Set INTERNAL_SECRET to the new secret.
  3. Click Add and create INTERNAL_SECRET_PREVIOUS with the old secret.
  4. Click Save.

catalog-api's variables with INTERNAL_SECRET and INTERNAL_SECRET_PREVIOUS highlighted and their values hidden

orders-api still sends the old secret, and orders keep working. Open catalog-api's Logs tab and search for previous secret: every internal call made with the old value is listed.

catalog-api's Logs tab filtered to lines saying internal call used the previous secret

B. orders-api switches to the new secret

  1. Open orders-api, Production, Settings, and click Edit.
  2. Set INTERNAL_SECRET to the new secret and click Save.

orders-api's INTERNAL_SECRET highlighted with the Save button

When orders-api has redeployed, place an order in the shop. Search catalog-api's logs for previous secret again. You should see no new lines: orders-api now uses the new secret.

C. catalog-api drops the old secret

  1. Back in catalog-api, Settings, click Edit.
  2. Click the x at the end of the INTERNAL_SECRET_PREVIOUS row.

The INTERNAL_SECRET_PREVIOUS row highlighted, ready to be removed

  1. Confirm with Remove, then click Save.

The Remove variable dialog for INTERNAL_SECRET_PREVIOUS with the Remove button highlighted

Once catalog-api has redeployed, the old secret no longer works. Put your old secret in the header (the example shows the one from Part 1). You should see Unauthorized:

$ curl -X POST https://main-catalog-api-yourworkspace.light-cloud.io/internal/products/1/reserve \
  -H "content-type: application/json" \
  -H "x-internal-secret: cf91d1e26971527e2ec7362e9452f5abb4670e767ab1a0af" \
  -d '{"quantity":1}'
{"error":"Unauthorized"}
PS> curl.exe -X POST https://main-catalog-api-yourworkspace.light-cloud.io/internal/products/1/reserve `
  -H "content-type: application/json" `
  -H "x-internal-secret: cf91d1e26971527e2ec7362e9452f5abb4670e767ab1a0af" `
  -d '{"quantity":1}'
{"error":"Unauthorized"}

The same request with the new secret succeeds (and reserves one item, so run it once). In my run, the old secret stopped working 70 seconds after saving, and no order failed during the whole rotation.

Step 4: Fail fast when catalog-api is down

A call between services can hang: the other service may be starting up, overloaded or misconfigured. Without a limit, the customer waits for as long as the connection does. orders-api now gives catalog-api 5 seconds and then answers with a clear error:

# How long to wait for catalog-api before giving up on an order.
CATALOG_TIMEOUT_SECONDS = float(os.environ.get("CATALOG_TIMEOUT_SECONDS", "5"))

try:
    async with httpx.AsyncClient(timeout=CATALOG_TIMEOUT_SECONDS) as client:
        reply = await client.post(
            f"{CATALOG_API_URL}/internal/products/{order.product_id}/reserve",
            json={"quantity": order.quantity},
            headers={"x-internal-secret": INTERNAL_SECRET, "x-request-id": request_id},
        )
except httpx.HTTPError as error:
    log(request_id, "catalog-api unreachable", error=type(error).__name__)
    raise HTTPException(status_code=503, detail="Catalog service unavailable, please try again")

When catalog-api cannot be reached, an order now returns:

{"detail": "Catalog service unavailable, please try again"}

with status 503, and orders-api's logs show catalog-api unreachable with the error type. To change the limit, add CATALOG_TIMEOUT_SECONDS to orders-api's environment variables. Keep it well below the time a customer is willing to wait for a button click.

Why 5 seconds? catalog-api scales to zero when idle, and its first request after a quiet period includes a cold start. Part 4 measures that cold start, so you can set the timeout from a number instead of a guess.

Troubleshooting

Access-Control-Allow-Origin is missing for localhost

WEB_ORIGIN must list the exact origin: scheme, host and port, and no slash at the end. http://localhost:5173/ and http://127.0.0.1:5173/ are different origins. Check that the value has no spaces around the comma, save, and wait about a minute for the redeploy.

Orders fail with 502 during the rotation

orders-api sent a secret catalog-api did not accept. Usually the order of the steps was swapped: orders-api got the new secret before catalog-api accepted it. Set INTERNAL_SECRET_PREVIOUS on catalog-api to whatever orders-api currently sends, and orders work again.

Orders return 503 "Catalog service unavailable, please try again"

orders-api could not reach catalog-api within CATALOG_TIMEOUT_SECONDS. Check CATALOG_API_URL on orders-api and that catalog-api's Production environment is Deployed. If catalog-api scaled to zero, the first order after a quiet period can take longer; Part 4 covers that.

FAQ

Can one API allow more than one CORS origin?

Yes. List every allowed origin in the API's CORS setting. In this tutorial WEB_ORIGIN holds a comma-separated list, for example the production site and http://localhost:5173/.

How do I change a shared secret between services without downtime?

Let the receiving service accept both the old and the new secret, switch the calling service to the new one, then remove the old one. At no point does a caller hold a secret the receiver refuses.

Do I need to redeploy after changing an environment variable on Light Cloud?

No. Saving environment variables redeploys the service on its own. In my run the new values were live in 50 to 70 seconds.

Why does only one service redeploy when I push to a monorepo?

Each Light Cloud app has a Root directory. A push only redeploys the apps whose folder the push changed.

What should a service do when another service does not answer?

Stop waiting after a short timeout and return a clear error, such as 503 Service Unavailable, instead of leaving the customer's request hanging.


This article first appeared on the Light Cloud Blog: https://blog.light-cloud.com/tutorials/microservices-on-light-cloud-part-2-connect-services.

view more: ‹ prev next ›