Design Principles¶
Terminology note: This document uses the canonical terms defined in UBIQUITOUS_LANGUAGE.md. When in doubt about a term, refer there.
Numbering note: Principles are numbered for reference. New principles are appended at the end; existing numbers are never reordered.
1. Usage is derived, not specified¶
Resource consumption flows from high-level parameters (traffic, frequency, payload size) through a dependency graph. Individual resource usage is never a free variable — it is computed from the graph topology and propagation rules.
Academic basis: "Workload derivation" (CostHat, UCC 2016); "Symbolic cost expressions" parameterized by input features (Skyler, ASPLOS 2026).
2. The model is a directed acyclic graph¶
Services, functions, or infrastructure components are nodes. Edges represent invocation or data dependencies, weighted by call frequency and data-size ratios. The graph must be a DAG — cycles in cost derivation are undefined.
Academic basis: "Service call graph" (CostHat); "Workload dependency graph" (GARMA); "Serverless Economic Graph" (Skyler).
3. Cost propagation is compositional¶
The cost of the system is the sum of derived costs for each node. Each node's cost depends on its incoming workload and its per-unit cost factors. Propagation is bottom-up: from entry points through internal dependencies.
Academic basis: Cost propagation in CostHat §2.4; CostHat Equation 2:
C(ζ) = Σ c(s,ζ).
4. Parameters are first-class citizens¶
Traffic rates, payload sizes, call frequencies, and other workload parameters are symbolic variables in the model. They can be varied for what-if analysis without re-deriving the graph structure.
Academic basis: "Input-parameter sensitive function models" (Eismann et al., ICPE 2020); "Performance Model Parameters" (GARMA/Palladio).
5. Two inputs, one engine¶
The cost engine joins two separate intermediate representations:
- Resource representation — what infrastructure exists (from parsing .tf files, Pulumi exports, or CDK synth). Contains resource types, configurations, regions, and attributes.
- Cost model representation — how infrastructure is used (from YAML, TypeScript, or Python SDK). Contains DAG topology, call frequencies, and per-node usage metrics.
Changing your infrastructure (add an RDS instance) changes the Resource representation. Changing your assumptions (double traffic, add a region) changes the Cost model representation. The engine joins them: Cost model representation × Resource representation → derived usage × pricing = cost.
This means you can run the same cost model against different environments (dev, staging, prod) or run different scenarios (1 user vs 100K) against the same infrastructure.
6. The model is provider-agnostic¶
The dependency graph and usage derivation are independent of pricing. Provider-specific pricing is plugged in as a separate layer: derived usage × unit price = cost. This separates the what you use problem from the what it costs problem.
Academic basis: Skyler's pluggable pricing plugins; CostHat's separation of workload model from cost model.
7. The model supports sensitivity analysis¶
Given the graph and parameters, the model must answer: "Which parameter changes affect cost the most?" and "What happens to total cost if traffic doubles?" This requires symbolic or parametric representation, not just point estimates.
Academic basis: CostHat's what-if analysis (§4); Skyler's cost prediction queries (§6); GARMA's bounded best/worst-case estimation.
8. The model accommodates large language model token costs¶
Token consumption in large language model API calls follows the same propagation pattern as traditional cloud resources: input tokens flow into a node, processing produces output tokens, and those flow to downstream nodes. The model must not assume only traditional cloud metrics.
This is a novel contribution; existing academic work does not explicitly model large language model token economics in DAG-based cost derivation.
9. DAG-first UX with flat overrides as escape hatch¶
The DAG is the primary interface. Flat per-resource overrides exist for migration and edge cases but are explicitly discouraged:
- The DAG syntax (
→ resource: rate) reads like traffic flow, not a spreadsheet - Entry point + frequency drives all derived usage — one number change cascades
- Flat overrides require specifying every resource's usage independently
- Validation warns when flat overrides conflict with a known call chain
- A
graphcommand renders the DAG visually
The anti-pattern is Infracost's
infracost-usage.yml, where every usage value is a manual override with no relationships between resources.
Always-on resources and per-metric fixed flags¶
Always-on infrastructure (a load balancer, NAT gateway, a reserved instance) has frequency-independent cost. Two mechanisms support it without abusing the flat-override escape hatch:
- A node carrying any fixed cost is always-on: it is costed without a synthetic incoming edge and is not reported as unreachable. You no longer add a meaningless edge just to make a fixed resource reachable.
- A usage metric may set
fixed: trueso its value is a flat monthly total that does not scale with the derived invocation count. This lets one node carry both a fixed dimension and a usage-driven dimension (NAT gateway hours + GB processed; ALB-hours + LCUs; Secrets Manager secret-month + per-call) instead of splitting into two nodes.flatOverride: trueremains the all-or-nothing shorthand for "every metric is fixed".
The flat-vs-derived conflict warning fires only for the genuine case: a fully-fixed node (flatOverride, or every metric marked fixed) that also receives incoming DAG edges. A node mixing fixed and usage-driven metrics legitimately consumes its edges and does not warn.
10. Type-safe SDK from infrastructure-as-code type generation¶
Status: not implemented yet. The engine has no code generator from infrastructure-as-code schemas (#430). The one generation step that runs today turns
cost-model.schema.jsonintosdk/ts/src/types.generated.ts. This section describes the intended design.
The SDK generates types from your infrastructure definition, so you cannot reference non-existent resources or attributes:
graph LR
TF[.tf files] --> SCHEMA[terraform providers schema -json]
PULS[Pulumi] --> EXPORT[pulumi stack export --json]
CDK[CDK] --> SYNTH[cdk synth]
SCHEMA --> CG[codegen]
EXPORT --> CG
SYNTH --> CG
CG --> TYPES[types.ts]
Generated types provide:
- ResourceAddress — union type of all resource addresses in your project (autocomplete + compile-time errors)
- Per-resource usage params — a Lambda node accepts
compute,memory,invocations; a DynamoDB node acceptsreadUnits,writeUnits,storageGb - Node type metadata — compute nodes can call things; storage nodes are leaves; routing nodes can call compute nodes
// ✅ Valid — resource exists, usage params match node type
api.calls("aws_lambda_function.get_user", [
{ to: "aws_dynamodb_table.users", rate: 1, type: "read" },
]);
// ❌ Compile error — resource doesn't exist in your .tf files
api.calls("aws_s3_bucket.oops", []);
// ❌ Compile error — storageGb is not a Lambda usage param
api.usage("aws_lambda_function.get_user", { storageGb: "10" });
Code generation is the standard approach (CDKTF, Pulumi, AWS CDK all do it). No compiler extensions needed.
11. Three surfaces, one schema¶
Users declare cost models through three interfaces, all producing the same Cost model representation:
| Interface | For | Strengths |
|---|---|---|
| YAML | Terraform users, quick sketches, CI pipelines | Familiar, diffable, easy in reviews |
| TypeScript SDK | Pulumi/CDK users, complex scenarios | Type-safe, programmatic what-if loops, integrates with infrastructure-as-code code |
| Python SDK | Data teams, Jupyter notebooks, automation | Scriptable, pandas integration, sensitivity analysis |
All three share a JSON Schema as the single source of truth. The YAML validates against it. The SDKs generate types from it. The engine consumes it.
graph TD
YAML[YAML file<br/>validates against schema]
TS[TypeScript SDK<br/>types from .tf files]
PY[Python SDK<br/>types from .tf files]
YAML --> CMR[cost model representation<br/>JSON Schema]
TS --> CMR
PY --> CMR
CMR --> Engine[Cost Engine<br/>Python]
Engine --> Analysis[Sensitivity & Analysis<br/>pandas · matplotlib · Jupyter]
The YAML syntax for DAG definition:
workflow: my-api
entry: aws_api_gateway_rest_api.my_api
frequency: 1000/min
calls:
aws_api_gateway_rest_api.my_api:
data_out: 50KB
→ aws_lambda_function.get_user: 0.8
→ aws_lambda_function.create_user: 0.2
aws_lambda_function.get_user:
compute: 200ms
memory: 256MB
→ aws_dynamodb_table.users:
rate: 1
type: read
aws_lambda_function.create_user:
compute: 350ms
memory: 512MB
→ aws_dynamodb_table.users:
rate: 1
type: write
The TypeScript SDK mirrors the YAML one-to-one:
const api = Workflow.fromTf("my-api", "./infra/", {
entry: "aws_api_gateway_rest_api.my_api",
frequency: perMinute(1000),
});
api.calls("aws_api_gateway_rest_api.my_api", [
{ to: "aws_lambda_function.get_user", rate: 0.8 },
{ to: "aws_lambda_function.create_user", rate: 0.2 },
]);
The Python SDK mirrors both:
api = Workflow.from_tf("my-api", "./infra/",
entry="aws_api_gateway_rest_api.my_api",
frequency=per_minute(1000),
)
api.calls("aws_api_gateway_rest_api.my_api", [
Call(to="aws_lambda_function.get_user", rate=0.8),
Call(to="aws_lambda_function.create_user", rate=0.2),
])
Same mental model, three surfaces. Learn one, know all three.
12. The engine is Python¶
The cost engine — directed acyclic graph traversal, workload derivation, pricing, and sensitivity analysis — is implemented in Python. This follows from the other principles:
- Sensitivity analysis requires data tooling (Principle 7). pandas, numpy, matplotlib, Jupyter — no other language matches this for running 10,000 what-if scenarios and plotting cost curves.
- Three surfaces share a JSON Schema contract (Principle 11). The Python engine reads the cost model representation; TypeScript and YAML produce it. Correctness is ensured by shared test fixtures, not shared code.
- The core computation is small (~200 lines). Reimplementing in TypeScript for browser or IDE use is straightforward because the JSON Schema is the source of truth.
- YAML is a first-class surface (Principle 11). Python has the most mature handling (ruamel.yaml preserves comments and formatting).
If the TypeScript surface later needs a standalone engine, it reimplements the same core logic. Both must pass the same test fixtures — this is a correctness contract, not a code-sharing contract.
13. Pricing is queried, not hardcoded¶
Cloud prices change monthly. Hard-coded prices in source code go stale. The pricing layer exposes a query interface and the engine never knows or cares where the data comes from.
The catalog must handle tiered pricing (e.g., Lambda GB-seconds has three price tiers), free tiers (first 1M requests free), and regional variation across all services and regions. Prices are fetched live from the pricing source and refresh on a schedule — cloud pricing changes monthly at most. The bundled seed price list is a test fixture and offline fallback, not a normal setup step.
A provider plugin architecture is premature. Use a normalized multi-cloud source until it doesn't cover a needed provider.
14. A model states the engine it needs¶
An engine too old for a model does not fail. It silently prices what it does not understand, and the resulting total reads as reasonable. A model written against the SaaS shape vocabulary therefore reports $0 for every shaped node on an engine from before that feature, and exits 0. The model author has no signal that a term went missing.
The model states its requirement, and the engine refuses rather than guessing:
version: "1.0"
requiresEngine: ">=0.2.0"
- The value is a PEP 440 specifier, evaluated with
packaging, the library pip itself uses. A minimum, an exact pin, and an excluded range all work. - The check runs in
CostEngine.compute, so every caller hits the same gate: the CLI, the SDK, and any script that builds an engine. A mismatch raisesEngineRequirementError, aValueErrorsubclass, matching what callers already catch. validatereports the same mismatch, so a reader finds it before pricing.- A model that declares no requirement behaves exactly as before. The field is opt-in, and every model written before it keeps working.
The engine version is infra_cost_model.__version__, and pyproject.toml reads it from there. One number, one place. Bump the minor version when a model written for the new engine would price differently, or not at all, on the previous one. That is what a pin is worth pinning against.
The pin cannot help a model run on an engine older than the pin itself. An engine predating this field drops the key along with any other key it does not recognise, which is the failure the field exists to make visible on engines that have it.