From 1d80f59453b3bce7ca190b94f226384e95c27ca7 Mon Sep 17 00:00:00 2001
From: "renovate[bot]" <29139614+renovate[bot]@users.noreply.github.com>
Date: Tue, 14 Jul 2026 09:49:30 +0200
Subject: [PATCH 01/15] Renovate: Update debian:trixie-slim Docker digest to
020c0d2 (#1041)
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
This PR contains the following updates:
| Package | Type | Update | Change |
|---|---|---|---|
| debian | | digest | `28de087` → `020c0d2` |
| debian | final | digest | `28de087` → `020c0d2` |
---
### Configuration
📅 **Schedule**: (in timezone Europe/Berlin)
- Branch creation
- At any time (no schedule defined)
- Automerge
- At any time (no schedule defined)
🚦 **Automerge**: Disabled by config. Please merge this manually once you
are satisfied.
♻ **Rebasing**: Whenever PR becomes conflicted, or you tick the
rebase/retry checkbox.
🔕 **Ignore**: Close this PR and you won't be reminded about these
updates again.
---
- [ ] If you want to rebase/retry this PR, check
this box
---
This PR was generated by [Mend Renovate](https://mend.io/renovate/).
View the [repository job
log](https://developer.mend.io/github/cobaltcore-dev/cortex).
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
---
postgres/Dockerfile | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/postgres/Dockerfile b/postgres/Dockerfile
index 483c4ff84..706109e9b 100644
--- a/postgres/Dockerfile
+++ b/postgres/Dockerfile
@@ -1,4 +1,4 @@
-FROM debian:trixie-slim@sha256:28de0877c2189802884ccd20f15ee41c203573bd87bb6b883f5f46362d24c5c2
+FROM debian:trixie-slim@sha256:020c0d20b9880058cbe785a9db107156c3c75c2ac944a6aa7ab59f2add76a7bd
# explicitly set user/group IDs
RUN set -eux; \
From 7e36a65cbe6d348b09f41f1c84cb12622c41db3f Mon Sep 17 00:00:00 2001
From: "renovate[bot]" <29139614+renovate[bot]@users.noreply.github.com>
Date: Tue, 14 Jul 2026 09:49:56 +0200
Subject: [PATCH 02/15] Renovate: Update External dependencies (#1040)
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
This PR contains the following updates:
| Package | Change |
[Age](https://docs.renovatebot.com/merge-confidence/) |
[Adoption](https://docs.renovatebot.com/merge-confidence/) |
[Passing](https://docs.renovatebot.com/merge-confidence/) |
[Confidence](https://docs.renovatebot.com/merge-confidence/) | Type |
Update |
|---|---|---|---|---|---|---|---|
|
[github.com/mattn/go-sqlite3](https://redirect.github.com/mattn/go-sqlite3)
| `v1.14.47` → `v1.14.48` |

|

|

|

| require | patch |
|
[kube-prometheus-stack](https://redirect.github.com/prometheus-operator/kube-prometheus)
([source](https://redirect.github.com/prometheus-community/helm-charts))
| `87.15.1` → `87.15.2` |

|

|

|

| | patch |
---
### Release Notes
mattn/go-sqlite3 (github.com/mattn/go-sqlite3)
###
[`v1.14.48`](https://redirect.github.com/mattn/go-sqlite3/releases/tag/v1.14.48):
1.14.48
[Compare
Source](https://redirect.github.com/mattn/go-sqlite3/compare/v1.14.47...v1.14.48)
#### What's Changed
- Add Serialize and Deserialize support by
[@otoolep](https://redirect.github.com/otoolep) in
[#1089](https://redirect.github.com/mattn/go-sqlite3/pull/1089)
- Replace namedValue with driver.NamedValue to avoid copying exec/query
args by [@charlievieth](https://redirect.github.com/charlievieth)
in
[#1128](https://redirect.github.com/mattn/go-sqlite3/pull/1128)
- Add go 1.20 to workflow matrix, remove 1.17 by
[@connyay](https://redirect.github.com/connyay) in
[#1136](https://redirect.github.com/mattn/go-sqlite3/pull/1136)
- Add build tags to support both x86 and ARM compilation on macOS by
[@Spaider](https://redirect.github.com/Spaider) in
[#1069](https://redirect.github.com/mattn/go-sqlite3/pull/1069)
- Fix virtual table example. by
[@andrzh](https://redirect.github.com/andrzh) in
[#1149](https://redirect.github.com/mattn/go-sqlite3/pull/1149)
- Update README.md by
[@parthokr](https://redirect.github.com/parthokr) in
[#1163](https://redirect.github.com/mattn/go-sqlite3/pull/1163)
- Update amalgamation code by
[@mattn](https://redirect.github.com/mattn) in
[#1166](https://redirect.github.com/mattn/go-sqlite3/pull/1166)
- Update amalgamation code by
[@mattn](https://redirect.github.com/mattn) in
[#1197](https://redirect.github.com/mattn/go-sqlite3/pull/1197)
- Fix docker job by [@itizir](https://redirect.github.com/itizir)
in
[#1201](https://redirect.github.com/mattn/go-sqlite3/pull/1201)
- Fix musl build
([#1164](https://redirect.github.com/mattn/go-sqlite3/issues/1164))
by [@leso-kn](https://redirect.github.com/leso-kn) in
[#1177](https://redirect.github.com/mattn/go-sqlite3/pull/1177)
- update go version to 1.19 by
[@mattn](https://redirect.github.com/mattn) in
[#1208](https://redirect.github.com/mattn/go-sqlite3/pull/1208)
- Update amalgamation code to 3.45.0 by
[@mattn](https://redirect.github.com/mattn) in
[#1207](https://redirect.github.com/mattn/go-sqlite3/pull/1207)
- Update amalgamation code to 3.45.1 by
[@mattn](https://redirect.github.com/mattn) in
[#1211](https://redirect.github.com/mattn/go-sqlite3/pull/1211)
- close channel by [@mattn](https://redirect.github.com/mattn) in
[#1213](https://redirect.github.com/mattn/go-sqlite3/pull/1213)
- fix: some typos by
[@pomadev](https://redirect.github.com/pomadev) in
[#1222](https://redirect.github.com/mattn/go-sqlite3/pull/1222)
- Add support for libsqlite3 on z/OS by
[@dustin-ward](https://redirect.github.com/dustin-ward) in
[#1239](https://redirect.github.com/mattn/go-sqlite3/pull/1239)
- Update amalgamation code to 3.46.1 by
[@mattn](https://redirect.github.com/mattn) in
[#1273](https://redirect.github.com/mattn/go-sqlite3/pull/1273)
- close statement when missing query arguments by
[@mattn](https://redirect.github.com/mattn) in
[#1281](https://redirect.github.com/mattn/go-sqlite3/pull/1281)
- Upgrade upload-artifact action by
[@jonstacks](https://redirect.github.com/jonstacks) in
[#1300](https://redirect.github.com/mattn/go-sqlite3/pull/1300)
- Remove suggestion that CGO isn't always needed by
[@samjewell](https://redirect.github.com/samjewell) in
[#1290](https://redirect.github.com/mattn/go-sqlite3/pull/1290)
- remove superfluous use of runtime.SetFinalizer on SQLiteRows by
[@charlievieth](https://redirect.github.com/charlievieth) in
[#1301](https://redirect.github.com/mattn/go-sqlite3/pull/1301)
- Fix sqlite3\_opt\_unlock\_notify with USE\_LIBSQLITE3 by
[@q66](https://redirect.github.com/q66) in
[#1262](https://redirect.github.com/mattn/go-sqlite3/pull/1262)
- Fix memory leak in callbackRetText function by
[@hionay](https://redirect.github.com/hionay) in
[#1259](https://redirect.github.com/mattn/go-sqlite3/pull/1259)
- docs: clarify GCP section by
[@justinsb](https://redirect.github.com/justinsb) in
[#1305](https://redirect.github.com/mattn/go-sqlite3/pull/1305)
- Add ability to set an int64 file control by
[@jonstacks](https://redirect.github.com/jonstacks) in
[#1298](https://redirect.github.com/mattn/go-sqlite3/pull/1298)
- Update amalgamation code to 3.49.1 by
[@mattn](https://redirect.github.com/mattn) in
[#1335](https://redirect.github.com/mattn/go-sqlite3/pull/1335)
- Update amalgamation code to 3.50.3 by
[@mattn](https://redirect.github.com/mattn) in
[#1343](https://redirect.github.com/mattn/go-sqlite3/pull/1343)
- Drop userauth implementation by
[@mattn](https://redirect.github.com/mattn) in
[#1344](https://redirect.github.com/mattn/go-sqlite3/pull/1344)
- fix syntax error by
[@eraytufan](https://redirect.github.com/eraytufan) in
[#1346](https://redirect.github.com/mattn/go-sqlite3/pull/1346)
- update amalgamation code by
[@mattn](https://redirect.github.com/mattn) in
[#1347](https://redirect.github.com/mattn/go-sqlite3/pull/1347)
- use quote include instead of angled include for sqlite3-binding.h by
[@nautaa](https://redirect.github.com/nautaa) in
[#1362](https://redirect.github.com/mattn/go-sqlite3/pull/1362)
- Upgrade SQLite to version
[`3051001`](https://redirect.github.com/mattn/go-sqlite3/commit/3051001)
by [@mattn](https://redirect.github.com/mattn) in
[#1366](https://redirect.github.com/mattn/go-sqlite3/pull/1366)
- Feat: add percentile extension option by
[@dsonck92](https://redirect.github.com/dsonck92) in
[#1364](https://redirect.github.com/mattn/go-sqlite3/pull/1364)
- Upgrade SQLite to version
[`3051002`](https://redirect.github.com/mattn/go-sqlite3/commit/3051002)
by [@mattn](https://redirect.github.com/mattn) in
[#1370](https://redirect.github.com/mattn/go-sqlite3/pull/1370)
- Use unsafe slice by [@mattn](https://redirect.github.com/mattn)
in
[#1373](https://redirect.github.com/mattn/go-sqlite3/pull/1373)
- Call sqlite3\_clear\_bindings() after sqlite3\_reset() in bind() by
[@mattn](https://redirect.github.com/mattn) in
[#1374](https://redirect.github.com/mattn/go-sqlite3/pull/1374)
- Upgrade SQLite to version
[`3051003`](https://redirect.github.com/mattn/go-sqlite3/commit/3051003)
by [@mattn](https://redirect.github.com/mattn) in
[#1375](https://redirect.github.com/mattn/go-sqlite3/pull/1375)
- Ensure Close always removes runtime finalizer to prevent memory leak
by [@mattn](https://redirect.github.com/mattn) in
[#1376](https://redirect.github.com/mattn/go-sqlite3/pull/1376)
- Fix json example by
[@Jaculabilis](https://redirect.github.com/Jaculabilis) in
[#1313](https://redirect.github.com/mattn/go-sqlite3/pull/1313)
- Add missing virtual table constraint op constants by
[@theimpostor](https://redirect.github.com/theimpostor) in
[#1379](https://redirect.github.com/mattn/go-sqlite3/pull/1379)
- Eliminate unnecessary bounds checks in hot paths by
[@mattn](https://redirect.github.com/mattn) in
[#1381](https://redirect.github.com/mattn/go-sqlite3/pull/1381)
- \[codex] optimize sqlite bind fast path by
[@mattn](https://redirect.github.com/mattn) in
[#1382](https://redirect.github.com/mattn/go-sqlite3/pull/1382)
- \[codex] batch row column fetches in Next by
[@mattn](https://redirect.github.com/mattn) in
[#1383](https://redirect.github.com/mattn/go-sqlite3/pull/1383)
- Raise minimum Go version to 1.21 by
[@mattn](https://redirect.github.com/mattn) in
[#1384](https://redirect.github.com/mattn/go-sqlite3/pull/1384)
- Reduce sqlite bind overhead by
[@mattn](https://redirect.github.com/mattn) in
[#1385](https://redirect.github.com/mattn/go-sqlite3/pull/1385)
- reduce CGO call overhead for exec and bind paths by
[@mattn](https://redirect.github.com/mattn) in
[#1386](https://redirect.github.com/mattn/go-sqlite3/pull/1386)
- \[codex] add opt-in statement cache by
[@mattn](https://redirect.github.com/mattn) in
[#1387](https://redirect.github.com/mattn/go-sqlite3/pull/1387)
- Fix panic when querying input with no SQL (only comments/whitespace)
by [@mattn](https://redirect.github.com/mattn) in
[#1392](https://redirect.github.com/mattn/go-sqlite3/pull/1392)
- evict least-recently-used stmt when cache is full by
[@mattn](https://redirect.github.com/mattn) in
[#1388](https://redirect.github.com/mattn/go-sqlite3/pull/1388)
- Upgrade SQLite to version
[`3053000`](https://redirect.github.com/mattn/go-sqlite3/commit/3053000)
by [@mattn](https://redirect.github.com/mattn) in
[#1394](https://redirect.github.com/mattn/go-sqlite3/pull/1394)
- add sqlite\_dbstat tag for the DBSTAT virtual table by
[@calmh](https://redirect.github.com/calmh) in
[#1338](https://redirect.github.com/mattn/go-sqlite3/pull/1338)
- avoid out of bounds write in unlock\_notify\_wait on 64 bit platforms
by [@calmh](https://redirect.github.com/calmh) in
[#1399](https://redirect.github.com/mattn/go-sqlite3/pull/1399)
- modernise reflect.SliceHeader to unsafe.Slice by
[@calmh](https://redirect.github.com/calmh) in
[#1400](https://redirect.github.com/mattn/go-sqlite3/pull/1400)
- guard oversized string length in ResultText by
[@dxbjavid](https://redirect.github.com/dxbjavid) in
[#1402](https://redirect.github.com/mattn/go-sqlite3/pull/1402)
- bind via sqlite3\_bind\_text64/blob64 to avoid 32-bit length
truncation by [@dxbjavid](https://redirect.github.com/dxbjavid)
in
[#1403](https://redirect.github.com/mattn/go-sqlite3/pull/1403)
- Upgrade SQLite to version
[`3053002`](https://redirect.github.com/mattn/go-sqlite3/commit/3053002)
by [@mattn](https://redirect.github.com/mattn) in
[#1404](https://redirect.github.com/mattn/go-sqlite3/pull/1404)
- guard oversized blob length in callbackRetBlob by
[@dxbjavid](https://redirect.github.com/dxbjavid) in
[#1405](https://redirect.github.com/mattn/go-sqlite3/pull/1405)
- preserve embedded NUL bytes in custom function text values by
[@dxbjavid](https://redirect.github.com/dxbjavid) in
[#1406](https://redirect.github.com/mattn/go-sqlite3/pull/1406)
- Follow documented call order for sqlite3\_value\_blob in
callbackArgString by [@mattn](https://redirect.github.com/mattn)
in
[#1407](https://redirect.github.com/mattn/go-sqlite3/pull/1407)
- Use atomic.Value for handle table and add concurrent lookup benchmark
by [@mattn](https://redirect.github.com/mattn) in
[#1412](https://redirect.github.com/mattn/go-sqlite3/pull/1412)
- cache column metadata for prepared and cached statements by
[@mattn](https://redirect.github.com/mattn) in
[#1413](https://redirect.github.com/mattn/go-sqlite3/pull/1413)
- free leaked schema string in GetFilename by
[@dxbjavid](https://redirect.github.com/dxbjavid) in
[#1408](https://redirect.github.com/mattn/go-sqlite3/pull/1408)
- Fix race in SQLiteStmt.Close by holding conn lock across cache check
by [@mattn](https://redirect.github.com/mattn) in
[#1416](https://redirect.github.com/mattn/go-sqlite3/pull/1416)
- Add CodeRabbit as a sponsor by
[@mattn](https://redirect.github.com/mattn) in
[#1417](https://redirect.github.com/mattn/go-sqlite3/pull/1417)
- Return error from vtable cursor open instead of ignoring it by
[@mattn](https://redirect.github.com/mattn) in
[#1419](https://redirect.github.com/mattn/go-sqlite3/pull/1419)
- Check sqlite3\_malloc64 result in Deserialize by
[@mattn](https://redirect.github.com/mattn) in
[#1420](https://redirect.github.com/mattn/go-sqlite3/pull/1420)
- Fix panic when registered functions return named types by
[@mattn](https://redirect.github.com/mattn) in
[#1421](https://redirect.github.com/mattn/go-sqlite3/pull/1421)
- Return error instead of silently ignoring unsupported bind types by
[@mattn](https://redirect.github.com/mattn) in
[#1422](https://redirect.github.com/mattn/go-sqlite3/pull/1422)
- Add CodeRabbit configuration by
[@mattn](https://redirect.github.com/mattn) in
[#1418](https://redirect.github.com/mattn/go-sqlite3/pull/1418)
- Close database on all error paths in Open by
[@mattn](https://redirect.github.com/mattn) in
[#1423](https://redirect.github.com/mattn/go-sqlite3/pull/1423)
- Check preupdate value fetch result to avoid NULL dereference by
[@mattn](https://redirect.github.com/mattn) in
[#1424](https://redirect.github.com/mattn/go-sqlite3/pull/1424)
- Use C.int in exported callbacks to match C declarations by
[@mattn](https://redirect.github.com/mattn) in
[#1425](https://redirect.github.com/mattn/go-sqlite3/pull/1425)
- Fix leak of extension load error message by
[@mattn](https://redirect.github.com/mattn) in
[#1426](https://redirect.github.com/mattn/go-sqlite3/pull/1426)
- Upgrade SQLite to version
[`3053003`](https://redirect.github.com/mattn/go-sqlite3/commit/3053003)
by [@mattn](https://redirect.github.com/mattn) in
[#1427](https://redirect.github.com/mattn/go-sqlite3/pull/1427)
#### New Contributors
- [@charlievieth](https://redirect.github.com/charlievieth) made
their first contribution in
[#1128](https://redirect.github.com/mattn/go-sqlite3/pull/1128)
- [@connyay](https://redirect.github.com/connyay) made their
first contribution in
[#1136](https://redirect.github.com/mattn/go-sqlite3/pull/1136)
- [@Spaider](https://redirect.github.com/Spaider) made their
first contribution in
[#1069](https://redirect.github.com/mattn/go-sqlite3/pull/1069)
- [@andrzh](https://redirect.github.com/andrzh) made their first
contribution in
[#1149](https://redirect.github.com/mattn/go-sqlite3/pull/1149)
- [@parthokr](https://redirect.github.com/parthokr) made their
first contribution in
[#1163](https://redirect.github.com/mattn/go-sqlite3/pull/1163)
- [@leso-kn](https://redirect.github.com/leso-kn) made their
first contribution in
[#1177](https://redirect.github.com/mattn/go-sqlite3/pull/1177)
- [@pomadev](https://redirect.github.com/pomadev) made their
first contribution in
[#1222](https://redirect.github.com/mattn/go-sqlite3/pull/1222)
- [@dustin-ward](https://redirect.github.com/dustin-ward) made
their first contribution in
[#1239](https://redirect.github.com/mattn/go-sqlite3/pull/1239)
- [@jonstacks](https://redirect.github.com/jonstacks) made their
first contribution in
[#1300](https://redirect.github.com/mattn/go-sqlite3/pull/1300)
- [@samjewell](https://redirect.github.com/samjewell) made their
first contribution in
[#1290](https://redirect.github.com/mattn/go-sqlite3/pull/1290)
- [@q66](https://redirect.github.com/q66) made their first
contribution in
[#1262](https://redirect.github.com/mattn/go-sqlite3/pull/1262)
- [@hionay](https://redirect.github.com/hionay) made their first
contribution in
[#1259](https://redirect.github.com/mattn/go-sqlite3/pull/1259)
- [@justinsb](https://redirect.github.com/justinsb) made their
first contribution in
[#1305](https://redirect.github.com/mattn/go-sqlite3/pull/1305)
- [@eraytufan](https://redirect.github.com/eraytufan) made their
first contribution in
[#1346](https://redirect.github.com/mattn/go-sqlite3/pull/1346)
- [@nautaa](https://redirect.github.com/nautaa) made their first
contribution in
[#1362](https://redirect.github.com/mattn/go-sqlite3/pull/1362)
- [@dsonck92](https://redirect.github.com/dsonck92) made their
first contribution in
[#1364](https://redirect.github.com/mattn/go-sqlite3/pull/1364)
- [@Jaculabilis](https://redirect.github.com/Jaculabilis) made
their first contribution in
[#1313](https://redirect.github.com/mattn/go-sqlite3/pull/1313)
- [@theimpostor](https://redirect.github.com/theimpostor) made
their first contribution in
[#1379](https://redirect.github.com/mattn/go-sqlite3/pull/1379)
- [@calmh](https://redirect.github.com/calmh) made their first
contribution in
[#1338](https://redirect.github.com/mattn/go-sqlite3/pull/1338)
- [@dxbjavid](https://redirect.github.com/dxbjavid) made their
first contribution in
[#1402](https://redirect.github.com/mattn/go-sqlite3/pull/1402)
**Full Changelog**:
prometheus-community/helm-charts
(kube-prometheus-stack)
###
[`v87.15.2`](https://redirect.github.com/prometheus-community/helm-charts/releases/tag/kube-prometheus-stack-87.15.2)
[Compare
Source](https://redirect.github.com/prometheus-community/helm-charts/compare/kube-prometheus-stack-87.15.1...kube-prometheus-stack-87.15.2)
kube-prometheus-stack collects Kubernetes manifests, Grafana dashboards,
and Prometheus rules combined with documentation and scripts to provide
easy to operate end-to-end Kubernetes cluster monitoring with Prometheus
using the Prometheus Operator.
#### What's Changed
- \[kube-prometheus-stack] fix externalUrl of alertmanager and
prometheus when route is configured by
[@DanielRaapDev](https://redirect.github.com/DanielRaapDev) in
[#7107](https://redirect.github.com/prometheus-community/helm-charts/pull/7107)
#### New Contributors
- [@DanielRaapDev](https://redirect.github.com/DanielRaapDev)
made their first contribution in
[#7107](https://redirect.github.com/prometheus-community/helm-charts/pull/7107)
**Full Changelog**:
---
### Configuration
📅 **Schedule**: (in timezone Europe/Berlin)
- Branch creation
- "after 6pm every weekday,every weekend,before 8am every weekday"
- Automerge
- At any time (no schedule defined)
🚦 **Automerge**: Enabled.
♻ **Rebasing**: Whenever PR is behind base branch, or you tick the
rebase/retry checkbox.
👻 **Immortal**: This PR will be recreated if closed unmerged. Get
[config
help](https://redirect.github.com/renovatebot/renovate/discussions) if
that's undesired.
---
- [ ] If you want to rebase/retry this PR, check
this box
---
This PR was generated by [Mend Renovate](https://mend.io/renovate/).
View the [repository job
log](https://developer.mend.io/github/cobaltcore-dev/cortex).
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
---
go.mod | 2 +-
go.sum | 4 ++--
helm/dev/cortex-prometheus-operator/Chart.yaml | 2 +-
3 files changed, 4 insertions(+), 4 deletions(-)
diff --git a/go.mod b/go.mod
index fcdc8942b..5d1606e51 100644
--- a/go.mod
+++ b/go.mod
@@ -73,7 +73,7 @@ require (
github.com/json-iterator/go v1.1.12 // indirect
github.com/kylelemons/godebug v1.1.0 // indirect
github.com/lib/pq v1.12.3
- github.com/mattn/go-sqlite3 v1.14.47
+ github.com/mattn/go-sqlite3 v1.14.48
github.com/moby/sys/user v0.4.0 // indirect
github.com/modern-go/concurrent v0.0.0-20180306012644-bacd9c7ef1dd // indirect
github.com/modern-go/reflect2 v1.0.3-0.20250322232337-35a7c28c31ee // indirect
diff --git a/go.sum b/go.sum
index d1011eae6..7d4fbb9a4 100644
--- a/go.sum
+++ b/go.sum
@@ -153,8 +153,8 @@ github.com/kylelemons/godebug v1.1.0 h1:RPNrshWIDI6G2gRW9EHilWtl7Z6Sb1BR0xunSBf0
github.com/kylelemons/godebug v1.1.0/go.mod h1:9/0rRGxNHcop5bhtWyNeEfOS8JIWk580+fNqagV/RAw=
github.com/lib/pq v1.12.3 h1:tTWxr2YLKwIvK90ZXEw8GP7UFHtcbTtty8zsI+YjrfQ=
github.com/lib/pq v1.12.3/go.mod h1:/p+8NSbOcwzAEI7wiMXFlgydTwcgTr3OSKMsD2BitpA=
-github.com/mattn/go-sqlite3 v1.14.47 h1:jOBI62gS7nKeZv+as1oGEy0+1qISgXwH/QBlR6KbfIo=
-github.com/mattn/go-sqlite3 v1.14.47/go.mod h1:6JTjA44L93a0QCyJef5YvlPoKXntQPjzWv5gtm9sB6w=
+github.com/mattn/go-sqlite3 v1.14.48 h1:7XHIgl0a8HwOaiK4E47ozLkST78rR9+OtNGx27D/TFs=
+github.com/mattn/go-sqlite3 v1.14.48/go.mod h1:6JTjA44L93a0QCyJef5YvlPoKXntQPjzWv5gtm9sB6w=
github.com/moby/docker-image-spec v1.3.1 h1:jMKff3w6PgbfSa69GfNg+zN/XLhfXJGnEx3Nl2EsFP0=
github.com/moby/docker-image-spec v1.3.1/go.mod h1:eKmb5VW8vQEh/BAr2yvVNvuiJuY6UIocYsFu/DxxRpo=
github.com/moby/sys/user v0.4.0 h1:jhcMKit7SA80hivmFJcbB1vqmw//wU61Zdui2eQXuMs=
diff --git a/helm/dev/cortex-prometheus-operator/Chart.yaml b/helm/dev/cortex-prometheus-operator/Chart.yaml
index ca8ea75f5..304a42a57 100644
--- a/helm/dev/cortex-prometheus-operator/Chart.yaml
+++ b/helm/dev/cortex-prometheus-operator/Chart.yaml
@@ -10,4 +10,4 @@ dependencies:
# CRDs of the prometheus operator, such as PrometheusRule, ServiceMonitor, etc.
- name: kube-prometheus-stack
repository: oci://ghcr.io/prometheus-community/charts
- version: 87.15.1
+ version: 87.15.2
From e06153f81457a66a1781eb91c5965e139a07d2ef Mon Sep 17 00:00:00 2001
From: "github-actions[bot]"
<41898282+github-actions[bot]@users.noreply.github.com>
Date: Tue, 14 Jul 2026 09:51:20 +0200
Subject: [PATCH 03/15] fix(postgres): rebuild image to resolve CVEs (#1033)
The daily CVE scan detected fixable vulnerabilities in the published
`cortex-postgres` image. A test rebuild confirms that rebuilding
reduces the CVE count (via `apt-get upgrade` picking up security
patches).
Merging this PR triggers the image rebuild and publish pipeline.
This PR was created automatically by the `rebuild-postgres` workflow.
Co-authored-by: umswmayj <140147670+umswmayj@users.noreply.github.com>
---
postgres/rebuild-trigger | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/postgres/rebuild-trigger b/postgres/rebuild-trigger
index 443bf050c..3fbb2fb47 100644
--- a/postgres/rebuild-trigger
+++ b/postgres/rebuild-trigger
@@ -1 +1 @@
-27331309223
+29312059711
From e720bca4df2614766e5e026fedac6dc7e11d4ce3 Mon Sep 17 00:00:00 2001
From: "renovate[bot]" <29139614+renovate[bot]@users.noreply.github.com>
Date: Wed, 15 Jul 2026 10:28:09 +0200
Subject: [PATCH 04/15] Renovate: Update kube-prometheus-stack Docker tag to
v87.16.0 (#1046)
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
This PR contains the following updates:
| Package | Update | Change |
|---|---|---|
|
[kube-prometheus-stack](https://redirect.github.com/prometheus-operator/kube-prometheus)
([source](https://redirect.github.com/prometheus-community/helm-charts))
| minor | `87.15.2` → `87.16.1` |
---
### Release Notes
prometheus-community/helm-charts
(kube-prometheus-stack)
###
[`v87.16.1`](https://redirect.github.com/prometheus-community/helm-charts/releases/tag/kube-prometheus-stack-87.16.1)
kube-prometheus-stack collects Kubernetes manifests, Grafana dashboards,
and Prometheus rules combined with documentation and scripts to provide
easy to operate end-to-end Kubernetes cluster monitoring with Prometheus
using the Prometheus Operator.
#### What's Changed
- \[kube-prometheus-stack] Update Helm release prometheus-node-exporter
to v4.56.1 by
[@renovate](https://redirect.github.com/renovate)\[bot] in
[#7114](https://redirect.github.com/prometheus-community/helm-charts/pull/7114)
**Full Changelog**:
###
[`v87.16.0`](https://redirect.github.com/prometheus-community/helm-charts/releases/tag/kube-prometheus-stack-87.16.0)
kube-prometheus-stack collects Kubernetes manifests, Grafana dashboards,
and Prometheus rules combined with documentation and scripts to provide
easy to operate end-to-end Kubernetes cluster monitoring with Prometheus
using the Prometheus Operator.
#### What's Changed
- \[kube-prometheus-stack] Update kube-prometheus-stack dependency
non-major updates by
[@renovate](https://redirect.github.com/renovate)\[bot] in
[#7110](https://redirect.github.com/prometheus-community/helm-charts/pull/7110)
**Full Changelog**:
---
### Configuration
📅 **Schedule**: (in timezone Europe/Berlin)
- Branch creation
- "after 6pm every weekday,every weekend,before 8am every weekday"
- Automerge
- At any time (no schedule defined)
🚦 **Automerge**: Enabled.
♻ **Rebasing**: Whenever PR is behind base branch, or you tick the
rebase/retry checkbox.
🔕 **Ignore**: Close this PR and you won't be reminded about this update
again.
---
- [ ] If you want to rebase/retry this PR, check
this box
---
This PR was generated by [Mend Renovate](https://mend.io/renovate/).
View the [repository job
log](https://developer.mend.io/github/cobaltcore-dev/cortex).
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
---
helm/dev/cortex-prometheus-operator/Chart.yaml | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/helm/dev/cortex-prometheus-operator/Chart.yaml b/helm/dev/cortex-prometheus-operator/Chart.yaml
index 304a42a57..250b1f1a7 100644
--- a/helm/dev/cortex-prometheus-operator/Chart.yaml
+++ b/helm/dev/cortex-prometheus-operator/Chart.yaml
@@ -10,4 +10,4 @@ dependencies:
# CRDs of the prometheus operator, such as PrometheusRule, ServiceMonitor, etc.
- name: kube-prometheus-stack
repository: oci://ghcr.io/prometheus-community/charts
- version: 87.15.2
+ version: 87.16.1
From c677c29ee6a665cc616fa6d8b1dfc085e04be740 Mon Sep 17 00:00:00 2001
From: Markus Wieland <44964229+SoWieMarkus@users.noreply.github.com>
Date: Wed, 15 Jul 2026 11:13:35 +0200
Subject: [PATCH 05/15] feat: enhance multicluster client 'no cluster matched'
error log messages (#1043)
- Add `extractClusterSelector` method to routers that get the attribute
to match on based on the provided object
- Use this method for the error log message
```txt
# Before
no cluster matched for GVK cortex.cloud/v1alpha1, Kind=History
# Now
no cluster matched for GVK cortex.cloud/v1alpha1, Kind=History: cluster selector `qa-ber-1` did not match any candidate
```
---
pkg/multicluster/client.go | 17 ++++-
pkg/multicluster/client_test.go | 36 +++++----
pkg/multicluster/routers.go | 125 ++++++++++++++++++++++++++++++++
3 files changed, 161 insertions(+), 17 deletions(-)
diff --git a/pkg/multicluster/client.go b/pkg/multicluster/client.go
index a8e597c8d..979258a44 100644
--- a/pkg/multicluster/client.go
+++ b/pkg/multicluster/client.go
@@ -245,9 +245,19 @@ func (c *Client) clusterForWrite(gvk schema.GroupVersionKind, obj any) (cluster.
return r.cluster, nil
}
}
+ // No remote match — fall back to home if the GVK is also configured there.
+ if c.homeGVKs[gvk] {
+ return c.HomeCluster, nil
+ }
+ // No match and no home fallback — return an error.
+ selector, err := router.extractClusterSelector(obj)
+ if err != nil {
+ return nil, fmt.Errorf("failed to extract cluster selector for GVK %s: %w", gvk, err)
+ }
+ return nil, &NoClusterMatchedError{GVK: gvk, ClusterSelector: selector}
}
- // If we couldn't find a matching remote cluster (not configured or not found) but the GVK is configured for home, return the home cluster.
+ // No remotes configured for this GVK — fall back to home if available.
if c.homeGVKs[gvk] {
return c.HomeCluster, nil
}
@@ -272,11 +282,12 @@ func IsDuplicateError(err error) bool {
// happens in multi-AZ setups where the resource targets an AZ that has no
// configured cluster.
type NoClusterMatchedError struct {
- GVK schema.GroupVersionKind
+ GVK schema.GroupVersionKind
+ ClusterSelector string
}
func (e *NoClusterMatchedError) Error() string {
- return fmt.Sprintf("no cluster matched for GVK %s", e.GVK)
+ return fmt.Sprintf("no cluster matched for GVK %s: cluster selector %q did not match any candidate", e.GVK, e.ClusterSelector)
}
// IsNoClusterMatchedError returns true if the error indicates that no
diff --git a/pkg/multicluster/client_test.go b/pkg/multicluster/client_test.go
index 6ed7215a3..6e965d440 100644
--- a/pkg/multicluster/client_test.go
+++ b/pkg/multicluster/client_test.go
@@ -132,29 +132,37 @@ func newTestScheme(t *testing.T) *runtime.Scheme {
// testRouter is a simple ResourceRouter for testing.
type testRouter struct{}
-// alwaysMatchRouter matches any object to any cluster.
-type alwaysMatchRouter struct{}
-
-func (r alwaysMatchRouter) Match(any, map[string]string) (bool, error) {
- return true, nil
-}
-
func (r testRouter) Match(obj any, labels map[string]string) (bool, error) {
- cm, ok := obj.(*corev1.ConfigMap)
- if !ok {
- return false, nil
- }
az, ok := labels["az"]
if !ok {
return false, nil
}
- objAZ, ok := cm.Labels["az"]
- if !ok {
- return false, nil
+ objAZ, err := r.extractClusterSelector(obj)
+ if err != nil {
+ return false, err
}
return objAZ == az, nil
}
+func (r testRouter) extractClusterSelector(obj any) (string, error) {
+ cm, ok := obj.(*corev1.ConfigMap)
+ if !ok {
+ return "", errors.New("object is not a ConfigMap")
+ }
+ return cm.Labels["az"], nil
+}
+
+// alwaysMatchRouter matches any object to any cluster.
+type alwaysMatchRouter struct{}
+
+func (r alwaysMatchRouter) Match(any, map[string]string) (bool, error) {
+ return true, nil
+}
+
+func (r alwaysMatchRouter) extractClusterSelector(obj any) (string, error) {
+ return "", nil
+}
+
var configMapGVK = schema.GroupVersionKind{Group: "", Version: "v1", Kind: "ConfigMap"}
var configMapListGVK = schema.GroupVersionKind{Group: "", Version: "v1", Kind: "ConfigMapList"}
var podGVK = schema.GroupVersionKind{Group: "", Version: "v1", Kind: "Pod"}
diff --git a/pkg/multicluster/routers.go b/pkg/multicluster/routers.go
index fdbd251dc..4432e2a02 100644
--- a/pkg/multicluster/routers.go
+++ b/pkg/multicluster/routers.go
@@ -28,11 +28,36 @@ var DefaultResourceRouters = map[schema.GroupVersionKind]ResourceRouter{
// by matching the resource content against the cluster's labels.
type ResourceRouter interface {
Match(obj any, labels map[string]string) (bool, error)
+ // extractClusterSelector extracts the routing key from the object (e.g. availability zone).
+ // Used to enrich error messages when no cluster matches.
+ extractClusterSelector(obj any) (string, error)
}
// HypervisorResourceRouter routes hypervisors to clusters based on availability zone.
type HypervisorResourceRouter struct{}
+func (h HypervisorResourceRouter) extractClusterSelector(obj any) (string, error) {
+ switch v := obj.(type) {
+ case *hv1.Hypervisor:
+ if v == nil {
+ return "", errors.New("object is nil")
+ }
+ az, ok := v.Labels[corev1.LabelTopologyZone]
+ if !ok {
+ return "", errors.New("hypervisor does not have availability zone label")
+ }
+ return az, nil
+ case hv1.Hypervisor:
+ az, ok := v.Labels[corev1.LabelTopologyZone]
+ if !ok {
+ return "", errors.New("hypervisor does not have availability zone label")
+ }
+ return az, nil
+ default:
+ return "", errors.New("object is not a Hypervisor")
+ }
+}
+
func (h HypervisorResourceRouter) Match(obj any, labels map[string]string) (bool, error) {
var hv hv1.Hypervisor
@@ -61,6 +86,26 @@ func (h HypervisorResourceRouter) Match(obj any, labels map[string]string) (bool
// ReservationsResourceRouter routes reservations to clusters based on availability zone.
type ReservationsResourceRouter struct{}
+func (r ReservationsResourceRouter) extractClusterSelector(obj any) (string, error) {
+ switch v := obj.(type) {
+ case *v1alpha1.Reservation:
+ if v == nil {
+ return "", errors.New("object is nil")
+ }
+ if v.Spec.AvailabilityZone == "" {
+ return "", errors.New("reservation does not have availability zone in spec")
+ }
+ return v.Spec.AvailabilityZone, nil
+ case v1alpha1.Reservation:
+ if v.Spec.AvailabilityZone == "" {
+ return "", errors.New("reservation does not have availability zone in spec")
+ }
+ return v.Spec.AvailabilityZone, nil
+ default:
+ return "", errors.New("object is not a Reservation")
+ }
+}
+
func (r ReservationsResourceRouter) Match(obj any, labels map[string]string) (bool, error) {
var res v1alpha1.Reservation
@@ -89,6 +134,26 @@ func (r ReservationsResourceRouter) Match(obj any, labels map[string]string) (bo
// CommittedResourceRouter routes committed resources to clusters based on availability zone.
type CommittedResourceRouter struct{}
+func (c CommittedResourceRouter) extractClusterSelector(obj any) (string, error) {
+ switch v := obj.(type) {
+ case *v1alpha1.CommittedResource:
+ if v == nil {
+ return "", errors.New("object is nil")
+ }
+ if v.Spec.AvailabilityZone == "" {
+ return "", errors.New("committed resource does not have availability zone in spec")
+ }
+ return v.Spec.AvailabilityZone, nil
+ case v1alpha1.CommittedResource:
+ if v.Spec.AvailabilityZone == "" {
+ return "", errors.New("committed resource does not have availability zone in spec")
+ }
+ return v.Spec.AvailabilityZone, nil
+ default:
+ return "", errors.New("object is not a CommittedResource")
+ }
+}
+
func (c CommittedResourceRouter) Match(obj any, labels map[string]string) (bool, error) {
var cr v1alpha1.CommittedResource
@@ -116,6 +181,26 @@ func (c CommittedResourceRouter) Match(obj any, labels map[string]string) (bool,
// FlavorGroupCapacityResourceRouter routes flavor group capacity CRDs to clusters based on availability zone.
type FlavorGroupCapacityResourceRouter struct{}
+func (f FlavorGroupCapacityResourceRouter) extractClusterSelector(obj any) (string, error) {
+ switch v := obj.(type) {
+ case *v1alpha1.FlavorGroupCapacity:
+ if v == nil {
+ return "", errors.New("object is nil")
+ }
+ if v.Spec.AvailabilityZone == "" {
+ return "", errors.New("flavor group capacity does not have availability zone in spec")
+ }
+ return v.Spec.AvailabilityZone, nil
+ case v1alpha1.FlavorGroupCapacity:
+ if v.Spec.AvailabilityZone == "" {
+ return "", errors.New("flavor group capacity does not have availability zone in spec")
+ }
+ return v.Spec.AvailabilityZone, nil
+ default:
+ return "", errors.New("object is not a FlavorGroupCapacity")
+ }
+}
+
func (f FlavorGroupCapacityResourceRouter) Match(obj any, labels map[string]string) (bool, error) {
var fgc v1alpha1.FlavorGroupCapacity
@@ -143,6 +228,26 @@ func (f FlavorGroupCapacityResourceRouter) Match(obj any, labels map[string]stri
// HistoryResourceRouter routes histories to clusters based on availability zone.
type HistoryResourceRouter struct{}
+func (h HistoryResourceRouter) extractClusterSelector(obj any) (string, error) {
+ switch v := obj.(type) {
+ case *v1alpha1.History:
+ if v == nil {
+ return "", errors.New("object is nil")
+ }
+ if v.Spec.AvailabilityZone == nil || *v.Spec.AvailabilityZone == "" {
+ return "", errors.New("history does not have availability zone in spec")
+ }
+ return *v.Spec.AvailabilityZone, nil
+ case v1alpha1.History:
+ if v.Spec.AvailabilityZone == nil || *v.Spec.AvailabilityZone == "" {
+ return "", errors.New("history does not have availability zone in spec")
+ }
+ return *v.Spec.AvailabilityZone, nil
+ default:
+ return "", errors.New("object is not a History")
+ }
+}
+
func (h HistoryResourceRouter) Match(obj any, labels map[string]string) (bool, error) {
var hist v1alpha1.History
@@ -170,6 +275,26 @@ func (h HistoryResourceRouter) Match(obj any, labels map[string]string) (bool, e
// ProjectQuotaResourceRouter routes project quotas to clusters based on availability zone.
type ProjectQuotaResourceRouter struct{}
+func (p ProjectQuotaResourceRouter) extractClusterSelector(obj any) (string, error) {
+ switch v := obj.(type) {
+ case *v1alpha1.ProjectQuota:
+ if v == nil {
+ return "", errors.New("object is nil")
+ }
+ if v.Spec.AvailabilityZone == "" {
+ return "", errors.New("project quota does not have availability zone in spec")
+ }
+ return v.Spec.AvailabilityZone, nil
+ case v1alpha1.ProjectQuota:
+ if v.Spec.AvailabilityZone == "" {
+ return "", errors.New("project quota does not have availability zone in spec")
+ }
+ return v.Spec.AvailabilityZone, nil
+ default:
+ return "", errors.New("object is not a ProjectQuota")
+ }
+}
+
func (p ProjectQuotaResourceRouter) Match(obj any, labels map[string]string) (bool, error) {
var pq v1alpha1.ProjectQuota
From a0ba342cee1583963fdf0503b86bc20d4a04437a Mon Sep 17 00:00:00 2001
From: Markus Wieland <44964229+SoWieMarkus@users.noreply.github.com>
Date: Wed, 15 Jul 2026 16:57:13 +0200
Subject: [PATCH 06/15] fix: persist pipeline results before history upsert
(#1047)
A refactor in the past led to the history creation being before the
result was written to the decision object. That caused all our histories
to report no host found.
I adjusted that and added some tests to catch something like this
earlier next time.
---
.../filter_weigher_pipeline_controller.go | 7 +++-
.../scheduling/lib/history_client_test.go | 22 ++++++++---
.../filter_weigher_pipeline_controller.go | 7 +++-
.../filter_weigher_pipeline_controller.go | 7 +++-
.../filter_weigher_pipeline_controller.go | 7 +++-
...filter_weigher_pipeline_controller_test.go | 37 +++++++++++++++++++
.../filter_weigher_pipeline_controller.go | 7 +++-
7 files changed, 83 insertions(+), 11 deletions(-)
diff --git a/internal/scheduling/cinder/filter_weigher_pipeline_controller.go b/internal/scheduling/cinder/filter_weigher_pipeline_controller.go
index 99abfc1ba..209e9e342 100644
--- a/internal/scheduling/cinder/filter_weigher_pipeline_controller.go
+++ b/internal/scheduling/cinder/filter_weigher_pipeline_controller.go
@@ -112,6 +112,12 @@ func (c *FilterWeigherPipelineController) process(ctx context.Context, decision
}
result, err := pipeline.Run(request)
+ // Persist the result on the decision before upserting history so that
+ // CreateOrUpdateHistory can observe the target host and ordered hosts.
+ // On error the result is empty/meaningless, so only set it on success.
+ if err == nil {
+ decision.Status.Result = &result
+ }
if !request.Options.SkipHistory {
if upsertErr := c.HistoryManager.CreateOrUpdateHistory(ctx, decision, nil, err); upsertErr != nil {
log.Error(upsertErr, "failed to create/update history")
@@ -121,7 +127,6 @@ func (c *FilterWeigherPipelineController) process(ctx context.Context, decision
log.Error(err, "failed to run pipeline")
return err
}
- decision.Status.Result = &result
log.Info("decision processed successfully", "duration", time.Since(startedAt))
return nil
}
diff --git a/internal/scheduling/lib/history_client_test.go b/internal/scheduling/lib/history_client_test.go
index 9cf866913..1e7f63559 100644
--- a/internal/scheduling/lib/history_client_test.go
+++ b/internal/scheduling/lib/history_client_test.go
@@ -215,12 +215,16 @@ func TestHistoryClient_CreateOrUpdateHistory(t *testing.T) {
// assertions on the history list (archived entries only).
expectHistoryLen int
// assertions on the current decision.
- expectTargetHost *string
- expectSuccessful bool
- expectCondStatus metav1.ConditionStatus
- expectReason string
- checkExplanation func(t *testing.T, explanation string)
- checkCurrentHosts func(t *testing.T, hosts []string)
+ expectTargetHost *string
+ expectSuccessful bool
+ expectCondStatus metav1.ConditionStatus
+ expectReason string
+ // expectMessageContains, if set, asserts the Ready condition message
+ // surfaces this substring (used to check the pipeline-error path tells
+ // the user why scheduling failed).
+ expectMessageContains string
+ checkExplanation func(t *testing.T, explanation string)
+ checkCurrentHosts func(t *testing.T, hosts []string)
}{
{
name: "create new history",
@@ -362,6 +366,9 @@ func TestHistoryClient_CreateOrUpdateHistory(t *testing.T) {
expectSuccessful: false,
expectCondStatus: metav1.ConditionFalse,
expectReason: v1alpha1.HistoryReasonPipelineRunFailed,
+ // The Ready condition message must tell the user the pipeline
+ // errored and why, so devops can troubleshoot from the CRD status.
+ expectMessageContains: "pipeline run failed: no hosts available",
checkExplanation: func(t *testing.T, explanation string) {
if !strings.Contains(explanation, "no hosts available") {
t.Errorf("expected explanation to contain error text, got: %q", explanation)
@@ -546,6 +553,9 @@ func TestHistoryClient_CreateOrUpdateHistory(t *testing.T) {
if readyCond.Reason != tt.expectReason {
t.Errorf("condition reason = %q, want %q", readyCond.Reason, tt.expectReason)
}
+ if tt.expectMessageContains != "" && !strings.Contains(readyCond.Message, tt.expectMessageContains) {
+ t.Errorf("condition message = %q, want it to contain %q", readyCond.Message, tt.expectMessageContains)
+ }
// Verify explanation.
if tt.checkExplanation != nil {
diff --git a/internal/scheduling/machines/filter_weigher_pipeline_controller.go b/internal/scheduling/machines/filter_weigher_pipeline_controller.go
index 0063010d8..2bcf5ee0d 100644
--- a/internal/scheduling/machines/filter_weigher_pipeline_controller.go
+++ b/internal/scheduling/machines/filter_weigher_pipeline_controller.go
@@ -135,6 +135,12 @@ func (c *FilterWeigherPipelineController) process(ctx context.Context, decision
// Execute the scheduling pipeline. Options not set: machine scheduling always records history.
request := ironcore.MachinePipelineRequest{Pools: pools.Items}
result, err := pipeline.Run(request)
+ // Persist the result on the decision before upserting history so that
+ // CreateOrUpdateHistory can observe the target host and ordered hosts.
+ // On error the result is empty/meaningless, so only set it on success.
+ if err == nil {
+ decision.Status.Result = &result
+ }
if !request.Options.SkipHistory {
if upsertErr := c.HistoryManager.CreateOrUpdateHistory(ctx, decision, nil, err); upsertErr != nil {
log.Error(upsertErr, "failed to create/update history")
@@ -144,7 +150,6 @@ func (c *FilterWeigherPipelineController) process(ctx context.Context, decision
log.V(1).Error(err, "failed to run scheduler pipeline")
return errors.New("failed to run scheduler pipeline")
}
- decision.Status.Result = &result
log.Info("decision processed successfully", "duration", time.Since(startedAt))
// Set the machine pool ref on the machine.
diff --git a/internal/scheduling/manila/filter_weigher_pipeline_controller.go b/internal/scheduling/manila/filter_weigher_pipeline_controller.go
index 6e00593a9..a0bc8b960 100644
--- a/internal/scheduling/manila/filter_weigher_pipeline_controller.go
+++ b/internal/scheduling/manila/filter_weigher_pipeline_controller.go
@@ -112,6 +112,12 @@ func (c *FilterWeigherPipelineController) process(ctx context.Context, decision
}
result, err := pipeline.Run(request)
+ // Persist the result on the decision before upserting history so that
+ // CreateOrUpdateHistory can observe the target host and ordered hosts.
+ // On error the result is empty/meaningless, so only set it on success.
+ if err == nil {
+ decision.Status.Result = &result
+ }
if !request.Options.SkipHistory {
if upsertErr := c.HistoryManager.CreateOrUpdateHistory(ctx, decision, nil, err); upsertErr != nil {
log.Error(upsertErr, "failed to create/update history")
@@ -121,7 +127,6 @@ func (c *FilterWeigherPipelineController) process(ctx context.Context, decision
log.Error(err, "failed to run pipeline")
return err
}
- decision.Status.Result = &result
log.Info("decision processed successfully", "duration", time.Since(startedAt))
return nil
}
diff --git a/internal/scheduling/nova/filter_weigher_pipeline_controller.go b/internal/scheduling/nova/filter_weigher_pipeline_controller.go
index e1ce5ff0f..0c6917e1a 100644
--- a/internal/scheduling/nova/filter_weigher_pipeline_controller.go
+++ b/internal/scheduling/nova/filter_weigher_pipeline_controller.go
@@ -199,6 +199,12 @@ func (c *FilterWeigherPipelineController) process(ctx context.Context, decision
}
result, err := pipeline.Run(request)
+ // Persist the result on the decision before upserting history so that
+ // CreateOrUpdateHistory can observe the target host and ordered hosts.
+ // On error the result is empty/meaningless, so only set it on success.
+ if err == nil {
+ decision.Status.Result = &result
+ }
if !request.Options.SkipHistory {
c.upsertHistory(ctx, decision, err)
}
@@ -206,7 +212,6 @@ func (c *FilterWeigherPipelineController) process(ctx context.Context, decision
log.Error(err, "failed to run pipeline")
return &request, err
}
- decision.Status.Result = &result
meta.SetStatusCondition(&decision.Status.Conditions, metav1.Condition{
Type: v1alpha1.DecisionConditionReady,
Status: metav1.ConditionTrue,
diff --git a/internal/scheduling/nova/filter_weigher_pipeline_controller_test.go b/internal/scheduling/nova/filter_weigher_pipeline_controller_test.go
index caf145491..ccd565f56 100644
--- a/internal/scheduling/nova/filter_weigher_pipeline_controller_test.go
+++ b/internal/scheduling/nova/filter_weigher_pipeline_controller_test.go
@@ -12,6 +12,7 @@ import (
"time"
corev1 "k8s.io/api/core/v1"
+ "k8s.io/apimachinery/pkg/api/meta"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"k8s.io/apimachinery/pkg/runtime"
"k8s.io/apimachinery/pkg/types"
@@ -406,6 +407,9 @@ func TestFilterWeigherPipelineController_ProcessNewDecisionFromAPI(t *testing.T)
expectHistoryCreated bool
expectUpdatedStatus bool
errorContains string
+ // verifyHistory, if set, is invoked with the persisted History CRD so
+ // tests can assert the decision was written through correctly.
+ verifyHistory func(t *testing.T, history *v1alpha1.History)
}{
{
name: "successful processing with decision creation enabled",
@@ -453,6 +457,36 @@ func TestFilterWeigherPipelineController_ProcessNewDecisionFromAPI(t *testing.T)
expectResult: true,
expectHistoryCreated: true,
expectUpdatedStatus: true,
+ // Regression guard: the History CRD must reflect the actual
+ // decision (target host selected, ordered hosts, Ready=True).
+ // Previously the result was written to the decision *after* the
+ // history upsert, so history always recorded a failed decision
+ // with no target host.
+ verifyHistory: func(t *testing.T, history *v1alpha1.History) {
+ cur := history.Status.Current
+ if !cur.Successful {
+ t.Errorf("expected history current.successful=true, got false")
+ }
+ if cur.TargetHost == nil {
+ t.Errorf("expected history current.targetHost to be set, got nil")
+ }
+ if len(cur.OrderedHosts) == 0 {
+ t.Errorf("expected history current.orderedHosts to be populated, got empty")
+ }
+ if cur.Explanation == "" {
+ t.Errorf("expected history current.explanation to be set, got empty")
+ }
+ ready := meta.FindStatusCondition(history.Status.Conditions, v1alpha1.HistoryConditionReady)
+ if ready == nil {
+ t.Fatalf("expected Ready condition to be set")
+ }
+ if ready.Status != metav1.ConditionTrue {
+ t.Errorf("expected Ready condition status True, got %s", ready.Status)
+ }
+ if ready.Reason != v1alpha1.HistoryReasonSchedulingSucceeded {
+ t.Errorf("expected Ready reason %q, got %q", v1alpha1.HistoryReasonSchedulingSucceeded, ready.Reason)
+ }
+ },
},
{
name: "successful processing with decision creation disabled",
@@ -730,6 +764,9 @@ func TestFilterWeigherPipelineController_ProcessNewDecisionFromAPI(t *testing.T)
}
time.Sleep(5 * time.Millisecond)
}
+ if tt.verifyHistory != nil {
+ tt.verifyHistory(t, &histories.Items[0])
+ }
} else {
var histories v1alpha1.HistoryList
if err := client.List(context.Background(), &histories); err != nil {
diff --git a/internal/scheduling/pods/filter_weigher_pipeline_controller.go b/internal/scheduling/pods/filter_weigher_pipeline_controller.go
index cfbd06315..898c5068d 100644
--- a/internal/scheduling/pods/filter_weigher_pipeline_controller.go
+++ b/internal/scheduling/pods/filter_weigher_pipeline_controller.go
@@ -149,6 +149,12 @@ func (c *FilterWeigherPipelineController) process(ctx context.Context, decision
// Execute the scheduling pipeline. Options not set: pod scheduling always records history.
request := pods.PodPipelineRequest{Nodes: nodes.Items, Pod: *pod}
result, err := pipeline.Run(request)
+ // Persist the result on the decision before upserting history so that
+ // CreateOrUpdateHistory can observe the target host and ordered hosts.
+ // On error the result is empty/meaningless, so only set it on success.
+ if err == nil {
+ decision.Status.Result = &result
+ }
if !request.Options.SkipHistory {
if upsertErr := c.HistoryManager.CreateOrUpdateHistory(ctx, decision, nil, err); upsertErr != nil {
log.Error(upsertErr, "failed to create/update history")
@@ -158,7 +164,6 @@ func (c *FilterWeigherPipelineController) process(ctx context.Context, decision
log.V(1).Error(err, "failed to run scheduler pipeline")
return errors.New("failed to run scheduler pipeline")
}
- decision.Status.Result = &result
log.Info("decision processed successfully", "duration", time.Since(startedAt))
// Assign the first node returned by the pipeline using a Binding.
From e1ec6e1ffd9126636593f6ed5ed8f10f67736f20 Mon Sep 17 00:00:00 2001
From: mblos <156897072+mblos@users.noreply.github.com>
Date: Thu, 16 Jul 2026 08:36:17 +0200
Subject: [PATCH 07/15] fix: account for CPU & mem as a binding constraint in
slot counting (#1044)
Fix capacity slot counting to use both memory and CPU as binding constraints, not memory alone.
---
Tiltfile | 6 +
api/v1alpha1/flavor_group_capacity_types.go | 10 +-
.../committed-resource-reservations.md | 371 +++++++-----------
.../cortex.cloud_flavorgroupcapacities.yaml | 31 +-
.../reservations/capacity/controller.go | 332 +++++++++-------
.../reservations/capacity/controller_test.go | 159 +++++++-
.../scheduling/reservations/capacity/split.go | 21 +-
.../reservations/capacity/split_test.go | 10 +-
8 files changed, 542 insertions(+), 398 deletions(-)
diff --git a/Tiltfile b/Tiltfile
index e836cdbc5..bc0c1c407 100644
--- a/Tiltfile
+++ b/Tiltfile
@@ -47,6 +47,12 @@ if len(env_set_overrides) > 0:
else:
print("=== No CORTEX_ environment variables found ===")
+region = os.getenv('OS_REGION_NAME')
+if region:
+ print("=== Deriving region-scoped URLs from OS_REGION_NAME=" + region + " ===")
+ env_set_overrides.append('openstack.url=https://identity-3.' + region + '.cloud.sap/v3')
+ env_set_overrides.append('prometheus.url=https://metrics-internal.scaleout.' + region + '.cloud.sap/')
+
load('ext://helm_resource', 'helm_resource', 'helm_repo')
helm_repo(
'Prometheus Community Helm Repo',
diff --git a/api/v1alpha1/flavor_group_capacity_types.go b/api/v1alpha1/flavor_group_capacity_types.go
index 7e9ee36c0..dbdb54277 100644
--- a/api/v1alpha1/flavor_group_capacity_types.go
+++ b/api/v1alpha1/flavor_group_capacity_types.go
@@ -113,11 +113,17 @@ type FlavorGroupCapacityStatus struct {
// +kubebuilder:object:root=true
// +kubebuilder:subresource:status
// +kubebuilder:resource:scope=Cluster
-// +kubebuilder:printcolumn:name="FlavorGroup",type="string",JSONPath=".spec.flavorGroup"
+// +kubebuilder:printcolumn:name="Group",type="string",JSONPath=".spec.flavorGroup"
// +kubebuilder:printcolumn:name="AZ",type="string",JSONPath=".spec.availabilityZone"
// +kubebuilder:printcolumn:name="Running",type="integer",JSONPath=".status.runningInstances"
-// +kubebuilder:printcolumn:name="LastReconcile",type="date",JSONPath=".status.lastReconcileAt"
+// +kubebuilder:printcolumn:name="Avail",type="integer",JSONPath=".status.exclusivelyFreeSlots"
// +kubebuilder:printcolumn:name="Ready",type="string",JSONPath=".status.conditions[?(@.type=='Ready')].status"
+// +kubebuilder:printcolumn:name="Reconciled",type="date",JSONPath=".status.lastReconcileAt"
+// +kubebuilder:printcolumn:name="Age",type="date",JSONPath=".metadata.creationTimestamp",priority=1
+// +kubebuilder:printcolumn:name="Free_Mem",type="string",JSONPath=".status.freeCapacity.memory",priority=1
+// +kubebuilder:printcolumn:name="Excl_Mem",type="string",JSONPath=".status.exclusivelyFreeCapacity.memory",priority=1
+// +kubebuilder:printcolumn:name="Free_CPU",type="string",JSONPath=".status.freeCapacity.cores",priority=1
+// +kubebuilder:printcolumn:name="Excl_CPU",type="string",JSONPath=".status.exclusivelyFreeCapacity.cores",priority=1
// FlavorGroupCapacity caches pre-computed capacity data for one flavor group in one AZ.
// One CRD exists per (flavor group × AZ) pair, updated by the capacity controller on a fixed interval.
diff --git a/docs/reservations/committed-resource-reservations.md b/docs/reservations/committed-resource-reservations.md
index 3775f34b2..4c3f448be 100644
--- a/docs/reservations/committed-resource-reservations.md
+++ b/docs/reservations/committed-resource-reservations.md
@@ -1,50 +1,34 @@
# Committed Resource Reservation System
-Cortex reserves hypervisor capacity for customers who pre-commit resources (committed resources, CRs), and exposes usage and capacity data via APIs.
+Cortex reserves hypervisor capacity for customers who pre-commit resources (committed resources, CRs), and exposes usage and capacity data to Limes via the LIQUID API.
+
+Implementation: `internal/scheduling/reservations/commitments/`
- [Committed Resource Reservation System](#committed-resource-reservation-system)
+ - [Architecture Overview](#architecture-overview)
+ - [Limes State → Cortex Action](#limes-state--cortex-action)
+ - [Resource Types](#resource-types)
+ - [Commitment Lifecycle](#commitment-lifecycle)
+ - [Reservation Lifecycle](#reservation-lifecycle)
+ - [Capacity Blocking](#capacity-blocking)
+ - [InFlightReservation](#inflightreservation)
+ - [APIs](#apis)
+ - [Change-Commitments](#change-commitments)
+ - [Quota](#quota)
+ - [Report-Usage](#report-usage)
+ - [Report-Capacity](#report-capacity)
+ - [Capacity Reporting Reference](#capacity-reporting-reference)
+ - [FlavorGroupCapacity CRD — per-flavor fields](#flavorgroupcapacity-crd--per-flavor-fields)
+ - [FlavorGroupCapacity CRD — group-level fields](#flavorgroupcapacity-crd--group-level-fields)
+ - [Prometheus metrics](#prometheus-metrics)
+ - [Report-Capacity REST endpoint](#report-capacity-rest-endpoint)
+ - [Syncer Task](#syncer-task)
+ - [Placement Observability](#placement-observability)
- [Configuration and Observability](#configuration-and-observability)
- - [Lifecycle Management](#lifecycle-management)
- - [State (CRDs)](#state-crds)
- - [CR Commitment Lifecycle](#cr-commitment-lifecycle)
- - [Resource types](#resource-types)
- - [CommittedResource Controller](#committedresource-controller)
- - [Reservation Lifecycle](#reservation-lifecycle)
- - [VM Lifecycle](#vm-lifecycle)
- - [Capacity Blocking](#capacity-blocking)
- - [InFlightReservation](#inflightreservation)
- - [Reservation Controller](#reservation-controller)
- - [Info API](#info-api)
- - [Change-Commitments API](#change-commitments-api)
- - [Quota API](#quota-api)
- - [Report-Usage API](#report-usage-api)
- - [Report-Capacity API](#report-capacity-api)
- - [Syncer Task](#syncer-task)
- - [Placement Observability (CRS Evaluation)](#placement-observability-crs-evaluation)
-
-The CR reservation implementation is located in `internal/scheduling/reservations/commitments/`. Key components include:
-- `CommittedResource` controller — acceptance, rejection, child Reservation CRUD (memory) or arithmetic headroom check (cores)
-- `Reservation` controller — placement, VM allocation verification
-- API endpoints (`api/`)
-- Capacity and usage calculation logic
-- Syncer for periodic state sync
-## Configuration and Observability
+## Architecture Overview
-**Configuration**: Helm values for intervals, API flags, and pipeline configuration are defined in `helm/bundles/cortex-nova/values.yaml`. Key configuration includes:
-- API endpoint toggles (change-commitments, report-usage, report-capacity) — each endpoint can be disabled independently
-- Reconciliation intervals (grace period, active monitoring)
-- Scheduling pipeline selection per flavor group
-- Per-flavor-group resource flags (`handlesCommitments`, `hasCapacity`, `hasQuota`) controlling which resource types are active for each group
-
-**Metrics and Alerts**: Defined in `helm/bundles/cortex-nova/templates/alerts.yaml` with prefixes:
-- `cortex_committed_resource_change_api_*`
-- `cortex_committed_resource_usage_api_*`
-- `cortex_committed_resource_capacity_api_*`
-
-## Lifecycle Management
-
-The system is organized around two CRD types and two controllers. `CommittedResource` CRDs represent customer commitments; `Reservation` CRDs represent individual hypervisor capacity slots. Each has its own controller with a well-defined responsibility boundary.
+The system is organized around two CRD types and two controllers. `CommittedResource` CRDs represent customer commitments; `Reservation` CRDs represent individual hypervisor capacity slots held on behalf of a commitment.
```mermaid
flowchart LR
@@ -80,71 +64,30 @@ flowchart LR
ResCtrl -->|update status| Res
```
-### State (CRDs)
-
-**`CommittedResource` CRD** — primary source of truth for a commitment accepted by Cortex. One CRD per commitment UUID. Spec holds the commitment identity (project, flavor group, resource type, amount, ...). Status holds the acceptance outcome (`Ready` condition with reason `Planned`/`Reserving`/`Rejected`/`Accepted`), the accepted amount, and usage fields populated by the usage reconciler: `AssignedInstances` (VM UUIDs deterministically assigned to this CR), `UsedResources` (total resource consumption of assigned VMs), `LastUsageReconcileAt`, and `UsageObservedGeneration`.
-
-**`Reservation` CRD** — a single reservation slot on a hypervisor, owned by a `CommittedResource`. One `CommittedResource` may drive multiple `Reservation` CRDs (one per flavor-sized slot). Only memory commitments create Reservation CRDs; cores commitments do not. See [./failover-reservations.md](./failover-reservations.md) for the failover reservation type.
-
-**`ProjectQuota` CRD** — per-project, per-AZ quota store. One CRD exists per (project × availability zone) pair, named `quota-{projectID}-{az}`. Written by the Quota API when Limes pushes quota (one CRD is created for each AZ in the request). The quota controller reconciles usage into the status: `TotalUsage` and `PaygUsage` are flat `map[string]int64` fields tracking per-resource consumption in that AZ. The controller watches CommittedResource and Hypervisor CRDs to maintain these values via periodic full reconciles, incremental HV diffs, and PaygUsage-only recomputes triggered by CommittedResource status changes.
-
-**`FlavorGroupCapacity` CRD** — per-flavor-group, per-AZ capacity snapshot maintained by the capacity controller (outside this subsystem). The Report-Capacity endpoint reads these to compute available capacity.
-
-### CR Commitment Lifecycle
-
-The CR commitment lifecycle covers everything from a commitment being accepted by Limes through to Cortex confirming or rejecting it. The `CommittedResource` CRD is the entry point; the `CommittedResource` controller owns the acceptance decision.
-
-**Limes state → Cortex action:**
-
-| Limes State | Meaning | Cortex action |
-|---|---|---|
-| `planned` | Future start, no guarantee yet | No capacity reserved |
-| `pending` | Limes asking for a yes/no decision now | One-shot acceptance attempt — accept or reject; no retry |
-| `guaranteed` / `confirmed` | Capacity must be honoured | Accept and keep in sync; see failure handling below |
-| `superseded` / `expired` | Commitment no longer active | Release all held capacity |
-
-#### Resource types
+`FlavorGroupCapacity` CRDs are maintained by the capacity controller (outside this subsystem) and read by the Report-Capacity endpoint. `ProjectQuota` CRDs are written by the Quota API and read by the Report-Usage endpoint.
-Cortex handles two resource types for committed resources, with different acceptance mechanisms:
+## Limes State → Cortex Action
-**Memory (`_ram`)** — Cortex creates and manages `Reservation` CRDs on specific hypervisors. Acceptance means Cortex can place the required number of reservation slots via the scheduling pipeline. If placement is impossible (no hosts with enough free memory), the commitment is rejected or retried depending on the commitment state and `AllowRejection` flag.
+| Limes State | Cortex action |
+|---|---|
+| `planned` | No capacity reserved |
+| `pending` | One-shot acceptance attempt — accept or reject; no retry |
+| `guaranteed` / `confirmed` | Accept and keep in sync; retry indefinitely unless `AllowRejection=true` |
+| `superseded` / `expired` | Release all held capacity |
-**CPU cores (`_cores`)** — No `Reservation` CRDs are created. Cortex checks whether sufficient CPU headroom exists by comparing the requested cores against the total CPU capacity for the flavor group and AZ (as reported by the `FlavorGroupCapacity` CRD) minus cores already committed by other active CRs. This is a lightweight arithmetic check that does not interact with the scheduling pipeline.
+`AllowRejection` mirrors the request's `RequiresConfirmation` flag. When set, the controller rejects and rolls back on failure rather than retrying. On any rejection, capacity is rolled back to the last successfully accepted amount (or fully released if never accepted).
-The two types share the same lifecycle states and the same acceptance/rejection semantics — they differ only in how capacity is verified and held.
+## Resource Types
-#### CommittedResource Controller
+Cortex handles two resource types with different acceptance mechanisms:
-The controller accepts or rejects commitments and keeps the allocated capacity in sync with what Limes expects.
+**Memory (`_ram`)** — Cortex creates `Reservation` CRDs on specific hypervisors. Acceptance requires the scheduler to place the required slots. If placement fails, the commitment is rejected or retried based on its state and `AllowRejection`.
-**`pending`** — Cortex is being asked for a yes/no answer. A single acceptance attempt is made. On failure, the commitment is rejected and all held capacity is released. No retry.
-
-**`guaranteed` / `confirmed`** — Cortex is expected to honour the commitment indefinitely. The default is to keep retrying on failure (`Ready=False, Reason=Reserving`). Callers that can tolerate rejection set `AllowRejection=true`; the controller then rejects on failure rather than retrying.
-
-**On rejection** — any capacity held for this CR is rolled back to the last successfully accepted amount (or fully released if never accepted).
-
-**Reconcile trigger flow:**
-
-```mermaid
-sequenceDiagram
- participant API as Change-Commitments API
- participant CRCtrl as CR Controller
- participant CRCRD as CommittedResource CRD
- participant ResCRD as Reservation CRD
- participant ResCtrl as Reservation Controller
-
- API->>CRCRD: write (create/update)
- CRCRD-->>CRCtrl: watch fires
- CRCtrl->>ResCRD: create/update child slots (memory only)
- ResCRD-->>ResCtrl: watch fires
- ResCtrl->>ResCRD: update (ObservedParentGeneration, Ready=True/False)
- ResCRD-->>CRCtrl: watch fires (Reservation→parent CR lookup)
- CRCtrl->>CRCRD: update status (Accepted / Reserving / Rejected)
-```
+**CPU cores (`_cores`)** — No `Reservation` CRDs are created. Cortex does an arithmetic headroom check: requested cores vs. total CPU capacity for the flavor group and AZ (from `FlavorGroupCapacity`) minus cores already held by active CRs. Lightweight, no scheduler interaction.
-For cores commitments the middle steps (Reservation CRUD, Reservation controller) are skipped — the CR controller updates the `CommittedResource` status directly after the arithmetic check.
+The two types share lifecycle states and acceptance/rejection semantics — they differ only in how capacity is verified and held.
-**CommittedResource status states:**
+## Commitment Lifecycle
```mermaid
stateDiagram-v2
@@ -170,22 +113,30 @@ stateDiagram-v2
Planned --> [*] : deleted
```
-### Reservation Lifecycle
+The reconcile trigger chain for memory commitments:
-*Applies to memory commitments only. Cores commitments do not create Reservations.*
+```mermaid
+sequenceDiagram
+ participant API as Change-Commitments API
+ participant CRCtrl as CR Controller
+ participant CRCRD as CommittedResource CRD
+ participant ResCRD as Reservation CRD
+ participant ResCtrl as Reservation Controller
-| Component | Event | Timing | Action |
-|-----------|-------|--------|--------|
-| **Reservation Controller** | `Reservation` created | Immediate (watch) | Find host via scheduler API, set `TargetHost` |
-| **Scheduling Pipeline** | VM Create, Migrate, Resize | Immediate | Add VM to `Spec.Allocations` |
-| **Reservation Controller** | Reservation CRD updated | `committedResourceRequeueIntervalGracePeriod` (default: 1 min) | Defer verification for new VMs still spawning; update `Status.Allocations` |
-| **Reservation Controller** | Hypervisor CRD updated (VM appeared/disappeared) | Immediate (event-driven) | Verify allocations via Hypervisor CRD; remove gone VMs from `Spec.Allocations` |
-| **Reservation Controller** | Periodic safety-net | `committedResourceRequeueIntervalActive` (default: 5 min) | Same as above; catches any missed events |
-| **Reservation Controller** | Optimize unused slots | >> minutes | Assign PAYG VMs or re-place reservations |
+ API->>CRCRD: write (create/update)
+ CRCRD-->>CRCtrl: watch fires
+ CRCtrl->>ResCRD: create/update child slots
+ ResCRD-->>ResCtrl: watch fires
+ ResCtrl->>ResCRD: update (Ready=True/False)
+ ResCRD-->>CRCtrl: watch fires
+ CRCtrl->>CRCRD: update status (Accepted / Reserving / Rejected)
+```
-#### VM Lifecycle
+## Reservation Lifecycle
-VM allocations are tracked within reservations:
+*Applies to memory commitments only.*
+
+A `Reservation` CRD represents one flavor-sized slot on a specific hypervisor. The Reservation controller uses the **Hypervisor CRD as the sole source of truth** for VM presence — no Nova API calls.
```mermaid
flowchart LR
@@ -201,177 +152,159 @@ flowchart LR
C -->|update Spec/Status.Allocations| Res
```
-**Allocation fields**:
-- `Spec.Allocations` — Expected VMs (written by the scheduling pipeline on placement)
-- `Status.Allocations` — Confirmed VMs (written by the controller after verifying the VM is on the expected host)
+VM allocation has two fields with distinct semantics: `Spec.Allocations` (expected — written by the scheduling pipeline) and `Status.Allocations` (confirmed — written by the controller after the VM is verified on the expected hypervisor). New VMs stay in `Spec` only during a grace period to allow for startup time. After the grace period, absence from the Hypervisor CRD removes the VM.
-**VM allocation state diagram**:
+When a VM is confirmed on a reservation for the first time, the controller proactively removes it from `Spec.Allocations` on all other candidate reservations. This frees phantom capacity blocks immediately rather than waiting for each candidate's grace period to expire.
-The controller uses the **Hypervisor CRD** as the sole source of truth for VM allocation verification:
+`MaxConcurrentReconciles=1` on the Reservation controller is intentional — parallel reconciles would allow concurrent placements to race and double-book a slot.
-```mermaid
-stateDiagram-v2
- direction LR
- state "Spec only (grace period)" as SpecOnly
- state "Spec + Status (on expected host)" as Confirmed
-
- [*] --> SpecOnly : placement (create, migrate, resize)
- SpecOnly --> SpecOnly : within grace period
- SpecOnly --> Confirmed : found on HV CRD after grace period
- SpecOnly --> [*] : not on HV CRD after grace period
- Confirmed --> [*] : not on HV CRD
-```
-
-**Candidate reservation cleanup**: When a VM is newly confirmed on a reservation (transitions from Spec-only to Spec+Status for the first time), the controller immediately removes that VM's UUID from `Spec.Allocations` on all other candidate reservations that still carry it. This proactive cleanup frees phantom capacity blocks on non-selected hosts immediately rather than waiting for each candidate reservation's own grace period expiry or periodic requeue to detect that the VM landed elsewhere.
-
-**Note**: VM allocations may not consume all resources of a reservation slot. A reservation with 128 GB may have VMs totaling only 96 GB if that fits the project's needs. Allocations may exceed reservation capacity (e.g., after VM resize).
-
-#### Capacity Blocking
+### Capacity Blocking
-**Blocking rules by allocation state:**
+Each active Reservation blocks capacity on its target hypervisor so the scheduler cannot double-allocate. The block is recalculated on every reconcile:
-| State | In HV Allocation? | Reservation must block? |
-|---|---|---|
-| No allocations | — | Full `Spec.Resources` |
-| Confirmed (Spec + Status) | Yes — already subtracted | No — subtract from reservation block |
-| Spec only (not yet running) | No — not yet on host | Yes — must remain in reservation block |
-
-**Formal calculation (stable state, `Spec.TargetHost == Status.Host`):**
+**Stable state (`Spec.TargetHost == Status.Host`):**
```
-confirmed = sum of resources for VMs in both Spec.Allocations and Status.Allocations
-spec_only_unblocked = sum of resources for VMs in Spec.Allocations only, NOT having an active pessimistic blocking reservation on this host
+confirmed = resources of VMs in both Spec and Status allocations
+spec_only_unblocked = resources of Spec-only VMs without an active InFlightReservation on this host
remaining = max(0, Spec.Resources - confirmed)
block = max(remaining, spec_only_unblocked)
```
-**Interaction with pessimistic blocking reservations:**
-
-When a VM is in flight (Nova choosing between candidates), a pessimistic blocking reservation exists on each candidate host. For any SpecOnly VM that has such a reservation on the same host, the pessimistic blocking reservation is the authority — the CR reservation must not double-count it. The `spec_only_unblocked` term excludes those VMs.
-
-See the [InFlightReservation](#inflightreservation) section below for how these reservations are managed.
-
-**Migration state (`Spec.TargetHost != Status.Host`):**
+The `spec_only_unblocked` term exists because an InFlightReservation on the same host already blocks those resources pessimistically — the CR reservation must not double-count them.
-When a reservation is being migrated to a new host, block the full `max(Spec.Resources, spec_only_unblocked)` on **both** hosts — no subtraction of confirmed VMs. VMs may be split across hosts mid-migration and the split is not reliably known from reservation data alone; conservatively blocking both hosts prevents overcommit during the transition. The over-blocking resolves once migration completes and `Spec.TargetHost == Status.Host` again.
+**Migration state (`Spec.TargetHost != Status.Host`):** Block full `max(Spec.Resources, spec_only_unblocked)` on **both** hosts. VMs may be split across hosts mid-migration; conservative blocking on both prevents overcommit until migration completes.
-**Corner cases:**
+**Corner cases worth noting:**
+- Confirmed VMs exceed reservation size (e.g. after resize): clamp `remaining` to 0, never negative
+- Spec-only VM larger than remaining slot: block `spec_only_unblocked` — those resources will land when the VM starts
+- Live migration within a reservation: handled implicitly by `hv.Status.Allocation`, which libvirt reports on both source and target during migration; no special logic needed
-- **Confirmed VMs exceed reservation size** (e.g., after VM resize): `Spec.Resources - confirmed` goes negative. Clamp to `0` — otherwise the filter would add capacity back to the host.
+### InFlightReservation
-- **Spec-only VM larger than remaining reservation** (e.g., confirmed VMs have consumed most of the slot, and a new VM awaiting startup is larger than what remains): `remaining < spec_only_unblocked`. Block `spec_only_unblocked` — the VM will consume those resources when it starts, and they are not yet in HV Allocation.
+A short-lived `InFlightReservation` CRD is created at the end of each VM placement run, one per candidate host returned to Nova. It pessimistically blocks capacity on every candidate while Nova decides where the VM lands — preventing a second concurrent placement from booking the same slot.
-- **VM live migration within a reservation** (VM moves away from the reservation's host): handled implicitly by `hv.Status.Allocation`. Libvirt reports resource consumption on both source and target during live migration, so both hosts' `hv.Status.Allocation` already reflects the in-flight state. No special filter logic needed. The reservation controller will eventually remove the VM from the reservation once it's confirmed on the wrong host past the grace period.
+Created by the scheduling pipeline; deleted once the VM is confirmed on a host or after a timeout. Skipped for non-VM-placement runs (reservation scheduling, capacity probes, failover — all set `SkipInflight`).
-#### InFlightReservation
+## APIs
-An `InFlightReservation` is a short-lived Reservation CRD (type `InFlightReservation`) that pessimistically blocks capacity on each candidate host while a VM is being scheduled. It prevents double-booking when multiple scheduling decisions are in flight concurrently.
+### Change-Commitments
-**Lifecycle:**
-- **Created** by the scheduling pipeline at the end of a successful placement run, one per candidate host returned to Nova. Creation is skipped when the `SkipInflight` pipeline option is set (used by reservation scheduling, capacity checks, and failover — any non-VM-placement run).
-- **Deleted** once the VM has been confirmed on a host (the in-flight reservation is no longer needed) or after a timeout if the VM never lands.
+`POST /commitments/v1/change-commitments`
-**Spec fields** (`InFlightReservationSpec`):
-- `VMID` — Nova server UUID of the VM being scheduled
-- `UserID` — owner of the VM
-- `ProjectID` — project/tenant of the VM
-- `Intent` — lifecycle operation that triggered the placement (e.g., create, migrate, resize)
+**Write-intent, watch-for-outcome**: the handler writes `CommittedResource` CRDs and polls their `Ready` condition until terminal. It does not interact with Reservation CRDs directly.
-**Interaction with CR reservations:** When computing how much capacity a CR reservation must block, Spec-only VMs that already have an InFlightReservation on the same host are excluded from the CR reservation's block calculation (the `spec_only_unblocked` term). This avoids double-counting resources that are already blocked by the pessimistic InFlightReservation.
+**All-or-nothing semantics**: if any commitment in a batch cannot be fulfilled, the entire request is rolled back. All modified CRDs are restored to their pre-request specs.
-#### Reservation Controller
+### Quota
-The `Reservation` controller watches `Reservation` CRDs and `Hypervisor` CRDs. `MaxConcurrentReconciles=1` prevents overbooking during concurrent placements.
+`PUT /commitments/v1/projects/:project_id/quota`
-**Placement** — finds hosts for new reservations (calls scheduler API). Placement requests include a `domain_name` scheduler hint resolved from the reservation's `DomainID` via Keystone. This allows the `filter_external_customer` pipeline filter to enforce host restrictions for external customer domains. Domain name resolution uses an in-process cache that stores names indefinitely (domain names are immutable in OpenStack). If the Keystone integration is not configured (`keystoneSecretRef` absent), the hint is omitted and domain-based host restrictions are not enforced.
+Persists Limes quota as `ProjectQuota` CRDs (one per project × AZ). The quota controller reconciles actual usage into each CRD's status. Writes are idempotent; concurrent writes are resolved with retry-on-conflict.
-**Allocation Verification** — tracks VM lifecycle on reservations. The controller uses the Hypervisor CRD as the sole source of truth, with two triggers:
-- New VMs (within `committedResourceAllocationGracePeriod`, default: 15 min): verification deferred — VM may still be spawning; requeued every `committedResourceRequeueIntervalGracePeriod` (default: 1 min)
-- Established VMs: verified reactively when the Hypervisor CRD changes (VM appeared or disappeared in `Status.Instances`), with `committedResourceRequeueIntervalActive` (default: 5 min) as a safety-net fallback
-- Missing unconfirmed VMs (in `Spec.Allocations` only): removed from `Spec.Allocations` when not found on the Hypervisor CRD after the grace period
-- Missing confirmed VMs (already present in `Status.Allocations`): bypass the grace period entirely — their disappearance from the Hypervisor CRD is treated as authoritative and they are removed immediately
+### Report-Usage
-**Reservation migration is not supported yet.**
+`POST /commitments/v1/projects/:project_id/report-usage`
-### Info API
+Reports current usage per flavor group (ram, cores, instances). VM-to-commitment assignment is **pre-computed** by a background usage reconciler that writes into `CommittedResource.Status` — it is not calculated inline at request time. This assignment is deterministic but may differ from Cortex's internal scheduling assignment.
-`GET /commitments/v1/info` — describes the full service to Limes: which flavor groups are active, what resource types each group exposes (ram, cores, instances), their units, LIQUID topologies, and whether each accepts commitments.
+For flavor groups with `HandlesCommitments=true`, the response includes per-AZ quota from `ProjectQuota` CRDs.
-- RAM resources with `HandlesCommitments=true` use `AZSeparatedTopology` — Limes treats quota as AZ-specific and sends per-AZ breakdowns in quota requests.
-- All other resources (cores, instances, and RAM without commitments) use `AZAwareTopology` — no per-AZ quota.
+### Report-Capacity
-Limes calls this endpoint once on startup and whenever the service description changes.
+`POST /commitments/v1/report-capacity`
-### Change-Commitments API
+Reports available capacity per flavor group and AZ, read from pre-computed `FlavorGroupCapacity` CRDs. If a CRD's `Ready` condition is stale, usage is omitted from the response (capacity is still reported) to avoid underreporting during a controller outage.
-The change-commitments API receives batched commitment changes from Limes and applies them using a **write-intent, watch-for-outcome** pattern: the handler creates or updates `CommittedResource` CRDs and polls their `Status.Conditions` until each reaches a terminal state — it does not interact with `Reservation` CRDs directly.
+### Capacity Reporting Reference
-**Request Semantics**: A request can contain multiple commitment changes across different projects and flavor groups. The semantic is **all-or-nothing** — if any commitment in the batch cannot be fulfilled (e.g., insufficient capacity), the entire request is rejected and rolled back.
+This section maps every reporting surface to the values it exposes, the resource dimensions it considers, and the cluster state it reflects.
-**Operations**:
-1. For each commitment in the batch, create or update a `CommittedResource` CRD. `Spec.AllowRejection` mirrors the request's `RequiresConfirmation` flag: `true` for changes where Limes needs a yes/no answer (new commitments, resizes), `false` for non-confirming changes (deletions, status-only transitions) where Limes doesn't act on the rejection reason
-2. Poll `CommittedResource.Status.Conditions[Ready]` until each reaches a terminal state: `Reason=Accepted` (success), `Reason=Planned` (deferred; accepted), or `Reason=Rejected` (failure) — only for confirming changes; non-confirming changes return immediately without polling
-3. On any failure or timeout, restore all modified `CommittedResource` CRDs to their pre-request specs (or delete newly-created ones)
+#### FlavorGroupCapacity CRD — per-flavor fields
-The `CommittedResource` controller handles all downstream work. `AllowRejection=true` tells it to reject and roll back on failure rather than retrying indefinitely.
+| Field | Dimensions | Cluster state | Notes |
+|---|---|---|---|
+| `TotalCapacityVMSlots` | Min(memory, CPU) | Empty datacenter | All reservation types ignored; competing groups not subtracted |
+| `TotalCapacityHosts` | Min(memory, CPU) | Empty datacenter | Host count for `TotalCapacityVMSlots` |
+| `PlaceableVMs` | Min(memory, CPU) | Current + reservations | If this flavor consumed all remaining capacity; competing groups not subtracted |
+| `PlaceableHosts` | Min(memory, CPU) | Current + reservations | Host count for `PlaceableVMs` |
-### Quota API
+#### FlavorGroupCapacity CRD — group-level fields
-`PUT /commitments/v1/projects/:project_id/quota` — receives the project's quota allocation from Limes and persists it as `ProjectQuota` CRDs, one per (project × availability zone) combination, named `quota-{projectID}-{az}`. For flavor groups with `HandlesCommitments=true`, Limes sends per-AZ quota breakdowns; each AZ gets its own CRD with a flat `Quota map[string]int64` holding per-resource quota values for that zone. The quota controller then reconciles usage into each CRD's status (`TotalUsage`, `PaygUsage`). Writes are idempotent; concurrent writes are resolved with retry-on-conflict.
+| Field | Dimensions | Cluster state | Notes |
+|---|---|---|---|
+| `FreeCapacity` | Memory + Cores (separate) | Current + reservations | Raw sum across candidate hosts; may double-count across groups sharing hosts |
+| `ExclusivelyFreeCapacity` | Memory + Cores (separate) | Current + reservations | Round-robin split result — sum across all groups never exceeds installed capacity |
+| `ExclusivelyFreeSlots` | Min(memory, CPU) → Memory | Current + reservations | `ExclusivelyFreeCapacity[memory] / smallestFlavorMemBytes`; the memory pool is CPU-gated: the round-robin excludes hosts where the flavor doesn't fit on CPU before summing bytes |
+| `TotalCapacity` | Memory + Cores (separate) | Empty datacenter | `max(TotalCapacityVMSlots × flavorResources)` over all flavors in the group |
+| `CommittedCapacity` | Memory (slot units) | — | Active CR accepted amounts in smallest-flavor slot units |
+| `RunningInstances` / `RunningResources` | Memory + Cores | — | Actual running VMs in this group × AZ |
-### Report-Usage API
+#### Prometheus metrics
-`POST /commitments/v1/projects/:project_id/report-usage` — reports current resource usage for a project.
+All metrics carry `flavor_group` and `az` labels; per-flavor metrics additionally carry `flavor_name`.
-For each flavor group `X` that accepts commitments, Cortex exposes three resource types:
-- `hw_version_X_ram` — RAM in units of the smallest flavor in the group (`HandlesCommitments=true`)
-- `hw_version_X_cores` — CPU cores (`HandlesCommitments=false`; derived from RAM via fixed ratio where applicable)
-- `hw_version_X_instances` — instance count (`HandlesCommitments=false`)
+| Metric suffix | Source field | Dimensions | Cluster state |
+|---|---|---|---|
+| `_vm_slots_empty_datacenter` | `TotalCapacityVMSlots` | Min(memory, CPU) | Empty datacenter |
+| `_vm_slots_placeable` | `PlaceableVMs` | Min(memory, CPU) | Current + reservations |
+| `_hosts_empty_datacenter` | `TotalCapacityHosts` | Min(memory, CPU) | Empty datacenter |
+| `_hosts_placeable` | `PlaceableHosts` | Min(memory, CPU) | Current + reservations |
+| `_free_capacity_gib` | `FreeCapacity[memory]` | Memory only | Current + reservations — may overlap across groups |
+| `_exclusively_free_capacity_gib` | `ExclusivelyFreeCapacity[memory]` | Memory only | Current + reservations |
+| `_exclusively_free_slots` | `ExclusivelyFreeSlots` | Min(memory, CPU) → Memory | Current + reservations |
+| `_committed_gib` | `CommittedCapacityBytes` | Memory | — |
+| `_committed_reservations` | `CommittedCapacity` | Memory (slot units) | — |
+| `_running_instances` | `RunningInstances` | — | — |
-For flavor groups with `HandlesCommitments=true`, the response includes per-AZ quota from the `ProjectQuota` CRDs (written by the Quota API).
+#### Report-Capacity REST endpoint
-VM-to-commitment assignment is read from pre-computed `CommittedResource.Status` fields rather than being calculated inline at request time. A dedicated **usage reconciler** (in `internal/scheduling/reservations/commitments/usage_reconciler.go`) watches `CommittedResource` and `Hypervisor` CRDs and periodically runs the deterministic assignment algorithm, writing `AssignedInstances`, `UsedResources`, `LastUsageReconcileAt`, and `UsageObservedGeneration` into each CommittedResource's status. The Report-Usage endpoint reads these status fields to determine which VMs belong to which commitment. If a CR has not yet been reconciled, its VMs appear as PAYG until the first usage reconcile completes.
+Capacity and usage are derived from `FlavorGroupCapacity` CRDs and reported per AZ for three resource types per group:
-For each VM, the API reports whether it accounts to a specific commitment or PAYG. This assignment is deterministic and may differ from the actual Cortex internal assignment used for scheduling.
+| Resource | Capacity formula | Usage formula | Notes |
+|---|---|---|---|
+| `_instances` | `runningInstances + ExclusivelyFreeSlots` | `runningInstances` | `ExclusivelyFreeSlots` is CPU-and-memory-gated (round-robin), final slot count via memory division |
+| `_ram` (fixed core ratio) | same as `_instances` | `runningInstances` | Slot count stands in for RAM |
+| `_ram` (variable) | `(runningMemBytes + ExclusivelyFreeCapacity[memory]) / ramUnitBytes` | `runningMemBytes / ramUnitBytes` | Both in declared units (e.g. GiB); `ramUnitBytes` configured per group |
+| `_cores` | `runningCoresCount + ExclusivelyFreeCapacity[cores]` | `runningCoresCount` | CPU-dimension-driven |
-### Report-Capacity API
+## Syncer Task
-`POST /commitments/v1/report-capacity` — reports available hypervisor capacity per flavor group and AZ. Capacity data is pre-computed by the capacity controller and stored in `FlavorGroupCapacity` CRDs; the endpoint aggregates these per-AZ values into the response. If a `FlavorGroupCapacity` CRD is stale (controller behind), the endpoint reports total capacity without subtracting usage to avoid underreporting.
+Runs periodically and reconciles local `CommittedResource` CRD state against Limes' view, correcting drift from missed API calls or restarts. Writes `CommittedResource` CRDs only — capacity management remains the controller's responsibility.
-### Syncer Task
+## Placement Observability
-The syncer task runs periodically and syncs local `CommittedResource` CRD state to match Limes' view of commitments, correcting drift from missed API calls or restarts. It writes `CommittedResource` CRDs only — capacity management is the controller's responsibility.
-
-### Placement Observability (CRS Evaluation)
-
-The `internal/scheduling/nova/crs/` package provides post-placement classification and Prometheus metrics for committed resource slot utilization. It answers the question: "For each VM placement (or no-host-found failure), what was the CR slot situation?"
-
-**Prometheus metrics:**
+The `internal/scheduling/nova/crs/` package classifies every placement decision by CR slot coverage and emits Prometheus metrics. This answers: "For each VM placement or no-host-found, what was the CR slot situation?"
| Metric | Labels | Description |
|--------|--------|-------------|
| `cortex_nova_no_host_found_total` | `cr_slot`, `flavor_group`, `intent` | No-host-found results classified by CR coverage |
| `cortex_nova_placement_total` | `flavor_group`, `intent`, `cr_slot` | Successful placements classified by CR slot outcome |
-PAYG placements (flavor not in any configured group) are not counted by either metric.
+PAYG placements (flavor not in any configured group) are not counted.
-**No-host-found classification (`cr_slot` label on `cortex_nova_no_host_found_total`):**
+**`cr_slot` values for no-host-found:**
-| Category | Meaning |
-|----------|---------|
+| Value | Meaning |
+|---|---|
| `no_cr` | Project has no active CommittedResources for the flavor group |
-| `cr_exhausted` | CommittedResources exist but are fully occupied (used >= capacity) |
-| `slot_exhausted` | CR has remaining capacity but no input host has a usable reservation slot |
-| `slot_blocked` | A usable slot exists on an input host but scheduling constraints excluded all such hosts |
+| `cr_exhausted` | CommittedResources exist but are fully occupied |
+| `slot_exhausted` | CR has remaining capacity but no candidate host has a usable reservation slot |
+| `slot_blocked` | A usable slot exists but scheduling constraints excluded all such hosts |
-**Placement classification (`cr_slot` label on `cortex_nova_placement_total`):**
+**`cr_slot` values for successful placements:**
-| Category | Meaning |
-|----------|---------|
+| Value | Meaning |
+|---|---|
| `no_cr` | No active CR or CR capacity fully exhausted |
| `slot_missed` | CR has remaining capacity but no candidate host has a slot with remaining memory > 0 |
| `slot_used` | CR has remaining capacity and at least one candidate host has a usable slot |
-**Slot evaluator:** The `SlotEvaluator` is built once per scheduling request from Hypervisor and Reservation CRDs (no further K8s reads during classification). It computes per-host free memory and indexes ready CR reservation slots by host. `HasUsableSlot` checks whether a host has a slot that can accommodate the VM under the overfill model: `slot.remaining + host.base_free >= vmMemBytes`.
+## Configuration and Observability
+
+**Configuration**: `helm/bundles/cortex-nova/values.yaml` — API endpoint toggles, reconciliation intervals, scheduling pipeline selection, and per-flavor-group resource flags.
-**Recorder:** The `Recorder` is called after each placement decision. On success (`slot_used`), it writes the VM UUID into the best-fit reservation slot (`PickSlot` selects the slot that maximises coverage with tightest-fit tiebreaking). On no-host-found, it classifies the failure and increments the counter.
+**Metrics and Alerts**: `helm/bundles/cortex-nova/templates/alerts.yaml`, prefixes:
+- `cortex_committed_resource_change_api_*`
+- `cortex_committed_resource_usage_api_*`
+- `cortex_committed_resource_capacity_api_*`
diff --git a/helm/library/cortex/files/crds/cortex.cloud_flavorgroupcapacities.yaml b/helm/library/cortex/files/crds/cortex.cloud_flavorgroupcapacities.yaml
index 952e15722..4102ea447 100644
--- a/helm/library/cortex/files/crds/cortex.cloud_flavorgroupcapacities.yaml
+++ b/helm/library/cortex/files/crds/cortex.cloud_flavorgroupcapacities.yaml
@@ -16,7 +16,7 @@ spec:
versions:
- additionalPrinterColumns:
- jsonPath: .spec.flavorGroup
- name: FlavorGroup
+ name: Group
type: string
- jsonPath: .spec.availabilityZone
name: AZ
@@ -24,12 +24,35 @@ spec:
- jsonPath: .status.runningInstances
name: Running
type: integer
- - jsonPath: .status.lastReconcileAt
- name: LastReconcile
- type: date
+ - jsonPath: .status.exclusivelyFreeSlots
+ name: Avail
+ type: integer
- jsonPath: .status.conditions[?(@.type=='Ready')].status
name: Ready
type: string
+ - jsonPath: .status.lastReconcileAt
+ name: Reconciled
+ type: date
+ - jsonPath: .metadata.creationTimestamp
+ name: Age
+ priority: 1
+ type: date
+ - jsonPath: .status.freeCapacity.memory
+ name: Free_Mem
+ priority: 1
+ type: string
+ - jsonPath: .status.exclusivelyFreeCapacity.memory
+ name: Excl_Mem
+ priority: 1
+ type: string
+ - jsonPath: .status.freeCapacity.cores
+ name: Free_CPU
+ priority: 1
+ type: string
+ - jsonPath: .status.exclusivelyFreeCapacity.cores
+ name: Excl_CPU
+ priority: 1
+ type: string
name: v1alpha1
schema:
openAPIV3Schema:
diff --git a/internal/scheduling/reservations/capacity/controller.go b/internal/scheduling/reservations/capacity/controller.go
index d3f5c8249..64a87befe 100644
--- a/internal/scheduling/reservations/capacity/controller.go
+++ b/internal/scheduling/reservations/capacity/controller.go
@@ -31,6 +31,7 @@ import (
"github.com/cobaltcore-dev/cortex/internal/knowledge/extractor/plugins/compute"
"github.com/cobaltcore-dev/cortex/internal/scheduling/reservations"
"github.com/cobaltcore-dev/cortex/pkg/multicluster"
+ "github.com/go-logr/logr"
)
var log = ctrl.Log.WithName("capacity-controller").WithValues("module", "capacity")
@@ -191,6 +192,29 @@ type vmUsage struct {
fresh bool
}
+// probeGroupResult holds the outcome of probing all flavors in one (group × AZ).
+type probeGroupResult struct {
+ groupName string
+ groupData compute.FlavorGroupFeature
+ flavors []v1alpha1.FlavorCapacityStatus
+ // allFresh is false if any scheduler probe failed; the group's CRD is left unchanged.
+ allFresh bool
+ smallestCandidates []string
+ committedCapacity int64
+}
+
+// flavorSlots returns the number of VM slots a resource map can fit for the given flavor.
+// It is the binding constraint across both memory and CPU: min(memSlots, cpuSlots).
+func flavorSlots(resources map[string]int64, flavorMemBytes, flavorVCPUs int64) int64 {
+ slots := resources[ResourceMemory] / flavorMemBytes
+ if flavorVCPUs > 0 {
+ if cpuSlots := resources[ResourceCores] / flavorVCPUs; cpuSlots < slots {
+ slots = cpuSlots
+ }
+ }
+ return slots
+}
+
// reconcileAll iterates all AZs, runs the round-robin split per AZ, then writes CRDs.
func (c *Reconciler) reconcileAll(ctx context.Context) error {
logger := LoggerFromContext(ctx)
@@ -222,22 +246,14 @@ func (c *Reconciler) reconcileAll(ctx context.Context) error {
usageByKey := c.computeVMUsage(ctx, flavorGroups, hvList.Items)
- var succeeded, failed int
for _, az := range azs {
- if err := c.reconcileAZ(ctx, az, flavorGroups, hvByName, blockedByReservations, usageByKey); err != nil {
- logger.Error(err, "failed to reconcile AZ", "az", az)
- failed++
- continue
- }
- succeeded += len(flavorGroups)
+ c.reconcileAZ(ctx, az, flavorGroups, hvByName, blockedByReservations, usageByKey)
}
logger.Info("capacity reconcile cycle completed",
"flavorGroups", len(flavorGroups),
"availabilityZones", len(azs),
"hypervisors", len(hvList.Items),
- "succeeded", succeeded,
- "failed", failed,
"duration", time.Since(startTime).String())
return nil
}
@@ -343,116 +359,100 @@ func hvRemainingResources(hv hv1.Hypervisor, blockedMemBytes int64) map[string]i
return result
}
-// reconcileAZ runs the round-robin capacity split for all flavor groups in one AZ,
-// then writes one FlavorGroupCapacity CRD per group that had all probes succeed.
-// Groups with failed probes are skipped — their CRDs retain the last good state.
-func (c *Reconciler) reconcileAZ(
+// probeGroup probes all flavors in a single group for one AZ and returns the result.
+// It preserves stale per-flavor values from the existing CRD on individual probe failures.
+func (c *Reconciler) probeGroup(
ctx context.Context,
+ groupName string,
+ groupData compute.FlavorGroupFeature,
az string,
- flavorGroups map[string]compute.FlavorGroupFeature,
hvByName map[string]hv1.Hypervisor,
blockedByReservations map[string]int64,
- usageByKey map[vmUsageKey]vmUsage,
-) error {
+) (probeGroupResult, error) {
logger := LoggerFromContext(ctx)
- type probeResult struct {
- groupName string
- groupData compute.FlavorGroupFeature
- flavors []v1alpha1.FlavorCapacityStatus
- // allFresh is false if any scheduler probe failed; the group's CRD is left unchanged.
- allFresh bool
- smallestCandidates []string
- committedCapacity int64
+ smallestFlavorBytes := int64(groupData.SmallestFlavor.MemoryMB) * 1024 * 1024 //nolint:gosec
+ if smallestFlavorBytes <= 0 {
+ return probeGroupResult{}, fmt.Errorf("smallest flavor %q has invalid memory %d MB",
+ groupData.SmallestFlavor.Name, groupData.SmallestFlavor.MemoryMB)
}
- results := make([]probeResult, 0, len(flavorGroups))
-
- groupNames := make([]string, 0, len(flavorGroups))
- for name := range flavorGroups {
- groupNames = append(groupNames, name)
+ // Load existing per-flavor data to preserve stale values on probe failure.
+ crdName := crdNameFor(groupName, az)
+ var existing v1alpha1.FlavorGroupCapacity
+ if err := c.client.Get(ctx, types.NamespacedName{Name: crdName}, &existing); err != nil && !apierrors.IsNotFound(err) {
+ return probeGroupResult{}, fmt.Errorf("failed to get FlavorGroupCapacity %s: %w", crdName, err)
+ }
+ existingByName := make(map[string]v1alpha1.FlavorCapacityStatus, len(existing.Status.Flavors))
+ for _, f := range existing.Status.Flavors {
+ existingByName[f.FlavorName] = f
}
- sort.Strings(groupNames)
- for _, groupName := range groupNames {
- groupData := flavorGroups[groupName]
+ // Probe all flavors. Sort for stable CRD output.
+ flavors := make([]compute.FlavorInGroup, len(groupData.Flavors))
+ copy(flavors, groupData.Flavors)
+ sort.Slice(flavors, func(i, j int) bool { return flavors[i].Name < flavors[j].Name })
- smallestFlavorBytes := int64(groupData.SmallestFlavor.MemoryMB) * 1024 * 1024 //nolint:gosec
- if smallestFlavorBytes <= 0 {
- logger.Error(fmt.Errorf("smallest flavor %q has invalid memory %d MB",
- groupData.SmallestFlavor.Name, groupData.SmallestFlavor.MemoryMB),
- "skipping flavor group", "flavorGroup", groupName)
- continue
- }
+ allFresh := true
+ newFlavors := make([]v1alpha1.FlavorCapacityStatus, 0, len(flavors))
+ var smallestCandidates []string
- // Probe all flavors. Sort for stable CRD output.
- flavors := make([]compute.FlavorInGroup, len(groupData.Flavors))
- copy(flavors, groupData.Flavors)
- sort.Slice(flavors, func(i, j int) bool { return flavors[i].Name < flavors[j].Name })
+ for _, flavor := range flavors {
+ cur := existingByName[flavor.Name]
+ cur.FlavorName = flavor.Name
- allFresh := true
- newFlavors := make([]v1alpha1.FlavorCapacityStatus, 0, len(flavors))
+ totalVMSlots, totalHosts, _, totalErr := c.probeScheduler(ctx, flavor, az, c.config.TotalPipeline, hvByName, true, nil)
+ placeableVMs, placeableHosts, candidates, placeableErr := c.probeScheduler(ctx, flavor, az, c.config.PlaceablePipeline, hvByName, false, blockedByReservations)
- // Load existing per-flavor data to preserve stale values on probe failure.
- crdName := crdNameFor(groupName, az)
- var existing v1alpha1.FlavorGroupCapacity
- if err := c.client.Get(ctx, types.NamespacedName{Name: crdName}, &existing); err != nil && !apierrors.IsNotFound(err) {
- return fmt.Errorf("failed to get FlavorGroupCapacity %s: %w", crdName, err)
+ if totalErr != nil {
+ allFresh = false
+ } else {
+ cur.TotalCapacityVMSlots = totalVMSlots
+ cur.TotalCapacityHosts = totalHosts
}
- existingByName := make(map[string]v1alpha1.FlavorCapacityStatus, len(existing.Status.Flavors))
- for _, f := range existing.Status.Flavors {
- existingByName[f.FlavorName] = f
+ if placeableErr != nil {
+ allFresh = false
+ } else {
+ cur.PlaceableVMs = placeableVMs
+ cur.PlaceableHosts = placeableHosts
}
+ // Capture candidates for the smallest flavor — used as split inputs.
+ if flavor.Name == groupData.SmallestFlavor.Name && placeableErr == nil {
+ smallestCandidates = candidates
+ }
+ newFlavors = append(newFlavors, cur)
+ }
- var smallestCandidates []string
- for _, flavor := range flavors {
- cur := existingByName[flavor.Name]
- cur.FlavorName = flavor.Name
-
- totalVMSlots, totalHosts, _, totalErr := c.probeScheduler(ctx, flavor, az, c.config.TotalPipeline, hvByName, true, nil)
- placeableVMs, placeableHosts, candidates, placeableErr := c.probeScheduler(ctx, flavor, az, c.config.PlaceablePipeline, hvByName, false, blockedByReservations)
+ committedCapacity, committedErr := c.sumCommittedCapacity(ctx, groupName, az, smallestFlavorBytes)
+ if committedErr != nil {
+ logger.Error(committedErr, "failed to sum committed capacity", "flavorGroup", groupName, "az", az)
+ committedCapacity = 0
+ }
- if totalErr != nil {
- allFresh = false
- } else {
- cur.TotalCapacityVMSlots = totalVMSlots
- cur.TotalCapacityHosts = totalHosts
- }
- if placeableErr != nil {
- allFresh = false
- } else {
- cur.PlaceableVMs = placeableVMs
- cur.PlaceableHosts = placeableHosts
- }
- // Capture candidates for the smallest flavor — used as split inputs.
- if flavor.Name == groupData.SmallestFlavor.Name && placeableErr == nil {
- smallestCandidates = candidates
- }
- newFlavors = append(newFlavors, cur)
- }
+ return probeGroupResult{
+ groupName: groupName,
+ groupData: groupData,
+ flavors: newFlavors,
+ allFresh: allFresh,
+ smallestCandidates: smallestCandidates,
+ committedCapacity: committedCapacity,
+ }, nil
+}
- committedCapacity, committedErr := c.sumCommittedCapacity(ctx, groupName, az, smallestFlavorBytes)
- if committedErr != nil {
- logger.Error(committedErr, "failed to sum committed capacity",
- "flavorGroup", groupName, "az", az)
- committedCapacity = 0
- }
+// buildSplitInputs constructs the HostState map and GroupInput slice needed by SplitCapacity.
+// Only groups where all probes succeeded are included.
+func buildSplitInputs(
+ results []probeGroupResult,
+ hvByName map[string]hv1.Hypervisor,
+ blockedByReservations map[string]int64,
+ az string,
+ logger logr.Logger,
+) (groupInputs []GroupInput, hosts map[string]HostState) {
- results = append(results, probeResult{
- groupName: groupName,
- groupData: groupData,
- flavors: newFlavors,
- allFresh: allFresh,
- smallestCandidates: smallestCandidates,
- committedCapacity: committedCapacity,
- })
- }
+ hosts = make(map[string]HostState)
+ groupInputs = make([]GroupInput, 0, len(results))
- // Build HostState and GroupInput for the round-robin split.
- // Only include groups where all probes succeeded.
- hosts := make(map[string]HostState)
- groupInputs := make([]GroupInput, 0, len(results))
for _, r := range results {
if !r.allFresh || r.smallestCandidates == nil {
continue
@@ -469,22 +469,17 @@ func (c *Reconciler) reconcileAZ(
continue
}
remaining := hvRemainingResources(hv, blockedByReservations[h])
- if remaining != nil {
- hosts[h] = HostState{Remaining: remaining}
- memSlots := remaining[ResourceMemory] / flavorMemBytes
- cpuSlots := remaining[ResourceCores] / flavorVCPUs
- usableSlots := memSlots
- if cpuSlots < usableSlots {
- usableSlots = cpuSlots
- }
- strandedMem := remaining[ResourceMemory] - usableSlots*flavorMemBytes
- strandedCPU := remaining[ResourceCores] - usableSlots*flavorVCPUs
- logger.V(1).Info("candidate host for capacity split",
- "az", az, "flavorGroup", r.groupName, "host", h,
- "usableSlots", usableSlots,
- "strandedMemoryGiB", strandedMem/(1024*1024*1024),
- "strandedCores", strandedCPU)
+ if remaining == nil {
+ continue
}
+ hosts[h] = HostState{Remaining: remaining}
+ usableSlots := flavorSlots(remaining, flavorMemBytes, flavorVCPUs)
+ strandedMem := remaining[ResourceMemory] - usableSlots*flavorMemBytes
+ strandedCPU := remaining[ResourceCores] - usableSlots*flavorVCPUs
+ logger.V(1).Info("candidate host slot details", "az", az, "flavorGroup", r.groupName, "host", h,
+ "usableSlots", usableSlots,
+ "strandedMemoryGiB", strandedMem/(1024*1024*1024),
+ "strandedCores", strandedCPU)
}
}
sort.Strings(candidateHosts) // stable order
@@ -497,15 +492,64 @@ func (c *Reconciler) reconcileAZ(
CandidateHosts: candidateHosts,
})
}
+ return groupInputs, hosts
+}
+
+// reconcileAZ probes all flavor groups in one AZ, splits capacity across groups,
+// and writes one FlavorGroupCapacity CRD per group that had all probes succeed.
+func (c *Reconciler) reconcileAZ(
+ ctx context.Context,
+ az string,
+ flavorGroups map[string]compute.FlavorGroupFeature,
+ hvByName map[string]hv1.Hypervisor,
+ blockedByReservations map[string]int64,
+ usageByKey map[vmUsageKey]vmUsage,
+) {
+
+ logger := LoggerFromContext(ctx)
+
+ groupNames := make([]string, 0, len(flavorGroups))
+ for name := range flavorGroups {
+ groupNames = append(groupNames, name)
+ }
+ sort.Strings(groupNames)
+
+ results := make([]probeGroupResult, 0, len(groupNames))
+ for _, groupName := range groupNames {
+ r, err := c.probeGroup(ctx, groupName, flavorGroups[groupName], az, hvByName, blockedByReservations)
+ if err != nil {
+ logger.Error(err, "skipping flavor group", "flavorGroup", groupName, "az", az)
+ continue
+ }
+ results = append(results, r)
+ }
+
+ groupInputs, hosts := buildSplitInputs(results, hvByName, blockedByReservations, az, logger)
+ freeResources, exclusiveResources, unassigned, strandedByHost := SplitCapacity(groupInputs, hosts)
- freeResources, exclusiveResources, unassigned := SplitCapacity(groupInputs, hosts)
if unassigned[ResourceMemory] > 0 || unassigned[ResourceCores] > 0 {
+ groupNames := make([]string, 0, len(groupInputs))
+ hostToGroups := make(map[string][]string)
+ for _, g := range groupInputs {
+ groupNames = append(groupNames, g.Name)
+ for _, h := range g.CandidateHosts {
+ hostToGroups[h] = append(hostToGroups[h], g.Name)
+ }
+ }
logger.Info("fragmented capacity not assigned to any group",
"az", az,
"unassignedMemoryGiB", unassigned[ResourceMemory]/(1024*1024*1024),
"unassignedCores", unassigned[ResourceCores],
"candidateHosts", len(hosts),
- "groups", len(groupInputs))
+ "groups", groupNames)
+ for host, res := range strandedByHost {
+ logger.V(1).Info("stranded host resources after split",
+ "az", az,
+ "host", host,
+ "strandedMemoryGiB", res[ResourceMemory]/(1024*1024*1024),
+ "strandedCores", res[ResourceCores],
+ "eligibleGroups", hostToGroups[host])
+ }
}
// Write one CRD per group. Skip groups with failed probes — their CRDs retain last good state.
@@ -523,7 +567,26 @@ func (c *Reconciler) reconcileAZ(
"flavorGroup", r.groupName, "az", az)
}
}
- return nil
+}
+
+// computeTotalCapacity returns the maximum memory bytes and CPU cores representable
+// by the flavor with the highest slot count in the group (empty-datacenter view).
+func computeTotalCapacity(newFlavors []v1alpha1.FlavorCapacityStatus, flavorSpecByName map[string]compute.FlavorInGroup) (maxMemBytes, maxCPUCores int64) {
+ for _, f := range newFlavors {
+ spec, ok := flavorSpecByName[f.FlavorName]
+ if !ok || f.TotalCapacityVMSlots <= 0 {
+ continue
+ }
+ memBytes := f.TotalCapacityVMSlots * int64(spec.MemoryMB) * 1024 * 1024 //nolint:gosec
+ cpuCores := f.TotalCapacityVMSlots * int64(spec.VCPUs) //nolint:gosec
+ if memBytes > maxMemBytes {
+ maxMemBytes = memBytes
+ }
+ if cpuCores > maxCPUCores {
+ maxCPUCores = cpuCores
+ }
+ }
+ return maxMemBytes, maxCPUCores
}
// writeCRD upserts one FlavorGroupCapacity CRD with fresh computed values.
@@ -558,28 +621,11 @@ func (c *Reconciler) writeCRD(
return fmt.Errorf("failed to get FlavorGroupCapacity %s: %w", crdName, err)
}
- // TotalCapacity: for each flavor multiply slot count by its resources; take the max
- // across all flavors independently. The flavor best matching the host's resource
- // ratio saturates more resources and produces a higher product.
flavorSpecByName := make(map[string]compute.FlavorInGroup, len(groupData.Flavors))
for _, f := range groupData.Flavors {
flavorSpecByName[f.Name] = f
}
- var maxMemBytes, maxCPUCores int64
- for _, f := range newFlavors {
- spec, ok := flavorSpecByName[f.FlavorName]
- if !ok || f.TotalCapacityVMSlots <= 0 {
- continue
- }
- memBytes := f.TotalCapacityVMSlots * int64(spec.MemoryMB) * 1024 * 1024 //nolint:gosec
- cpuCores := f.TotalCapacityVMSlots * int64(spec.VCPUs) //nolint:gosec
- if memBytes > maxMemBytes {
- maxMemBytes = memBytes
- }
- if cpuCores > maxCPUCores {
- maxCPUCores = cpuCores
- }
- }
+ maxMemBytes, maxCPUCores := computeTotalCapacity(newFlavors, flavorSpecByName)
patch := client.MergeFrom(existing.DeepCopy())
existing.Status.Flavors = newFlavors
@@ -598,8 +644,11 @@ func (c *Reconciler) writeCRD(
}
existing.Status.FreeCapacity = resMapToQuantity(freeRes)
existing.Status.ExclusivelyFreeCapacity = resMapToQuantity(exclusiveRes)
+ var exclusivelyFreeSlots int64
if flavorMemBytes := int64(groupData.SmallestFlavor.MemoryMB) * 1024 * 1024; flavorMemBytes > 0 { //nolint:gosec
- existing.Status.ExclusivelyFreeSlots = exclusiveRes[ResourceMemory] / flavorMemBytes
+ flavorVCPUs := int64(groupData.SmallestFlavor.VCPUs) //nolint:gosec
+ exclusivelyFreeSlots = flavorSlots(exclusiveRes, flavorMemBytes, flavorVCPUs)
+ existing.Status.ExclusivelyFreeSlots = exclusivelyFreeSlots
}
existing.Status.LastReconcileAt = metav1.Now()
@@ -633,6 +682,7 @@ func (c *Reconciler) probeScheduler(
if flavorBytes <= 0 {
return 0, 0, nil, fmt.Errorf("flavor %q has invalid memory %d MB", flavor.Name, flavor.MemoryMB)
}
+ flavorVCPUs := int64(flavor.VCPUs) //nolint:gosec
// Build EligibleHosts from all known hypervisors so that novaLimitHostsToRequest
// (which filters the response to hosts present in the request) does not zero out
@@ -680,7 +730,7 @@ func (c *Reconciler) probeScheduler(
if !ok {
continue
}
- var capBytes int64
+ var resources map[string]int64
if ignoreAllocations {
effCap := hv.Status.EffectiveCapacity
if effCap == nil {
@@ -693,15 +743,17 @@ func (c *Reconciler) probeScheduler(
if !ok {
continue
}
- capBytes = memCap.Value()
+ resources = map[string]int64{ResourceMemory: memCap.Value()}
+ if cpuCap, ok := effCap[hv1.ResourceCPU]; ok {
+ resources[ResourceCores] = cpuCap.Value()
+ }
} else {
- remaining := hvRemainingResources(hv, blockedByReservations[hostName])
- if remaining == nil {
+ resources = hvRemainingResources(hv, blockedByReservations[hostName])
+ if resources == nil {
continue
}
- capBytes = remaining[ResourceMemory]
}
- if slots := capBytes / flavorBytes; slots > 0 {
+ if slots := flavorSlots(resources, flavorBytes, flavorVCPUs); slots > 0 {
capacity += slots
candidateHosts = append(candidateHosts, hostName)
}
diff --git a/internal/scheduling/reservations/capacity/controller_test.go b/internal/scheduling/reservations/capacity/controller_test.go
index 0ba4b8f5c..d01110af0 100644
--- a/internal/scheduling/reservations/capacity/controller_test.go
+++ b/internal/scheduling/reservations/capacity/controller_test.go
@@ -226,11 +226,9 @@ func TestReconcileAZ_CreatesCRD(t *testing.T) {
}
hvByName := map[string]hv1.Hypervisor{"host-1": *hv}
- if err := ctrl.reconcileAZ(context.Background(), az,
+ ctrl.reconcileAZ(context.Background(), az,
map[string]compute.FlavorGroupFeature{groupName: groupData},
- hvByName, map[string]int64{}, map[vmUsageKey]vmUsage{}); err != nil {
- t.Fatalf("reconcileAZ failed: %v", err)
- }
+ hvByName, map[string]int64{}, map[vmUsageKey]vmUsage{})
var crd v1alpha1.FlavorGroupCapacity
if err := fakeClient.Get(context.Background(), types.NamespacedName{Name: crdNameFor(groupName, az)}, &crd); err != nil {
@@ -300,11 +298,9 @@ func TestReconcileAZ_SkipsCRDWriteOnSchedulerError(t *testing.T) {
Flavors: []compute.FlavorInGroup{smallFlavor},
}
- if err := ctrl.reconcileAZ(context.Background(), az,
+ ctrl.reconcileAZ(context.Background(), az,
map[string]compute.FlavorGroupFeature{groupName: groupData},
- map[string]hv1.Hypervisor{}, map[string]int64{}, map[vmUsageKey]vmUsage{}); err != nil {
- t.Fatalf("reconcileAZ failed: %v", err)
- }
+ map[string]hv1.Hypervisor{}, map[string]int64{}, map[vmUsageKey]vmUsage{})
// Stale probes → CRD must NOT be written; last good state is preserved.
var list v1alpha1.FlavorGroupCapacityList
@@ -362,13 +358,9 @@ func TestReconcileAZ_IdempotentUpdate(t *testing.T) {
groups := map[string]compute.FlavorGroupFeature{groupName: groupData}
// First call
- if err := ctrl.reconcileAZ(context.Background(), az, groups, hvByName, map[string]int64{}, map[vmUsageKey]vmUsage{}); err != nil {
- t.Fatalf("first reconcileAZ failed: %v", err)
- }
+ ctrl.reconcileAZ(context.Background(), az, groups, hvByName, map[string]int64{}, map[vmUsageKey]vmUsage{})
// Second call — should not error on the already-existing CRD.
- if err := ctrl.reconcileAZ(context.Background(), az, groups, hvByName, map[string]int64{}, map[vmUsageKey]vmUsage{}); err != nil {
- t.Fatalf("second reconcileAZ failed: %v", err)
- }
+ ctrl.reconcileAZ(context.Background(), az, groups, hvByName, map[string]int64{}, map[vmUsageKey]vmUsage{})
var crd v1alpha1.FlavorGroupCapacity
if err := fakeClient.Get(context.Background(), types.NamespacedName{Name: crdName}, &crd); err != nil {
@@ -594,12 +586,9 @@ func TestReconcileAZ_ZeroMemoryFlavorSkipped(t *testing.T) {
SmallestFlavor: compute.FlavorInGroup{Name: "bad-flavor", MemoryMB: 0},
}
// reconcileAZ logs and skips groups with zero memory; it does not return an error.
- err := c.reconcileAZ(context.Background(), "az-a",
+ c.reconcileAZ(context.Background(), "az-a",
map[string]compute.FlavorGroupFeature{"hana-v2": groupData},
nil, nil, nil)
- if err != nil {
- t.Errorf("reconcileAZ should not return error for zero-memory flavor, got: %v", err)
- }
// No CRD should have been created.
var list v1alpha1.FlavorGroupCapacityList
@@ -611,6 +600,136 @@ func TestReconcileAZ_ZeroMemoryFlavorSkipped(t *testing.T) {
}
}
+func TestFlavorSlots(t *testing.T) {
+ const (
+ mem4GiB = 4 * 1024 * 1024 * 1024
+ mem8GiB = 8 * 1024 * 1024 * 1024
+ mem32GiB = 32 * 1024 * 1024 * 1024
+ )
+ tests := []struct {
+ name string
+ memRemaining int64
+ coresRemaining int64
+ flavorMem int64
+ flavorCPUs int64
+ want int64
+ }{
+ {
+ name: "memory is binding constraint",
+ // 8 GiB available, flavor needs 4 GiB and 2 cores; 64 cores available → 2 mem-slots, 32 cpu-slots
+ memRemaining: mem8GiB, coresRemaining: 64, flavorMem: mem4GiB, flavorCPUs: 2, want: 2,
+ },
+ {
+ name: "CPU is binding constraint",
+ // 32 GiB available (fits 8 slots), only 3 cores available (fits 1 slot at 2 vcpus)
+ memRemaining: mem32GiB, coresRemaining: 3, flavorMem: mem4GiB, flavorCPUs: 2, want: 1,
+ },
+ {
+ name: "both constraints equal",
+ memRemaining: mem8GiB, coresRemaining: 4, flavorMem: mem4GiB, flavorCPUs: 2, want: 2,
+ },
+ {
+ name: "zero VCPUs — CPU dimension ignored",
+ memRemaining: mem8GiB, coresRemaining: 0, flavorMem: mem4GiB, flavorCPUs: 0, want: 2,
+ },
+ {
+ name: "not enough memory for even one slot",
+ memRemaining: mem4GiB - 1, coresRemaining: 64, flavorMem: mem4GiB, flavorCPUs: 2, want: 0,
+ },
+ {
+ name: "not enough CPU for even one slot",
+ memRemaining: mem32GiB, coresRemaining: 1, flavorMem: mem4GiB, flavorCPUs: 2, want: 0,
+ },
+ }
+ for _, tt := range tests {
+ t.Run(tt.name, func(t *testing.T) {
+ resources := map[string]int64{
+ ResourceMemory: tt.memRemaining,
+ ResourceCores: tt.coresRemaining,
+ }
+ got := flavorSlots(resources, tt.flavorMem, tt.flavorCPUs)
+ if got != tt.want {
+ t.Errorf("flavorSlots() = %d, want %d", got, tt.want)
+ }
+ })
+ }
+}
+
+func TestComputeTotalCapacity(t *testing.T) {
+ mb := func(mb int64) int64 { return mb * 1024 * 1024 }
+ tests := []struct {
+ name string
+ flavors []v1alpha1.FlavorCapacityStatus
+ specs map[string]compute.FlavorInGroup
+ wantMemBytes int64
+ wantCPU int64
+ }{
+ {
+ name: "single flavor",
+ flavors: []v1alpha1.FlavorCapacityStatus{
+ {FlavorName: "small", TotalCapacityVMSlots: 10},
+ },
+ specs: map[string]compute.FlavorInGroup{
+ "small": {Name: "small", MemoryMB: 4096, VCPUs: 2},
+ },
+ wantMemBytes: 10 * mb(4096),
+ wantCPU: 20,
+ },
+ {
+ name: "picks flavor with most total memory, not most slots",
+ // mem: large wins (2×32GiB=64GiB > 10×4GiB=40GiB); CPU: small wins (10×2=20 > 2×8=16)
+ flavors: []v1alpha1.FlavorCapacityStatus{
+ {FlavorName: "small", TotalCapacityVMSlots: 10},
+ {FlavorName: "large", TotalCapacityVMSlots: 2},
+ },
+ specs: map[string]compute.FlavorInGroup{
+ "small": {Name: "small", MemoryMB: 4096, VCPUs: 2},
+ "large": {Name: "large", MemoryMB: 32768, VCPUs: 8},
+ },
+ wantMemBytes: 2 * mb(32768),
+ wantCPU: 20,
+ },
+ {
+ name: "zero slots excluded",
+ flavors: []v1alpha1.FlavorCapacityStatus{
+ {FlavorName: "small", TotalCapacityVMSlots: 0},
+ {FlavorName: "large", TotalCapacityVMSlots: 3},
+ },
+ specs: map[string]compute.FlavorInGroup{
+ "small": {Name: "small", MemoryMB: 4096, VCPUs: 2},
+ "large": {Name: "large", MemoryMB: 8192, VCPUs: 4},
+ },
+ wantMemBytes: 3 * mb(8192),
+ wantCPU: 12,
+ },
+ {
+ name: "all zero slots",
+ flavors: []v1alpha1.FlavorCapacityStatus{{FlavorName: "small", TotalCapacityVMSlots: 0}},
+ specs: map[string]compute.FlavorInGroup{"small": {MemoryMB: 4096, VCPUs: 2}},
+ wantMemBytes: 0,
+ wantCPU: 0,
+ },
+ {
+ name: "empty input",
+ flavors: nil,
+ specs: map[string]compute.FlavorInGroup{},
+ wantMemBytes: 0,
+ wantCPU: 0,
+ },
+ }
+ for _, tt := range tests {
+ t.Run(tt.name, func(t *testing.T) {
+ gotMem, gotCPU := computeTotalCapacity(tt.flavors, tt.specs)
+ if gotMem != tt.wantMemBytes {
+ t.Errorf("maxMemBytes = %d, want %d", gotMem, tt.wantMemBytes)
+ }
+ if gotCPU != tt.wantCPU {
+ t.Errorf("maxCPUCores = %d, want %d", gotCPU, tt.wantCPU)
+ }
+ })
+ }
+}
+
// Verify that the module-level log variable from reservations package doesn't
// collide with the one in this package.
func TestPackageLogVar(t *testing.T) {
@@ -789,9 +908,7 @@ func TestComputeVMUsage_ZerosOutWhenAllVMsRemoved(t *testing.T) {
}
// Now run reconcileAZ to verify the CRD gets zeroed out.
- if err := ctrl.reconcileAZ(context.Background(), az, groups, hvByName, map[string]int64{}, usageByKey); err != nil {
- t.Fatalf("reconcileAZ failed: %v", err)
- }
+ ctrl.reconcileAZ(context.Background(), az, groups, hvByName, map[string]int64{}, usageByKey)
var crd v1alpha1.FlavorGroupCapacity
if err := fakeClient.Get(context.Background(), types.NamespacedName{Name: crdName}, &crd); err != nil {
diff --git a/internal/scheduling/reservations/capacity/split.go b/internal/scheduling/reservations/capacity/split.go
index ebeefb5bc..6ba73ef0b 100644
--- a/internal/scheduling/reservations/capacity/split.go
+++ b/internal/scheduling/reservations/capacity/split.go
@@ -196,27 +196,34 @@ func allocateRoundRobin(states []groupState, hostRes map[string]map[string]int64
}
}
-// computeUnassigned sums remaining resources on candidate hosts after allocation.
+// computeUnassigned sums remaining resources on candidate hosts after allocation
+// and returns per-host stranded resources for operator visibility.
// Non-candidate hosts are excluded — their leftover is not fragmentation.
-func computeUnassigned(groups []GroupInput, hostRes map[string]map[string]int64) map[string]int64 {
+func computeUnassigned(groups []GroupInput, hostRes map[string]map[string]int64) (unassigned map[string]int64, strandedByHost map[string]map[string]int64) {
candidateSet := make(map[string]struct{})
for _, g := range groups {
for _, h := range g.CandidateHosts {
candidateSet[h] = struct{}{}
}
}
- unassigned := make(map[string]int64)
+ unassigned = make(map[string]int64)
+ strandedByHost = make(map[string]map[string]int64)
for h, res := range hostRes {
if _, isCandidate := candidateSet[h]; !isCandidate {
continue
}
+ hasStranded := false
for r, remaining := range res {
if remaining > 0 {
unassigned[r] += remaining
+ hasStranded = true
}
}
+ if hasStranded {
+ strandedByHost[h] = res
+ }
}
- return unassigned
+ return
}
// collectExclusiveResources builds the exclusive allocation map from group assigned counts.
@@ -248,12 +255,12 @@ func collectExclusiveResources(states []groupState) map[string]map[string]int64
//
// The caller divides exclusiveResources[group][ResourceMemory] by the group's flavor memory
// to obtain the slot count meaningful to that group.
-func SplitCapacity(groups []GroupInput, hosts map[string]HostState) (freeResources, exclusiveResources map[string]map[string]int64, unassigned map[string]int64) {
+func SplitCapacity(groups []GroupInput, hosts map[string]HostState) (freeResources, exclusiveResources map[string]map[string]int64, unassigned map[string]int64, strandedByHost map[string]map[string]int64) {
states := initGroupStates(groups, hosts)
freeResources = computeFreeResources(groups, hosts)
hostRes := copyHostResources(hosts)
allocateRoundRobin(states, hostRes)
- unassigned = computeUnassigned(groups, hostRes)
+ unassigned, strandedByHost = computeUnassigned(groups, hostRes)
exclusiveResources = collectExclusiveResources(states)
- return freeResources, exclusiveResources, unassigned
+ return
}
diff --git a/internal/scheduling/reservations/capacity/split_test.go b/internal/scheduling/reservations/capacity/split_test.go
index a4b1ff212..d5588d0d7 100644
--- a/internal/scheduling/reservations/capacity/split_test.go
+++ b/internal/scheduling/reservations/capacity/split_test.go
@@ -234,7 +234,7 @@ func TestComputeUnassigned(t *testing.T) {
}
for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) {
- got := computeUnassigned(tc.groups, tc.hostRes)
+ got, _ := computeUnassigned(tc.groups, tc.hostRes)
for r, want := range tc.wantUnassigned {
if got[r] != want {
t.Errorf("unassigned[%s] = %d, want %d", r, got[r], want)
@@ -449,7 +449,7 @@ func TestSplitCapacity(t *testing.T) {
for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) {
- free, assigned, unassigned := SplitCapacity(tc.groups, tc.hosts)
+ free, assigned, unassigned, _ := SplitCapacity(tc.groups, tc.hosts)
for groupName, wantMem := range tc.wantAssignedMem {
if got := assigned[groupName][ResourceMemory]; got != wantMem {
@@ -492,7 +492,7 @@ func TestSplitCapacity_SumNeverExceedsTotal(t *testing.T) {
"h3": host(24*GiB, 12),
}
- _, assigned, _ := SplitCapacity(groups, hosts)
+ _, assigned, _, _ := SplitCapacity(groups, hosts)
var totalInstalled, totalAssigned int64
for _, hs := range hosts {
@@ -518,9 +518,9 @@ func TestSplitCapacity_Deterministic(t *testing.T) {
"h2": host(8*GiB, 4),
}
- _, first, firstUnassigned := SplitCapacity(groups, hosts)
+ _, first, firstUnassigned, _ := SplitCapacity(groups, hosts)
for i := range 10 {
- _, got, gotUnassigned := SplitCapacity(groups, hosts)
+ _, got, gotUnassigned, _ := SplitCapacity(groups, hosts)
for _, g := range groups {
if got[g.Name][ResourceMemory] != first[g.Name][ResourceMemory] {
t.Errorf("run %d: assigned[%s][memory] = %d, want %d (non-deterministic)",
From fc47858efaae7393208936640e9bbe6019b05db5 Mon Sep 17 00:00:00 2001
From: Markus Wieland <44964229+SoWieMarkus@users.noreply.github.com>
Date: Thu, 16 Jul 2026 09:09:02 +0200
Subject: [PATCH 08/15] fix: broken workflows with large gh runner (#1050)
Signed-off-by: Markus Wieland
---
.github/workflows/codeql.yaml | 2 +-
.github/workflows/push-images.yaml | 2 +-
2 files changed, 2 insertions(+), 2 deletions(-)
diff --git a/.github/workflows/codeql.yaml b/.github/workflows/codeql.yaml
index 47b415216..c85947d68 100644
--- a/.github/workflows/codeql.yaml
+++ b/.github/workflows/codeql.yaml
@@ -24,7 +24,7 @@ permissions:
jobs:
analyze:
name: CodeQL
- runs-on: large_runner_16core_64gb
+ runs-on: ubuntu-latest #large_runner_16core_64gb
steps:
- name: Check out code
uses: actions/checkout@v7
diff --git a/.github/workflows/push-images.yaml b/.github/workflows/push-images.yaml
index ef3c39e9c..ce71706e1 100644
--- a/.github/workflows/push-images.yaml
+++ b/.github/workflows/push-images.yaml
@@ -17,7 +17,7 @@ jobs:
packages: write
attestations: write
id-token: write
- runs-on: large_runner_16core_64gb
+ runs-on: ubuntu-latest #large_runner_16core_64gb
steps:
- uses: actions/checkout@v7
- name: Set up QEMU
From 5e0cb88ae838b9eba8af1b27f1e51a1498f9528f Mon Sep 17 00:00:00 2001
From: Markus Wieland <44964229+SoWieMarkus@users.noreply.github.com>
Date: Thu, 16 Jul 2026 09:39:37 +0200
Subject: [PATCH 09/15] feat: add configurable HTTP User-Agent for service
identity (#1045)
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
Cortex previously made outgoing HTTP requests without identifying
itself, so receiving services couldn't tell who was calling them. This
PR adds a unified, configurable User-Agent (e.g.
`cortex-nova/sha-70af93a8`) to all outgoing requests, set per deployment
via values.yaml.
The tricky part is that cortex has two kinds of HTTP clients. Most
requests go through the shared http.DefaultTransport (via
http.DefaultClient or plain &http.Client{}), which we cover in one shot
at startup with
`httpext.WrapTransport(&http.DefaultTransport).SetOverrideUserAgent(...)`
(from `sapcc/go-bits`). The SSO clients however, build their own
transport and therefore don't share the default one — so they need a
separate mechanism. For those, pkg/sso sets the header on its own
transport via sso.SetUserAgent(...), with sso.WrapUserAgent(rt) for
hand-built transports.
The User-Agent is configured as two fields (component + version) rather
than a single string, so they map cleanly onto the component/version
format of the go-bits lib without any join-then-split. The Helm library
chart defaults them to the release name and chart appVersion (the
deployed image tag), and bundles can override as needed.
---
cmd/manager/main.go | 21 +++++
cmd/shim/main.go | 25 ++++++
.../cortex/templates/manager/manager.yaml | 8 ++
helm/library/cortex/values.yaml | 7 ++
internal/shim/placement/shim.go | 2 +-
pkg/sso/sso.go | 45 +++++++++-
pkg/sso/sso_test.go | 83 ++++++++++++++++++-
7 files changed, 185 insertions(+), 6 deletions(-)
diff --git a/cmd/manager/main.go b/cmd/manager/main.go
index d8508d1c7..a7ae683d3 100644
--- a/cmd/manager/main.go
+++ b/cmd/manager/main.go
@@ -67,6 +67,7 @@ import (
"github.com/cobaltcore-dev/cortex/pkg/conf"
"github.com/cobaltcore-dev/cortex/pkg/monitoring"
"github.com/cobaltcore-dev/cortex/pkg/multicluster"
+ "github.com/cobaltcore-dev/cortex/pkg/sso"
"github.com/cobaltcore-dev/cortex/pkg/task"
hv1 "github.com/cobaltcore-dev/openstack-hypervisor-operator/api/v1"
"github.com/sapcc/go-bits/httpext"
@@ -90,6 +91,15 @@ func init() {
// +kubebuilder:scaffold:scheme
}
+// UserAgentConfig identifies this cortex deployment to the services it talks
+// to. Rendered by helm from the release name and chart version.
+type UserAgentConfig struct {
+ // Component is the service name, e.g. "cortex-nova".
+ Component string `json:"component,omitempty"`
+ // Version is the deployed version, e.g. "sha-70af93a8".
+ Version string `json:"version,omitempty"`
+}
+
type MainConfig struct {
// ID used to identify leader election participants.
LeaderElectionID string `json:"leaderElectionID,omitempty"`
@@ -97,6 +107,8 @@ type MainConfig struct {
EnabledControllers []string `json:"enabledControllers"`
// List of enabled tasks.
EnabledTasks []string `json:"enabledTasks"`
+ // User-Agent sent with all outgoing HTTP requests.
+ UserAgent UserAgentConfig `json:"userAgent,omitempty"`
}
//nolint:gocyclo
@@ -105,6 +117,15 @@ func main() {
mainConfig := conf.GetConfigOrDie[MainConfig]()
restConfig := ctrl.GetConfigOrDie()
+ // Identify this cortex deployment to the services it talks to via the
+ // User-Agent header, before any HTTP requests are made. The shared
+ // http.DefaultTransport covers http.DefaultClient and any http.Client
+ // without its own transport; SSO clients build their own transport and
+ // are handled separately by pkg/sso.
+ httpext.WrapTransport(&http.DefaultTransport).
+ SetOverrideUserAgent(mainConfig.UserAgent.Component, mainConfig.UserAgent.Version)
+ sso.SetUserAgent(mainConfig.UserAgent.Component, mainConfig.UserAgent.Version)
+
// Custom entrypoint for scheduler e2e tests.
// Usage: /main [json-override]
// The optional json-override is merged on top of the ConfigMap config, e.g.:
diff --git a/cmd/shim/main.go b/cmd/shim/main.go
index 45a6938de..29865e0c5 100644
--- a/cmd/shim/main.go
+++ b/cmd/shim/main.go
@@ -17,6 +17,7 @@ import (
"github.com/cobaltcore-dev/cortex/pkg/conf"
"github.com/cobaltcore-dev/cortex/pkg/monitoring"
"github.com/cobaltcore-dev/cortex/pkg/multicluster"
+ "github.com/cobaltcore-dev/cortex/pkg/sso"
hv1 "github.com/cobaltcore-dev/openstack-hypervisor-operator/api/v1"
"github.com/sapcc/go-bits/httpext"
"k8s.io/apimachinery/pkg/runtime"
@@ -50,11 +51,35 @@ func init() {
utilruntime.Must(hv1.AddToScheme(scheme)) // Hypervisor crd
}
+// UserAgentConfig identifies this cortex deployment to the services it talks
+// to. Rendered by helm from the release name and chart version.
+type UserAgentConfig struct {
+ // Component is the service name, e.g. "cortex-nova".
+ Component string `json:"component,omitempty"`
+ // Version is the deployed version, e.g. "sha-70af93a8".
+ Version string `json:"version,omitempty"`
+}
+
+type MainConfig struct {
+ // User-Agent sent with all outgoing HTTP requests.
+ UserAgent UserAgentConfig `json:"userAgent,omitempty"`
+}
+
func main() {
ctx := ctrl.SetupSignalHandler()
+ mainConfig := conf.GetConfigOrDie[MainConfig]()
restConfig := ctrl.GetConfigOrDie()
+ // Identify this cortex deployment to the services it talks to via the
+ // User-Agent header, before any HTTP requests are made. The shared
+ // http.DefaultTransport covers http.DefaultClient and any http.Client
+ // without its own transport; SSO clients build their own transport and
+ // are handled separately by pkg/sso.
+ httpext.WrapTransport(&http.DefaultTransport).
+ SetOverrideUserAgent(mainConfig.UserAgent.Component, mainConfig.UserAgent.Version)
+ sso.SetUserAgent(mainConfig.UserAgent.Component, mainConfig.UserAgent.Version)
+
var metricsAddr string
var apiBindAddr string
var metricsCertPath, metricsCertName, metricsCertKey string
diff --git a/helm/library/cortex/templates/manager/manager.yaml b/helm/library/cortex/templates/manager/manager.yaml
index 052b99170..190232d8c 100644
--- a/helm/library/cortex/templates/manager/manager.yaml
+++ b/helm/library/cortex/templates/manager/manager.yaml
@@ -133,6 +133,14 @@ data:
{{- if .Values.conf }}
{{- $mergedConf = mergeOverwrite .Values.conf $mergedConf }}
{{- end }}
+ {{- /* Default the outgoing HTTP User-Agent so every deployment is */ -}}
+ {{- /* identifiable by the services it talks to: component from the */ -}}
+ {{- /* release name, version from the chart appVersion (the image tag). */ -}}
+ {{- $userAgent := dict "component" .Release.Name "version" .Chart.AppVersion }}
+ {{- if hasKey $mergedConf "userAgent" }}
+ {{- $userAgent = mergeOverwrite $userAgent $mergedConf.userAgent }}
+ {{- end }}
+ {{- $_ := set $mergedConf "userAgent" $userAgent }}
{{ toJson $mergedConf }}
---
apiVersion: v1
diff --git a/helm/library/cortex/values.yaml b/helm/library/cortex/values.yaml
index 49f6cb227..7627d86a5 100644
--- a/helm/library/cortex/values.yaml
+++ b/helm/library/cortex/values.yaml
@@ -109,6 +109,13 @@ conf:
schedulingDomain: cortex
# Used to differentiate different cortex deployments in the same cluster (e.g. leader election ID)
leaderElectionID: cortex-unknown
+ # User-Agent sent with all outgoing HTTP requests, identifying this cortex
+ # deployment to the services it talks to. When unset, it defaults to the
+ # release name and chart appVersion (e.g. component "cortex-nova",
+ # version "sha-70af93a8"), producing "cortex-nova/sha-70af93a8".
+ # userAgent:
+ # component: cortex-nova
+ # version: v1.2.3
enabledControllers:
# The explanation controller is available for all decision resources.
- explanation-controller
diff --git a/internal/shim/placement/shim.go b/internal/shim/placement/shim.go
index 77ff72a66..5731e8793 100644
--- a/internal/shim/placement/shim.go
+++ b/internal/shim/placement/shim.go
@@ -379,7 +379,7 @@ func (s *Shim) initHTTPClient(ctx context.Context) error {
transport.ResponseHeaderTimeout = 60 * time.Second
transport.ExpectContinueTimeout = 1 * time.Second
transport.IdleConnTimeout = 90 * time.Second
- s.httpClient = &http.Client{Transport: transport, Timeout: 60 * time.Second}
+ s.httpClient = &http.Client{Transport: sso.WrapUserAgent(transport), Timeout: 60 * time.Second}
setupLog.Info("Testing connection to placement API", "url", s.config.PlacementURL)
req, err := http.NewRequestWithContext(ctx, http.MethodGet, s.config.PlacementURL, http.NoBody)
diff --git a/pkg/sso/sso.go b/pkg/sso/sso.go
index 5f8241bea..ac86e089c 100644
--- a/pkg/sso/sso.go
+++ b/pkg/sso/sso.go
@@ -16,6 +16,27 @@ import (
"sigs.k8s.io/controller-runtime/pkg/client"
)
+// userAgent is the User-Agent value set on every outgoing HTTP request made
+// through clients created by this package, identifying cortex as the caller.
+// It defaults to "cortex" and can be overridden once at process startup via
+// SetUserAgent (e.g. "cortex-nova/sha-70af93a8").
+var userAgent = "cortex"
+
+// SetUserAgent sets the User-Agent that this package's HTTP clients send. The
+// component and version are combined as "component/version"; if version is
+// empty, only the component is used. An empty component leaves the default
+// ("cortex") unchanged.
+func SetUserAgent(component, version string) {
+ if component == "" {
+ return
+ }
+ if version == "" {
+ userAgent = component
+ return
+ }
+ userAgent = component + "/" + version
+}
+
// Configuration for single-sign-on (SSO).
type SSOConfig struct {
Cert string `json:"cert,omitempty"`
@@ -36,6 +57,28 @@ func (lrt *requestLogger) RoundTrip(req *http.Request) (*http.Response, error) {
return lrt.T.RoundTrip(req)
}
+// Custom HTTP round tripper that sets the cortex User-Agent on every request,
+// identifying cortex as the caller. SSO clients build their own transport and
+// therefore need this, as they don't share http.DefaultTransport.
+type userAgentTransport struct {
+ T http.RoundTripper
+}
+
+// RoundTrip sets the User-Agent header to the cortex user agent. The request
+// is cloned so the caller's request is not mutated.
+func (uat *userAgentTransport) RoundTrip(req *http.Request) (*http.Response, error) {
+ req = req.Clone(req.Context())
+ req.Header.Set("User-Agent", userAgent)
+ return uat.T.RoundTrip(req)
+}
+
+// WrapUserAgent wraps the given RoundTripper so that every request carries the
+// cortex User-Agent header. Use it for hand-built transports that neither go
+// through NewHTTPClient nor the shared http.DefaultTransport.
+func WrapUserAgent(rt http.RoundTripper) http.RoundTripper {
+ return &userAgentTransport{T: rt}
+}
+
// Kubernetes connector which initializes the sso connection from a secret.
type Connector struct{ client.Client }
@@ -110,5 +153,5 @@ func NewHTTPClient(conf SSOConfig) (*http.Client, error) {
if conf.Cert == "" {
slog.Debug("making http requests without SSO")
}
- return &http.Client{Transport: &requestLogger{T: transport}}, nil
+ return &http.Client{Transport: &userAgentTransport{T: &requestLogger{T: transport}}}, nil
}
diff --git a/pkg/sso/sso_test.go b/pkg/sso/sso_test.go
index cf8551392..ca3a453a7 100644
--- a/pkg/sso/sso_test.go
+++ b/pkg/sso/sso_test.go
@@ -8,6 +8,76 @@ import (
"testing"
)
+// captureTransport records the last request it saw and returns a canned response.
+type captureTransport struct{ last *http.Request }
+
+func (c *captureTransport) RoundTrip(req *http.Request) (*http.Response, error) {
+ c.last = req
+ return &http.Response{StatusCode: http.StatusOK, Body: http.NoBody, Header: make(http.Header)}, nil
+}
+
+func TestSetUserAgent(t *testing.T) {
+ old := userAgent
+ defer func() { userAgent = old }()
+
+ tests := []struct {
+ name string
+ component string
+ version string
+ want string
+ }{
+ {name: "ComponentAndVersion", component: "cortex-nova", version: "sha-abc", want: "cortex-nova/sha-abc"},
+ {name: "ComponentOnly", component: "cortex-nova", version: "", want: "cortex-nova"},
+ {name: "EmptyComponentKeepsDefault", component: "", version: "sha-abc", want: "cortex"},
+ }
+ for _, tt := range tests {
+ t.Run(tt.name, func(t *testing.T) {
+ userAgent = "cortex"
+ SetUserAgent(tt.component, tt.version)
+ if userAgent != tt.want {
+ t.Errorf("userAgent = %q, want %q", userAgent, tt.want)
+ }
+ })
+ }
+}
+
+func TestUserAgentTransport(t *testing.T) {
+ old := userAgent
+ userAgent = "cortex-nova/test"
+ defer func() { userAgent = old }()
+
+ tests := []struct {
+ name string
+ incomingUA string
+ }{
+ {name: "NoExistingUA", incomingUA: ""},
+ {name: "OverridesExistingUA", incomingUA: "gophercloud/2.0"},
+ }
+ for _, tt := range tests {
+ t.Run(tt.name, func(t *testing.T) {
+ capture := &captureTransport{}
+ uat := &userAgentTransport{T: capture}
+ req, err := http.NewRequest(http.MethodGet, "https://example.com", http.NoBody)
+ if err != nil {
+ t.Fatalf("NewRequest() error = %v", err)
+ }
+ if tt.incomingUA != "" {
+ req.Header.Set("User-Agent", tt.incomingUA)
+ }
+ if _, err := uat.RoundTrip(req); err != nil {
+ t.Fatalf("RoundTrip() error = %v", err)
+ }
+ if got := capture.last.Header.Get("User-Agent"); got != "cortex-nova/test" {
+ t.Errorf("User-Agent = %q, want %q", got, "cortex-nova/test")
+ }
+ // The original request must not be mutated.
+ if got := req.Header.Get("User-Agent"); got != tt.incomingUA {
+ t.Errorf("original request User-Agent mutated: got %q, want %q", got, tt.incomingUA)
+ }
+ })
+ }
+}
+
func TestNewHTTPClient(t *testing.T) {
tests := []struct {
name string
@@ -154,14 +224,19 @@ nyCru8FaKdd+A5MBMSTb8MX0LcnWvdQ=
t.Fatalf("NewHTTPClient() error = %v, want no error", err)
}
- transport, ok := client.Transport.(*requestLogger)
+ transport, ok := client.Transport.(*userAgentTransport)
+ if !ok {
+ t.Fatalf("Expected transport to be of type *userAgentTransport, got %T", client.Transport)
+ }
+
+ logger, ok := transport.T.(*requestLogger)
if !ok {
- t.Fatalf("Expected transport to be of type *requestLogger, got %T", client.Transport)
+ t.Fatalf("Expected transport to be of type *requestLogger, got %T", transport.T)
}
- httpTransport, ok := transport.T.(*http.Transport)
+ httpTransport, ok := logger.T.(*http.Transport)
if !ok {
- t.Fatalf("Expected inner transport to be of type *http.Transport, got %T", transport.T)
+ t.Fatalf("Expected inner transport to be of type *http.Transport, got %T", logger.T)
}
if httpTransport.TLSClientConfig == nil {
From d0fa4713f2c74afbee1e134605692af708edb506 Mon Sep 17 00:00:00 2001
From: "github-actions[bot]"
<41898282+github-actions[bot]@users.noreply.github.com>
Date: Thu, 16 Jul 2026 10:02:52 +0200
Subject: [PATCH 10/15] bump app version [skip ci] (#1037)
bump app version [skip ci]
```
bumped cortex: sha-ee9cd485 -> sha-fc47858e
bumped cortex-shim: sha-ee9cd485 -> sha-7e36a65c
bumped cortex-postgres: sha-af707446 -> sha-e06153f8
```
Co-authored-by: SoWieMarkus <44964229+SoWieMarkus@users.noreply.github.com>
---
helm/library/cortex-postgres/Chart.yaml | 2 +-
helm/library/cortex-shim/Chart.yaml | 2 +-
helm/library/cortex/Chart.yaml | 2 +-
3 files changed, 3 insertions(+), 3 deletions(-)
diff --git a/helm/library/cortex-postgres/Chart.yaml b/helm/library/cortex-postgres/Chart.yaml
index b64823e29..6b7edd809 100644
--- a/helm/library/cortex-postgres/Chart.yaml
+++ b/helm/library/cortex-postgres/Chart.yaml
@@ -6,4 +6,4 @@ name: cortex-postgres
description: Postgres setup for Cortex.
type: application
version: 0.6.8
-appVersion: "sha-af707446"
+appVersion: "sha-e06153f8"
diff --git a/helm/library/cortex-shim/Chart.yaml b/helm/library/cortex-shim/Chart.yaml
index 3ff6405b7..846787384 100644
--- a/helm/library/cortex-shim/Chart.yaml
+++ b/helm/library/cortex-shim/Chart.yaml
@@ -3,6 +3,6 @@ name: cortex-shim
description: A Helm chart to distribute cortex shims.
type: application
version: 0.1.6
-appVersion: "sha-ee9cd485"
+appVersion: "sha-7e36a65c"
icon: "https://example.com/icon.png"
dependencies: []
diff --git a/helm/library/cortex/Chart.yaml b/helm/library/cortex/Chart.yaml
index 64433287e..338b0289c 100644
--- a/helm/library/cortex/Chart.yaml
+++ b/helm/library/cortex/Chart.yaml
@@ -3,6 +3,6 @@ name: cortex
description: A Helm chart to distribute cortex.
type: application
version: 0.3.0
-appVersion: "sha-ee9cd485"
+appVersion: "sha-fc47858e"
icon: "https://example.com/icon.png"
dependencies: []
From ee4cb62230d9ec5b5e443b689ac5403e05cfbfcb Mon Sep 17 00:00:00 2001
From: Markus Wieland <44964229+SoWieMarkus@users.noreply.github.com>
Date: Thu, 16 Jul 2026 10:15:42 +0200
Subject: [PATCH 11/15] fix: add signed-off-by line to bump app version commit
message (#1052)
Signed-off-by: Markus Wieland
---
.github/workflows/update-appversion.yml | 2 ++
1 file changed, 2 insertions(+)
diff --git a/.github/workflows/update-appversion.yml b/.github/workflows/update-appversion.yml
index a89de4203..a83623dfb 100644
--- a/.github/workflows/update-appversion.yml
+++ b/.github/workflows/update-appversion.yml
@@ -74,6 +74,8 @@ jobs:
bump app version [skip ci]
${{ steps.summary.outputs.log }}
+
+ Signed-off-by: ${{ github.actor }} <${{ github.actor_id }}+${{ github.actor }}@users.noreply.github.com>
title: bump app version [skip ci]
body: |
bump app version [skip ci]
From 9ba9a62bce7ff4c50ad07e8128c5c32f7b7bb3ab Mon Sep 17 00:00:00 2001
From: "cortex-ai-agents[bot]"
<279748396+cortex-ai-agents[bot]@users.noreply.github.com>
Date: Thu, 16 Jul 2026 12:53:25 +0200
Subject: [PATCH 12/15] Add changelog entry for release PR #1051 (#1055)
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
## Summary
- Add changelog entry documenting all component releases included in PR
#1051
- Merge after #1051
## Test plan
- [ ] Verify CHANGELOG.md renders correctly on GitHub
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-authored-by: cortex-ai-agents[bot] <279748396+cortex-ai-agents[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7
---
CHANGELOG.md | 50 ++++++++++++++++++++++++++++++++++++++++++++++++++
1 file changed, 50 insertions(+)
diff --git a/CHANGELOG.md b/CHANGELOG.md
index 4a491ea09..bd7cc17f7 100644
--- a/CHANGELOG.md
+++ b/CHANGELOG.md
@@ -1,5 +1,55 @@
# Changelog
+## 2026-07-16 — [#1051](https://github.com/cobaltcore-dev/cortex/pull/1051)
+
+### cortex v0.3.1 (sha-fc47858e)
+
+Non-breaking changes:
+- Add configurable HTTP User-Agent for service identity in SSO/HTTP client layer ([#1045](https://github.com/cobaltcore-dev/cortex/pull/1045))
+- Enhance multicluster client 'no cluster matched' error log messages with richer diagnostic context ([#1043](https://github.com/cobaltcore-dev/cortex/pull/1043))
+- Account for CPU & memory as a binding constraint in slot counting ([#1044](https://github.com/cobaltcore-dev/cortex/pull/1044))
+- Persist pipeline results before history upsert to prevent data loss on error ([#1047](https://github.com/cobaltcore-dev/cortex/pull/1047))
+- Fix broken CI workflows with large GitHub runner configuration ([#1050](https://github.com/cobaltcore-dev/cortex/pull/1050))
+- Add missing `SkipCommittedResourceTracking` to Options table and failover docs ([#1034](https://github.com/cobaltcore-dev/cortex/pull/1034))
+- Update external dependencies ([#1040](https://github.com/cobaltcore-dev/cortex/pull/1040), [#1032](https://github.com/cobaltcore-dev/cortex/pull/1032))
+
+### cortex-shim v0.1.7 (sha-7e36a65c)
+
+Includes updated image sha-7e36a65c.
+
+### cortex-postgres v0.6.9 (sha-e06153f8)
+
+Non-breaking changes:
+- Rebuild image to resolve CVEs ([#1033](https://github.com/cobaltcore-dev/cortex/pull/1033))
+
+### cortex-nova v0.0.81
+
+Includes updated charts cortex v0.3.1, cortex-postgres v0.6.9.
+
+### cortex-cinder v0.0.81
+
+Includes updated charts cortex v0.3.1, cortex-postgres v0.6.9.
+
+### cortex-manila v0.0.81
+
+Includes updated charts cortex v0.3.1, cortex-postgres v0.6.9.
+
+### cortex-crds v0.0.81
+
+Includes updated chart cortex v0.3.1.
+
+### cortex-ironcore v0.0.81
+
+Includes updated chart cortex v0.3.1.
+
+### cortex-pods v0.0.81
+
+Includes updated chart cortex v0.3.1.
+
+### cortex-placement-shim v0.1.7
+
+Includes updated chart cortex-shim v0.1.7.
+
## 2026-07-13 — [#1036](https://github.com/cobaltcore-dev/cortex/pull/1036)
### cortex v0.3.0 (sha-ee9cd485)
From c6322ff32d9e8a67e3f97b258a82b3ebe5949779 Mon Sep 17 00:00:00 2001
From: "cortex-ai-agents[bot]"
<279748396+cortex-ai-agents[bot]@users.noreply.github.com>
Date: Thu, 16 Jul 2026 12:53:47 +0200
Subject: [PATCH 13/15] Bump chart versions for release PR #1051 (#1053)
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
## Summary
- Bump helm chart versions for release PR #1051
- Bumped: cortex 0.3.0→0.3.1, cortex-shim 0.1.6→0.1.7, cortex-postgres
0.6.8→0.6.9, cortex-prometheus-operator 0.2.1→0.2.2, bundles
0.0.80→0.0.81, cortex-placement-shim 0.1.6→0.1.7
- This PR must be merged before #1051
## Test plan
- [ ] Verify all Chart.yaml version bumps are correct
- [ ] Ensure CI passes with updated chart versions
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-authored-by: cortex-ai-agents[bot] <279748396+cortex-ai-agents[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7
---
helm/bundles/cortex-cinder/Chart.yaml | 8 ++++----
helm/bundles/cortex-crds/Chart.yaml | 4 ++--
helm/bundles/cortex-ironcore/Chart.yaml | 4 ++--
helm/bundles/cortex-manila/Chart.yaml | 8 ++++----
helm/bundles/cortex-nova/Chart.yaml | 8 ++++----
helm/bundles/cortex-placement-shim/Chart.yaml | 4 ++--
helm/bundles/cortex-pods/Chart.yaml | 4 ++--
helm/dev/cortex-prometheus-operator/Chart.yaml | 2 +-
helm/library/cortex-postgres/Chart.yaml | 2 +-
helm/library/cortex-shim/Chart.yaml | 2 +-
helm/library/cortex/Chart.yaml | 2 +-
11 files changed, 24 insertions(+), 24 deletions(-)
diff --git a/helm/bundles/cortex-cinder/Chart.yaml b/helm/bundles/cortex-cinder/Chart.yaml
index 3f954cd0b..858510351 100644
--- a/helm/bundles/cortex-cinder/Chart.yaml
+++ b/helm/bundles/cortex-cinder/Chart.yaml
@@ -5,23 +5,23 @@ apiVersion: v2
name: cortex-cinder
description: A Helm chart deploying Cortex for Cinder.
type: application
-version: 0.0.80
+version: 0.0.81
appVersion: 0.1.0
dependencies:
# from: file://../../library/cortex-postgres
- name: cortex-postgres
repository: oci://ghcr.io/cobaltcore-dev/cortex/charts
- version: 0.6.8
+ version: 0.6.9
# from: file://../../library/cortex
- name: cortex
repository: oci://ghcr.io/cobaltcore-dev/cortex/charts
- version: 0.3.0
+ version: 0.3.1
alias: cortex-knowledge-controllers
# from: file://../../library/cortex
- name: cortex
repository: oci://ghcr.io/cobaltcore-dev/cortex/charts
- version: 0.3.0
+ version: 0.3.1
alias: cortex-scheduling-controllers
# Owner info adds a configmap to the kubernetes cluster with information on
diff --git a/helm/bundles/cortex-crds/Chart.yaml b/helm/bundles/cortex-crds/Chart.yaml
index 13eb15db9..c8f73b1c7 100644
--- a/helm/bundles/cortex-crds/Chart.yaml
+++ b/helm/bundles/cortex-crds/Chart.yaml
@@ -5,13 +5,13 @@ apiVersion: v2
name: cortex-crds
description: A Helm chart deploying Cortex CRDs.
type: application
-version: 0.0.80
+version: 0.0.81
appVersion: 0.1.0
dependencies:
# from: file://../../library/cortex
- name: cortex
repository: oci://ghcr.io/cobaltcore-dev/cortex/charts
- version: 0.3.0
+ version: 0.3.1
# Owner info adds a configmap to the kubernetes cluster with information on
# the service owner. This makes it easier to find out who to contact in case
diff --git a/helm/bundles/cortex-ironcore/Chart.yaml b/helm/bundles/cortex-ironcore/Chart.yaml
index 4c0beecb2..5e42a57aa 100644
--- a/helm/bundles/cortex-ironcore/Chart.yaml
+++ b/helm/bundles/cortex-ironcore/Chart.yaml
@@ -5,13 +5,13 @@ apiVersion: v2
name: cortex-ironcore
description: A Helm chart deploying Cortex for IronCore.
type: application
-version: 0.0.80
+version: 0.0.81
appVersion: 0.1.0
dependencies:
# from: file://../../library/cortex
- name: cortex
repository: oci://ghcr.io/cobaltcore-dev/cortex/charts
- version: 0.3.0
+ version: 0.3.1
# Owner info adds a configmap to the kubernetes cluster with information on
# the service owner. This makes it easier to find out who to contact in case
diff --git a/helm/bundles/cortex-manila/Chart.yaml b/helm/bundles/cortex-manila/Chart.yaml
index b7da9c9e0..3be82952a 100644
--- a/helm/bundles/cortex-manila/Chart.yaml
+++ b/helm/bundles/cortex-manila/Chart.yaml
@@ -5,23 +5,23 @@ apiVersion: v2
name: cortex-manila
description: A Helm chart deploying Cortex for Manila.
type: application
-version: 0.0.80
+version: 0.0.81
appVersion: 0.1.0
dependencies:
# from: file://../../library/cortex-postgres
- name: cortex-postgres
repository: oci://ghcr.io/cobaltcore-dev/cortex/charts
- version: 0.6.8
+ version: 0.6.9
# from: file://../../library/cortex
- name: cortex
repository: oci://ghcr.io/cobaltcore-dev/cortex/charts
- version: 0.3.0
+ version: 0.3.1
alias: cortex-knowledge-controllers
# from: file://../../library/cortex
- name: cortex
repository: oci://ghcr.io/cobaltcore-dev/cortex/charts
- version: 0.3.0
+ version: 0.3.1
alias: cortex-scheduling-controllers
# Owner info adds a configmap to the kubernetes cluster with information on
diff --git a/helm/bundles/cortex-nova/Chart.yaml b/helm/bundles/cortex-nova/Chart.yaml
index 27e9b455a..7df970361 100644
--- a/helm/bundles/cortex-nova/Chart.yaml
+++ b/helm/bundles/cortex-nova/Chart.yaml
@@ -5,23 +5,23 @@ apiVersion: v2
name: cortex-nova
description: A Helm chart deploying Cortex for Nova.
type: application
-version: 0.0.80
+version: 0.0.81
appVersion: 0.1.0
dependencies:
# from: file://../../library/cortex-postgres
- name: cortex-postgres
repository: oci://ghcr.io/cobaltcore-dev/cortex/charts
- version: 0.6.8
+ version: 0.6.9
# from: file://../../library/cortex
- name: cortex
repository: oci://ghcr.io/cobaltcore-dev/cortex/charts
- version: 0.3.0
+ version: 0.3.1
alias: cortex-knowledge-controllers
# from: file://../../library/cortex
- name: cortex
repository: oci://ghcr.io/cobaltcore-dev/cortex/charts
- version: 0.3.0
+ version: 0.3.1
alias: cortex-scheduling-controllers
# Owner info adds a configmap to the kubernetes cluster with information on
diff --git a/helm/bundles/cortex-placement-shim/Chart.yaml b/helm/bundles/cortex-placement-shim/Chart.yaml
index dfc035bda..36905ed39 100644
--- a/helm/bundles/cortex-placement-shim/Chart.yaml
+++ b/helm/bundles/cortex-placement-shim/Chart.yaml
@@ -5,13 +5,13 @@ apiVersion: v2
name: cortex-placement-shim
description: A Helm chart deploying the Cortex placement shim.
type: application
-version: 0.1.6
+version: 0.1.7
appVersion: 0.1.0
dependencies:
# from: file://../../library/cortex-shim
- name: cortex-shim
repository: oci://ghcr.io/cobaltcore-dev/cortex/charts
- version: 0.1.6
+ version: 0.1.7
# Owner info adds a configmap to the kubernetes cluster with information on
# the service owner. This makes it easier to find out who to contact in case
# of issues. See: https://github.com/sapcc/helm-charts/pkgs/container/helm-charts%2Fowner-info
diff --git a/helm/bundles/cortex-pods/Chart.yaml b/helm/bundles/cortex-pods/Chart.yaml
index f58ac7cd3..864e88447 100644
--- a/helm/bundles/cortex-pods/Chart.yaml
+++ b/helm/bundles/cortex-pods/Chart.yaml
@@ -5,13 +5,13 @@ apiVersion: v2
name: cortex-pods
description: A Helm chart deploying Cortex for Pods.
type: application
-version: 0.0.80
+version: 0.0.81
appVersion: 0.1.0
dependencies:
# from: file://../../library/cortex
- name: cortex
repository: oci://ghcr.io/cobaltcore-dev/cortex/charts
- version: 0.3.0
+ version: 0.3.1
# Owner info adds a configmap to the kubernetes cluster with information on
# the service owner. This makes it easier to find out who to contact in case
diff --git a/helm/dev/cortex-prometheus-operator/Chart.yaml b/helm/dev/cortex-prometheus-operator/Chart.yaml
index 250b1f1a7..fc8798062 100644
--- a/helm/dev/cortex-prometheus-operator/Chart.yaml
+++ b/helm/dev/cortex-prometheus-operator/Chart.yaml
@@ -5,7 +5,7 @@ apiVersion: v2
name: cortex-prometheus-operator
description: Prometheus operator setup for Cortex.
type: application
-version: 0.2.1
+version: 0.2.2
dependencies:
# CRDs of the prometheus operator, such as PrometheusRule, ServiceMonitor, etc.
- name: kube-prometheus-stack
diff --git a/helm/library/cortex-postgres/Chart.yaml b/helm/library/cortex-postgres/Chart.yaml
index 6b7edd809..b5abe1d7c 100644
--- a/helm/library/cortex-postgres/Chart.yaml
+++ b/helm/library/cortex-postgres/Chart.yaml
@@ -5,5 +5,5 @@ apiVersion: v2
name: cortex-postgres
description: Postgres setup for Cortex.
type: application
-version: 0.6.8
+version: 0.6.9
appVersion: "sha-e06153f8"
diff --git a/helm/library/cortex-shim/Chart.yaml b/helm/library/cortex-shim/Chart.yaml
index 846787384..5e5507ed7 100644
--- a/helm/library/cortex-shim/Chart.yaml
+++ b/helm/library/cortex-shim/Chart.yaml
@@ -2,7 +2,7 @@ apiVersion: v2
name: cortex-shim
description: A Helm chart to distribute cortex shims.
type: application
-version: 0.1.6
+version: 0.1.7
appVersion: "sha-7e36a65c"
icon: "https://example.com/icon.png"
dependencies: []
diff --git a/helm/library/cortex/Chart.yaml b/helm/library/cortex/Chart.yaml
index 338b0289c..bc45750c5 100644
--- a/helm/library/cortex/Chart.yaml
+++ b/helm/library/cortex/Chart.yaml
@@ -2,7 +2,7 @@ apiVersion: v2
name: cortex
description: A Helm chart to distribute cortex.
type: application
-version: 0.3.0
+version: 0.3.1
appVersion: "sha-fc47858e"
icon: "https://example.com/icon.png"
dependencies: []
From a6c0bb3b92b7a1a77b0e2cc282823318372bd510 Mon Sep 17 00:00:00 2001
From: mblos <156897072+mblos@users.noreply.github.com>
Date: Thu, 16 Jul 2026 16:34:51 +0200
Subject: [PATCH 14/15] feat: update release command, with easier pr structure,
updating release PR (#1058)
- simplifies the /release workflow from three deliverables (separate
bump PR + changelog PR + description update) to two: a single prep PR
combining chart bumps and changelog
- updates the release PR description and title
---------
Signed-off-by: mblos
---
.claude/commands/release.md | 71 +++++++++++++++++++++----------------
1 file changed, 40 insertions(+), 31 deletions(-)
diff --git a/.claude/commands/release.md b/.claude/commands/release.md
index 083d9ad14..d1388c576 100644
--- a/.claude/commands/release.md
+++ b/.claude/commands/release.md
@@ -1,15 +1,14 @@
---
allowed-tools: Read, Write, Edit, Bash(*), Agent
-description: Release orchestrator — opens a chart-bump PR, opens a changelog PR, and rewrites the release PR description to reference both. Usage: /release PR_NUMBER
+description: Release orchestrator — opens a single release-prep PR combining changelog and chart bumps, and rewrites the release PR description. Usage: /release PR_NUMBER
---
# Release Orchestrator
-You orchestrate the release process for a given release PR. Three deliverables, in order:
+You orchestrate the release process for a given release PR. Two deliverables, in order:
-1. A bump PR for helm chart versions (`release/bump-charts-`).
-2. A changelog PR with the release notes (`release/changelog-`), using the bumped versions.
-3. The release PR description updated with the changelog and references to both PRs.
+1. A single prep PR (`release/prepare-`) combining the changelog entry and helm chart version bumps.
+2. The release PR description updated with the changelog and a reference to the prep PR.
You are the only mutator. The investigator subagents — `release-digest`, `release-bump-planner`, `release-changelog-writer` — are read-only by construction. They return text; you apply edits, run git, push branches, and dispatch `pull-request-creator` to open PRs. Never call `gh pr create` directly.
@@ -17,9 +16,18 @@ You are the only mutator. The investigator subagents — `release-digest`, `rele
## Phase 1: Setup
-Read `AGENTS.md`. Capture `` from the user's invocation. Then:
+Read `AGENTS.md`. Capture `` from the user's invocation. If no number was provided, find the open PR targeting `main` whose head branch matches a release pattern:
+```sh
+gh pr list --state open --base main --json number,title,headRefName | \
+ jq '.[] | select(.headRefName | test("release|bump-app-version"; "i"))'
```
+
+If exactly one candidate is found, use it and tell the user which PR was detected. If none or multiple, abort and ask the user to specify the PR number explicitly.
+
+Then:
+
+```sh
git fetch origin main
git status --porcelain
git rev-parse --abbrev-ref HEAD
@@ -59,6 +67,7 @@ Save its full output as ``. From the plan extract:
- The `### Bundle dependency updates` block — likewise.
- The `### Bundle self-bumps` block — likewise.
- The single `### Bumped Versions Summary` line — the only piece you forward to Phase 5. Save it as ``.
+- The new cortex library version (the `` side of `cortex →`) — save as ``.
---
@@ -70,14 +79,7 @@ Starting from `main` with a clean tree, apply the plan to the working tree:
- For each line in `### Bundle dependency updates`, use `Edit` on the named `helm/bundles//Chart.yaml` to change the `version:` field of the dependency entry at the given index. Anchor your `Edit` on the specific old version string plus the dependency's `name:` and any `alias:` line so the match is unique.
- For each line in `### Bundle self-bumps`, use `Edit` on the named bundle's Chart.yaml to change the top-level `version:`. Anchor on the chart's `name: ` plus the version line to disambiguate from dependency `version:` entries.
-Dispatch **`pull-request-creator`** with:
-
-- `branch`: `release/bump-charts-`
-- `commit_message`: `Bump chart versions for release PR #`
-- `motivation`: `Bump helm chart versions for release PR #. Bumped: . This PR must be merged before #.`
-- `assign_reviewers`: `false` (release-mechanics PRs route to the release owner regardless of code area)
-
-Capture `` and `` from its report. The agent leaves the working tree clean on `release/bump-charts-` — switch back yourself with `git checkout main` before the next phase.
+Do NOT commit yet — leave the edits uncommitted in the working tree.
---
@@ -100,40 +102,48 @@ Produce the changelog entry.
Save its full output as ``.
----
+Read the first 100 lines of `CHANGELOG.md` (or the full file if it does not exist) to check whether an entry for `#` is already present (use `grep -F "[#]"`):
-## Phase 6: Apply the changelog
+- **First run** (no existing entry): prepend `` (followed by a blank line) directly under the `# Changelog` header, before any existing entries.
+- **Update run** (entry already present): replace the entire existing entry for `#` — from its `##` heading line down to (but not including) the next `##` heading or end of file — with ``. Do not prepend a second entry.
-If `CHANGELOG.md` does not exist, write it with `# Changelog\n\n` followed by ``. Otherwise, read the file and prepend `` (followed by a blank line) directly under the `# Changelog` header, before any existing entries.
+If `CHANGELOG.md` does not exist at all, create it with `# Changelog\n\n` followed by ``.
+
+Do NOT commit yet — both `helm/` edits and `CHANGELOG.md` remain uncommitted in the working tree.
+
+---
+
+## Phase 6: Open the prep PR
Dispatch **`pull-request-creator`** with:
-- `branch`: `release/changelog-`
-- `commit_message`: `Add changelog entry for release PR #`
-- `motivation`: `Add changelog entry for release PR #. Merge after #.`
+- `branch`: `release/prepare-`
+- `commit_message`: `Release cortex `
+- `motivation`: `Release prep for #: changelog entry and helm chart version bumps. Merge this before merging #.`
- `assign_reviewers`: `false`
-Capture `` and ``. The agent leaves the working tree clean on `release/changelog-` — `git checkout main` yourself before Phase 7.
+Capture `` and `` from its report. The agent leaves the working tree clean on `release/prepare-` — switch back yourself with `git checkout main` before Phase 7.
+
+`pull-request-creator`'s idempotency handles the "update" case: if the branch already exists with only bot commits, it resets and force-pushes automatically.
---
## Phase 7: Update the release PR description
-Build the new release PR description: `` followed by a Dependencies footer linking the bump PR and the changelog PR. Write it to a tempfile and pass `--body-file` to avoid shell quoting issues.
+Build the new release PR description: `` followed by a Dependencies footer. Write it to a tempfile and pass `--body-file` to avoid shell quoting issues.
-```
+```sh
TMP=$(mktemp)
cat > "$TMP" <<'BODY'
-## Changelog
+## Release cortex
## Dependencies
-- Bump PR: # (must be merged before this PR)
-- Changelog PR: # (merge after this PR)
+- Prep PR: # (must be merged before this PR)
BODY
-gh pr edit --body-file "$TMP"
+gh pr edit --title "Release cortex " --body-file "$TMP"
rm "$TMP"
```
@@ -148,9 +158,8 @@ Print:
```
## Release # Post-Open Summary
-- Bump PR: # ()
-- Changelog PR: # ()
-- Release PR #: description updated with changelog and PR references
+- Prep PR: # ()
+- Release PR #: description updated with changelog and prep PR reference
- Bumped:
```
@@ -161,5 +170,5 @@ If any phase aborted, list which phase and why, and skip the remaining phases
## Critical rules
- Phases 2 → 7 strictly in order. Each depends on the previous.
-- Never read chart files or `CHANGELOG.md` for analysis — that is what the investigator agents do. You read those files only for the mechanical `Edit` and prepend in Phases 4 and 6.
+- Never read chart files or `CHANGELOG.md` for analysis — that is what the investigator agents do. You read those files only for the mechanical `Edit` in Phase 4 and the mechanical prepend/replace in Phase 5 (reading the first 100 lines to detect an existing entry is explicitly permitted).
- All PR creation flows through `pull-request-creator`. Do not call `gh pr create` directly. The agent owns branch reset, commit, force-push, the human-commit guard, and clean-tree postcondition — you only stage the working-tree edits.
From 9dce62c834d7265ad2523864d3eddd59db66bbbd Mon Sep 17 00:00:00 2001
From: mblos <156897072+mblos@users.noreply.github.com>
Date: Fri, 17 Jul 2026 09:16:40 +0200
Subject: [PATCH 15/15] Revert "fix: broken workflows with large gh runner"
(#1060)
Reverts cobaltcore-dev/cortex#1050
Signed-off-by: mblos
---
.github/workflows/codeql.yaml | 2 +-
.github/workflows/push-images.yaml | 2 +-
2 files changed, 2 insertions(+), 2 deletions(-)
diff --git a/.github/workflows/codeql.yaml b/.github/workflows/codeql.yaml
index c85947d68..47b415216 100644
--- a/.github/workflows/codeql.yaml
+++ b/.github/workflows/codeql.yaml
@@ -24,7 +24,7 @@ permissions:
jobs:
analyze:
name: CodeQL
- runs-on: ubuntu-latest #large_runner_16core_64gb
+ runs-on: large_runner_16core_64gb
steps:
- name: Check out code
uses: actions/checkout@v7
diff --git a/.github/workflows/push-images.yaml b/.github/workflows/push-images.yaml
index ce71706e1..ef3c39e9c 100644
--- a/.github/workflows/push-images.yaml
+++ b/.github/workflows/push-images.yaml
@@ -17,7 +17,7 @@ jobs:
packages: write
attestations: write
id-token: write
- runs-on: ubuntu-latest #large_runner_16core_64gb
+ runs-on: large_runner_16core_64gb
steps:
- uses: actions/checkout@v7
- name: Set up QEMU