Skip to content

PrometheusMonitor: use oc exec instead of oc rsh to avoid 16KB truncation - #31705

Open
ptalgulk01 wants to merge 1 commit into
openshift:mainfrom
ptalgulk01:ptalgulk01/fix-prometheus-monitor-rsh-truncation
Open

ptalgulk01 wants to merge 1 commit into
openshift:mainfrom
ptalgulk01:ptalgulk01/fix-prometheus-monitor-rsh-truncation

Conversation

@ptalgulk01

@ptalgulk01 ptalgulk01 commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Summary

compat_otp.Monitor's Prometheus-querying methods (InstantQuery/SimpleQuery, RangeQuery, queryRules/GetAllRules/GetAlertRules/GetRecordRules, and GetAlerts) all fetch their response by running curl inside the prometheus-k8s pod via RemoteShPod, which executes the command with oc rsh.

Direct experimentation (see discussion on openshift/openshift-tests-private#30147) found that oc rsh-captured output gets hard-truncated at 16,384 bytes, while oc exec returns the complete response regardless of size:

Command Bytes Truncated?
oc rsh -n openshift-monitoring prometheus-k8s-0 sh -c "curl ..." 0 Yes (empty)
oc rsh -n openshift-monitoring prometheus-k8s-0 curl ... 16,384 Yes (cut at 16KB)
oc exec -n openshift-monitoring prometheus-k8s-0 -- sh -c "curl ..." 55,182 No (complete)
oc exec -n openshift-monitoring prometheus-k8s-0 -- curl ... 55,182 No (complete)

On a cluster with enough concurrently firing/pending alerts to push /api/v1/alerts past 16KB, this makes consumers of GetAlerts() fail with an unparseable/truncated JSON response, deterministically, on every call — not an intermittent flake.

Fix

  • Added RemoteExecPod() in pods.go: same signature shape as RemoteShPod, but runs the command via oc exec -n <namespace> <pod> -- <cmd...> instead of oc rsh.
  • Switched all four Monitor/PrometheusMonitor query methods (InstantQuery, RangeQuery, queryRules, GetAlerts) to use RemoteExecPod instead of RemoteShPod.
  • RemoteShPod itself is untouched, so this doesn't change behavior for any other caller.

Test plan

  • go build ./test/extended/util/compat_otp/... passes
  • Re-run an MCO alert test (e.g. TC 63865/63866 in openshift-tests-private) on a cluster with multiple concurrent alerts firing, to confirm GetAlerts() no longer truncates

Summary by CodeRabbit

  • Bug Fixes
    • Prometheus queries and alert retrieval now preserve complete command output, including results larger than 16 KB.

@openshift-merge-bot

Copy link
Copy Markdown
Contributor

Pipeline controller notification

This PR uses the pipeline controller for second-stage tests. Selection and triggering follow the repository configuration.

Use /test ? to list jobs, /pipeline remaining to request missing second-stage tests, or /pipeline required to rerun the selected second-stage set.

@coderabbitai

coderabbitai Bot commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

Note

Currently processing new changes in this PR. This may take a few minutes, please wait...

⚙️ Run configuration
  • Configuration used: Repository YAML (base), Central YAML (inherited)
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 7422473d-a832-4a91-87f0-ad9bbcfd1a95
📥 Commits

Reviewing files that changed from the base of the PR and between c664cb7 and 876fe2e.

📒 Files selected for processing (2)
  • test/extended/util/compat_otp/pods.go
  • test/extended/util/compat_otp/prometheus_monitoring.go
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create a new PR

Comment @coderabbitai help to get the list of available commands.

@openshift-ci

openshift-ci Bot commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

[APPROVALNOTIFIER] This PR is NOT APPROVED

This pull-request has been approved by: ptalgulk01
Once this PR has been reviewed and has the lgtm label, please assign tomasdavidorg for approval. For more information see the Code Review Process.

The full list of commands accepted by this bot can be found here.

Details Needs approval from an approver in each of these files:

Approvers can indicate their approval by writing /approve in a comment
Approvers can cancel approval by writing /approve cancel in a comment

@ptalgulk01 ptalgulk01 reopened this Oct 7, 2026
@openshift-merge-bot

Copy link
Copy Markdown
Contributor

Scheduling tests matching the pipeline_run_if_changed or not excluded by pipeline_skip_if_only_changed parameters:
/test e2e-aws-ovn-fips
/test e2e-aws-ovn-microshift
/test e2e-aws-ovn-microshift-serial
/test e2e-aws-ovn-serial-1of2
/test e2e-aws-ovn-serial-2of2
/test e2e-gcp-ovn
/test e2e-gcp-ovn-upgrade
/test e2e-metal-ipi-ovn-ipv6

@openshift-ci

openshift-ci Bot commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

@ptalgulk01: The following tests failed, say /retest to rerun all failed tests or /retest-required to rerun all mandatory failed tests:

Test name Commit Details Required Rerun command
ci/prow/e2e-aws-ovn-microshift 876fe2e link true /test e2e-aws-ovn-microshift
ci/prow/e2e-aws-ovn-microshift-serial 876fe2e link true /test e2e-aws-ovn-microshift-serial
ci/prow/e2e-aws-ovn-serial-2of2 876fe2e link true /test e2e-aws-ovn-serial-2of2
ci/prow/e2e-aws-ovn-serial-1of2 876fe2e link true /test e2e-aws-ovn-serial-1of2

Full PR test history. Your PR dashboard.

Details

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository. I understand the commands that are listed here.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant