Skip to content

feat(vertexai): add google_vertex_ai_online_evaluator resource - #18921

Merged
roaks3 merged 7 commits into
GoogleCloudPlatform:mainfrom
iliesicatrinel:online-evaluator-resource
Oct 5, 2026
Merged

roaks3 merged 7 commits into
GoogleCloudPlatform:mainfrom
iliesicatrinel:online-evaluator-resource

Conversation

@iliesicatrinel

@iliesicatrinel iliesicatrinel commented Sep 8, 2026 •

Copy link
Copy Markdown
Contributor

Adds a new resource google_vertex_ai_online_evaluator (GA + beta) for Vertex AI
online evaluation. It periodically samples traces/sessions from an agent and
evaluates them using configured metric sources, supporting both inline predefined
metrics and references to google_vertex_ai_evaluation_metric (via
metric_resource_name).

Release Notes

`google_vertex_ai_online_evaluator`

@modular-magician modular-magician added the awaiting-approval Pull requests that need reviewer's approval to run presubmit tests label Sep 8, 2026
@github-actions
github-actions Bot requested a review from roaks3 September 8, 2026 15:09
@github-actions

github-actions Bot commented Sep 8, 2026

Copy link
Copy Markdown

Googlers: For automatic test runs see go/terraform-auto-test-runs.

@roaks3, a repository maintainer, has been assigned to review your changes. If you have not received review feedback within 2 business days, please leave a comment on this PR asking them to take a look.

You can help make sure that review is quick by doing a self-review and by running impacted tests locally.

@iliesicatrinel
iliesicatrinel marked this pull request as draft September 8, 2026 15:25
@iliesicatrinel
iliesicatrinel marked this pull request as ready for review September 10, 2026 13:24
Inline metric config, provided as a JSON-formatted string using
camelCase field names to match the API format (see the
[Metric API documentation](https://cloud.google.com/vertex-ai/docs/reference/rest/v1/Metric)).
Suggested predefined `metricSpecName` values:

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think it would be better to list the metrics divided by the scope so that users don't specify trace scope metric and declare session scope ingestion only.
Btw what will happen in such case, will the metric silently not work, or will the apply fail with a pertinent error ?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I agree, added a clarification. For predefined metrics scope mismatch isn't silent (we have validation on our control plane) and apply fails with INVALID_ARGUMENT (e.g. Metric 'multi_turn_trajectory_quality_v1' is not a valid predefined metric. Only [tool_use_quality_v1, final_response_quality_v1, hallucination_v1, safety_v1] are supported). For a registered metric_resource_name only the name format is checked on create, and invalid metrics are surfaced at runtime (we update OnlineEvaluator state and write a log in the project).

@github-actions

Copy link
Copy Markdown

@roaks3 This PR has been waiting for review for 3 weekdays. Please take a look! Use the label disable-review-reminders to disable these notifications.

@roaks3

roaks3 commented Sep 16, 2026

Copy link
Copy Markdown
Contributor

We’re currently requesting a change in process for Googlers (or individuals on behalf of Googlers) contributing changes to the Terraform Provider for Google Cloud. Please see go/terraform-ssp-adjustment for details.

@roaks3 roaks3 closed this Sep 16, 2026
@roaks3 roaks3 reopened this Oct 1, 2026

@roaks3 roaks3 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, just a couple small comments

create_url: projects/{{project}}/locations/{{region}}/onlineEvaluators
update_verb: PATCH
update_mask: true
generate_list_resource: true

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Was this done intentionally? (is there value in users querying all evaluators?)

Just checking

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, the OnlineEvaluators API has a List method, so enumerating evaluators is a supported use case.

type: String
description: The region of the OnlineEvaluator (e.g. us-central1).
url_param_only: true
required: true

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Often we don't make region required because it can be set on the provider level, and normally users expect that setting to be used. From a quick look, it seems like most other vertexai resources make this optional.

Is there a reason region needs to be required here?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No strong reason, made it optional, matching the common vertexai pattern. Thanks

service/aiplatform-evaluation:
resources:
- google_vertex_ai_evaluation_.*
- google_vertex_ai_online_evaluator

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

FYI go/terraform-resource-owners-enrollment#manage-enrollment, the change here won't stick, it is reflective of an internal source

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, dropped the enrolled_teams.yml change. I'll add ourselves to internal proto instead.

@modular-magician modular-magician added service/aiplatform-evaluation and removed awaiting-approval Pull requests that need reviewer's approval to run presubmit tests labels Oct 1, 2026
@modular-magician

modular-magician commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

Hi there, I'm the Modular magician. I've detected the following information about your changes for commit 991212e:

Diff report

Your PR generated the following diffs in downstream repositories:

Repository Diff Link Changes
google provider View Diff 9 files changed, 3887 insertions(+)
google-beta provider View Diff 9 files changed, 3887 insertions(+)
terraform-google-conversion View Diff 1 file changed, 853 insertions(+)

Missing test report

Your PR includes resource fields which are not covered by any test.

Resource: google_vertex_ai_online_evaluator (5 total tests)
Please add an acceptance test which includes these fields. The test should include the following:

resource "google_vertex_ai_online_evaluator" "primary" {
  cloud_observability {
    log_view = # value needed
    session_scope {
      filter {
        duration {
          comparison_operator = # value needed
          value               = # value needed
        }
        model_call_errors {
          comparison_operator = # value needed
          value               = # value needed
        }
        model_calls {
          comparison_operator = # value needed
          value               = # value needed
        }
        tool_call_errors {
          comparison_operator = # value needed
          value               = # value needed
        }
        tool_calls {
          comparison_operator = # value needed
          value               = # value needed
        }
        user_turns {
          comparison_operator = # value needed
          value               = # value needed
        }
      }
    }
    trace_scope {
      filter {
        total_token_usage {
          comparison_operator = # value needed
          value               = # value needed
        }
      }
    }
    trace_view = # value needed
  }
}

Test report

Analytics

Total Tests Passed Skipped Affected
136 119 11 6
Affected Service Packages
  • vertexai

Learn how VCR tests work


Step 1: Replaying Mode

Action taken

Found 6 affected test(s) by replaying old test recordings. Starting RECORDING based on the most recent commit.

Click here to see the affected tests
  • TestAccVertexAIFeatureOnlineStoreFeatureview_vertexAiFeatureonlinestoreFeatureview_featureRegistry_updated
  • TestAccVertexAIOnlineEvaluatorListQuery_generated
  • TestAccVertexAIOnlineEvaluator_update
  • TestAccVertexAIOnlineEvaluator_vertexAiOnlineEvaluatorBasicExample
  • TestAccVertexAIOnlineEvaluator_vertexAiOnlineEvaluatorFullExample
  • TestAccVertexAIReasoningEngine_apiParity

View the replaying VCR build log


Step 2: Recording Mode

Recording Mode Findings Test Name
✅ Log - TestAccVertexAIOnlineEvaluatorListQuery_generated
✅ Log - TestAccVertexAIOnlineEvaluator_update
✅ Log - TestAccVertexAIOnlineEvaluator_vertexAiOnlineEvaluatorBasicExample
✅ Log - TestAccVertexAIOnlineEvaluator_vertexAiOnlineEvaluatorFullExample
❌ Error · Log ⚪ Nightly fails 100% of 24 TestAccVertexAIFeatureOnlineStoreFeatureview_vertexAiFeatureonlinestoreFeatureview_featureRegistry_updated
❌ Error · Log ⚪ Nightly fails 100% of 24 TestAccVertexAIReasoningEngine_apiParity

Caution

Issues requiring attention before PR completion

🔴 Initial Recording Failed: Some tests failed during the recording step. See the table above for details.

🟡 Known Nightly Failures: 2 of the tests that failed in recording also fail or are flaky in the last 30 days of nightly runs, so they may be unrelated to this PR. In nightly findings, ⚪ means the test already fails in nightly, 🟡 that it is flaky there, and 🔴 that it is healthy in nightly and so more likely broken by this PR.

Please address these issues to complete your PR. If you believe these detections are incorrect or unrelated to your change, please raise the concern with your reviewer.

View the recording VCR build log or the debug logs folder for detailed results.

@iliesicatrinel, @maxgasztych VCR tests complete for 991212e!

@modular-magician modular-magician added the awaiting-approval Pull requests that need reviewer's approval to run presubmit tests label Oct 2, 2026
@github-actions
github-actions Bot requested a review from roaks3 October 2, 2026 06:47

@roaks3 roaks3 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, is there a reason we can't test the fields flagged by the check? It wouldn't be a strict blocker, but testing each field is generally important for us to confirm the fields are configured correctly, and it helps for flagging regressions.

@github-actions
github-actions Bot requested a review from roaks3 October 5, 2026 07:59
@iliesicatrinel

Copy link
Copy Markdown
Contributor Author

LGTM, is there a reason we can't test the fields flagged by the check? It wouldn't be a strict blocker, but testing each field is generally important for us to confirm the fields are configured correctly, and it helps for flagging regressions.

Sure, added coverage for all the flagged fields.

@roaks3

roaks3 commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

/gcbrun

@modular-magician modular-magician removed the awaiting-approval Pull requests that need reviewer's approval to run presubmit tests label Oct 5, 2026
@modular-magician

modular-magician commented Oct 5, 2026 •

Copy link
Copy Markdown
Collaborator

Hi there, I'm the Modular magician. I've detected the following information about your changes for commit 7ab1fc8:

Diff report

Your PR generated the following diffs in downstream repositories:

Repository Diff Link Changes
google provider View Diff 9 files changed, 3970 insertions(+)
google-beta provider View Diff 9 files changed, 3970 insertions(+)
terraform-google-conversion View Diff 1 file changed, 853 insertions(+)

Missing service labels

The following new resources do not have corresponding service labels:

  • google_vertex_ai_online_evaluator

If you believe this detection to be incorrect please raise the concern with your reviewer. Googlers: This error is safe to ignore once you've completed go/fix-missing-service-labels.
An override-missing-service-label label can be added to allow merging.

Test report

Analytics

Total Tests Passed Skipped Affected
136 121 11 4
Affected Service Packages
  • vertexai

Learn how VCR tests work


Step 1: Replaying Mode

Action taken

Found 4 affected test(s) by replaying old test recordings. Starting RECORDING based on the most recent commit.

Click here to see the affected tests
  • TestAccVertexAIFeatureOnlineStoreFeatureview_vertexAiFeatureonlinestoreFeatureview_featureRegistry_updated
  • TestAccVertexAIOnlineEvaluator_update
  • TestAccVertexAIOnlineEvaluator_vertexAiOnlineEvaluatorFullExample
  • TestAccVertexAIReasoningEngine_apiParity

View the replaying VCR build log


Step 2: Recording Mode

Recording Mode Findings Test Name
✅ Log - TestAccVertexAIOnlineEvaluator_update
✅ Log - TestAccVertexAIOnlineEvaluator_vertexAiOnlineEvaluatorFullExample
❌ Error · Log ⚪ Nightly fails 100% of 23 TestAccVertexAIFeatureOnlineStoreFeatureview_vertexAiFeatureonlinestoreFeatureview_featureRegistry_updated
❌ Error · Log ⚪ Nightly fails 100% of 23 TestAccVertexAIReasoningEngine_apiParity

Caution

Issues requiring attention before PR completion

🔴 Initial Recording Failed: Some tests failed during the recording step. See the table above for details.

🟡 Known Nightly Failures: 2 of the tests that failed in recording also fail or are flaky in the last 30 days of nightly runs, so they may be unrelated to this PR. In nightly findings, ⚪ means the test already fails in nightly, 🟡 that it is flaky there, and 🔴 that it is healthy in nightly and so more likely broken by this PR.

Please address these issues to complete your PR. If you believe these detections are incorrect or unrelated to your change, please raise the concern with your reviewer.

View the recording VCR build log or the debug logs folder for detailed results.

@iliesicatrinel, @roaks3, @maxgasztych VCR tests complete for 7ab1fc8!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants