Skip to content

Add Databricks SQL warehouse lifecycle operators - #70088

Open
Vamsi-klu wants to merge 5 commits into
apache:mainfrom
Vamsi-klu:agent/databricks-warehouse-lifecycle
Open

Add Databricks SQL warehouse lifecycle operators#70088
Vamsi-klu wants to merge 5 commits into
apache:mainfrom
Vamsi-klu:agent/databricks-warehouse-lifecycle

Conversation

@Vamsi-klu

@Vamsi-klu Vamsi-klu commented Jul 19, 2026

Copy link
Copy Markdown
Contributor

This adds first-class operators for starting and stopping existing Databricks SQL warehouses, including optional polling until the requested lifecycle state is reached.

related: #21377

Problem

Airflow's Databricks provider can execute SQL against a warehouse, but it has no first-class way to manage an existing warehouse's start/stop lifecycle. Dag authors currently need custom REST calls around their SQL tasks.

What changed

  • Add DatabricksHook methods for retrieving, starting, and stopping a warehouse through the Databricks SQL Warehouses API.
  • Add a validated WarehouseState model for the six documented lifecycle states.
  • Add DatabricksStartWarehouseOperator and DatabricksStopWarehouseOperator with idempotent pre-checks, optional waiting, monotonic deadlines, and explicit terminal-state errors.
  • Register the operators in provider metadata and add a how-to guide plus a system-test example with unconditional stop cleanup.

Warehouse IDs are embedded in the documented REST paths; no request-body workaround or new dependency is introduced. The implementation follows the Databricks SQL Warehouses API.

Scope

This is the Phase 1 scope proposed on #21377: get/state/start/stop plus synchronous waiting. Create/delete, edit, warehouse-by-name resolution, async hooks, and deferrable operators remain outside this PR so the initial contribution stays reviewable and independently useful.

Behavior and compatibility

  • Starting an already RUNNING warehouse and stopping an already STOPPED warehouse are no-ops.
  • Existing STARTING/STOPPING transitions are reused instead of issuing duplicate requests.
  • A start accepted while the warehouse still reports STOPPING continues polling; Databricks API transition rejections propagate unchanged.
  • Waiting uses time.monotonic(), starts no new poll after the configured deadline, and still honors a target or deletion state returned by a poll that began before the deadline.
  • The new operator validates templated warehouse IDs at execution time and remains import-compatible across supported Airflow versions through common.compat.sdk.

Validation

  • breeze run pytest providers/databricks/tests/unit/databricks/operators/test_databricks_warehouse.py -xvs — 23 passed.
  • breeze testing providers-tests --test-type "Providers[databricks]" — 842 passed, 12 skipped.
  • breeze testing providers-tests --test-type "Providers[amazon,common.compat,common.sql,databricks,google,openlineage]" — 11,858 passed, 185 skipped.
  • breeze run mypy providers/databricks/src/airflow/providers/databricks/exceptions.py providers/databricks/src/airflow/providers/databricks/hooks/databricks.py providers/databricks/src/airflow/providers/databricks/operators/databricks_warehouse.py — success, no issues.
  • Explicit nine-file prek pre-commit checks — passed.
  • Explicit nine-file prek manual checks — passed, including the providers mypy hook.
  • breeze build-docs --docs-only --clean-build databricks — documentation build successful; the generated guide contains both lifecycle examples.
  • breeze run pytest providers/databricks/tests/system/databricks/example_databricks_sql_warehouse.py --collect-only -q — 1 system test collected.
  • breeze ci selective-check --commit-ref HEAD — selected provider unit/compatibility tests, provider mypy, docs, Python scans, and the system-test path; no UI tests selected.

Reviewer evidence

This PR has no UI surface, so before/after screenshots and browser validation are not applicable. The REST boundary is covered with autospecced request assertions, operator behavior is covered with a specced hook, and the system example is import-validated. No Databricks workspace credentials were used or required for these deterministic lifecycle tests.

The narrow Phase 1 scope was posted on the issue before implementation: #21377 (comment)


Was generative AI tooling used to co-author this PR?
  • Yes — Codex (GPT-5)

Generated-by: Codex (GPT-5) following the guidelines

@Vamsi-klu

Vamsi-klu commented Jul 19, 2026

Copy link
Copy Markdown
Contributor Author

Local validation evidence for the Databricks SQL warehouse lifecycle implementation:

  • Focused operator suite: 21 passed.
  • Full Databricks provider suite: 838 passed, 12 skipped.
  • Selective six-provider dependency matrix: 11,858 passed, 185 skipped.
  • Changed-source mypy: Success: no issues found in 3 source files.
  • Explicit pre-commit and manual prek checks: passed; the manual run included the providers mypy hook.
  • Databricks docs build: successful; generated output contains both start and stop examples.
  • System-test example: one test_run collected successfully.
  • Selective-check analysis selected the expected provider unit/compatibility, provider mypy, docs, Python scan, and system-test jobs; it selected no UI work.

The tests assert the exact Databricks Warehouses API paths, idempotent start/stop behavior, transition-in-progress behavior, terminal failure states, templated-ID validation, and strict monotonic timeout handling.

There are no UI changes in this PR, so screenshots would not add reviewer signal. No Databricks credentials were used: REST calls are mocked at the hook boundary, while the system-test Dag is import-validated. Live workspace execution can be added later if a reviewer specifically requests it.


Drafted-by: Codex (GPT-5)

@Vamsi-klu

Copy link
Copy Markdown
Contributor Author

@eladkal @moomindani Can i get some feedack/Stamp for the PR please? Thanks!

@Vamsi-klu
Vamsi-klu marked this pull request as ready for review July 19, 2026 07:01
@potiuk potiuk added the ready for maintainer review Set after triaging when all criteria pass. label Jul 20, 2026

@moomindani moomindani left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for this — nice, self-contained contribution.

Conventions I checked and found consistent: _DatabricksWarehouseBaseOperator sharing start/stop mirrors GCP's _DataprocStartStopClusterBaseOperator; WarehouseState follows the existing RunState / SQLStatementState shape in the hook; and wait_for_termination / polling_period_seconds / databricks_retry_* match DatabricksSQLStatementsOperator — note these differ from AWS's wait_for_completion, but matching the provider is the right precedence. time.monotonic(), spec/autospec mocks, and the all_done cleanup task in the system example are all correct.

I validated the lifecycle behaviour against a real workspace (2X-Small serverless warehouse, auto_stop_mins=10) rather than only reading the code:

Probe Observed
POST /start from STOPPED, then tight-poll GET 3/3 trials flipped STOPPED -> STARTING within 0.45-0.46s
POST /stop from RUNNING reached STOPPED in ~2s
Start requested while stopping STARTING at t=3s, RUNNING at t=8s, no API rejection

Two things I'd like maintainer input on, then some small cleanups.

1. STOPPED as a start-path failure state is racy. The first poll after start_warehouse() runs with no time.sleep() in between, so a single lagging GET fails the task even though the start succeeded. Sub-second in my probes, but structural — details and a reproduction inline.

2. Shipping these without a deferrable mode is the part I'd most like a second opinion on. I know the PR body scopes deferrable out of Phase 1, and I understand wanting to keep the first contribution reviewable. But start/stop are multi-minute waits that hold a worker slot for their whole duration, which is the canonical case for deferrable operators — and the comparable operators elsewhere all have one:

  • AWS RedshiftResumeClusterOperator / RedshiftPauseClusterOperatordeferrable + dedicated triggers
  • GCP DataprocStartClusterOperator / DataprocStopClusterOperatordeferrable
  • This provider's own DatabricksRunNowOperator, DatabricksSQLStatementsOperator, and sensors — all deferrable

So these two operators would be the only blocking-poll operators in the Databricks provider. My concern is less "please add it now" and more that deferring it has a compatibility cost: once released, wait_for_termination and timeout are public API, and retrofitting deferrable around them is awkward — DatabricksSQLStatementsOperator needs its "wait_timeout": "0s" trick precisely to make one set of parameters serve both paths. Doing it up front is cheaper than reconciling it later.

The groundwork is mostly there: the hook already has _a_do_api_call, and the existing a_get_cluster_state / a_get_sql_statement_state are only a handful of lines each, so a_get_warehouse_state plus a DatabricksWarehouseStateTrigger alongside the two existing triggers looks like a modest addition rather than a redesign.

I'm not blocking on this — it is a scope judgement that belongs to the committers, not to me, and "merge Phase 1 now, add deferrable in Phase 2" is a legitimate answer if the parameter surface is settled deliberately. I'd just rather it be an explicit decision than an omission noticed after release.

Also ran locally:

  • pytest test_databricks_warehouse.py test_databricks.py — 149 passed.
  • prek --stage pre-commit — the only two failures are Update providers build files and Validate provider.yaml files, both from Docker not running on my machine, not from your diff.

Drafted-by: Claude Code (Opus 5)

Comment thread providers/databricks/tests/system/databricks/example_databricks_sql_warehouse.py Outdated
@Vamsi-klu
Vamsi-klu force-pushed the agent/databricks-warehouse-lifecycle branch from c6a5489 to c6f018d Compare July 27, 2026 05:26
@Vamsi-klu

Copy link
Copy Markdown
Contributor Author

@moomindani, thanks for the thorough review and for validating the lifecycle behavior against a real workspace.

I pushed c6f018d and addressed all four inline comments:

  • Start polling now treats a STOPPED response immediately after the start request as potentially stale and continues until RUNNING, deletion, or timeout. The regression test covers STOPPED pre-check → start request → stale STOPPED poll → RUNNING.
  • The shared polling path now uses WarehouseState.is_deleted as the source of truth for terminal deletion states.
  • Hook construction is inlined into the cached _hook property.
  • The unused ENV_ID assignment is removed from the system-test example.

I also corrected the PR description's timeout wording: no new poll starts after the deadline, while a target or deletion state returned by an already-started poll is still honored.

On deferrable execution: I agree it would be valuable, but I am deliberately keeping it in Phase 2 rather than broadening this Phase 1 PR. That scope was recorded on #21377 and in the PR description before implementation. Since the warehouse start/stop endpoints return immediately, a future deferrable path can remain additive while preserving wait_for_termination, polling_period_seconds, and timeout. That follow-up will need the async state hook, serialized trigger and timeout behavior, operator completion path, and compatibility tests. If a committer considers deferrable execution a pre-merge requirement, I can revisit the scope here.

Validation on the rebased branch:

  • Full Databricks provider suite: 842 passed, 12 skipped.
  • Focused operator and warehouse-hook suites: 35 passed.
  • System-test example: 1 test collected.
  • Provider mypy: Success: no issues found.
  • Branch-level pre-commit and manual prek checks: passed.

Drafted-by: Codex (GPT-5); reviewed by @Vamsi-klu before posting

@moomindani moomindani left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Verified c6f018d by running it rather than reading the summary — all four are correctly addressed.

The race fix is the right shape: dropping the failure_states parameter and keying terminal detection off state.is_deleted fixes finding 1 and 2 in one move. I re-ran my original reproduction and the behaviour is now:

Scenario Before Now
STOPPED (pre-check) → stale STOPPEDRUNNING error reaches RUNNING
Warehouse never leaves STOPPED error (misleading) timeout, last state: STOPPED
DELETING mid-wait error error (unchanged)

That is exactly the trade I hoped for — a genuine never-starts now surfaces as a timeout with the last observed state in the message, which is more diagnosable than the old immediate failure.

The test updates are what I'd have asked for: parametrizing test_starts_then_waits_until_running over ["STARTING", "STOPPED"] pins the regression, and re-pointing the start leg of test_execute_raises_on_failure_state from STOPPED to DELETING keeps the terminal-state assertion meaningful instead of just deleting it. 150 passed locally. prek --stage pre-commit is clean apart from the two Docker-dependent hooks that fail on my machine regardless of the diff. Diff against current main is your 9 files only.

On deferrable: that's a reasonable answer, and recording it explicitly is all I was after. My concern was an unexamined omission, not the choice itself — you've now stated the Phase 2 plan and the parameter-compatibility reasoning, so a committer can weigh it deliberately. No objection from me to merging Phase 1 as scoped.

Nothing further from my side.


Drafted-by: Claude Code (Opus 5)

@eladkal

eladkal commented Jul 27, 2026

Copy link
Copy Markdown
Contributor

So @moomindani if I get it right you are approving the change?

A slow final status request can finish after the deadline even when it confirms the requested state. Treating that observation as a timeout can fail an otherwise successful Dag.
A warehouse can still report STOPPED immediately after the start request because the API response is eventually consistent. Keep polling until RUNNING or deletion/timeout so valid starts do not fail spuriously.
@Vamsi-klu
Vamsi-klu force-pushed the agent/databricks-warehouse-lifecycle branch from c6f018d to a486797 Compare July 31, 2026 06:51
@Vamsi-klu

Vamsi-klu commented Jul 31, 2026

Copy link
Copy Markdown
Contributor Author

Hi @eladkal All four findings are addressed and retested per your review. @moomindani the review is marked COMMENTED. Can i get maintainer approval please? Thanks!

@eladkal
eladkal self-requested a review August 6, 2026 05:48

@eladkal eladkal left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I see there is an open point around defer support. What is the plan here?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

About file name. I know this is following what we have now in the provider but this doesn't align with project conventions.
it should be warehouse.py

I will raise followup PR to fix the other files soon

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Renamed in 727bfc3. It is operators/warehouse.py and tests/unit/databricks/operators/test_warehouse.py now, and I updated everything that pointed at the old name: provider.yaml, the regenerated get_provider_info, the two :class: references in the docs, the mock.patch targets in the unit tests and the system example import. git grep databricks_warehouse across the provider comes back empty.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

same comment about file name

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Renamed alongside the operator module in 727bfc3, see the reply on the other file-name thread for the full list of what got updated.

Comment on lines +310 to +315
def to_json(self) -> str:
return json.dumps(self.__dict__)

@classmethod
def from_json(cls, data: str) -> WarehouseState:
return WarehouseState(**json.loads(data))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't see any core using these. What are they for?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nothing uses them, you are right. I added them for symmetry with RunState and SQLStatementState, whose to_json and from_json are used by the deferrable triggers, but WarehouseState has no trigger yet so they were dead code from day one. Removed both along with the round-trip test in 727bfc3. They will come back with DatabricksWarehouseStateTrigger when the deferrable work lands. is_deleted stays, since _wait_for_state uses it now.

Comment thread providers/databricks/docs/operators/sql_warehouse.rst Outdated

@moomindani moomindani left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@eladkal sorry for the slow reply to your question from Jul 27 — yes, as far as my own review goes I have no remaining objections. All four findings I raised were addressed, and I re-verified on a486797 rather than re-reading the summary: stale STOPPED after start now reaches RUNNING, a warehouse that never leaves STOPPED times out with the last observed state, and 150 tests pass. The warehouse code is byte-identical to the c6f018d I checked earlier.

I am deliberately not marking this approved, though, because your four points from Aug 6 are still open and three of them need code changes. I checked them and they all hold:

  • File naming (warehouse.py) — amazon uses athena.py / ec2.py, google uses bigquery.py; none repeat the provider name, so the existing databricks_*.py files are the deviation.
  • Unused hook methods — confirmed. WarehouseState.to_json / from_json have no production caller; only test_databricks.py:1619 round-trips them against themselves. (is_deleted is now used by _wait_for_state, so that one is fine.)
  • Docs — agreed, drop the system-test framing at sql_warehouse.rst:49 and just describe what the trigger rule does.

Separately: the branch is 186 commits behind main. Diffed against main it looks like this PR reverts #70130 and #69442 — that is a stale-base artifact, not a real revert (against the merge base it is only the 9 files of this PR). Worth rebasing so CI runs on current main and that diff stops misleading reviewers.


Drafted-by: Claude Code (Opus 5)

Vamsi-klu and others added 2 commits August 8, 2026 04:50
…rehouse-lifecycle

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Rename the operator module to warehouse.py to follow provider naming
conventions, drop the unused WarehouseState.to_json/from_json helpers
that have no production caller until the deferrable trigger lands, and
describe the all_done trigger rule behavior in the docs without
referring to system tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@Vamsi-klu

Copy link
Copy Markdown
Contributor Author

@eladkal pushed 727bfc3, which covers all four of your points: the module is renamed to warehouse.py, the unused WarehouseState serializers are gone, and the docs no longer mention system tests.

On defer support, my plan was to add it as a follow-up rather than fold it in here. Phase 2 adds a_get_warehouse_state to the hook plus a DatabricksWarehouseStateTrigger next to the two triggers the provider already has, keeping wait_for_termination, polling_period_seconds and timeout exactly as they are, so deferrable lands as an additive change rather than a parameter redesign. @moomindani reviewed that reasoning and had no objection to merging phase 1 as scoped. If you would rather see deferrable in this PR before it merges, say so and I will extend it here instead.

@moomindani I also merged latest main in the same push, so the diff no longer looks like it reverts #70130 and #69442.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants