cancel
Showing results for 
Search instead for 
Did you mean: 
Get Started Discussions
Start your journey with Databricks by joining discussions on getting started guides, tutorials, and introductory topics. Connect with beginners and experts alike to kickstart your Databricks experience.
cancel
Showing results for 
Search instead for 
Did you mean: 

Serverless compute returns 403 on S3 buckets in the account-regional namespace

alexgfowler_ld
New Contributor
SUMMARY
Serverless SQL warehouses and serverless notebook/jobs compute cannot read S3 buckets created in the AWS account-regional namespace (names ending -{accountId}-{region}-an). Classic compute reads the same buckets through the same IAM role without trouble. The failure is a 403 Forbidden with a null S3 request ID, raised by the bundled aws-sdk-java/1.12.681.

Measured across two days, same region, same role, same credential, same file, with the reads interleaved seconds apart:
Global namespace buckets: 15 of 15 reads succeeded
Account-regional namespace buckets: 0 of 14 reads succeeded

ENVIRONMENT
Cloud: AWS
Workspace and UC metastore: ap-southeast-1 (Singapore)
Affected compute: Serverless SQL warehouse (PRO, 2X-Small), serverless notebook/jobs compute
Unaffected compute: Classic job clusters, non-serverless PRO SQL warehouse
Client named in errors: Hadoop 3.5.0, aws-sdk-java/1.12.681
Storage access: Unity Catalog storage credential over a cross-account IAM role, external locations, managed and external tables

THE SYMPTOM
Any serverless read of an object in an account-regional bucket fails at the first
getFileStatus call:
 
[UNAUTHORIZED_ACCESS] Unauthorized access:
s3://my-bucket-ACCOUNTID-ap-southeast-1-an/bronze/.../_delta_log:
getFileStatus on s3://my-bucket-ACCOUNTID-ap-southeast-1-an/bronze/.../_delta_log:
shaded.databricks.awssdk.com.amazonaws.services.s3.model.AmazonS3Exception: Forbidden;
request: HEAD https://my-bucket-ACCOUNTID-ap-southeast-1-an.s3-ap-southeast-1.amazonaws.com
Hadoop 3.5.0, aws-sdk-java/1.12.681 ...
credentials-provider: ...BasicSessionCredentials
credential-header: AWS4-HMAC-SHA256 Credential=ASIA...../20260919/ap-southeast-1/s3/aws4_request
signature-present: true
(Service: Amazon S3; Status Code: 403; Error Code: 403 Forbidden;
Request ID: null; S3 Extended Request ID: null; Proxy: 192.168.200.20)
SQLSTATE: 42501

 

Two details are worth noting immediately.

1. Request ID: null. A genuine IAM denial from S3 returns a real request ID. We have one in our records from a deliberately under-permissioned probe. A null request ID is the signature of this problem and the fastest way to tell the two apart.
 
2. signature-present: true, with a real ASIA session credential correctly scoped to the
region. Credentials were vended and the request was signed. This is not a
missing-credential problem.

FINDING 1: ACCOUNT-REGIONAL BUCKETS FAIL, GLOBAL-NAMESPACE BUCKETS SUCCEED

AWS lets you create a general-purpose bucket in either the shared global namespace or your account-regional namespace. The latter requires the name to end -{accountId}-{region}-an and is created with --bucket-namespace account-regional. It is a security feature: the name can only ever be claimed by your account, so it cannot be re-registered by a stranger after deletion.

Serverless cannot read them.

Evidence: 
Two buckets created 90 seconds apart in ap-southeast-1, holding a byte-identical file,
accessed through the same IAM role, the same Unity Catalog storage credential, and the
same external-location and volume mechanism. The only difference is the namespace.

Measured across two separate days, on cold warehouse sessions, with the two reads
interleaved seconds apart:
Global namespace: 15 of 15 successful reads
Account-regional namespace: 0 of 14 successful reads

Both credentials were tested against both buckets, in both directions, and the result never moved. The same holds for real Delta tables, not just probe files. A catalog whose managed location is a global-namespace bucket serves a 1,634-row Delta table to serverless without issue, while the identical data in an account-regional bucket returns 403 in the same session.

What this is not
We eliminated these by test, not by argument.
  • IAM permissions. Classic compute reads and writes the same bucket through the same role. The two bucket policies involved are structurally identical, same actions including s3:ListBucket, differing only in the bucket ARN.
  • The storage credential. Two different UC storage credentials, both over the same role, give identical results on both buckets. Neither credential is impaired: both succeed elsewhere.
  • Unity Catalog configuration. A brand-new external location, volume, schema and catalog reproduce it. A brand-new external location over a working bucket succeeds. Object age is not a factor.
  • Bucket configuration. Identical AES256 default encryption, no bucket policy, BucketOwnerEnforced, identical public access block on both.
  • Workspace identity. Reproduced on a second, independently created workspace with its own VPC.
  • Access method. Fails through external tables, managed tables and UC volumes alike.

FINDING 2: WHY ONE REGION IS UNAFFECTED, AND THE SDK ANGLE

ap-southeast-5 (Malaysia, GA August 2024) is unaffected even for account-regional buckets. 
Same account, same role, same credential. It simply works, and still does. 
 
The bundled client named in every error is aws-sdk-java/1.12.681
 
Extracting com/amazonaws/partitions/endpoints.json from that exact artifact shows its partition metadata is dated 2024-03-15, five months before ap-southeast-5 became generally available, and the region is absent from it entirely.

That gives a coherent, testable hypothesis: regions known to the pinned SDK route through
a legacy client path that does not handle account-regional bucket addressing, and regions
unknown to it fall through to a different path that does. The AWS SDK for Java v1 final
release (1.12.797, December 2025) does contain ap-southeast-5, so if that is the
mechanism, refreshing the dependency would move ap-southeast-5 into the failing set rather
than fixing anything.

We cannot verify this from outside. We cannot inspect which client handles the request,
and no supported Spark configuration on serverless lets us influence the storage path. We
offer it as a starting point, not a conclusion. The measured facts are that one region
works, others fail, and the difference tracks that SDK release date.

CONFOUNDING FACTORS, PLEASE READ BEFORE REPRODUCING

These cost us several wrong conclusions.

1. A newly created S3 bucket is unreadable from serverless for up to about 2 hours

Regardless of namespace, and with the identical 403. We measured a brand-new bucket failing at 2 minutes, 27 minutes and 42 minutes old, then succeeding consistently from about 2 hours. Create two buckets, test them immediately, and both fail, leading to the confident and wrong conclusion that namespace is irrelevant. Three of our own experiments were voided this way. It is consistent with S3 virtual-hosted-style DNS propagation, given the failing requests use the legacy s3-REGION.amazonaws.com form.
 
Wait at least 2 hours after bucket creation before drawing any conclusion.

2. Control-plane credential validation passes regardless
 
Validating the storage credential returns READ, LIST and PATH_EXISTS all PASS on locations
that serverless compute cannot read. Validation exercises the role without the compute
session scoped-down policy. A green validation is not evidence.

3. Direct file queries swallow the real error

Querying a path directly returns a generic FAILED_TO_CREATE_PLAN_FOR_DIRECT_QUERY. Use an
external table, a managed table, or a UC volume. Otherwise the underlying S3 exception
never surfaces and you cannot tell this apart from anything else.

4. Single trials are not enough

We once observed a global-namespace bucket succeed twice and then fail 90 seconds later,
at the tail of the propagation window above. Run at least three trials per configuration
and treat any inconsistency as a result in itself.

5. S3 server access logging will not answer whether the request reached S3

We enabled it and planted a known-good control request. The control never appeared in the
delivered logs either, so the absence of the failing requests proves nothing. Server access
logging is best-effort by design. Use CloudTrail S3 data events if you need this question
answered.
 
6. Unity Catalog itself is fine

Listing catalogs, schemas, tables, external locations and storage credentials all work on
serverless, and lineage is captured normally. Only the reading of objects out of S3 fails.
Do not go hunting a Unity Catalog misconfiguration.
 
REPRODUCTION

Step 1. Create two buckets that differ only in namespace.
 
ACCOUNT_ID=your-aws-account-id
REGION=ap-southeast-1
AN=repro-$ACCOUNT_ID-$REGION-an
GL=repro-global-$ACCOUNT_ID

aws s3api create-bucket --bucket $AN --region $REGION \
--create-bucket-configuration LocationConstraint=$REGION \
--bucket-namespace account-regional

aws s3api create-bucket --bucket $GL --region $REGION \
--create-bucket-configuration LocationConstraint=$REGION
Any region that predates roughly March 2024 should show the behaviour.

Step 2. Put an identical file in each.
printf 'id,label\n1,test\n' > /tmp/c.csv
aws s3 cp /tmp/c.csv s3://$AN/files/c.csv --region $REGION
aws s3 cp /tmp/c.csv s3://$GL/files/c.csv --region $REGION

 

Step 3. Grant your existing Unity Catalog role read access to both buckets: s3:GetObject
on the object ARN, plus s3:ListBucket and s3:GetBucketLocation on the bucket ARN.

Step 4. In Unity Catalog, create one external location per bucket using the same storage
credential, and an EXTERNAL volume over each.

Step 5. Wait at least 2 hours. Then, on a serverless SQL warehouse or serverless notebook:
 
SELECT count(*) FROM read_files('/Volumes/CAT/SCH/global_vol/', format => 'csv');
SELECT count(*) FROM read_files('/Volumes/CAT/SCH/account_regional_vol/', format => 'csv');
The first returns 1. The second returns 403 with Request ID: null.

 

Control: point the same two queries at a non-serverless warehouse or a classic cluster. Both succeed. That is what localises the fault to serverless compute rather than to IAM, the bucket, or Unity Catalog.

WORKAROUNDS WHILE THIS IS UNPATCHED

1. Use a global-namespace bucket for anything serverless must read.

This is the only real fix available to customers today. Verified end to end: a catalog with
its managed location on a global-namespace bucket in the same region serves Delta tables to
both serverless SQL warehouses and serverless notebooks, with data loaded and written from
serverless too.

The cost is real and worth stating plainly. The account-regional namespace exists so that a
bucket name can never be claimed by another AWS account. Moving to the global namespace
gives that up. The namespace also cannot be changed in place, so this is a data migration,
not a rename.

 

2. Use classic compute for existing account-regional buckets.
 
Classic job clusters and non-serverless PRO SQL warehouses read them without trouble. The
trade-offs are minutes rather than seconds to start, a 10-minute minimum idle timeout on
non-serverless warehouses, and cluster management you may not want.

One trap worth knowing: a workflow task that omits its job-cluster setting silently falls
back to serverless and then fails at run time. Assert the setting in whatever you use to
define jobs.

 

3. Keep an eye on fine-grained access control.

Databricks documents a dedicated classic cluster as delegating row filters, column masks
and dynamic views to the workspace serverless compute. If that path applies, a table with a
row filter could fail on a classic cluster that otherwise works. We have no such filters and
have not tested this. Flagging it because classic compute is otherwise the safe ground.

A SEPARATE ISSUE WE HIT ALONG THE WAY: TRIAL ACCOUNTS AND SAME-REGION S3

Independent of the above, and worth recording because the two are easy to conflate. We lost time doing exactly that.

While our account was on a trial, serverless could not read any customer bucket in the workspace own region, global-namespace buckets included. Cross-region buckets were unaffected:
 
ap-southeast-5, account-regional, customer bucket A -> works
ap-southeast-5, global, customer bucket B (new) -> works
ap-southeast-1, global, customer bucket C -> 403
ap-southeast-1, account-regional, customer bucket D -> 403
 
Bucket C, a global-namespace bucket in the metastore own region, returned 403 during the trial. After the trial concluded, the same bucket, unchanged, reads reliably. Bucket D still fails, which is Finding 1.

Databricks documents serverless as reaching same-region S3 through a Databricks-owned private gateway, and different-region S3 out through NAT. A trial-level network restriction on the private gateway path would produce exactly the table above while leaving the NAT path alone. Two staff posts on this forum state that trial accounts carry network restrictions regardless of configured policy settings, and the serverless SQL warehouse requirements state the account must not be on a free trial.

This is inference, not measurement. We cannot put the account back on trial to test it. The before and after on bucket C is measured, and the trial end date is confirmed in writing. The mechanism is not.

Note the shape carefully. It is not that trial accounts cannot use customer S3 buckets. Cross-region customer buckets worked fine throughout. Stating it too broadly sent us chasing the wrong mechanism for days.

WHAT WOULD HELP FROM DATABRICKS

1. Confirm whether account-regional namespace buckets are supported by serverless compute. If not, this belongs in the documented limitations. The namespace is a first-class AWS S3 feature and nothing currently warns against it.

2. Identify the component returning the 403. The request is signed, the credential is valid, and the same address answers correctly when the request comes from other compute, but the response carries no S3 request ID.

3. Comment on the SDK hypothesis in Finding 2, which we cannot test from outside.

4. Confirm whether a trial account restricts same-region serverless S3 access. If it does, that belongs in the published trial limitations. We could find nothing stating it.
1 REPLY 1

Louis_Frolio
Databricks Employee
Databricks Employee

Greetings @alexgfowler_ld, I did some digging and here is what I found.

First, this is one of the most carefully built-out reports I've seen on this forum. Your A/B test already rules out the usual suspects: same role, credential, data, and region work from classic compute, and serverless fails only for the account-regional bucket. That's not an IAM or Unity Catalog setup problem. I can't see inside the serverless network path (nobody outside engineering can), but I can confirm what's confirmable and point you at the right door.

What AWS says

Account regional namespaces launched March 12, 2026, and the S3 user guide says these buckets support every S3 feature with no application changes required. So from the AWS side, nothing about the bucket should behave differently. This is something in the path between serverless compute and S3, not an S3 limitation.

Where the 403 is likely coming from

Your read on the null request ID is right. S3 stamps every response it generates, errors included, with x-amz-request-id and x-amz-id-2. Both are null and the exception names a proxy, so the simplest explanation is that the proxy in the serverless network path minted the 403 and the request never reached S3. Same-region S3 traffic from serverless rides a VPC gateway endpoint inside Databricks-managed VPCs (the firewall docs describe this), and cross-region traffic takes a different route, which is why ap-southeast-5 kept working during your trial.

On the SDK hypothesis

It holds together. The failing request went to the legacy dash-style hostname (bucket.s3-ap-southeast-1.amazonaws.com), which older partition data hands out for regions it knows, while an unknown region falls back to s3.REGION.amazonaws.com. So the two populations really do take different hostname shapes. I wouldn't call it root cause yet, though, and I'd hold off on the conclusion that S3 refuses these buckets on the legacy hostname: if S3 were saying no, you'd have a request ID. More likely a component in the serverless path matches on hostname and handles the two shapes differently. One detail to put in front of support: these bucket names now contain a region string and an account ID, so anything that parses names or hostnames to work out region or ownership has a fresh way to get it wrong. Also, the AWS SDK for Java 1.x reached end of support on December 31, 2025, so a client pinned to it was never going to learn about a bucket feature that shipped in March 2026.

Your other questions

Trial accounts: the serverless SQL warehouse requirements state plainly that the account must not be on a free trial. The mechanism isn't documented, so I can't confirm it's the same-region path specifically, but your before-and-after on bucket C is clean evidence and I agree it belongs in the published trial limitations.

Fine-grained access control: your caution is correct. On dedicated access mode compute, any query touching a table with row filters or column masks, or a dynamic view, is handed to the workspace's serverless compute for filtering, so it would fail even though the cluster itself can read the bucket. Standard access mode compute and pro SQL warehouses evaluate those controls themselves, so that's the safer ground in the meantime. Your note about tasks with no compute setting falling back to serverless is also correct.

Support status: the serverless limitations page says nothing about bucket namespaces either way. I don't know the answer, and the public docs don't establish it as unsupported.

Workarounds and next step

Your two workarounds are the right ones: classic compute for existing account-regional buckets, or a global-namespace bucket for anything serverless must read, after weighing the security trade-off you already described. Then please open a support case at help.databricks.com. Your post is most of the ticket already. Include the repro, the null request ID observation, the ap-southeast-5 contrast, the compute types and timestamps, and the exact hostnames in the failing versus working requests (that last one is most likely to shorten the investigation). Redact the account ID in bucket names. Only the serverless team can see what the proxy is doing with these requests, and if it turns out to be a gap, you've already written the documentation ticket.

References:

Regards, Louis.