<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic Best Practices and Questions for Configuring High QPS AI Search Endpoints in Databricks in Get Started Discussions</title>
    <link>https://community.databricks.com/t5/get-started-discussions/best-practices-and-questions-for-configuring-high-qps-ai-search/m-p/164365#M11945</link>
    <description>&lt;P class=""&gt;Hello Databricks Community,&lt;/P&gt;&lt;P class=""&gt;I am exploring the configuration and scaling of high QPS (Queries Per Second) endpoints for Databricks AI Search, as described in the&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;A class="" title="/aws/en/ai-search/high-qps" href="https://docs.databricks.com/aws/en/ai-search/high-qps" target="_self"&gt;official documentation&lt;/A&gt;. I have several questions and would appreciate insights from anyone who has experience with high-throughput search workloads:&lt;/P&gt;&lt;OL class=""&gt;&lt;LI&gt;What is the maximum QPS supported for AI Search endpoints, and how can I request higher QPS for my application?&lt;/LI&gt;&lt;LI&gt;How does Databricks provision infrastructure to meet the target QPS, and are there any guarantees for sustained throughput?&lt;/LI&gt;&lt;LI&gt;What factors influence the QPS limits for a given AI Search index (e.g., index size, data type, concurrency)?&lt;/LI&gt;&lt;LI&gt;How can I monitor and troubleshoot QPS-related issues, such as 429 (Too Many Requests) errors or latency degradation?&lt;/LI&gt;&lt;LI&gt;Are there recommended best practices for scaling AI Search endpoints to support real-time applications with high QPS requirements?&lt;/LI&gt;&lt;LI&gt;Can I dynamically adjust the target QPS after an endpoint is created, and what is the impact on performance and cost?&lt;/LI&gt;&lt;LI&gt;What are the differences between standard and high QPS endpoints, and how do I choose the right configuration for my workload?&lt;/LI&gt;&lt;LI&gt;Is there a way to estimate the required QPS for my application based on expected traffic and query complexity?&lt;/LI&gt;&lt;LI&gt;How does QPS scaling interact with index updates or sync operations—are there any limitations or considerations?&lt;/LI&gt;&lt;LI&gt;Are there any additional costs associated with requesting or maintaining high QPS endpoints?&lt;/LI&gt;&lt;/OL&gt;&lt;P class=""&gt;If you have practical experience, tips, or documentation links, please share! I believe this discussion will help others planning to scale their AI Search workloads.&lt;/P&gt;&lt;P class=""&gt;Thank you!&lt;/P&gt;</description>
    <pubDate>Wed, 29 Jul 2026 06:38:30 GMT</pubDate>
    <dc:creator>pj-celebal-tech</dc:creator>
    <dc:date>2026-07-29T06:38:30Z</dc:date>
    <item>
      <title>Best Practices and Questions for Configuring High QPS AI Search Endpoints in Databricks</title>
      <link>https://community.databricks.com/t5/get-started-discussions/best-practices-and-questions-for-configuring-high-qps-ai-search/m-p/164365#M11945</link>
      <description>&lt;P class=""&gt;Hello Databricks Community,&lt;/P&gt;&lt;P class=""&gt;I am exploring the configuration and scaling of high QPS (Queries Per Second) endpoints for Databricks AI Search, as described in the&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;A class="" title="/aws/en/ai-search/high-qps" href="https://docs.databricks.com/aws/en/ai-search/high-qps" target="_self"&gt;official documentation&lt;/A&gt;. I have several questions and would appreciate insights from anyone who has experience with high-throughput search workloads:&lt;/P&gt;&lt;OL class=""&gt;&lt;LI&gt;What is the maximum QPS supported for AI Search endpoints, and how can I request higher QPS for my application?&lt;/LI&gt;&lt;LI&gt;How does Databricks provision infrastructure to meet the target QPS, and are there any guarantees for sustained throughput?&lt;/LI&gt;&lt;LI&gt;What factors influence the QPS limits for a given AI Search index (e.g., index size, data type, concurrency)?&lt;/LI&gt;&lt;LI&gt;How can I monitor and troubleshoot QPS-related issues, such as 429 (Too Many Requests) errors or latency degradation?&lt;/LI&gt;&lt;LI&gt;Are there recommended best practices for scaling AI Search endpoints to support real-time applications with high QPS requirements?&lt;/LI&gt;&lt;LI&gt;Can I dynamically adjust the target QPS after an endpoint is created, and what is the impact on performance and cost?&lt;/LI&gt;&lt;LI&gt;What are the differences between standard and high QPS endpoints, and how do I choose the right configuration for my workload?&lt;/LI&gt;&lt;LI&gt;Is there a way to estimate the required QPS for my application based on expected traffic and query complexity?&lt;/LI&gt;&lt;LI&gt;How does QPS scaling interact with index updates or sync operations—are there any limitations or considerations?&lt;/LI&gt;&lt;LI&gt;Are there any additional costs associated with requesting or maintaining high QPS endpoints?&lt;/LI&gt;&lt;/OL&gt;&lt;P class=""&gt;If you have practical experience, tips, or documentation links, please share! I believe this discussion will help others planning to scale their AI Search workloads.&lt;/P&gt;&lt;P class=""&gt;Thank you!&lt;/P&gt;</description>
      <pubDate>Wed, 29 Jul 2026 06:38:30 GMT</pubDate>
      <guid>https://community.databricks.com/t5/get-started-discussions/best-practices-and-questions-for-configuring-high-qps-ai-search/m-p/164365#M11945</guid>
      <dc:creator>pj-celebal-tech</dc:creator>
      <dc:date>2026-07-29T06:38:30Z</dc:date>
    </item>
    <item>
      <title>Re: Best Practices and Questions for Configuring High QPS AI Search Endpoints in Databricks</title>
      <link>https://community.databricks.com/t5/get-started-discussions/best-practices-and-questions-for-configuring-high-qps-ai-search/m-p/164630#M11949</link>
      <description>&lt;P&gt;Answers to your questions:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;STRONG&gt;Max QPS:&lt;/STRONG&gt;&amp;nbsp;Standard endpoints default to &lt;STRONG&gt;20–200 QPS&lt;/STRONG&gt; depending on index size, and with high QPS you can scale to &lt;STRONG&gt;1,000+ QPS&lt;/STRONG&gt; on standard endpoints.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;How to request more:&lt;/STRONG&gt; Set &lt;CODE&gt;target_qps&lt;/CODE&gt; when creating or updating a &lt;STRONG&gt;standard&lt;/STRONG&gt; endpoint via UI, SDK, or REST. Databricks then provisions extra capacity automatically.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Guarantees:&lt;/STRONG&gt; It’s &lt;STRONG&gt;best-effort, not guaranteed&lt;/STRONG&gt;; actual throughput depends on workload and traffic shape.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;What affects QPS:&lt;/STRONG&gt; index size, vector dimensionality, query complexity, filter usage, &lt;CODE&gt;num_results&lt;/CODE&gt;, and whether you use managed embeddings / optimized query route.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Monitoring/troubleshooting:&lt;/STRONG&gt; Check endpoint &lt;CODE&gt;scaling_info.state&lt;/CODE&gt;, use the &lt;STRONG&gt;index URL&lt;/STRONG&gt; with &lt;STRONG&gt;service principal OAuth&lt;/STRONG&gt;, and look for &lt;CODE&gt;SCALING_CHANGE_IN_PROGRESS&lt;/CODE&gt;, 429s, or rising latency.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Best practices:&lt;/STRONG&gt; Use service principals, reuse the index object, keep &lt;CODE&gt;num_results&lt;/CODE&gt; modest, load test, and split/replicate across endpoints if you need more total throughput.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Dynamic changes:&lt;/STRONG&gt; Yes — you can update &lt;CODE&gt;target_qps&lt;/CODE&gt; after creation. While scaling is in progress, capacity changes aren’t immediate; avoid changing it again until the state becomes applied.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Standard vs high QPS:&lt;/STRONG&gt; High QPS is for &lt;STRONG&gt;standard endpoints only&lt;/STRONG&gt;; storage-optimized endpoints don’t support it. Standard is the right choice when sustained query throughput matters.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Estimating needs:&lt;/STRONG&gt; Use expected traffic + query shape, then load test. Databricks also provides a GenAI calculator for rough cost/capacity estimates.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Sync/update interaction:&lt;/STRONG&gt; Changing &lt;CODE&gt;target_qps&lt;/CODE&gt; does &lt;STRONG&gt;not&lt;/STRONG&gt; require a sync to take effect, but the new capacity applies only after provisioning completes.&lt;/LI&gt;
&lt;LI&gt;&lt;STRONG&gt;Cost:&lt;/STRONG&gt; Higher target QPS increases endpoint cost, and you’re charged for the provisioned capacity even if traffic is lower.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Docs:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;A href="https://docs.databricks.com/aws/en/ai-search/high-qps" target="_blank"&gt;High QPS guide&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://docs.databricks.com/aws/en/ai-search/best-practices" target="_blank"&gt;AI Search performance guide&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;A href="https://docs.databricks.com/aws/en/ai-search/cost-management" target="_blank"&gt;Cost management guide&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;</description>
      <pubDate>Fri, 31 Jul 2026 17:28:12 GMT</pubDate>
      <guid>https://community.databricks.com/t5/get-started-discussions/best-practices-and-questions-for-configuring-high-qps-ai-search/m-p/164630#M11949</guid>
      <dc:creator>Lu_Wang_ENB_DBX</dc:creator>
      <dc:date>2026-07-31T17:28:12Z</dc:date>
    </item>
    <item>
      <title>Re: Best Practices and Questions for Configuring High QPS AI Search Endpoints in Databricks</title>
      <link>https://community.databricks.com/t5/get-started-discussions/best-practices-and-questions-for-configuring-high-qps-ai-search/m-p/165288#M11985</link>
      <description>&lt;P class=""&gt;For high-QPS AI Search endpoints in Databricks, I’d start by separating the problem into &lt;STRONG&gt;latency, throughput, scaling, and cost&lt;/STRONG&gt;. Make sure you know your target QPS, acceptable p95/p99 latency, query size, embedding/model latency, and whether traffic arrives in bursts or stays steady.&lt;/P&gt;&lt;P&gt;A few practical best practices:&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;&lt;STRONG&gt;Load-test with realistic traffic&lt;/STRONG&gt; rather than relying on a small number of concurrent requests.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Monitor p50/p95/p99 latency&lt;/STRONG&gt;, errors, throttling, CPU/memory utilization, and queueing.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Tune endpoint capacity and autoscaling&lt;/STRONG&gt; based on sustained QPS as well as traffic spikes.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Keep retrieval efficient&lt;/STRONG&gt; by limiting unnecessary result sizes and avoiding expensive filtering where possible.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Cache repeated queries or embeddings&lt;/STRONG&gt; when the workload allows it.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Use retries carefully&lt;/STRONG&gt; with exponential backoff so a temporary slowdown doesn’t turn into a retry storm.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Separate ingestion/update workloads from serving traffic&lt;/STRONG&gt; when possible so index updates don’t negatively affect query performance.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Measure cost per request&lt;/STRONG&gt; alongside latency; simply adding capacity may solve QPS problems but create an unnecessarily expensive architecture.&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;The main questions I’d ask before choosing a configuration are: What QPS are you targeting, what p95 latency do you need, how large is the index, how frequently does it change, and are requests mostly identical or highly variable?&lt;/P&gt;&lt;P&gt;If you’re also looking for a straightforward web-based option to explore, &lt;A href="https://playeggycar.io/" target="_self"&gt;playeggycar.io&lt;/A&gt; is another site you can check out.&lt;/P&gt;</description>
      <pubDate>Mon, 10 Aug 2026 23:30:05 GMT</pubDate>
      <guid>https://community.databricks.com/t5/get-started-discussions/best-practices-and-questions-for-configuring-high-qps-ai-search/m-p/165288#M11985</guid>
      <dc:creator>Brodybenson</dc:creator>
      <dc:date>2026-08-10T23:30:05Z</dc:date>
    </item>
  </channel>
</rss>

