### Page: https://forum.scylladb.com/t/about-the-announcements-category/1 Title: About the Announcements category - Announcements - ScyllaDB Community NoSQL Forum Meta Description: News and information related to products, the ScyllaDB community, events, webinars, and so on Language: en Canonical URL: https://forum.scylladb.com/t/about-the-announcements-category/1 ## Headings Structure: H1: About the Announcements category H3: Related topics ## Main Content: H1: About the Announcements category H3: Related topics News and information related to products, the ScyllaDB community, events, webinars, and so on --- ### Page: https://forum.scylladb.com/t/about-the-database-community-category/3 Title: About the Database Community category - Database Community - ScyllaDB Community NoSQL Forum Meta Description: Have fun and engage with your fellow community members. Feel free to introduce yourself here, add feature requests, and provide feedback. Also, the place to post ScyllaDB related job opportunities. How are you using S… Language: en Canonical URL: https://forum.scylladb.com/t/about-the-database-community-category/3 ## Headings Structure: H1: About the Database Community category H3: Related topics ## Main Content: H1: About the Database Community category H3: Related topics Have fun and engage with your fellow community members. Feel free to introduce yourself here, add feature requests, and provide feedback. Also, the place to post ScyllaDB related job opportunities. How are you using ScyllaDB? --- ### Page: https://forum.scylladb.com/t/welcome-to-the-scylladb-community-nosql-forum/7 Title: Welcome to the ScyllaDB Community NoSQL Forum - Database Community - ScyllaDB Community NoSQL Forum Meta Description: Explore NoSQL topics as you learn from the ScyllaDB community Before posting, please search and make sure the question hasn't been asked before. To ask a question, [log in](https://forum.scylladb.com/login) first. A fe… Language: en Canonical URL: https://forum.scylladb.com/t/welcome-to-the-scylladb-community-nosql-forum/7 ## Headings Structure: H1: Welcome to the ScyllaDB Community NoSQL Forum H2: Explore NoSQL topics as you learn from the ScyllaDB community H3: Related topics ## Main Content: H1: Welcome to the ScyllaDB Community NoSQL Forum H2: Explore NoSQL topics as you learn from the ScyllaDB community H3: Related topics A few helpful resources: Also, here’s a blog with details on why we launched this forum and how it relates to Slack and our other community channels: Introducing the ScyllaDB Community Forum - ScyllaDB --- ### Page: https://forum.scylladb.com/t/about-the-scylladb-category/12 Title: About the ScyllaDB category - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: For general questions related to ScyllaDB products. Topics include Troubleshooting, Benchmarks, Data modeling, Drivers, 3rd party integrations, etc. Language: en Canonical URL: https://forum.scylladb.com/t/about-the-scylladb-category/12 ## Headings Structure: H1: About the ScyllaDB category H3: Related topics ## Main Content: H1: About the ScyllaDB category H3: Related topics For general questions related to ScyllaDB products. Topics include Troubleshooting, Benchmarks, Data modeling, Drivers, 3rd party integrations, etc. Hi, I studied you peoples forum and earlier also checked one webinar, any chance to join you people to learn something more. --- ### Page: https://forum.scylladb.com/t/about-the-university-and-training-category/13 Title: About the University and Training category - University and Training - ScyllaDB Community NoSQL Forum Meta Description: For topics regarding specific ScyllaDB University courses, lessons and training events. Language: en Canonical URL: https://forum.scylladb.com/t/about-the-university-and-training-category/13 ## Headings Structure: H1: About the University and Training category H3: Related topics ## Main Content: H1: About the University and Training category H3: Related topics For topics regarding specific ScyllaDB University courses, lessons and training events. --- ### Page: https://forum.scylladb.com/t/scylladb-summit-2023-call-for-speakers-is-open/30 Title: ScyllaDB Summit 2023 Call for Speakers is open! - Announcements - ScyllaDB Community NoSQL Forum Meta Description: Call for Speakers for ScyllaDB Summit 2023 is now open! If you have war stories of deployments or awesome tools and integrations we’d love to hear! Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-summit-2023-call-for-speakers-is-open/30 ## Headings Structure: H1: ScyllaDB Summit 2023 Call for Speakers is open! H3: Related topics ## Main Content: H1: ScyllaDB Summit 2023 Call for Speakers is open! H3: Related topics Call for Speakers for ScyllaDB Summit 2023 is now open! If you have war stories of deployments or awesome tools and integrations we’d love to hear! --- ### Page: https://forum.scylladb.com/t/high-performance-nosql-masterclass-register-now-for-09-nov-2022/31 Title: High Performance NoSQL Masterclass — register now for 09 Nov 2022 - Announcements - ScyllaDB Community NoSQL Forum Meta Description: Register now for our latest Masterclass, and read the agenda. Hosted by ScyllaDB and our friends at Pythian. If you’re registered, sound off below! Language: en Canonical URL: https://forum.scylladb.com/t/high-performance-nosql-masterclass-register-now-for-09-nov-2022/31 ## Headings Structure: H1: High Performance NoSQL Masterclass — register now for 09 Nov 2022 H3: Related topics ## Main Content: H1: High Performance NoSQL Masterclass — register now for 09 Nov 2022 H3: Related topics Register now for our latest Masterclass, and read the agenda. Hosted by ScyllaDB and our friends at Pythian. If you’re registered, sound off below! --- ### Page: https://forum.scylladb.com/t/engaging-with-the-scylladb-community/32 Title: Engaging with the ScyllaDB Community - Database Community - ScyllaDB Community NoSQL Forum Meta Description: There are already a number of ways to engage with the ScyllaDB Community. Let’s just go over a few: ScyllaDB User Slack (http://slack.scylladb.com/) ScyllaDB User Google Group “scylladb-users” (https://groups.google.co… Language: en Canonical URL: https://forum.scylladb.com/t/engaging-with-the-scylladb-community/32 ## Headings Structure: H1: Engaging with the ScyllaDB Community H3: Related topics ## Main Content: H1: Engaging with the ScyllaDB Community H3: Related topics There are already a number of ways to engage with the ScyllaDB Community. Let’s just go over a few: So why would we want another? Because users appreciate choice, and different people like to engage via different media. That said, there are usually strengths and weaknesses to different media which would gravitate certain conversations to different platforms. Online forums are nothing new, but community forums, discussion boards, persist because they have certain advantages: If people are happy to be using their existing media of choice to communicate with the ScyllaDB community and are getting the answers they need, then please stick with it! These forums are our way of adding more choice, and to hopefully to get better re-use of the knowledge you all share with the community day in and day out. As the Director of Technical Advocacy here at ScyllaDB, we hope to use this forum to communicate with you more actively, interactively and directly. A big shout-out to Guy Shtub who championed the project into fruition! Thanks Peter! A bit more about when to use slack and when to use this forum: Community Forum This forum is a great place to ask questions, get detailed answers, have in-depth discussions, and search for all the previously answered questions. Benefits: User Slack Our community Slack channel provides real-time chat with ScyllaDB experts and other ScyllaDB users. Benefits: Great work! Lets send it to the open world… An official Discord server would be so cool tohave for Scylla. --- ### Page: https://forum.scylladb.com/t/say-hello-how-are-you-using-scylladb/35 Title: Say Hello, How are you using ScyllaDB? - Database Community - ScyllaDB Community NoSQL Forum Meta Description: Welcome to the forum! Please go ahead and introduce yourself. ScyllaDB is used for many different use cases. Any interesting projects you’d like to share with your fellow community members? Are there any tips/tricks y… Language: en Canonical URL: https://forum.scylladb.com/t/say-hello-how-are-you-using-scylladb/35 ## Headings Structure: H1: Say Hello, How are you using ScyllaDB? H3: Related topics ## Main Content: H1: Say Hello, How are you using ScyllaDB? H3: Related topics Welcome to the forum! Please go ahead and introduce yourself. ScyllaDB is used for many different use cases. Any interesting projects you’d like to share with your fellow community members? Are there any tips/tricks you learned along the way? How did you first hear about ScyllaDB? I’ll go first My name is Guy Shtub, and I’m Head of Training at ScyllaDB. I joined Scylla over four years ago to build ScyllaDB University. My tip is that we have lots of topics covered in ScyllaDB University and in the Documentation. If you have any questions or issues, you can ask them here. Also, feel free to suggest new topics/lessons/hands-on labs! Hiii My name is Mahdi Fardkohan and I have been working as a DBA & BI Engineer in a software company for several years. I’m so passionate about Data Mesh Architecture, and I found ScyllaDB when I was researching high-performance open-source distributed databases for Big Data to use in a Retail startup and I was amazed by its capabilities. I’m honored to be here with great people like you in the ScyllaDB community, and I’m also very grateful to the venerable Scylla and its wonderful experts who put this power and knowledge into the hands of the world. Thanks for your kind feedback, Mahdi! My name is Yaniv and I’m the VP R&D of Scylla, very happy to be here and read the forum notes! My name is Diego and I’m the support engineer at Scylla, interested in the forum notes, thanks! Hi, my name is Fabio, and I’m a software engineer in testing leader. I’m into details, and passionate about performance, so happy to be here. Hi Everyone! Great to be here! This is my first time ever to join a cool community like ScyllaDB My name is Trinh, an expat working in Germany as Data Engineer. I have no experience with NoSQL in general, but using SQL and Distributed DB (ex: HDFS) a lot. I’d like to fullfill my knowledge in this area so that I can complete my picture of Data industry. Happy to meet you in my Twitter https://twitter.com/bkincities Hi, My name is Raya and I am a a software engineer in testing manager at ScyllaDB. Thrilled to be here Hi, My name is Orane Gabrielovitch, and I’m the Israel office admin. Happy to be here Hi all, I’m happy to be part of our ScyllaDB community Inna B. Operations & Admin Manager at ScyllaDB Hello ScyllaDB community. I’m happy to be here. I don’t use ScyllaDB but I’m curious to learn more. Hi there. I’m Suraj Vijayan, Undergraduate student from India. I’m a Full Stack web and app developer (started coding at 16) and currently into Rust and Cassandra. Learnt Cassandra (through DataStax workhops and their academy) and got certified as Developer Associate. I had always been looking into Discord’s tech stack and came to know the fact that they moved from Cassandra to ScyllaDB. Though I studied a lot about it in blogs, I so curios (excited too like - “Whoa!! Damn! Awesome!! why not give a try?” ). And this how I ended here. Looking forward to learn ScyllaDB. Talking about projects I have a made a lot, but the one I like a lot is my portfolio site. As its an full stack site and uses Cassandra (AstraDB) to collect website stats (yes a lot, no google analytics), people can comment, like and share my works and achievments. Why not give a try to it - surajvijayan.me and also an honorable mention to my AstraDB Navigator - A low code environment to manage you AstraDB instances. I guess this is an pretty huge message. but yeah thank you if you came this far. Hi, My name is Clive Attard, and I’m a software engineer learning new technologies. Excited to be here Hello, This is kartikeya. Currently started using scylla db in our org Hi, I am Abdulshakur. Excited to learn Scylladb Hi, I am Gary, looking forward to learn Scylladb! Hi , My name is Amir , I am a Software engineer. Hi Scylladb community I’m Charan working as a Database reliability engineer. Currently I’m learning and implementing scylla on my environments. Happy to connect here. --- ### Page: https://forum.scylladb.com/t/about-the-knowledge-base-category/40 Title: About the Knowledge Base category - Knowledge Base - ScyllaDB Community NoSQL Forum Meta Description: Topics for understanding and troubleshooting ScyllaDB. These are frequently asked questions and general topics. Language: en Canonical URL: https://forum.scylladb.com/t/about-the-knowledge-base-category/40 ## Headings Structure: H1: About the Knowledge Base category H3: Related topics ## Main Content: H1: About the Knowledge Base category H3: Related topics Topics for understanding and troubleshooting ScyllaDB. These are frequently asked questions and general topics. --- ### Page: https://forum.scylladb.com/t/what-is-the-difference-between-clustering-primary-partition-and-composite-or-compound-keys-in-scylladb/41 Title: What is the difference between Clustering, Primary, Partition, and Composite (or Compound) Keys in ScyllaDB? - Knowledge Base - ScyllaDB Community NoSQL Forum Meta Description: In ScyllaDB (and Apache Cassandra for that matter) A Primary Key is defined within a table. It is one or more columns used to identify a row. All tables must include a definition for a Primary Key. For example, in the ta… Language: en Canonical URL: https://forum.scylladb.com/t/what-is-the-difference-between-clustering-primary-partition-and-composite-or-compound-keys-in-scylladb/41 ## Headings Structure: H1: What is the difference between Clustering, Primary, Partition, and Composite (or Compound) Keys in ScyllaDB? H3: Related topics ## Main Content: H1: What is the difference between Clustering, Primary, Partition, and Composite (or Compound) Keys in ScyllaDB? H3: Related topics In ScyllaDB (and Apache Cassandra for that matter) A Primary Key is defined within a table. It is one or more columns used to identify a row. All tables must include a definition for a Primary Key. For example, in the table: The Primary Key is a single column – the pet_chip_id. If a Primary Key is made up of a single column, it is called a Simple Primary Key. It’s also possible to define the Primary Key to include more than one column, in which case it is called a Composite (or Compound) key. For example: In this case, the first part of the Primary Key is called the Partition Key (pet_chip_id in the above example) and the second part is called the Clustering Key (time). The Partition Key is responsible for data distribution across the nodes. It determines which node will store a given row. It can be one or more columns. The Clustering Key is responsible for sorting the rows within the partition. It can be zero or more columns. If a table has more than one column defined as the Primary Key, for example: In this case, the Partition Key includes two columns: pet_chip_id and time, and the Clustering Key is pet_name. Every query must include all the columns defined in the Partition Key (pet_chip_id and time) in this case. Look at another example: If there is more than one column in the Clustering Key (pet_name and heart_rate in the example above), the order of these columns defines the clustering order. For a given partition, all the rows are physically ordered inside ScyllaDB by the clustering order. This order determines what select queries you can efficiently run on this partition. In this example, the ordering is first by pet_name and then by heart_rate. In addition to the Partition Key columns, a query may include the Clustering Key. If it does include the Clustering Key columns, they must be used in the same order as they were defined. Additional Resources: --- ### Page: https://forum.scylladb.com/t/how-many-connections-should-i-open-from-each-scylladb-client-application/43 Title: How many connections should I open from each ScyllaDB client application? - Knowledge Base - ScyllaDB Community NoSQL Forum Meta Description: As a rule of thumb, for Scylla’s best performance, each client needs at least 1-3 connections per Scylla core. For example, in a cluster with three nodes, each node with 16 cores, each client application should open 32 (… Language: en Canonical URL: https://forum.scylladb.com/t/how-many-connections-should-i-open-from-each-scylladb-client-application/43 ## Headings Structure: H1: How many connections should I open from each ScyllaDB client application? H3: Related topics ## Main Content: H1: How many connections should I open from each ScyllaDB client application? H3: Related topics As a rule of thumb, for Scylla’s best performance, each client needs at least 1-3 connections per Scylla core. For example, in a cluster with three nodes, each node with 16 cores, each client application should open 32 (2x16) connections to each Scylla node. Additional Resources: --- ### Page: https://forum.scylladb.com/t/whats-the-best-way-to-count-scan-a-table-in-scylladb-full-table-scan/48 Title: What's the best way to count / scan a table in ScyllaDB? (full table scan) - Knowledge Base - ScyllaDB Community NoSQL Forum Meta Description: If you wish to perform a full table scan, please look into the following blog posts on how to do an efficient full table scan: The theory behind an efficient full table scan Blog by Avi (our CTO) Follow up Blog with co… Language: en Canonical URL: https://forum.scylladb.com/t/whats-the-best-way-to-count-scan-a-table-in-scylladb-full-table-scan/48 ## Headings Structure: H1: What's the best way to count / scan a table in ScyllaDB? (full table scan) H3: Related topics ## Main Content: H1: What's the best way to count / scan a table in ScyllaDB? (full table scan) H3: Related topics If you wish to perform a full table scan, please look into the following blog posts on how to do an efficient full table scan: Note: There’s also a python version of that efficient full table scan script. --- ### Page: https://forum.scylladb.com/t/scylladb-version-upgrade/49 Title: ScyllaDB Version Upgrade - Knowledge Base - ScyllaDB Community NoSQL Forum Meta Description: Can I upgrade from ScyllaDB Version X to version Y? Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-version-upgrade/49 ## Headings Structure: H1: ScyllaDB Version Upgrade H3: Related topics ## Main Content: H1: ScyllaDB Version Upgrade H3: Related topics Can I upgrade from ScyllaDB Version X to version Y? If you’re using ScyllaDB Cloud, you don’t have to upgrade, as it’s a fully managed service. It deploys the latest ScyllaDB Enterprise version, and all upgrades are performed by ScyllaDB. If you’re using ScyllaDB Open Source or Enterprise, you can upgrade to a newer version following these upgrade guides: A ScyllaDB upgrade is a rolling procedure - it does not require full cluster shutdown and is performed without any downtime or disruption of service. To ensure a successful upgrade and avoid breaking anything, you should perform your upgrades consecutively - to each successive version. For example, to upgrade from version 4.4 to 5.0, you should first upgrade ScyllaDB to version 4.5, next from 4.5 to 4.6, and finally, from 4.6 to 5.0, without skipping any major version. --- ### Page: https://forum.scylladb.com/t/scylladb-is-using-up-all-of-my-memory-what-s-going-on/50 Title: ScyllaDB is using up all of my memory. What’s going on? - Knowledge Base - ScyllaDB Community NoSQL Forum Meta Description: ScyllaDB is designed to use all the memory it has and to put it to good use. Most notably to cache data. By default, when ScyllaDB starts up, it inspects the node’s hardware configuration and claims all memory to itself,… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-is-using-up-all-of-my-memory-what-s-going-on/50 ## Headings Structure: H1: ScyllaDB is using up all of my memory. What’s going on? H3: Related topics ## Main Content: H1: ScyllaDB is using up all of my memory. What’s going on? H3: Related topics ScyllaDB is designed to use all the memory it has and to put it to good use. Most notably to cache data. By default, when ScyllaDB starts up, it inspects the node’s hardware configuration and claims all memory to itself, leaving some reserve for the operating system. The assumption is that ScyllaDB does not run in a shared environment. If for some reason you do want to give scyllaDB less memory, say for testing or development, you can do so: Additional Resources: --- ### Page: https://forum.scylladb.com/t/scylladb-university-live-1st-of-december/52 Title: ScyllaDB University LIVE - 1st of December - Announcements - ScyllaDB Community NoSQL Forum Meta Description: The next ScyllaDB University LIVE training event will take place on the 1st of December. It’s a half-day free online training with some of our best engineers and experts. Learn more here, hope to see you there. Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-university-live-1st-of-december/52 ## Headings Structure: H1: ScyllaDB University LIVE - 1st of December H3: Related topics ## Main Content: H1: ScyllaDB University LIVE - 1st of December H3: Related topics The next ScyllaDB University LIVE training event will take place on the 1st of December. It’s a half-day free online training with some of our best engineers and experts. Learn more here, hope to see you there. Check out my blog post ScyllaDB University LIVE, Fall 203: From Getting Started to Expert Tips & Tricks for more details, see you there! This is happening tomorrow (Wednesday). You can still save your spot here. --- ### Page: https://forum.scylladb.com/t/i-m-having-performance-issues-what-should-i-do/53 Title: I’m having performance issues. What should I do? - Knowledge Base - ScyllaDB Community NoSQL Forum Meta Description: ScyllaDB auto-tunes for optimal performance, however, users still need to apply best practices in order to get the most out of their ScyllaDB deployments. Lower than expected performance can be a result of many factors,… Language: en Canonical URL: https://forum.scylladb.com/t/i-m-having-performance-issues-what-should-i-do/53 ## Headings Structure: H1: I’m having performance issues. What should I do? H3: Related topics ## Main Content: H1: I’m having performance issues. What should I do? H3: Related topics ScyllaDB auto-tunes for optimal performance, however, users still need to apply best practices in order to get the most out of their ScyllaDB deployments. Lower than expected performance can be a result of many factors, from hardware (storage, CPU, network) to data modeling to the application layer. As a first step, make sure you have ScyllaDB Monitoring in place. Looking at the monitoring dashboards is the best way to look for bottlenecks and understand the cause of the issues. Here are some other tips: Additional Resources: --- ### Page: https://forum.scylladb.com/t/scylladb-monitoring-has-no-data-in-the-charts/54 Title: ScyllaDB Monitoring has no data in the charts - Knowledge Base - ScyllaDB Community NoSQL Forum Meta Description: The most common reason for seeing no data about your cluster after installing the ScyllaDB Monitoring stack is an issue with the connection to Prometheus. To check this: Login to the Prometheus console by pointing your… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-monitoring-has-no-data-in-the-charts/54 ## Headings Structure: H1: ScyllaDB Monitoring has no data in the charts H3: Related topics ## Main Content: H1: ScyllaDB Monitoring has no data in the charts H3: Related topics The most common reason for seeing no data about your cluster after installing the ScyllaDB Monitoring stack is an issue with the connection to Prometheus. To check this: Other things to check: Make sure you are not using the local network for the local IP range when using Docker containers. By default, the local IP range (127.0.0.X) is inside the Docker container. If you are trying to connect to a target via the local IP range from a Docker container, you need to use the -l flag to enable the local network stack. Verify that Prometheus is pointing to the correct target by checking prometheus/scylla_servers.yml. Make sure that your dashboard and Scylla versions are aligned. If, for example, you are running Scylla 5.1, you can specify a specific version with the -v flag when starting the monitoring stack: ./start-all.sh -v 5.1 Additional Resources: --- ### Page: https://forum.scylladb.com/t/when-to-use-filtering-and-when-not/55 Title: When to use filtering - and when not - Knowledge Base - ScyllaDB Community NoSQL Forum Meta Description: Sometimes you want to be able to query by different columns, but you’re not interested in creating secondary indexes. Filtering is one more way of allowing such queries. The mechanism is really simple – the coordinator w… Language: en Canonical URL: https://forum.scylladb.com/t/when-to-use-filtering-and-when-not/55 ## Headings Structure: H1: When to use filtering - and when not H3: Related topics ## Main Content: H1: When to use filtering - and when not H3: Related topics Sometimes you want to be able to query by different columns, but you’re not interested in creating secondary indexes. Filtering is one more way of allowing such queries. The mechanism is really simple – the coordinator will fetch all of the results specified by the key restrictions, and then filter out rows that do not match the rest of the restrictions. But, there’s a catch. Filtering can be very performance-heavy, it can even result in fetching all rows from the table and then filter out just a few rows. Because of that, queries that involve filtering must be explicitly allowed to do so, by appending CQL ALLOW FILTERING keyword to each such query. Here are some popular resources about CQL ALLOW FILTERING --- ### Page: https://forum.scylladb.com/t/how-do-you-use-scylladb-with-spring-boot/56 Title: How do you use ScyllaDB with Spring Boot? - Knowledge Base - ScyllaDB Community NoSQL Forum Meta Description: This blog explains how to use Spring Boot apps with ScyllaDB for time series data, taking advantage of shard-aware drivers and prepared statements: Using Spring Boot, ScyllaDB and Time Series Data - ScyllaDB Language: en Canonical URL: https://forum.scylladb.com/t/how-do-you-use-scylladb-with-spring-boot/56 ## Headings Structure: H1: How do you use ScyllaDB with Spring Boot? H3: Related topics ## Main Content: H1: How do you use ScyllaDB with Spring Boot? H3: Related topics This blog explains how to use Spring Boot apps with ScyllaDB for time series data, taking advantage of shard-aware drivers and prepared statements: Using Spring Boot, ScyllaDB and Time Series Data - ScyllaDB --- ### Page: https://forum.scylladb.com/t/about-the-blog-posts-category/58 Title: About the Blog Posts category - Blog Posts - ScyllaDB Community NoSQL Forum Meta Description: (Replace this first paragraph with a brief description of your new category. This guidance will appear in the category selection area, so try to keep it below 200 characters.) Use the following paragraphs for a longer description, or to establish category guidelines or rules: Why should people use this category? What is it for? How exactly is this different than the other categories we already have? What should topics in this category generally contain? Do we need this category? Can... Language: en Canonical URL: https://forum.scylladb.com/t/about-the-blog-posts-category/58 ## Headings Structure: H1: About the Blog Posts category H3: Related topics ## Main Content: H1: About the Blog Posts category H3: Related topics (Replace this first paragraph with a brief description of your new category. This guidance will appear in the category selection area, so try to keep it below 200 characters.) Use the following paragraphs for a longer description, or to establish category guidelines or rules: Why should people use this category? What is it for? How exactly is this different than the other categories we already have? What should topics in this category generally contain? Do we need this category? Can we merge with another category, or subcategory? Honestly, this category feels too vague. before making a new one, ask yourself: who is it for, what belongs here, and what doesn’t. most new categories fail because they overlap with existing ones or nobody posts there. If you can’t clearly say “post this here, not there,” just use tags or a subcategory instead. keeps the forum clean and less confusing. --- ### Page: https://forum.scylladb.com/t/how-scylladb-helped-an-adtech-company-focus-on-core-business/60 Title: How ScyllaDB Helped an AdTech Company Focus on Core Business - Blog Posts - ScyllaDB Community NoSQL Forum Meta Description: [image] Language: en Canonical URL: https://forum.scylladb.com/t/how-scylladb-helped-an-adtech-company-focus-on-core-business/60 ## Headings Structure: H1: How ScyllaDB Helped an AdTech Company Focus on Core Business H3: How ScyllaDB Helped an AdTech Company Focus on Core Business H3: Related topics ## Main Content: H1: How ScyllaDB Helped an AdTech Company Focus on Core Business H3: How ScyllaDB Helped an AdTech Company Focus on Core Business H3: Related topics AdTech innovator GumGum wanted to escape the maintenance issues connected with Apache Cassandra. This podcast shares how they moved to a database-as-a-service (ScyllaDB Cloud DBaaS). --- ### Page: https://forum.scylladb.com/t/experimental-features/61 Title: Experimental features - Knowledge Base - ScyllaDB Community NoSQL Forum Meta Description: Experimental features are still under development and are not stable enough to be used in production. Their design is not finalized, and their API will likely change, breaking backward or forward compatibility. Experimen… Language: en Canonical URL: https://forum.scylladb.com/t/experimental-features/61 ## Headings Structure: H1: Experimental features H3: Related topics ## Main Content: H1: Experimental features H3: Related topics Experimental features are still under development and are not stable enough to be used in production. Their design is not finalized, and their API will likely change, breaking backward or forward compatibility. Experimental features are available for ScyllaDB Open Source users to test and provide feedback. ScyllaDB Enterprise and ScyllaDB Cloud do not support experimental features. To list all the experimental features available in your ScyllaDB version, run scylla --help. To enable an experimental feature, you can do one of the following: --- ### Page: https://forum.scylladb.com/t/which-driver-should-i-use/62 Title: Which driver should I use? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Which driver should I use? Can I use Apache Cassandra drivers? Language: en Canonical URL: https://forum.scylladb.com/t/which-driver-should-i-use/62 ## Headings Structure: H1: Which driver should I use? H3: Related topics ## Main Content: H1: Which driver should I use? H3: Related topics Which driver should I use? Can I use Apache Cassandra drivers? ScyllaDB is compatible with Apache Cassandra on the protocol level, so that every Cassandra Driver will work with Scylla out of the box. That being said, for many drivers (Java, Python, Go, Rust…) there are Scylla forks that take advantage of Scylla features, like shared per core, to allow better performance. It is recommended to use this fork when available. See Scylla CQL Drivers | Scylla Docs for a list of drivers. --- ### Page: https://forum.scylladb.com/t/database-security-authentication-authorization-rba-encyrption-and-audit/63 Title: Database Security: Authentication, Authorization, RBA, Encyrption and Audit - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: We’re concerned with data breaches. How secure is ScyllaDB? Language: en Canonical URL: https://forum.scylladb.com/t/database-security-authentication-authorization-rba-encyrption-and-audit/63 ## Headings Structure: H1: Database Security: Authentication, Authorization, RBA, Encyrption and Audit H3: Scylla Security Checklist | Scylla Docs H3: Related topics ## Main Content: H1: Database Security: Authentication, Authorization, RBA, Encyrption and Audit H3: Scylla Security Checklist | Scylla Docs H3: Related topics We’re concerned with data breaches. How secure is ScyllaDB? Security is a big topic that includes many features like: Encryption (on rest, in transit), Authentication and Authorization, role-based access control, auditing, and more. An excellent place to start is Scylla Security Checklist: Scylla is an Apache Cassandra-compatible NoSQL data store that can handle 1 million transactions per second on a single server. --- ### Page: https://forum.scylladb.com/t/can-the-compression-strategy-be-changed-from-stcs-to-ics-live-or-is-a-node-restart-required/64 Title: Can the compression strategy be changed from STCS to ICS live, or is a node restart required? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Yes, you can change the compaction strategy live - without restarting any nodes. Note that there might be increased compaction activity for a while after you make the change, so schedule it to off-peak hours if possible… Language: en Canonical URL: https://forum.scylladb.com/t/can-the-compression-strategy-be-changed-from-stcs-to-ics-live-or-is-a-node-restart-required/64 ## Headings Structure: H1: Can the compression strategy be changed from STCS to ICS live, or is a node restart required? H3: Related topics ## Main Content: H1: Can the compression strategy be changed from STCS to ICS live, or is a node restart required? H3: Related topics Yes, you can change the compaction strategy live - without restarting any nodes. Note that there might be increased compaction activity for a while after you make the change, so schedule it to off-peak hours if possible. See How to Change Compaction Strategy for details. --- ### Page: https://forum.scylladb.com/t/scylladb-5-1-release-candidate/65 Title: ScyllaDB 5.1 release candidate - Announcements - ScyllaDB Community NoSQL Forum Meta Description: The Scylla team is pleased to announce ScyllaDB Open Source 5.1 RC4, a Release Candidate for the Scylla Open Source 5.1 minor release. We encourage you to run ScyllaDB 5.1 release candidates on your test environments; t… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-5-1-release-candidate/65 ## Headings Structure: H1: ScyllaDB 5.1 release candidate H3: Related topics ## Main Content: H1: ScyllaDB 5.1 release candidate H3: Related topics The Scylla team is pleased to announce ScyllaDB Open Source 5.1 RC4, a Release Candidate for the Scylla Open Source 5.1 minor release. We encourage you to run ScyllaDB 5.1 release candidates on your test environments; this will help ensure that an upgrade to ScyllaDB 5.1 General Availability will proceed smoothly with your workload. Use the release candidate with caution; RC4 is not production-ready yet. You can help stabilize Scylla Open Source 5.1 by reporting bugs here. Only the last two minor releases of the ScyllaDB Open Source project are supported. Once Scylla Open Source 5.1 is officially released, ScyllaDB Open Source 5.1 and 5.0 will be supported, and ScyllaDB 4.6 will be retired. For a complete description of ScyllaDB 5.1 see ScyllaDB 5.1. ScyllaDB 5.1 RC1 , RC2, RC3 Get ScyllaDB Open Source 5.1 (under “More Versions” for each distro) Upgrade from ScyllaDB 5.0 to 5.1 Updates and bug fixes since 5.1 RC3 Stability: Memtable flush aborts the node if it fails due to out-of storage space #11245 The Scylla team is pleased to announce ScyllaDB Open Source 5.1 RC4, a Release Candidate for the Scylla Open Source 5.1 minor release. We encourage you to run ScyllaDB 5.1 release candidates on your test environments; this will help ensure that an upgrade to ScyllaDB 5.1 General Availability will proceed smoothly with your workload. Use the release candidate with caution; RC4 is not production-ready yet. You can help stabilize Scylla Open Source 5.1 by reporting bugs here. Only the last two minor releases of the ScyllaDB Open Source project are supported. Once Scylla Open Source 5.1 is officially released, ScyllaDB Open Source 5.1 and 5.0 will be supported, and ScyllaDB 4.6 will be retired. For a complete description of ScyllaDB 5.1 see ScyllaDB 5.1. ScyllaDB 5.1 RC1 , RC2, RC3 Get ScyllaDB Open Source 5.1 (under “More Versions” for each distro) Upgrade from ScyllaDB 5.0 to 5.1 Updates and bug fixes since 5.1 RC3 The Scylla team is pleased to announce ScyllaDB Open Source 5.1 RC5, a Release Candidate for the Scylla Open Source 5.1 minor release. We encourage you to run ScyllaDB 5.1 release candidates on your test environments; this will help ensure that an upgrade to ScyllaDB 5.1 General Availability will proceed smoothly with your workload. Use the release candidate with caution; RC5 is not production-ready yet. You can help stabilize Scylla Open Source 5.1 by reporting bugs here. Only the last two minor releases of the ScyllaDB Open Source project are supported. Once Scylla Open Source 5.1 is officially released, ScyllaDB Open Source 5.1 and 5.0 will be supported, and ScyllaDB 4.6 will be retired. For a complete description of ScyllaDB 5.1 see ScyllaDB 5.1. ScyllaDB 5.1 RC1 , RC2, RC3, RC4 Get ScyllaDB Open Source 5.1 (under “More Versions” for each distro) Upgrade from ScyllaDB 5.0 to 5.1 Updates and bug fixes since 5.1 RC4 --- ### Page: https://forum.scylladb.com/t/sharechat-s-path-to-high-performance-nosql-q-a-with-geetish-nayak/67 Title: ShareChat’s Path to High-Performance NoSQL: Q&A with Geetish Nayak - Blog Posts - ScyllaDB Community NoSQL Forum Meta Description: [800x400-blog-sharechat-webinar-ondemand] Language: en Canonical URL: https://forum.scylladb.com/t/sharechat-s-path-to-high-performance-nosql-q-a-with-geetish-nayak/67 ## Headings Structure: H1: ShareChat’s Path to High-Performance NoSQL: Q&A with Geetish Nayak H3: ShareChat’s Path to High-Performance NoSQL: Q&A with Geetish Nayak H3: Related topics ## Main Content: H1: ShareChat’s Path to High-Performance NoSQL: Q&A with Geetish Nayak H3: ShareChat’s Path to High-Performance NoSQL: Q&A with Geetish Nayak H3: Related topics How India's social media unicorn achieves microsecond P99 latency with 1.2M op/sec – for 180M monthly active users expecting real-time engagement with 2.5B posts per month. --- ### Page: https://forum.scylladb.com/t/scylladb-support/69 Title: ScyllaDB Support - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: How does ScyllaDB Support work? Is it 24/7? Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-support/69 ## Headings Structure: H1: ScyllaDB Support H3: Related topics ## Main Content: H1: ScyllaDB Support H3: Related topics How does ScyllaDB Support work? Is it 24/7? ScyllaDB protects every customer with professional, top-notch technical support. 24/7 and around the globe our professional team is ready to deliver the knowledge and expertise needed to resolve any issue, answer any questions, and meet our customer satisfaction commitment. OSS users will get community support, official documentation and free access to ScyllaDB university. Scylla Enterprise and Scylla Cloud customers subscriptions include a ScyllaDB Enterprise license, tested and certified binaries, software updates, hot fixes, and technical NoSQL support. More information can be found here --- ### Page: https://forum.scylladb.com/t/migrating-to-scylladb-from-cassandra-mongodb-dynamodb/70 Title: Migrating to ScyllaDB from Cassandra MongoDB DynamoDB - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: What does the typical migration process from (Cassandra/DynamoDB/Mongo) look like? How long does it take? Language: en Canonical URL: https://forum.scylladb.com/t/migrating-to-scylladb-from-cassandra-mongodb-dynamodb/70 ## Headings Structure: H1: Migrating to ScyllaDB from Cassandra MongoDB DynamoDB H3: Related topics ## Main Content: H1: Migrating to ScyllaDB from Cassandra MongoDB DynamoDB H3: Related topics What does the typical migration process from (Cassandra/DynamoDB/Mongo) look like? How long does it take? This is a very wide question (you can learn more in a Scylla University dedicated lesson). You need to consider whether you’ll be doing a Hot/Live or a Cold/Offline migration. There’s the live traffic aspect and the historical data aspect. If you’ll be doing a HOT migration, here is the sequence you should follow: Now let’s talk about historical data migration. The following tools can be used to migrate historical data from Cassandra: The following tools can be used to migrate historical data from DynamoDB to ScyllaDB (CQL API / DynamoDB compatible API, a.k.a Alternator): In order to migrate from MongoDB you’ll 1st need to design your schema / data model. A nice tool that can help you with that is Hackolade As for data migration tools: Thanks @Tomersan! Very helpful. I’ll add that if you’re doing a migration, get in touch with us, we’re happy to help. Also, in the upcoming LIVE event, we’ll have a dedicated session on migrating from DynamoDB. Hate to necro this thread but it’s been about a year - has any of the information above changed substantially? We’re considering moving from our Cass 3.11 cluster (12 nodes, 2 DCs) to Scylla 5.1. Any recommendations or tips are welcome. No substantial change that I’m aware of. Later today, we’ll host another LIVE training event, and there will be a specific session on migrating from DynamoDB to ScyllaDB, including a hands-on example. Hope to see you there. @Guy Can I see recordings of such events? If not, think about it, it might be useful for newcomers. These events are available live only. We’re working on updating the Migration lesson on ScyllaDB University. Also, stay tuned for the next training event (we have one every few weeks) as well as for the upcoming ScyllaDB Summit. --- ### Page: https://forum.scylladb.com/t/scylladb-and-large-partitions/71 Title: ScyllaDB and Large Partitions - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Are large partitions still an issue and if so, how can I deal with it? What is the maximal partition size? Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-and-large-partitions/71 ## Headings Structure: H1: ScyllaDB and Large Partitions H3: Related topics ## Main Content: H1: ScyllaDB and Large Partitions H3: Related topics Are large partitions still an issue and if so, how can I deal with it? What is the maximal partition size? Large partitions are an anti-pattern in Scylla - one should aim to have data spread evenly among the partitions in the cluster. The presence of large partitions create issues such as higher shard latency (because it has a hot spot for data) or oversized allocation warnings/errors on logs. In any case, you can use Scylla’s system.large_partitions virtual table (doc) to help you track them, along with log lines that indicate their occurrence. Check this Scylla University lesson covering this topic. As to how avoid them on the first place, you should wisely plan your data modeling in a way that leads to a high cardinality distribution of data among partitions. Check this blog post for some ideas around that. --- ### Page: https://forum.scylladb.com/t/is-scylladb-right-for-my-application-scylladbs-sweet-spot/72 Title: Is ScyllaDB Right for My Application? ScyllaDB's Sweet Spot - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This is a question that I often get: I’m currently evaluating different databases for my application. What is the sweet spot for ScyllaDB? Language: en Canonical URL: https://forum.scylladb.com/t/is-scylladb-right-for-my-application-scylladbs-sweet-spot/72 ## Headings Structure: H1: Is ScyllaDB Right for My Application? ScyllaDB's Sweet Spot H3: Related topics ## Main Content: H1: Is ScyllaDB Right for My Application? ScyllaDB's Sweet Spot H3: Related topics This is a question that I often get: I’m currently evaluating different databases for my application. What is the sweet spot for ScyllaDB? I’d say it’s a need for HIGH volume throughput with HIGH cardinality and LOW (15ms or even single digit) tail (p95/99) latency. Many times our I/O scheduler will do most of the prioritization for you. Application not optimized enough? reach out to our Solution architects for advise on better data modeling. Want more optimization tips? Make sure to check Scylla Monitoring: Have cardinality problems? check this out. I agree with what Tomer wrote about performance (High throughput, low latency). I’d also add to the sweet spot High Availability and Big Data: There are other reasons that teams choose ScyllaDB: I’d recommend a relational database and not ScyllaDB if: I’m going to add a few more elements that make for a sweet spot for ScyllaDB: Transactional (OLTP) vs. Analytical (OLAP) — while ScyllaDB does have workload prioritization for balancing various workloads on the same cluster, and certain kinds of analytics can be run on ScyllaDB, we’re far more focused on transactional/operational workloads. ScyllaDB is a row-oriented data store, vs. a columnar database that is designed for analytics. Single-digit Millisecond P99 Latencies — ScyllaDB is optimized for locally-attached NVMe SSD. While we have a built-in row-based in-memory cache, we’re not a pure-play in-memory database or data grid. So you will find that ScyllaDB is faster than a database connected to block storage, but not as expensive as a RAM-based system. It works in a “goldilocks” zone optimizing performance and price. High Throughput — This is a subjective term, but generally ScyllaDB is the database to look at when you scale to tens of thousands, hundreds of thousands and millions of operations per second (OPS). ScyllaDB uses immutable LSM tree based storage, so it is optimized for fast writes. And because it has a built-in row-based cache we are also good for fast reads. Multi-datacenter Replication — ScyllaDB automagically does replication across multiple sites, so if you are looking to deploy your data in the region of your users we will take care of that distribution for you. You don’t need to do all four of these. Any one of these make ScyllaDB a good fit. But the more of these criteria apply to your use case, the more you should be adding ScyllaDB to a technology short list for consideration. This page provides some additional guidance: Fit - ScyllaDB --- ### Page: https://forum.scylladb.com/t/scylladb-on-aws-one-large-node-or-more-smaller-ones/73 Title: ScyllaDB on AWS, One Large Node or More Smaller Ones? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I’m looking into running a ScyllaDB cluster on AWS. How do I determine sizing? Is it better to run a few large instances or many smaller ones? Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-on-aws-one-large-node-or-more-smaller-ones/73 ## Headings Structure: H1: ScyllaDB on AWS, One Large Node or More Smaller Ones? H3: On-Demand Webinar: Does it still make sense to do Big Data with Small Nodes? H3: Related topics ## Main Content: H1: ScyllaDB on AWS, One Large Node or More Smaller Ones? H3: On-Demand Webinar: Does it still make sense to do Big Data with Small Nodes? H3: Related topics I’m looking into running a ScyllaDB cluster on AWS. How do I determine sizing? Is it better to run a few large instances or many smaller ones? Larger nodes are easier to manage and their resources aggregate much more performance and speed. Think of applying a patch or performing a rolling restart in a hundred small nodes, also the number of failures would be higher so you’d expect more often maintenance. A disadvantage would be: since for AWS the storage size increases proportionately the larger the instance type, more powerful instances would become an expensive option in cases when the user doesn’t have a huge dataset. But still it may be recommended depending on the throughput and latency requirements. A benchmark would help to find the sweet spot, as well as other details that can be found in a nicely explicative Webinar about this exact subject: In the world of Big Data, scaling out is the norm. However, many Big Data deployments are trapped in a sea of small box clusters. Join us to learn the pros and cons of large nodes, and explore why people resist using big machines. I recommend watching it! --- ### Page: https://forum.scylladb.com/t/scylladb-cloud-or-enterprise/74 Title: ScyllaDB Cloud or Enterprise? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: What are the benefits of using Scylla Cloud vs. Scylla Enterprise? Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-cloud-or-enterprise/74 ## Headings Structure: H1: ScyllaDB Cloud or Enterprise? H3: Related topics ## Main Content: H1: ScyllaDB Cloud or Enterprise? H3: Related topics What are the benefits of using Scylla Cloud vs. Scylla Enterprise? By choosing ScyllaDB Cloud you’ll get the latest Enterprise version of Scylla DB as a service. Our experts will take care of everything, leaving you to focus on your data and applications. With flexible pricing and sizing choices, ScyllaDB Cloud may be used to suit your needs right away and grow alongside you in the future. Our solutions team will continuously check and manage your database to make sure it is always fully operational and operating at peak efficiency. You can read more here --- ### Page: https://forum.scylladb.com/t/what-are-the-different-scylladb-flavors-when-should-i-use-each-one/75 Title: What are the different ScyllaDB flavors? When should I use each one? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: ScyllaDB comes in three different variants: Open Source: a free, open-source, community-supported version. This is licensed under the AGPL. It offers great performance, a low node count, reduced complexity, and a low … Language: en Canonical URL: https://forum.scylladb.com/t/what-are-the-different-scylladb-flavors-when-should-i-use-each-one/75 ## Headings Structure: H1: What are the different ScyllaDB flavors? When should I use each one? H3: Related topics ## Main Content: H1: What are the different ScyllaDB flavors? When should I use each one? H3: Related topics ScyllaDB comes in three different variants: Open Source: a free, open-source, community-supported version. This is licensed under the AGPL. It offers great performance, a low node count, reduced complexity, and a low total cost of ownership. If you’re technical, know your way around source code, and like to be the first one to work with the latest features, this might be the version for you. Enterprise: This variant is based on the ScyllaDB open-source project. It includes tested and certified production-ready binaries, software updates, and hotfixes. Additionally, it includes mission-critical technical support that guarantees that you have access to the engineers who developed ScyllaDB. ScyllaDB Enterprise has a commercial license. This version also includes the ScyllaDB Manager, which provides centralized cluster administration and recurrent task automation, automation of periodic repair, as well as other features which are exclusive to ScyllaDB Enterprise customers. Cloud: This is the fastest and most affordable NoSQL database as a service. You get fully managed ScyllaDB Enterprise clusters. Using ScyllaDB Cloud spares your team from database administrative tasks, giving you access to ScyllaDB clusters with automatic backup, repairs, performance optimization, security hardening, and 24/7 maintenance and support. This is offered for a fraction of the cost of other database as a service offerings, without any vendor lock-in. --- ### Page: https://forum.scylladb.com/t/difference-between-reshape-and-compaction/76 Title: Difference between reshape and compaction - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: What is the difference between reshape and compaction in scylla? Even after checking the related documentation, I don’t quite understand. Is it correct for reshape to do the following? reshape: To make compaction work… Language: en Canonical URL: https://forum.scylladb.com/t/difference-between-reshape-and-compaction/76 ## Headings Structure: H1: Difference between reshape and compaction H3: ScyllaDB Open Source 4.2 H3: Related topics ## Main Content: H1: Difference between reshape and compaction H3: ScyllaDB Open Source 4.2 H3: Related topics What is the difference between reshape and compaction in scylla? Even after checking the related documentation, I don’t quite understand. ScyllaDB Open Source 4.2 provides new features: a far more efficient binary search algorithm, and improvements to our DynamoDB compatible interface. Is it correct for reshape to do the following? reshape: To make compaction work well per shard, reshape sstables to fit critica before bootstrap. In the case of stcs, set sstable count for 4 to hit min_threshold , and when using LCS adjusts the sstable size to 160mb. If the node is restarted after draining it, the reshape operation occurs before node join, and compaction seems to operate after node join. Why does this work like that? Reshape is a Rewrite of a set of SSTables to satisfy a compaction strategy’s criteria. For example, restoring data from an old backup or before the strategy update. Reshape happens when the data is wildly out-of-shape and regular compaction would require a lot of work to get it back into shape. It is an offline operation which allows the system to devote all resources to the work. Regular compaction happens when the data gets somewhat out of shape (e.g. a new sstable is added when a memtable is flushed). It’s an online operation that requires mild resource use. --- ### Page: https://forum.scylladb.com/t/gui-visualization-for-scylladb/77 Title: GUI Visualization for ScyllaDB - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello, Team! I am a Scylladb newbie. Does ScyllaDB have a GUI? Language: en Canonical URL: https://forum.scylladb.com/t/gui-visualization-for-scylladb/77 ## Headings Structure: H1: GUI Visualization for ScyllaDB H3: Related topics ## Main Content: H1: GUI Visualization for ScyllaDB H3: Related topics Hello, Team! I am a Scylladb newbie. Does ScyllaDB have a GUI? You can manage ScyllaDB with the ScyllaDB Manager. Also have a look at the Monitoring Stack. If you are talking about the tools such as pgAdmin (for PostgreSQL) or SQL Server Management Studio, then this may help: I am using JetBrains Rider for my .NET C# code development on my Linux desktop, and I was surprised to find out that that IDE’s support for databases is almost on par with dedicated and paid-for tools, while it’s only a plugin to Rider. I am working with DDL and DML code for my ScyllaDB using JetBrains Rider. It’s a paid-for tool, but I already have it anyway, so I might as well use it for the DB work. I also tried their database tool (called “DataGrip”), but I did not find any functionality in it which would be useful for my needs and that is not already in the Rider plugin. DBeaver now supports ScyllaDB in the Enterprise version. More than just a GUI, it’s a universal database manager for SQL and NoSQL databases. You can read more about it in this blog post. It provides step-by-step instructions on how to connect DBeaver to ScyllaDB. Once connected, users can run CQL queries and view result tables directly within DBeaver. If you try this out, please share your experience here. --- ### Page: https://forum.scylladb.com/t/running-repair-after-changing-the-replication-factor/78 Title: Running Repair after changing the Replication Factor - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: After I change the replication settings of a keyspace, should I run a repair on all nodes? Or will Scylla rescale the whole structure on its own? Language: en Canonical URL: https://forum.scylladb.com/t/running-repair-after-changing-the-replication-factor/78 ## Headings Structure: H1: Running Repair after changing the Replication Factor H3: Related topics ## Main Content: H1: Running Repair after changing the Replication Factor H3: Related topics After I change the replication settings of a keyspace, should I run a repair on all nodes? Or will Scylla rescale the whole structure on its own? ScyllaDB doesn’t automatically stream data to new replicas after changing the RF. You have to run a repair for that to happen. --- ### Page: https://forum.scylladb.com/t/kubernetes-memory-limit-for-scylladb-pod/79 Title: Kubernetes memory limit for ScyllaDB pod - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I set up Kubernetes for a limit of 64GB RAM for a Scylla pod. What if Scylla hits 64GB? Will my system crash? Or will it auto-free memory? Also, if I increase the memory limit to 96GB, will this improve performance? Language: en Canonical URL: https://forum.scylladb.com/t/kubernetes-memory-limit-for-scylladb-pod/79 ## Headings Structure: H1: Kubernetes memory limit for ScyllaDB pod H3: Related topics ## Main Content: H1: Kubernetes memory limit for ScyllaDB pod H3: Related topics I set up Kubernetes for a limit of 64GB RAM for a Scylla pod. What if Scylla hits 64GB? Will my system crash? Or will it auto-free memory? Also, if I increase the memory limit to 96GB, will this improve performance? ScyllaDB grabs all available memory and manages it internally. It is completely normal to have ScyllaDB at the memory limit. This is by design. Regarding increasing the memory limit, of course, the more memory, the better. The excess memory – that is not strictly needed for normal operations – is used to cache data from the disk, greatly increasing access speed for cached rows. More memory also allows for supporting more storage (ScyllaDB aims for a 1:100 memory:disk ratio). --- ### Page: https://forum.scylladb.com/t/running-a-large-number-of-keyspaces-in-a-cluster/80 Title: Running a large number of keyspaces in a cluster - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: What’s the operational experience like with large numbers of keyspaces in a cluster? Are there any functional or operational limits to be aware of? Language: en Canonical URL: https://forum.scylladb.com/t/running-a-large-number-of-keyspaces-in-a-cluster/80 ## Headings Structure: H1: Running a large number of keyspaces in a cluster H3: Related topics ## Main Content: H1: Running a large number of keyspaces in a cluster H3: Related topics What’s the operational experience like with large numbers of keyspaces in a cluster? Are there any functional or operational limits to be aware of? It’s recommended to limit the number of keyspaces/tables in a single cluster to a few hundreds. Using more can be tricky in some cases as Scylla needs many files, and Memtable flush is per table. Additionally, the amount of RAM available and how homogenous those keyspaces are homogenous also have an impact. --- ### Page: https://forum.scylladb.com/t/calculating-per-record-memory-in-a-cluster/81 Title: Calculating "per record" memory in a cluster - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I have been tasked to come up with “some number” of how much memory it takes to store one record in Scylla (and related overhead if any). We had commands in Redis Cache that produced these metrics. Is that even possible … Language: en Canonical URL: https://forum.scylladb.com/t/calculating-per-record-memory-in-a-cluster/81 ## Headings Structure: H1: Calculating "per record" memory in a cluster H3: Related topics ## Main Content: H1: Calculating "per record" memory in a cluster H3: Related topics I have been tasked to come up with “some number” of how much memory it takes to store one record in Scylla (and related overhead if any). We had commands in Redis Cache that produced these metrics. Is that even possible to do in Scylla? I have scyllatop running and also the monitoring stack. One way to accomplish this task is to create a million “average” records, execute “nodetool flush” to make sure they get flushed to disk (which will also place them in cache), then look at the cache metrics and divide the cache memory usage by the total rows in cache. Thanks, a couple of follow-up questions. When you say “cache metrics,” do you imply ones provided by scyllatop ? I presume this is where I get cache memory usage from? Also, would that work in a clustered environment (I have 8 nodes spread over four DCs). I know cache is over cluster, but what happens if not all records are in the cache? How would I find out how many have been put into cache? The metrics are available by scyllatop, though a nicer way to see them is with Prometheus or Grafana (see scylla-monitoring.git). The metrics contain the number of partitions in cache, rows in cache, and memory in cache, so from there you compute any statistic you want. About the rows not in cache, well they don’t have any cache footprint. --- ### Page: https://forum.scylladb.com/t/which-cassandra-version-is-scylladb-it-compatible-with/84 Title: Which Cassandra version is ScyllaDB it compatible with? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: ScyllaDB is a drop-in replacement for Apache Cassandra 3.11, but it has some additional features from Apache Cassandra 4.0. See ScyllaDB and Apache Cassandra Compatibility for details. Language: en Canonical URL: https://forum.scylladb.com/t/which-cassandra-version-is-scylladb-it-compatible-with/84 ## Headings Structure: H1: Which Cassandra version is ScyllaDB it compatible with? H3: Related topics ## Main Content: H1: Which Cassandra version is ScyllaDB it compatible with? H3: Related topics ScyllaDB is a drop-in replacement for Apache Cassandra 3.11, but it has some additional features from Apache Cassandra 4.0. See ScyllaDB and Apache Cassandra Compatibility for details. --- ### Page: https://forum.scylladb.com/t/data-model-with-a-lot-of-empty-columns-collections/88 Title: Data model with a lot of empty columns, collections - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello, I’m trying to find a good practice for column vs. Map column. For example, if I have 200 columns in a table and I usually use only 30% of them, am I better off using a Map columns (collection) instead of having em… Language: en Canonical URL: https://forum.scylladb.com/t/data-model-with-a-lot-of-empty-columns-collections/88 ## Headings Structure: H1: Data model with a lot of empty columns, collections H3: Related topics ## Main Content: H1: Data model with a lot of empty columns, collections H3: Related topics Hello, I’m trying to find a good practice for column vs. Map column. For example, if I have 200 columns in a table and I usually use only 30% of them, am I better off using a Map columns (collection) instead of having empty columns? I’ve read here that storage is not affected but memory is. Any advice on this? Yes, storage is not affected by empty columns. They are simply not stored if empty. It is similar in memory. It used to be that we used a different container for columns storage based on the number of columns in the schema: we used a vector (very efficient lookups but empty columns also take memory) by default and switched to a set for larger column counts (less efficient lookups but empty columns don’t use memory). We now uniformly switched to a compact radix tree, which should also not use any memory for empty columns. So overall, I think you are better off with columns. Although if you are not yet using clustering keys, you might consider refactoring your schema, so some of these maybe-empty columns are separate rows. Thank you for your answer! We are effectively organizing our schema to migrate our data from PostgreSQL. Can you elaborate a little more about the idea of an empty column as a separate row? Like giving me an example and why it will be better that way. I meant that if there is a pattern of some columns being empty in certain partitions, you can maybe organize your schema such that these are separate clustering rows instead, part of the same partition. Can’t really write an example without knowing more about the schema. But having partitions with lots of columns is also completely fine. thanks for the awesome information. --- ### Page: https://forum.scylladb.com/t/running-nodetool-upgradesstables-after-version-upgrade/89 Title: Running nodetool upgradesstables after version upgrade - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: After an upgrade from 4.6 to 5.0, do we need to run the nodetool upgradesstables ? Language: en Canonical URL: https://forum.scylladb.com/t/running-nodetool-upgradesstables-after-version-upgrade/89 ## Headings Structure: H1: Running nodetool upgradesstables after version upgrade H3: Related topics ## Main Content: H1: Running nodetool upgradesstables after version upgrade H4: apache/cassandra/blob/trunk/src/java/org/apache/cassandra/io/sstable/format/big/BigFormat.java#L369 H3: Related topics After an upgrade from 4.6 to 5.0, do we need to run the nodetool upgradesstables ? The short answer is no. The long answer is: even if a new ScyllaDB version introduces a new sstable version, you don’t have to manually upgrade sstables via nodetool upgradesstables ,this will happen automatically as the sstables are compacted, and in general new sstables are written in the new version. Gradually all sstables are migrated to the new version. This might take some time, but this is fine. There is no rush to get all sstables to the same version. In some cases one might be in a hurry to get some of the optimizations a new sstables format version has. This is a quick summary of version content: For example me adds host id information to sstable, I’m not sure it’s for preformece optimizations --- ### Page: https://forum.scylladb.com/t/point-in-time-recovery-in-scylladb/90 Title: Point in Time Recovery in ScyllaDB - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: User Question: I’m a n00b in ScyllaDB and reading the docs I understand that there is no such thing as Point-In-Time-Recovery. So you can only restore whatever your last backup has. Am I right or did I misunderstand thi… Language: en Canonical URL: https://forum.scylladb.com/t/point-in-time-recovery-in-scylladb/90 ## Headings Structure: H1: Point in Time Recovery in ScyllaDB H3: Related topics ## Main Content: H1: Point in Time Recovery in ScyllaDB H3: Related topics User Question: I’m a n00b in ScyllaDB and reading the docs I understand that there is no such thing as Point-In-Time-Recovery. So you can only restore whatever your last backup has. Am I right or did I misunderstand things ? Answer: Correct. It is not something to worry about, however, as even if you lose a full AZ you can still serve requests and have your data fully considering you issue queries with LOCAL_QUORUM. And - of course - you can always go multi-DC for a true DR scenario. User Question: Indeed, but I am thinking about the scenario where the application has a new release and there is some bug or some operator executes the wrong DDL and blows things up. If that happened, I guess the only way to restore things up to the point of failure would be by replaying the transactions (Kafka) or by doing a PITR with an RDBMS and then dump all the data again onto ScyllaDB. Am I wrong ? Answer: As you are correct in the first question, then of course you are also correct in the last one. You can always snapshot before such changes. And for TRUNCATE/DROP table statements, Scylla will always snapshot the table in question (unless you told it not to) when these DDL statements are run --- ### Page: https://forum.scylladb.com/t/lost-connection-to-remote-peer-exception-when-querying-data-synchronously/91 Title: “Lost connection to remote peer” exception when querying data synchronously - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: User Question: I met “Lost connection to remote peer” exception when query data from scylla synchronously which will break the main client process. At the time of exception happened, seems one node became down but other… Language: en Canonical URL: https://forum.scylladb.com/t/lost-connection-to-remote-peer-exception-when-querying-data-synchronously/91 ## Headings Structure: H1: “Lost connection to remote peer” exception when querying data synchronously H3: Related topics ## Main Content: H1: “Lost connection to remote peer” exception when querying data synchronously H3: Related topics User Question: I met “Lost connection to remote peer” exception when query data from scylla synchronously which will break the main client process. At the time of exception happened, seems one node became down but other nodes were working well. Do I need to switch to asynchronous access? CQL deriver version is 4.14.1 Answer: Well, it seems inflight queries threw an exception which you didn’t handle. As a result, your main thread died. Handle it and retry the failed queries, and it should work picking up other coordinator nodes. --- ### Page: https://forum.scylladb.com/t/counter-columns-alternative-and-overcoming-limitations/94 Title: Counter columns alternative and overcoming limitations - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Question: According to the official documentation, among the restrictions on counter columns in ScyllaDB are: The only other columns in a table with a counter column can be columns of the primary key (which cannot be u… Language: en Canonical URL: https://forum.scylladb.com/t/counter-columns-alternative-and-overcoming-limitations/94 ## Headings Structure: H1: Counter columns alternative and overcoming limitations H3: Related topics ## Main Content: H1: Counter columns alternative and overcoming limitations H3: Related topics According to the official documentation, among the restrictions on counter columns in ScyllaDB are: In my table design I have hundreds of rows where each row needs to be linked to a counter column. The table has the following structure: Where K is the partition key, C is the cluster key, V is a value and COUNTER is a counter column. Generally, the above table is queried as follows: This results in about 500 rows being returned per query. As seen, each row is linked to a COUNTER column. Specifically, each combination of K1, K2, C1, C2, C3 is linked to a different COUNTER value. How am I suppose to model this table, if the COUNTER has to be moved to an entirely different table? If I understand correctly, the only way to do this, is to define another table → table_counter without any of the values (V): However, I have several issues with this approach: 1) It seems extremely inelegant to break up a logically cohesive table like this 2) It means that whenever I want to execute the previous query I would essentially need to execute two queries, instead of one, just to get the counter information linked to each K1 K2 C1 C2 C3 combination 3) I would also be forced to combine the results of the two queries above into a single data structure on the client side (for it to be useful) Is this correct? If yes, is there an alternative to COUNTER column where I could add the COUNTER column to the first table? One approach I was thinking of is to use a regular INTEGER as the counter column. Whenever the counter column needs to be updated I can read the current counter (integer) value and increment it on the client side and then write the new value back to the database. I understand that I won’t be protected from concurrent reads/writes so that if two clients read the counter value at the same time, and they both increment/update it, one write (update) will be lost (e.g, only the last write will be preserved). However, I can live with the occasional lost write as we are not tracking anything critical where a single (or handful) of lost writes will make a major difference. I also understand that this would require a read and then a write (two operations) each time I want to update the counter column, however, it allows me to keep the counter column as part of the original table as well as reduce the querying from two tables to one. Additionally, there would be no need to combine the results of querying two tables on the client side with this design. Seems more efficient than using a counter column and much more elegant. Is this approach a viable alternative to the COUNTER column? Are there any pitfalls I missed? Is there another approach that might work better in my example? *The question was asked on Stack Overflow by S.O.S Answer: An Alternative to the “classic” DRDT-inspired (but not quite) counter column is to use Scylla’s lightweight transactions (LWT) - basically your counter becomes a normal integer column, which you can read normally if you wish, but writes use a conditional update (UPDATE … IF …). For example to modify value V and also increment the counter you can: This pattern is known as “optimistic locking” - step 1-2 are “optimistic” in assuming that the new item they build will be able to be written, but if some other concurrent update beat us, step 1-2 will need to be repeated. This is to contrast with pessimistic locking approaches, where the client “takes a lock”, and only after knowing it is holding the lock, bothers to calculate the new value of the item (step 2). LWT is much more powerful than counters, and you can do with it a lot more than you can do with counters. Reads can be as efficient as regular reads, but note that writes do become slower. Scylla is working on a next-generation LWT implementation based on Raft (the current implementation is based on Paxos), so you can expect improvements in LWT write performance in the future. *The answer was provided on Stack Overflow by Nadav Har’El Thanks for your excellent response. Just to be clear, using LWT still requires the client to submit two transactions - one to read the current value and update the value on the client’s side and a second transaction to write the new value to db. In other words, it’s not possible to combine reading and updating into a single transaction with LWT. In other words, LWT only ensures that a concurrent write is not lost but it still requires two transactions from the client. In a case where I don’t mind losing the occasional write a regular integer column without LWT would suffice. Do you concur? Also, when would you recommend using LWT instead of custom COLUMN counter? Only when we expect few writes on the column or always? Thanks! CQL does not currently have the syntax to increment a non-counter column - doing UPDATE ... X = X + 1 will be an error if X is not a counter column. Both with LWT and without LWT. So you need to do a separate read - even though the underlying LWT implementation could have handle an atomic increment just fine without a read. Remember, though, that if the counter is not stand-alone and is used for optimistic locking (as I explained above) you would need that extra read anyway. And yes, if you don’t care about write isolation and missed increments, you can just use an unsafe read+write. --- ### Page: https://forum.scylladb.com/t/check-if-an-item-is-inside-a-list-or-collection/95 Title: Check if an item is inside a list or collection - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Question: I’d like to check if an item is inside a list. How can I do that? Language: en Canonical URL: https://forum.scylladb.com/t/check-if-an-item-is-inside-a-list-or-collection/95 ## Headings Structure: H1: Check if an item is inside a list or collection H3: Related topics ## Main Content: H1: Check if an item is inside a list or collection H3: Related topics Question: I’d like to check if an item is inside a list. How can I do that? Answer: To do that, you use the CONTAINS operator. This operator can be used on lists, sets, and maps. In the case of maps, CONTAINS applies to the map values. Also, see the Documentation. --- ### Page: https://forum.scylladb.com/t/is-there-an-api-for-scylla-nodetool/96 Title: Is there an API for Scylla Nodetool? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Question: Is there an API for nodetool? Especially nodetool tablestats? I saw there is scylladb/api/api-doc at master · scylladb/scylladb · GitHub. Is this the right place to look for APIs? *The question was asked on … Language: en Canonical URL: https://forum.scylladb.com/t/is-there-an-api-for-scylla-nodetool/96 ## Headings Structure: H1: Is there an API for Scylla Nodetool? H3: Related topics ## Main Content: H1: Is there an API for Scylla Nodetool? H4: scylladb/scylla-tools-java/blob/master/src/java/org/apache/cassandra/tools/nodetool/TableStats.java H4: scylladb/scylla-tools-java/blob/master/src/java/org/apache/cassandra/tools/nodetool/stats/TableStatsHolder.java#L117 H4: scylladb/scylla-jmx/blob/master/src/main/java/org/apache/cassandra/db/ColumnFamilyStore.java H3: Related topics Question: Is there an API for nodetool? Especially nodetool tablestats? I saw there is scylladb/api/api-doc at master · scylladb/scylladb · GitHub. Is this the right place to look for APIs? *The question was asked on Stack Overflow by SilentCanon Answer: The Scylla server indeed has a REST API and it’s documentation is at the URL you point to. You can find a Swagger UI when you start the Scylla server as well: https://docs.scylladb.com/operating-scylla/rest/. But, please note that nodetool, for example, does not use the API directly. Instead, it talks to the Scylla JMX proxy, which is a Java process that implements Cassandra-compatible JMX API. You can still use the REST API directly, but you have to figure out the mapping between the JMX operations and the REST API yourself. For something like nodetool tablestats, the first step is to check what JMX APIs nodetool uses: The command delegates to TableStatsHolder class: which uses the ColumnFamilyStoreMBean JMX API for querying table statistics. You can find the implementation of the JMX API in the scylla-jmx project by finding the ColumnFamilyStore class (without the MBean suffix): From that class, you can see that, for example, the ColumnFamilyStore.getSSTableCountPerLevel() method delegates to the column_family/sstables/per_level/ REST API URL. *The answer was provided on Stack Overflow by Pekka Enberg Hey, thanks for the answer but i’d like to ask more. If you know, or could point me at the direction of how to divide the actual scylla and the scylla-jmx. Meaning we wanted to get rid of cassandra on hosts just because it uses java, and it’s a huge package almost 30mb in size, while our goal was to make a tiniest distro possible. Is it possible to, let’s say, start scylla on the host and jmx nodetool services in a container on the same host, so API ports are reachable from inside the container? This way we would have no need for java in the initial distro and therefore on hosts We have not done that, but we’ll be looking at this direction in the future - makes total sense to separate them. Thank you for the quick response. Is there an approximate date you may look in to it, maybe in your backlog? So i would know when to get back to the though of replacing cassandra again? It’s not urgent. Besides being large and perhaps the need to update more often, it’s not a hassle. Splitting it has consequences that we’ll have to deal with. For example, running multiple containers side-by-side. Have heard you. Thanks anyway! --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-156-2022-11-27/113 Title: Last week in scylladb.git master (issue #156; 2022-11-27) - Database Community - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 41629e97de…996eac9569 range are covered. There were 90 non-merge commits from 12 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-156-2022-11-27/113 ## Headings Structure: H1: Last week in scylladb.git master (issue #156; 2022-11-27) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #156; 2022-11-27) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 41629e97de…996eac9569 range are covered. There were 90 non-merge commits from 12 authors in that period. Some notable commits: A regression where CQL ignored some WHERE clause components when a multi-column restriction ((col_a, col_b) < (0, 0)) was present was fixed. The bundled cqlsh now uses the ScyllaDB Python driver (rather than the generic Cassandra driver) and supports Scylla Cloud connection bundles. When a query completes a page, ScyllaDB caches the query activity as an inactive read. When the client requests the next page, ScyllaDB re-activates the read and continues where it left off. A bug in this mechanism that could cause crashes has been fixed. There is now documentation for the (mostly automatic) procedure to upgrade a cluster to use Raft, and for the manual procedure to recover in case of problems. A bug in alternator WHERE condition for global secondary index range key, that does not appear to have any user visible impact, has been fixed. The task manager is now aware of repair-related tasks. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-4-6-10/115 Title: [RELEASE] ScyllaDB 4.6.10 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 4.6.10, a bugfix release of the ScyllaDB 4.6 stable branch. Please note the latest ScyllaDB stable release is ScyllaDB 5.0, and you are encouraged to upgrade to it. Scyl… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-4-6-10/115 ## Headings Structure: H1: [RELEASE] ScyllaDB 4.6.10 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 4.6.10 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 4.6.10, a bugfix release of the ScyllaDB 4.6 stable branch. Please note the latest ScyllaDB stable release is ScyllaDB 5.0, and you are encouraged to upgrade to it. ScyllaDB Open Source 4.6.10, like all past and future 4.x.y releases, is backward compatible and supports rolling upgrades. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/scylladb-university-live-nosql-roundtable-avi-kivity-dor-laor-tzach-livyatan/116 Title: ScyllaDB University Live NoSQL Roundtable: Avi Kivity, Dor Laor & Tzach Livyatan - Blog Posts - ScyllaDB Community NoSQL Forum Meta Description: [800x400-blog-scylladb-u-roundtable-fall-22] Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-university-live-nosql-roundtable-avi-kivity-dor-laor-tzach-livyatan/116 ## Headings Structure: H1: ScyllaDB University Live NoSQL Roundtable: Avi Kivity, Dor Laor & Tzach Livyatan H3: ScyllaDB University Live NoSQL Roundtable: Avi Kivity, Dor Laor & Tzach... H3: Related topics ## Main Content: H1: ScyllaDB University Live NoSQL Roundtable: Avi Kivity, Dor Laor & Tzach Livyatan H3: ScyllaDB University Live NoSQL Roundtable: Avi Kivity, Dor Laor & Tzach... H3: Related topics A look at the expert NoSQL roundtables that are a part of every ScyllaDB University LIVE event – including the fall session on December 1. --- ### Page: https://forum.scylladb.com/t/scylladb-return-inconsistent-data-after-node-full-rebuild/119 Title: Scylladb return inconsistent data after node full rebuild - Database Community - ScyllaDB Community NoSQL Forum Meta Description: Hi! I am testing node rebuild after loosing the data volume in docker (like described here Rebuild a Node After Losing the Data Volume | Scylla Docs). My steps Create new cluster with static ip’s(3 DC, 2 nodes in each… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-return-inconsistent-data-after-node-full-rebuild/119 ## Headings Structure: H1: Scylladb return inconsistent data after node full rebuild H3: Related topics ## Main Content: H1: Scylladb return inconsistent data after node full rebuild H3: Related topics Hi! I am testing node rebuild after loosing the data volume in docker (like described here Rebuild a Node After Losing the Data Volume | Scylla Docs). Image version: scylladb/scylla:4.6.8 Hello, scylla replace operations fetches data from nodes in the same DC. In case RF = 1, it is expected the replacing node will miss some of the data after replace. You can run a cross DC repair to fix the data. In addition, I would recommend using RF more than 1 per DC. --- ### Page: https://forum.scylladb.com/t/i-wonder-how-much-cpu-share-each-workload-prioritization-type-oltp-olap-of-scylladb-has/121 Title: I wonder how much cpu share each Workload Prioritization type (oltp, olap) of scylladb has - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: There are two workload types in scylladb other than the default value, but it doesn’t seem to be specified in the documentation how much cpu share this option has. interactive - workload sensitive to latency, expected… Language: en Canonical URL: https://forum.scylladb.com/t/i-wonder-how-much-cpu-share-each-workload-prioritization-type-oltp-olap-of-scylladb-has/121 ## Headings Structure: H1: I wonder how much cpu share each Workload Prioritization type (oltp, olap) of scylladb has H3: Related topics ## Main Content: H1: I wonder how much cpu share each Workload Prioritization type (oltp, olap) of scylladb has H4: scylladb/scylladb/blob/master/docs/dev/service_levels.md H4: scylladb/scylladb/blob/b551cd254c6daf07014267950fcdd17aaf19acf0/transport/server.cc#L590-L598 H3: Related topics There are two workload types in scylladb other than the default value, but it doesn’t seem to be specified in the documentation how much cpu share this option has. Even if I check the related PR, it is not specified how much cpu share the corresponding workload type takes. May I know the relevant code? Hi @Stewart_Han , It is not about cpu shares, the workload type is used by ScyllaDB to decide on the proper action during excesive load on the system. For batch workloads , where the concurrency is bounded, we can throttle the load by delaying answers to the client, this in turn will make him send less requests per second for each thread. For interactive workload, throteling doesn’t make much sense since the requests are not queued on the client side so the only resort is to fail early and signal to the client that we are overloaded (OverloadedException), this behaviour is a little bit speculative since we need to predict in advance our ability to serve a specific requests withing the timeout limits. Here is a pointer to the code where we drop requests for interactive workload if we predict that we will not be able to meet the timeout: If the workload is not interactive scylla will just continue normally, assuming that the workload will converge to the maximal throughput possible. thanks for your reply! --- ### Page: https://forum.scylladb.com/t/meet-scylladb-s-new-vp-of-r-d-yaniv-kaul/123 Title: Meet ScyllaDB’s New VP of R&D, Yaniv Kaul - Blog Posts - ScyllaDB Community NoSQL Forum Meta Description: [800x400-blog-intro-yaniv-kaul (1)] Language: en Canonical URL: https://forum.scylladb.com/t/meet-scylladb-s-new-vp-of-r-d-yaniv-kaul/123 ## Headings Structure: H1: Meet ScyllaDB’s New VP of R&D, Yaniv Kaul H3: Meet ScyllaDB’s New VP of R&D, Yaniv Kaul H3: Related topics ## Main Content: H1: Meet ScyllaDB’s New VP of R&D, Yaniv Kaul H3: Meet ScyllaDB’s New VP of R&D, Yaniv Kaul H3: Related topics Get to know Yaniv Kaul, who just joined ScyllaDB as VP of Research and Development, in this quick Q & A. You'll hear about his fascination with tackling complex challenges: from distributed system engineering to photographing bees. Does anyone here have questions for Yaniv? --- ### Page: https://forum.scylladb.com/t/release-scylla-5-1-0-part-1/126 Title: [RELEASE] Scylla 5.1.0 - part 1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of ScyllaDB Open Source 5.1, a production-ready release of our open-source NoSQL database. ScyllaDB 5.1 introduces Partition level rate limit, Distributed select coun… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-5-1-0-part-1/126 ## Headings Structure: H1: [RELEASE] Scylla 5.1.0 - part 1 H2: New Features H3: Distributed SELECT COUNT(*) H3: Limit partition access rate H3: Load and stream H3: Materialized Views: Prune H3: Alternator updates H3: Materialized Views: Synchronous Mode - Experimental H3: Performance: Eliminate exceptions from the read and write path H3: Raft Updates H3: Web Assembly (WASM) based UDA/UDF - Experimental H1: [RELEASE] Scylla 5.1.0 - part 2 H2: Updates in this Release H3: Deployment and Packaging H3: CQL API updates H3: Stability and Performance Improvements H3: Tooling H3: Storage H3: Configuration H3: Monitoring and Tracing H3: Bug Fixes H3: Related topics ## Main Content: H1: [RELEASE] Scylla 5.1.0 - part 1 H2: New Features H3: Distributed SELECT COUNT(*) H3: Limit partition access rate H3: Load and stream H3: Materialized Views: Prune H3: Alternator updates H3: Materialized Views: Synchronous Mode - Experimental H3: Performance: Eliminate exceptions from the read and write path H3: Raft Updates H3: Web Assembly (WASM) based UDA/UDF - Experimental H1: [RELEASE] Scylla 5.1.0 - part 2 H2: Updates in this Release H3: Deployment and Packaging H3: CQL API updates H3: Stability and Performance Improvements H3: Tooling H3: Storage H3: Configuration H3: Monitoring and Tracing H3: Bug Fixes H3: Related topics The ScyllaDB team is pleased to announce the release of ScyllaDB Open Source 5.1, a production-ready release of our open-source NoSQL database. ScyllaDB 5.1 introduces Partition level rate limit, Distributed select count, and more functional, performance and stability improvements. Only the last two minor releases of the ScyllaDB Open Source project are supported. As ScyllaDB Open Source 5.1 is officially released, only ScyllaDB Open Source 5.1 and ScyllaDB 5.0 will be supported; ScyllaDB 4.6 will be retired. Users are encouraged to upgrade to the latest and greatest release. Upgrade instructions are here. Get ScyllaDB Open Source 5.1 as binary packages (RPM/DEB), AWS AMI, GCP Image and Docker image Upgrade from ScyllaDB 5.0 to ScyllaDB 5.1 ScyllaDB will now automatically run SELECT COUNT(*) statements on all nodes and all shards in parallel, which brings a considerable speedup, even 100X in larger clusters. This feature is limited to queries that do not use GROUP BY or filtering. The implementation includes a new level of coordination. A Super-Coordinator node splits aggregation queries into sub-queries, distributes them across some group of coordinators, and merges results. Like a regular coordinator, the Super-Coordinator is a per operation function. A 3 node cluster setup on powerful desktops (3x32 vCPU) Filled the cluster with ~2 * 10^8 rows using scylla-bench and run: time cqlsh --request-timeout=3600 -e “select count(*) from scylla_bench.test using timeout 1h;” Before Distributed Select: 68s After Distributed Select: 2s You can disable this feature by setting enable_parallelized_aggregation config parameter to false. It is now possible to limit read rates and writes rates into a partition with a new WITH per_partition_rate_limit clause for the CREATE TABLE and ALTER TABLE statements. This is useful to prevent hot-partition problems when high rate reads or writes are bogus (for example, arriving from spam bots). #4703 Limits are configured separately for reads and writes. Some examples: ALTER TABLE t WITH per_partition_rate_limit = { ‘max_reads_per_second’: 100, ‘max_writes_per_second’: 200 Limit reads only, no limit for writes: ALTER TABLE t WITH per_partition_rate_limit = { ‘max_reads_per_second’: 200 This feature extends nodetool refresh to allow loading arbitrary sstables that do not belong to a particular node into the cluster. It loads the sstables from disk, calculates the data’s owning nodes, and automatically streams the data to the owning nodes. In particular this is useful when restoring a cluster from backup. For example, say the old cluster has 6 nodes and the new cluster has 3 nodes. One can copy the sstables from the old cluster to the new nodes and trigger the load and stream process. This can make restores and migrations much easier: Load_and_stream option also updates the relevant Materialized Views #9205 curl -X POST "http://{ip}:10000/storage_service/sstables/{keyspace}?cf={table}&load_and_stream=true Note there is an open bug, #282 , for the Nodetool refresh --load-and-stream operation. Until it is fixed, use the REST API above. A new CQL extension PRUNE MATERIALIZED VIEW statement can now be used to remove inconsistent rows from materialized views. A special statement is dedicated for pruning ghost rows from materialized views. A ghost row is an inconsistency issue which manifests itself by having rows in a materialized view which do not correspond to any base table rows. Such inconsistencies should be prevented altogether and ScyllaDB strives to avoid them, but if they happen, this statement can be used to restore a materialized view to a fully consistent state without rebuilding it from scratch. PRUNE MATERIALIZED VIEW my_view; PRUNE MATERIALIZED VIEW my_view WHERE token(v) > 7 AND token(v) < 1535250; PRUNE MATERIALIZED VIEW my_view WHERE v = 19; Alternator, ScyllaDB’s implementation of the DynamoDB API, now provides the following improvements: There is now a synchronous mode for materialized views. In ordinary, asynchronous mode materialized views operations return before the view is updated. In synchronous mode materialized views operations do not return until the view is updated. This enhances consistency but reduces availability as in some situations all nodes might be required to be functional. CREATE MATERIALIZED VIEW main.mv AS SELECT * FROM main.t WITH synchronous_updates = true; ALTER MATERIALIZED VIEW main.mv WITH synchronous_updates = true; When a coordinator times out, it generates an exception which is then caught in a higher layer and converted to a protocol message. Since exceptions are slow, this can make a node that experiences timeouts become even slower. To prevent that, the coordinator write path and read path has been converted not to use exceptions for timeout cases, treating them as another kind of result value instead. Further work on the read path and on the replica reduces the cost of timeouts, so that goodput is preserved while a node is overloaded. Improvement results below While strong consistent schema management remains experimental in 5.1, the work on the Raft consensus algorithm continues toward more use cases, such as safe topology updates, improving traceability and stability. Here are selected updates in this release: ScyllaDB 5.1 brings experimental support for Wasm-based User Defined Functions (UDFs) and User Defined Aggregates (UDAs). The CQL syntax is compatible with Apache Cassandra. Examples: CREATE FUNCTION sample ( arg int ) …; CREATE FUNCTION sample ( arg text ) …; A full example of using Rust to create a UDF will be shared soon. To enable WASM UDF in ScyllaDB 5.1: –enable-user-defined-functions true --experimental-features udf experimental_features: Issues fixed in this release: More update in Part 2 below! A list of CQL bug fix and extensions: The LIKE operator on descending order clustering keys now works. #10183 ScyllaDB would incorrectly use an index with some IN queries, leading to incorrect results. This is now fixed. In CREATE AGGREGATE statements, the INITCOND and FINALFUNC clauses are now optional (defaulting to NULL and the identity function respectively). CREATE KEYSPACE now has a WITH STORAGE clause, allowing to customize where data is stored. For now, this is only a placeholder for future extensions. When talking to drivers using the older v3 protocol, ScyllaDB did not serialize timeout exceptions correctly, resulting in the driver complaining about protocol violations. This is now fixed. #5610 The CQL grammar was relaxed to allow bind markers in collection literals, e.g. UPDATE tab SET my_set = { ?, ‘foobar’, :variable }. ScyllaDB now validates collections for NULLs more carefully. #10580 After this change, the following query INSERT INTO ks.t (list_column) VALUES (?); And the driver sending a list with null inside as the bound value, something like [1, 2, null, 4] Would result in an invalid_request_exception instead of an ugly marshaling error. The tool can be used to list the different API functions and their parameters, and to print detailed help for each function. Then, when invoking any function, scylla-api-cli performs basic validation on the function arguments and prints the result to the standard output. Note that json results msy be pretty-printed using commonly available command line utilities. It is recommended to use scylla-api-cli for interactive usage of the REST API over plain http tools, like curl, to prevent human errors. It is now possible to limit, and control in real time, the bandwidth of streaming and compaction. These and more configuration updates below: Below are a list of monitoring and tracing related work in this release: For a full list of fixed issues see git log and 5.1 release candidates notes. --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-0-5/128 Title: [RELEASE] ScyllaDB 5.0.5 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.0.5, a bugfix release of the ScyllaDB 5.0 stable branch. ScyllaDB Open Source 5.0.5, like all past and future 5.x.y releases, is backward compatible and supports rolling… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-0-5/128 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.0.5 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.0.5 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.0.5, a bugfix release of the ScyllaDB 5.0 stable branch. ScyllaDB Open Source 5.0.5, like all past and future 5.x.y releases, is backward compatible and supports rolling upgrades. Note that the latest stable branch of ScyllaDB is 5.1, and you are encouraged to upgrade. Issues fixed in this release: ScyllaDB 5.0.5 was acutally released at Oct, resent by mistake. --- ### Page: https://forum.scylladb.com/t/what-table-options-should-be-used-for-fast-writes-reads-and-no-deletes/129 Title: What table options should be used for fast writes & reads and no deletes - Database Community - ScyllaDB Community NoSQL Forum Meta Description: ( New user to Scylladb … excuse my naivety ) I am going to use ScyllaDB cluster ( single DC ) for key value pair , key will be the queue-id and value will be the job data . Key size 32 bytes , values varying avg 50kb -… Language: en Canonical URL: https://forum.scylladb.com/t/what-table-options-should-be-used-for-fast-writes-reads-and-no-deletes/129 ## Headings Structure: H1: What table options should be used for fast writes & reads and no deletes H3: Related topics ## Main Content: H1: What table options should be used for fast writes & reads and no deletes H3: Related topics ( New user to Scylladb … excuse my naivety ) I am going to use ScyllaDB cluster ( single DC ) for key value pair , key will be the queue-id and value will be the job data . Key size 32 bytes , values varying avg 50kb - max 10 Mb There will by 800 million writes every day at peaks of 4 k writes per second ( which might grow) No batch inserts all single . No updates . Reads will happen exactly once per record To avoid any tombstones I will use 1 table per day and drop entire tables after 2 days There are my doubts What table type should I used . I think Caching enabled=true ? What is the ideal concurrency I should design my writers and readers ? If there are no updates or deletes can I use this to speed up my read/writes by setting up some config ? You can use default table, however I suggest to use TimeWindow Compaction Strategy instead of dropping and creating the table with TTL of X days (make sure you properly then set the window unit in TWCS, don’t create within your TTL more than 12-13 windows please) then tombstones won’t be an issue, Scylla will effectively throw them away (thanks to TTL and windows) However a warning here is - you won’t overwrite old data here (outside of current window), if you will, then this effective removing of old data will get broken. (and in such case default ICS with TTL should also work with either SAG or periodic major compaction) I don’t know what will be your latency SLAs, but 50k-10M payloads are huuuge range, where the rows around 50k will be quite fast, but processing of 10M payload might bottleneck the cpu(shard). Limits we suggest to keep are here: scylladb/config.cc at master · scylladb/scylladb · GitHub , so for you if we assume single row partition, then it’s about cell size, which is 1 MB, you will have 10MB, so I’d expect not single digit ms latencies, but worst case 10x more Above largely depends on how many reads per second will you do and how many cpus will be there (and how good will be your PK distribution). Concurrency depends on distribution and how many cpus you will have and how many reads/writes (with ideally percentiles for that data, since that payload size range of yours is big). Some guidance is in Sizing Up Your ScyllaDB Cluster - ScyllaDB , or check sizing calculators (take them as guidance, not as a rule of thumb, cassandra-stress is your best friend here to see how much 1 cpu will be able to handle with your RF). You can also read Maximizing Performance via Concurrency While Minimizing Timeouts in Distributed Databases - ScyllaDB . If there are no updates and deletes, then this is perfect TTL + TWCS situation assuming you really want to drop all your data older than 2 days. And that is basically your best tuning - Compaction | ScyllaDB Docs (but do check other strategies, too) --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-0-6/133 Title: [RELEASE] ScyllaDB 5.0.6 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.0.6, a bugfix release of the ScyllaDB 5.0 stable branch. ScyllaDB Open Source 5.06, like all past and future 5.x.y releases, is backward compatible and supports rolling … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-0-6/133 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.0.6 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.0.6 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.0.6, a bugfix release of the ScyllaDB 5.0 stable branch. ScyllaDB Open Source 5.06, like all past and future 5.x.y releases, is backward compatible and supports rolling upgrades. Note that the latest stable branch of ScyllaDB is 5.1, and you are encouraged to upgrade. Issues fixed in this release: --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-157-2022-12-04/134 Title: Last week in scylladb.git master (issue #157; 2022-12-04) - Database Community - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 996eac9569…b551cd254c range are covered. There were 149 non-merge commits from 19 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-157-2022-12-04/134 ## Headings Structure: H1: Last week in scylladb.git master (issue #157; 2022-12-04) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #157; 2022-12-04) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 996eac9569…b551cd254c range are covered. There were 149 non-merge commits from 19 authors in that period. Some notable commits: A crash was fixed during an illegal lightweight transaction INSERT with NULL clustering key Alternator, ScyllaDB’s implementation of the DynamoDB API, now considers the time-to-live implementation stable and no longer experimental. Transient errors in Alternator TTL scanning (where alternator looks for expired items) no longer cause Alternator to abort the scan; instead it continues. ScyllaDB caches rows and (since 4.6) index entries in a single unified cache. It was observed that in some small-partition workloads index caching causes a performance regression, so index caching is now disabled by default. It can still be enabled for workloads that benefit from it. We plan to re-enable it when the regression is fixed. The bundled cqlsh now considers system_distributed_everywhere a system keyspace. Last week’s change to the scylla driver in cqlsh was reverted, as it causes regressions around the USE statement. Evaluation of Boolean binary operators (e.g. “=”) has been refactored to use the same expression evaluation code as other expressions, paving the way for relaxation of the CQL grammar to be more similar to SQL. WebAssembly (WASM) has been re-enabled for aarch64 (ARM). The topology management code is more relaxed about unknown endpoints to prevent crashes in tests that check for edge cases. This fixes a recent regression. Hinted handoff now checks that a node exists in topology before doing anything; this helps with a recent regression due to topology refactoring. In SELECT JSON statements, column names can be given aliases (just as with traditional SELECT). However, ScyllaDB ignored those aliases. It will now honor them. The container (docker) image now uses the C locale to reduce image size. The Raft protocol implementation now supports changing a node’s IP address without changing its identity. Raft failure detection now uses a separate RPC verb from gossip failure detection. The task manager now controls repair tasks with shard granularity. A crash in the task manager related to repair tasks has been fixed. A crash while fetching repaid ids from the repair history table was fixed. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/io-uring-what-is-it/136 Title: Io_uring: what is it? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I’ve recently been seeing io_uring come up in locations. Understanding it has something to do with input and output, what exactly is it? *The question was originally asked on Stack Overflow by Amani Language: en Canonical URL: https://forum.scylladb.com/t/io-uring-what-is-it/136 ## Headings Structure: H1: Io_uring: what is it? H3: Related topics ## Main Content: H1: Io_uring: what is it? H3: Related topics I’ve recently been seeing io_uring come up in locations. Understanding it has something to do with input and output, what exactly is it? *The question was originally asked on Stack Overflow by Amani io_uring is a (new as of mid-2019) Linux kernel interface to efficiently allows you to send and receive data asynchronously. It was originally designed to target block devices and files but has since gained the ability to work with things like network sockets. Unlike something like epoll(), it is built around a completion model rather than a readiness model. This is desirable because other operating systems have used the completion model successfully for some time. io_uring provides something competitive and complete for Linux without the drawbacks the previous Linux AIO interface has. The author of io_uring has written a PDF document titled Efficient IO with io_uring, which technically discusses its usage. A gentler introduction is provided by the Lord of the io_uring guide. You can read ScyllaDB developer Glauber Costa proselytize it in How io_uring and eBPF Will Revolutionize Programming in Linux. Lastly, LWN.net has written about io_uring many times. *The answer was originally provided on Stack Overflow by Anon --- ### Page: https://forum.scylladb.com/t/what-are-the-differences-between-column-families-in-cassandras-data-model-compared-to-bigtable/137 Title: What are the differences between column families in Cassandra's data model compared to Bigtable? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I am learning about Cassandra’s data model and its relation to Bigtable but have some things I still don’t understand regarding the Column Family concept. Is the column-family based data model of Cassandra the same as t… Language: en Canonical URL: https://forum.scylladb.com/t/what-are-the-differences-between-column-families-in-cassandras-data-model-compared-to-bigtable/137 ## Headings Structure: H1: What are the differences between column families in Cassandra's data model compared to Bigtable? H3: Related topics ## Main Content: H1: What are the differences between column families in Cassandra's data model compared to Bigtable? H3: Related topics I am learning about Cassandra’s data model and its relation to Bigtable but have some things I still don’t understand regarding the Column Family concept. Is the column-family based data model of Cassandra the same as the column-family based data model of Google’s BigTable? Firstly I’ve read the Bigtable paper, including the part about its data model, that is, how data is stored. As far as I understood, each table in Bigtable relies on a multi-dimensional sparse map with the dimensions row, column, and time. The map is sorted by rows. Columns can be grouped with the name convention family:qualifier to a column family. Therefore, a single row can contain multiple-column families. Although it is stated that Cassandra relies on the Bigtable data model, I read multiple times that in Cassandra, a column family contains multiple rows and is, to some extent, comparable to a table in relational data stores. Isn’t this contrary to Bigtable’s approach, where a row could contain multiple column families? What comes first, the column family or row? Are these concepts even comparable? *Based on a question originally asked on Stack Overflow by OxideNt When Cassandra started, its data model was indeed based on BigTable’s. A row of data could include any number of columns, each of these columns has a name and a value. A row could have a thousand different columns, and a different row could have a thousand other columns - rows do not have to have the same columns. Such a database is called “schema-less”, because there is no schema that each row needs to adhere to. But Toto, we’re not in Kansas anymore - and Cassandra’s model changed in focus (though not in essence) since, and I’ll try to explain how and why: As Cassandra matured, its developers started to realize that schema-less isn’t as great as they once thought it was. Schemas are valuable in ensuring application correctness. Moreover, one doesn’t normally get to 1000 columns in a single row just because there are 1000 individually-named fields in one record. Rather, the more common case is that the record actually contains 200 entries, each with 5 fields. The schema should fix these 5 fields that every one of these entries should have, and what defines each of these separate entries is called a “clustering key”. So around the time of Cassandra 0.8, these ideas were introduced to Cassandra as the “CQL” (Cassandra Query Language). For example, in CQL, one declares that a column-family (which was dutifully renamed “table”) has a schema, with a known list of fields: This schema says that each wide row in the table (now, in modern Cassandra, this was renamed a “partition”) with the key “groupname” is a possibly long list of users, each with username, email, and age fields. The first name in the “PRIMARY KEY” specifier is the partition key (it determines the key of the wide rows), and the second is called the clustering key (it determines the key of the small rows that together make up the wide rows). Despite the new CQL dressup, Cassandra continued to implement these new concepts using the good-old-BigTable-wide-row-without-schema implementation. For example, consider that our data has a group “mygroup” with two people, (john, john@somewhere.com, 27) and (joe, joe@somewhere.com, 38). Cassandra adds the following four column names->values to the wide row: Note how we ended up with a wide row with 4 columns - 2 non-key fields per row (email and age), multiplied by the number of rows in the partition (2). The clustering key field “username” no longer appears anywhere as the value, but rather as part of the column’s name! So If we have two username values “john” and “joe”, We have some columns prefixed “john” and some columns prefixed “joe”, and when we read the column “joe:email” we know this is the value of the email field of the row which has username=joe. Cassandra still has this internal duality - converting the user-facing CQL rows and clustering keys into old-style wide rows. Previously, Cassandra’s on-disk format known as “SSTables” was still schema-less and used composite names as shown above for column names. I wrote a detailed description of the SSTable format on Scylla’s site SSTables Data File · scylladb/scylladb Wiki · GitHub (Scylla is a more efficient C++ re-implementation of Cassandra to which I contribute). However, column names are very inefficient in this format so Cassandra, in version 3.0, switched to a different file format, which for the first time, accepts clustering keys and schema-full rows as first-class citizens. This was the last nail in the coffin of the schema-less Cassandra from 13 years ago. Cassandra is now schema-full, all the way. *Based on an answer originally on Stack Overflow by Nadav Har’El --- ### Page: https://forum.scylladb.com/t/error-message-key-cartesian-product-size-is-greater-than-maximum/138 Title: Error Message: key cartesian product size {} is greater than maximum {} - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I’m seeing this error while reading data from ScyllaDB RequestHandler: ip:9042 replied with server error (clustering key cartesian product size 600 is greater than maximum 100), defuncting connection. Any idea? *Based … Language: en Canonical URL: https://forum.scylladb.com/t/error-message-key-cartesian-product-size-is-greater-than-maximum/138 ## Headings Structure: H1: Error Message: key cartesian product size {} is greater than maximum {} H3: Related topics ## Main Content: H1: Error Message: key cartesian product size {} is greater than maximum {} H3: Related topics I’m seeing this error while reading data from ScyllaDB RequestHandler: ip:9042 replied with server error (clustering key cartesian product size 600 is greater than maximum 100), defuncting connection. Any idea? *Based on a question originally asked on Stack Overflow by rohan-vadje This error is returned to prevent too large restriction sets from being generated, which may put a strain on your server. If you’re aware of the risks and know a reasonable upper bound of the number of restrictions for your queries, you can manually change the maximum in scylla.yaml, e.g. max_clustering_key_restrictions_per_query: 650. Note, however, that this option has a warning in its description, and it should be acknowledged: In particular, setting this flag above a couple of hundred is risky - 600 should be alright, but at this point, you could also consider rephrasing your query so that they have less values in their IN restrictions - perhaps splitting some queries into multiple smaller ones? Source from Scylla tracker: https://github.com/scylladb/scylla/pull/4797 *Based on an answer originally on Stack Overflow by Piotr Sarna --- ### Page: https://forum.scylladb.com/t/scylladbs-read-path-vs-cassandras-and-performance-when-using-hdd-vs-ssd/139 Title: ScyllaDB's read path vs. Cassandra's and performance when using HDD vs SSD - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: What is the difference between ScyllaDB’s read path and Cassandra’s read path? When I stress test Cassandra and ScyllaDB, ScyllaDB’s read performance is poorer by five times compared to Cassandra’s using 16 cores and a n… Language: en Canonical URL: https://forum.scylladb.com/t/scylladbs-read-path-vs-cassandras-and-performance-when-using-hdd-vs-ssd/139 ## Headings Structure: H1: ScyllaDB's read path vs. Cassandra's and performance when using HDD vs SSD H3: Related topics ## Main Content: H1: ScyllaDB's read path vs. Cassandra's and performance when using HDD vs SSD H3: Related topics What is the difference between ScyllaDB’s read path and Cassandra’s read path? When I stress test Cassandra and ScyllaDB, ScyllaDB’s read performance is poorer by five times compared to Cassandra’s using 16 cores and a normal HDD. I expect better read performance on ScyllaDB compared to Cassandra when using a normal HDD. Can someone please confirm, if it’s possible to achieve better read performance using a normal HDD? If so, what changes are required to Scylla’s config? Please guide me! *Based on a question originally asked on Stack Overflow by sateesh There can be various reasons why you are not getting the most out of your Scylla Cluster. I strongly recommend reading the following docs that will provide you with more insights: *Based on an answer on Stack Overflow by TomerSan Both Cassandra and ScyllaDB utilize the same disk storage architecture (LSM). That means that they have relatively the same disk access patterns because the algorithms are largely the same. The LSM trees were built with the idea in mind that it is not necessary to do instant in-place updates. It consists of immutable data buckets that are large continuous pieces of data on disk. That means less random IO, and more sequential IO for which the HDD works great (not counting utilized parallelism by modern database implementations). All the above means that the difference that you see is not induced by the difference in how those databases use a disk. It must be related to the configuration differences and what happens underneath. Maybe ScyllaDB tries to utilize more parallelism or more aggressively do compaction. It depends. To be able to say anything specific, please share your tests, envs, and configurations. *Based on an answer on Stack Overflow by Ivan Prisyazhnyy Both databases use an LSM tree, but Scylla has thread-per-core architecture on top plus we use O_Direct while C* uses the page cache. Scylla also has a sophisticated IO scheduler that makes sure not to overload the disk, and thus scylla_setup runs a benchmark automatically to tune. Check your output of it in io.conf. There are far more things to review, an more information from you is required to look into it. In general, Scylla should perform better in this case as well, but your disk is likely to be the bottleneck in both cases. *Based on an answer on Stack Overflow by dor laor Some other responses focused on write performance, but this isn’t what you asked about - you asked about reads. Uncached read performance on HDDs is bound to be poor in Cassandra and Scylla, because reads from disk each require several seeks on the HDD, and even the best HDD cannot do more than, say, 200 of those seeks per second. Even with a RAID of several of these disks, you will rarely be able to do more than, say, 1000 requests per second. Since a modern multi-core can do orders of magnitude more CPU work than 1000 requests per second, in both Scylla and Cassandra cases, you’ll likely see free CPU. So Scylla’s main benefit of using much less CPU per request will not even matter when the disk is the performance bottleneck. In such cases, I would expect Scylla’s and Cassandra’s performance (I am assuming that you’re measuring throughput when you talk about performance?) should be roughly the same. If still, you’re seeing better throughput from Cassandra than Scylla, several details may explain why, beyond the general client misconfiguration issues raised in other responses: I don’t know which of these differences - or something else - is causing the performance of your use-case to be lower in Scylla, but please keep in mind that whatever you fix, your performance is always going to be bad with HDDs. With SDDs, we’ve measured in the past more than a million random-access read requests per second on a single node. HDDs cannot come to anything close. If you really need optimum performance or performance per dollar, SDDs are the way to go. *Based on an answer on Stack Overflow by Nadav Har’El --- ### Page: https://forum.scylladb.com/t/column-design-and-udt-recommendations-for-a-specific-problem/141 Title: Column design and UDT recommendations for a specific problem - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, I’m learning about Cassandra/Scylla and I have a question. Lets say I’m storing books and each book has N fields, like title, alt title, synopsis, etc… Each of this fields is a string and has an associated language (… Language: en Canonical URL: https://forum.scylladb.com/t/column-design-and-udt-recommendations-for-a-specific-problem/141 ## Headings Structure: H1: Column design and UDT recommendations for a specific problem H3: Related topics ## Main Content: H1: Column design and UDT recommendations for a specific problem H3: Related topics Hi, I’m learning about Cassandra/Scylla and I have a question. Lets say I’m storing books and each book has N fields, like title, alt title, synopsis, etc… Each of this fields is a string and has an associated language (or “unknown”/null). What would be the most efficient way to store this kind of data? I was thinking about two solutions: The first one consists of a UDT column with: title map>, desc map>... the key is the language, for example “en” or “English” and the set is all variants (a language may have a couple). I would add a field into the UDT for every possible field I can think of (title and desc are 2 of 16) like the example below: raw_scraps.fields( title frozen>>, alt_title frozen>>, synopsis_or_description frozen>>, background_or_context frozen>>, status frozen>>, publication frozen>>, author frozen>>, artist frozen>>, serialization frozen>>, tag_genre frozen>>, tag_theme frozen>>, tag_rating frozen>>, tag_demographic frozen>>, tag_format frozen>>, tag_uncategorized frozen>>, content_url frozen>>, original_language frozen>) The other idea consists of a single map where the text is the language and fields is a UDT with 16 title set, desc set... inside. A naive approach would be to store the language string for every possible field, so a UDT with 16 pairs of language+field, but then this would consume a lot of disk. A book may have 10-20 fields in 1-4 languages. Notice that I’m mostly considering disk storage, since I expect scylla to use less resources for the language string if I group fields by their language in an associative collection, but I want to know the general approach for performance, either cpu and disk and how would you implement this. My first preference us to use clustering keys, not collections. This places just once instance of attribute_language per (book, language) pair, rather then once per (book, language, attribute) triplet. It also allows selecting just one language, but that may not be in your requirements. Also: prefer set<> to list<>. Use frozen<> where you can, but be aware it’s less suitable for updating. Your suggestion might not work for my usecase since not all fields have the same set of languages. For example, I may have 2 titles in english, 1 in spanish and 4 in chinese but 1 synopsis in english 3 in french and 1 in russian. So your suggestion can be suitable but kind of painful to work with specially when doing queries: first I’ll change title and synopsis to set and for every entry I would need N (being N the number of different languages in total) rows. For every field in a row I would be storing the subset of texts for that language, and I would use null/empty set when there are no texts for a given field in a language. This works perfectly if each field uses the same or almost the same languages, but fails when that can’t be guaranteed like in my case :C maybe with this new information you might have any other suggestion? PD, I’ve read about set being suggested in place of list but for my usecase I wont be updating the data. Only insert, read and delete, (afaik list degrades on update), not sure about disk space on list vs set. Note sure what the problem is - if there are to titles in some language, leave it NULL. It consumes almost no space. list vs. set are almost the same. Lists only cause trouble if unfrozen. Anyway, your original proposal is fine too, it just duplicates language for each field and leads to messy query results, but that’s not a problem for a machine. --- ### Page: https://forum.scylladb.com/t/release-scylla-monitoring-stack-4-1/146 Title: [RELEASE] Scylla Monitoring Stack 4.1 - Announcements - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of ScyllaDB Monitoring Stack 4.1.0 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB Enterprise and ScyllaDB Open Source, based on Prometheus … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-monitoring-stack-4-1/146 ## Headings Structure: H1: [RELEASE] Scylla Monitoring Stack 4.1 H3: Related topics ## Main Content: H1: [RELEASE] Scylla Monitoring Stack 4.1 H3: Related topics The ScyllaDB team is pleased to announce the release of ScyllaDB Monitoring Stack 4.1.0 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB Enterprise and ScyllaDB Open Source, based on Prometheus and Grafana. ScyllaDB Monitoring Stack 4.1.0 supports: This release brings new panels and graphs, bug fixes, stability improvements and performance enhancements. Versions updates for Scylla Monitoring Stack 4.1.00 New Information in ScyllaDB Dashboards Don’t configure Grafana loki data source if Loki is disable #1818 --- ### Page: https://forum.scylladb.com/t/announcing-scylladb-open-source-5-1/149 Title: Announcing ScyllaDB Open Source 5.1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: ScyllaDB Open Source 5.1 is now available. It is a feature-based minor release for the ScyllaDB Open Source 5.0 major release of earlier this year. It also resolves over 100 issues, bringing improved capabilities, sta… Language: en Canonical URL: https://forum.scylladb.com/t/announcing-scylladb-open-source-5-1/149 ## Headings Structure: H1: Announcing ScyllaDB Open Source 5.1 H3: Related topics ## Main Content: H1: Announcing ScyllaDB Open Source 5.1 H3: Related topics ScyllaDB Open Source 5.1 is now available. It is a feature-based minor release for the ScyllaDB Open Source 5.0 major release of earlier this year. It also resolves over 100 issues, bringing improved capabilities, stability and performance to our NoSQL database server and its APIs. READ THE FULL BLOG: Announcing ScyllaDB Open Source 5.1 - ScyllaDB --- ### Page: https://forum.scylladb.com/t/release-scylla-manager-3-0-2-and-2-6-5/156 Title: [RELEASE] Scylla Manager 3.0.2 and 2.6.5 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces the release of ScyllaDB Manager 3.0.2 and 2.6.5, production-ready ScyllaDB Manager patch releases of the stable ScyllaDB Manager 3.0 and ScyllaDB Manager 2.6 branches. As always, ScyllaDB Man… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-manager-3-0-2-and-2-6-5/156 ## Headings Structure: H1: [RELEASE] Scylla Manager 3.0.2 and 2.6.5 H3: Related topics ## Main Content: H1: [RELEASE] Scylla Manager 3.0.2 and 2.6.5 H3: Related topics The ScyllaDB team announces the release of ScyllaDB Manager 3.0.2 and 2.6.5, production-ready ScyllaDB Manager patch releases of the stable ScyllaDB Manager 3.0 and ScyllaDB Manager 2.6 branches. As always, ScyllaDB Manager customers and users are encouraged to upgrade in coordination with the ScyllaDB support team. ScyllaDB Manager is a centralized cluster administration and recurrent tasks automation tool. Note that this upgrade affects both ScyllaDB Manager Server and Manager Agent; it’s recommended to upgrade both. These releases fix #3235, a compatibility issue with ScyllaDB, introduced in ScyllaDB Open Source 5.0. --- ### Page: https://forum.scylladb.com/t/register-now-for-scylladb-summit-2023/158 Title: Register Now for ScyllaDB Summit 2023 - Blog Posts - ScyllaDB Community NoSQL Forum Meta Description: [800x400-blog-summit-2023-registration-768x384] Language: en Canonical URL: https://forum.scylladb.com/t/register-now-for-scylladb-summit-2023/158 ## Headings Structure: H1: Register Now for ScyllaDB Summit 2023 H3: Register Now for ScyllaDB Summit 2023 H3: Related topics ## Main Content: H1: Register Now for ScyllaDB Summit 2023 H3: Register Now for ScyllaDB Summit 2023 H3: Related topics “Database monsters of the world, connect!” Once again we will be hosting the next annual ScyllaDB Summit on February 15th and 16th, 2023. It will be, as usual, all free and all online. We’ll have dozens of speakers, from your professional peers from... --- ### Page: https://forum.scylladb.com/t/whats-new-w-scylladb-december-2022/159 Title: What's New w/ ScyllaDB - December 2022 - Announcements - ScyllaDB Community NoSQL Forum Meta Description: DECEMBER 2022 NEWSLETTER Learn from Discord, Hulu & Strava & More at ScyllaDB Summit 2023 ScyllaDB Summit is coming to you (free + virtual) February 15-16. Engineers from Discord, Hulu, Strava, ShareChat, and more will… Language: en Canonical URL: https://forum.scylladb.com/t/whats-new-w-scylladb-december-2022/159 ## Headings Structure: H1: What's New w/ ScyllaDB - December 2022 H3: Related topics ## Main Content: H1: What's New w/ ScyllaDB - December 2022 H3: Related topics DECEMBER 2022 NEWSLETTER Learn from Discord, Hulu & Strava & More at ScyllaDB Summit 2023 ScyllaDB Summit is coming to you (free + virtual) February 15-16. Engineers from Discord, Hulu, Strava, ShareChat, and more will be sharing NoSQL strategies, practical tips, and all the latest performance, resilience, and ecosystem innovations impacting the world of data-intensive applications. It’s also a great opportunity to learn about ScyllaDB’s latest innovations and best practices that will help your team get the most out of ScyllaDB. Register now: Register Now for ScyllaDB Summit 2023 - ScyllaDB Introducing the ScyllaDB Community Forum Introducing the ScyllaDB Community Forum, an open platform where users can learn from one another’s experiences with ScyllaDB. Learn more: Introducing the ScyllaDB Community Forum - ScyllaDB R&D Team Bonuses: 10-Second Survey Many companies offer year-end bonuses to entice R&D teams to significantly reduce their cloud spend. Does yours? Share your experience and get a chance to win a $250 Amazon gift card! Take 10-second survey: R & D Team Bonuses Survey ShareChat’s Path to High-Performance NoSQL: Q&A with Geetish Nayak How India’s social media unicorn achieves microsecond P99 latency with 1.2M op/sec – for 180M monthly active users expecting real-time engagement with 2.5B posts per month. Read more: ShareChat’s Path to High-Performance NoSQL: Q&A with Geetish Nayak - ScyllaDB How ScyllaDB Helped an AdTech Company Focus on Core Business AdTech innovator GumGum wanted to escape Cassandra’s maintenance burden. Learn why and how they moved to a database-as-a-service (ScyllaDB Cloud DBaaS). Read more: How ScyllaDB Helped an AdTech Company Focus on Core Business - ScyllaDB ScyllaDB Ranked as one of the Fastest-Growing Companies in North America on the 2022 Deloitte Technology Fast 500™ With 800% growth, ScyllaDB ranked as North America’s 185th fastest-growing company on the 2022 Deloitte Technology Fast 500 – the only database on the list. Read more: ScyllaDB Ranked as one of the Fastest-Growing Companies in North America on the 2022 Deloitte Technology Fast 500™ - ScyllaDB Message Queue and NoSQL with 10x Performance Boost Byte Byte Go’s take on how ScyllaDB achieves at least an order of magnitude improvement in performance vs. Cassandra by taking full advantage of recent trends in computer architecture. Read more: Message Queue and NoSQL with 10x Performance Boost ScyllaDB Internal R&D Summit: Together to the Top Hear what happened when ScyllaDB’s Research & Development team recently gathered in Tbilisi, Georgia to share knowledge, learn from each other, and have fun together. Read more: ScyllaDB Internal R&D Summit: Together to the Top - ScyllaDB EVENTS AND MORE! ScyllaDB Virtual Workshop December 15 | 10am PT | 1pm ET | 5pm GMT Ready to try out ScyllaDB and want to make sure you’re “doing it right?” Spend an hour with our architects for a crash course in what ScyllaDB is all about, the core concepts you need to know, and a step-by-step demonstration of how to get started. Register here: Scylla Virtual Workshop ScyllaDB Summit 2023 February 15-16, 2023 | Online Join us at ScyllaDB Summit 2023 to explore what’s needed to power instantaneous experiences with massive distributed datasets for this next tech cycle. Register now: ScyllaDB Summit 2023 | ScyllaDB NEW SCYLLADB RELEASES AND UPDATES Side note: We’d love your feedback on this format. Any suggestions for making it more valuable? --- ### Page: https://forum.scylladb.com/t/how-does-scylladb-find-the-node-containing-the-data-i-want/161 Title: How Does ScyllaDB Find the Node Containing the Data I Want? - Knowledge Base - ScyllaDB Community NoSQL Forum Meta Description: The driver can connect to any Scylla node and perform a query. That node will be designated as the coordinator node for the given query. The coordinator node can be the replica node (the one holding the data), but it doe… Language: en Canonical URL: https://forum.scylladb.com/t/how-does-scylladb-find-the-node-containing-the-data-i-want/161 ## Headings Structure: H1: How Does ScyllaDB Find the Node Containing the Data I Want? H3: Related topics ## Main Content: H1: How Does ScyllaDB Find the Node Containing the Data I Want? H3: Related topics The driver can connect to any Scylla node and perform a query. That node will be designated as the coordinator node for the given query. The coordinator node can be the replica node (the one holding the data), but it doesn’t have to be. In ScyllaDB (and Apache Cassandra) Each node in the cluster is responsible for a set of tokens. The coordinator node hashes the Partition Key, using the Partition Hash Function to determine which nodes are responsible for that data. Because the partition hash function is known to the client, token-aware drivers can optimize the performance by choosing the coordinator node as one of the replica nodes. This is efficient and as a result the number of network hops is lower and the cluster internal load gets reduced. Scylla shard-aware drivers further increase performance by routing the query not only to the right replica node but also to the right shard (or CPU core) within that node. Additional Resources: Scylla Architecture - Fault Tolerance on ScyllaDB Docs *The question was originally asked on the user slack channel --- ### Page: https://forum.scylladb.com/t/best-way-to-fetch-n-rows-in-scylladb-count-limit-or-paging/162 Title: Best way to Fetch N rows in ScyllaDB: Count, Limit or Paging - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I have a use case where I utilize ScyllaDB to limit users’ actions in the past 24h. Let’s say the user is only allowed to make an order 3 times in the last 24h. I am using ScyllaDB’s ttl and making a count on the number … Language: en Canonical URL: https://forum.scylladb.com/t/best-way-to-fetch-n-rows-in-scylladb-count-limit-or-paging/162 ## Headings Structure: H1: Best way to Fetch N rows in ScyllaDB: Count, Limit or Paging H3: Related topics ## Main Content: H1: Best way to Fetch N rows in ScyllaDB: Count, Limit or Paging H3: Related topics I have a use case where I utilize ScyllaDB to limit users’ actions in the past 24h. Let’s say the user is only allowed to make an order 3 times in the last 24h. I am using ScyllaDB’s ttl and making a count on the number of records in the table to achieve this. I am also using https://github.com/spaolacci/murmur3 to get the hash for the partition key. However, I would like to know what is the most efficient way to query the table. So I have a few queries in which I’d like to understand better and compare the behavior(please correct me if any of my statement is wrong): count() will implement a full-scan query, meaning that it may query more than necessary records into the table. SELECT COUNT(1) FROM orders WHERE hash_id=? AND user_id=?; limit will only limit the number of records returned to the client. Meaning it will still query all records that match its predicates but only limit the ones returned. SELECT user_id FROM orders WHERE hash_id=? AND user_id=? LIMIT ?; I’m a bit new to this, but if I read the docs correctly, it should only query the up until it received the first N records without having to query the whole table. So if I limit the page size to the number of records I want to fetch and only query the first page, would it work correctly? and will it have a consistent result? docs: Paging | ScyllaDB Docs my query is still using limit, but utilizing the driver to achieve this with https://github.com/gocql/gocql iter := conn.Query( “SELECT user_id FROM orders WHERE hash_id=? AND user_id=? LIMIT ?”, hashID, userID,3 ).PageSize(3).PageState(nil).Iter() Please let me know if my analysis was correct and which method would be best to choose *The question was asked on Stack Overflow by Radhian Amri Your client should always use paging - otherwise, you risk adding pressure to the query coordinator, which may introduce latency and memory fragmentation. If you use the Scylla Monitoring stack (and you should if you don’t!), refer to the CQL Optimization dashboard and - more specifically - to the Paged Queries panel. Now, to your question. It seems to be that your example is a bit minimalist for what you are actually wanting to achieve, and - even then - should it not be, we have to consider such set-up at scale. Eg, There may be a tenant allowed which is allowed to place 3 orders within a day, but another tenant allowed to place 1 million orders within a week? If the above assumption is correct - and with the options at hand you have given - you are better off using LIMIT with paging. The reason is that there are some particular problems with the description you’ve given at hand: Of course, it may be that I am wrong, but I’d like to suggest different approaches which may or may not apply to you and to your use case. SELECT something FROM orders WHERE hash_id=? AND user_id=? AND ts >= ? AND ts < ?; SELECT count FROM counter_table WHERE hash_id=? AND user_id=? AND date=?; *The answer was provided on Stack Overflow by Felipe Mendes I have a few points I want to add to what Felipe wrote already: First, you don’t need to hash the partition key yourself. You can use anything you want for the partition key, even consecutive numbers. The partition key doesn’t need to be random-looking. Scylla will internally hash the partition key to improve the load balancing. You don’t need to know or care which hashing algorithm ScyllaDB uses, but interestingly, it’s a variant of murmur3 too (which is not identical to the one you used - it’s a modified algorithm originally picked by the Cassandra developers). Second, you should know - and decide whether you care - that the limit you are trying to enforce is not a hard limit when faced with concurrent operations: Imagine that the given partition already has two records - and now two concurrent record addition requests come in. Both can check that there are just two records, decide it’s fine to add the third - and then when both add their record - and you end up with four records. You’ll need to decide whether this is fine for you that a user can get in 4 requests in a day if they are lucky, or it’s a disaster. Note that theoretically, you can get even more than 4 - if the user manages to send N requests at exactly the same time, they may be able to get 2+N records in the database (but in the usual case, they won’t manage to get many superfluous records). If you’ll want 3 to be a hard limit, you’ll probably need to change your solution - perhaps to one based on LWT and not use TTL. Third, I want to note that there is not an important performance difference between COUNT and LIMIT when you know a-priori that there will only be up to 3 (or perhaps, as explained above, 4 or some other similarly small number) results. If you assume that the SELECT only yields three or less results, and it can never be a thousand results, then it doesn’t really matter if you just retrieve them or count them - you should just do whichever is convenient for you. In any case, I think that paging is not a good solution for your need. For such short results and you can just use the default page size, and you’ll never reach it anyway, and also paging hints to the server that you will likely continue reading on the next page - and it caches the buffers it needs to do that - while in this case, you know that you’ll never continue after the first three results. So in short, don’t use any special paging setup here - just use the default page size (which is 1MB), and it will never be reached anyway. *The answer was provided on Stack Overflow by Nadav Har’El --- ### Page: https://forum.scylladb.com/t/running-cql-query-on-different-nodes-in-the-cluster-gives-different-results/163 Title: Running CQL query on different nodes in the cluster gives different results - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I have a cluster of 3 nodes in a single DC. I have followed the instructions given in https://docs.scylladb.com/operating-scylla/procedures/cluster-management/add-dc-to-existing-dc/ to add 3 new nodes in a new DC. After … Language: en Canonical URL: https://forum.scylladb.com/t/running-cql-query-on-different-nodes-in-the-cluster-gives-different-results/163 ## Headings Structure: H1: Running CQL query on different nodes in the cluster gives different results H3: Related topics ## Main Content: H1: Running CQL query on different nodes in the cluster gives different results H3: Related topics I have a cluster of 3 nodes in a single DC. I have followed the instructions given in https://docs.scylladb.com/operating-scylla/procedures/cluster-management/add-dc-to-existing-dc/ to add 3 new nodes in a new DC. After the nodes in the new DC are started, I check nodetool status and ensure all are up and running. Now since all nodes are part of the same cluster, I assume the query results should be the same irrespective of which node I run the cql query on, isn’t it? But I see that the data is different when the query is run on different nodes. In fact the query results are different when the query is run on different nodes of the same DC too! The following differences are observed (this is not a complete list though): My keyspace previously used SimpleStrategy with a replication factor of 2. While adding the new DC, as part of the steps described in the documentation, I have modified it to use NetworkTopologyStrategy with a replication factor of 2 in both DCs: Why is this difference? What am I missing? This is a sample keyspace and table definition: This query gives different results on different runs (even within the same minute). Sometimes it shows the same result on all nodes. The data is not getting written to this keyspace anymore. Why am I seeing this difference on multiple runs? The 2 runs were just a few seconds apart. Why is there a difference? Can someone throw some light on this please? *The question was asked on Stack Overflow by Shobhana Sriram You didn’t mention when was the last time you ran nodetool repair (or using Scylla Manager to run repair) on this cluster. ScyllaDB (as well as Cassandra) uses eventual consistency, which means your write request will be satisfied when the Consistency level (CL) of that request was achieved. If you used CL=ONE for your writes then only 1 replica needs to ACK for the application to consider this successful. The replication to the 2nd replica will be done a-synchronically (and can also fail for various reasons). Here comes the anti-entropy mechanism, which you can read more about here: https://docs.scylladb.com/architecture/anti-entropy/ You must make sure your cluster completed cluster-wide repair before the table’s gc_grace_seconds value (default 10 days) You should have also fully repaired your cluster before adding the 2nd DC, or at least done it after you’ve added the 2nd DC. This is also written in our docs. On top of all the above, you are running a CQL query, and unless you changed the CQL query’s CL, it’s using the default CL=ONE. This means that EVERY single replica in the cluster (from either of the 2 DCs) can respond to your read request, and as explained above, the data is most likely inconsistent currently. Read more about Architecture → Ring Architecture / CL here: https://docs.scylladb.com/architecture/console-CL-full-demo/ https://docs.scylladb.com/architecture/ringarchitecture/ I highly recommend you to visit Scylla University and learn more about all that I wrote here and much more: https://university.scylladb.com/courses/scylla-essentials-overview/lessons/architecture/ *The answer was provided on Stack Overflow by TomerSan --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-158-2022-12-11/164 Title: Last week in scylladb.git master (issue #158; 2022-12-11) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the b551cd254c…e47794ed98 range are covered. There were 92 non-merge commits from 16 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-158-2022-12-11/164 ## Headings Structure: H1: Last week in scylladb.git master (issue #158; 2022-12-11) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #158; 2022-12-11) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the b551cd254c…e47794ed98 range are covered. There were 92 non-merge commits from 16 authors in that period. Some notable commits: Secondary indexes on static columns are now supported. Usually repair can compare and update the same shard in different nodes, for example shard 3 in one node is compared against shard 3 in another. When the number of shards in nodes is dissimilar, this doesn’t work and each shard compares against data from multiple shards in other nodes. This is now made more efficient by reducing sstable reader thrashing for this dissimilar shard count case. The algorithm for removing nodes from the token ring was corrected and made more efficient. It’s not known that this had any user impact. Raft’s error handling when an ex-cluster-member tries to modify configuration was improved; it is now rejected with a permanent error. The CQL server will now only run requests that benefit from concurrency (e.g. QUERY and EXECUTE) in parallel. Configuration and authentication related requests will be serialized, reducing the chance for errors in those code paths. Alternator, ScyllaDB’s implementation of the DynamoDB API, implements Time-To-Live (TTL) by scanning data. Some shutdown-related hangs in the scanning process were fixed. The bundled Java driver was updated to version 3.11.2.4. When a materialized view processes updates to the base table, it locks the partition and clustering key. In some rare cases involving one of the locks timing out but the other not, this can cause a crash. This was fixed by acquiring the locks sequentially. The bundled scylla types tool can now serialize a value to the sstable binary format. The node replace procedure, used to replace a dead node, will now work through Raft (when enabled). A rare bug involving an allocation failure while updating cached rows was fixed. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/reddit-like-app-use-single-or-multiple-tables/168 Title: Reddit like app - use single or multiple tables? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello there, I am evaluating ScyllaDB for purpose of creating an application that would be in Reddit style. So, I have like Subreddits (Topics), Posts (Questions) and Comments (answers for that specific post). Do I nee… Language: en Canonical URL: https://forum.scylladb.com/t/reddit-like-app-use-single-or-multiple-tables/168 ## Headings Structure: H1: Reddit like app - use single or multiple tables? H3: Related topics ## Main Content: H1: Reddit like app - use single or multiple tables? H3: Related topics Hello there, I am evaluating ScyllaDB for purpose of creating an application that would be in Reddit style. So, I have like Subreddits (Topics), Posts (Questions) and Comments (answers for that specific post). Do I need to separate tables for Topics, Questions and Comments? Do they need to have some between them with relations? Performance-wise, what is the best model? What do you think about this proposal? And should I have relations between them. The best model depends on what you expect to be the most common query. You should start with listing queries. e.g. all posts for a subreddit, all comments within a post, and weight them by frequency and importance. The table structure will be derived from that. The Data Modeling and Application Development course on ScyllaDB University, starting with Basic Data Modeling, might also be of interest. --- ### Page: https://forum.scylladb.com/t/migrating-from-dynamodb-to-scylladb/173 Title: Migrating from DynamoDB to ScyllaDB - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Agenda- Migrate DynamoDB tables to ScyllaDB (Schema as well as data) Does the Scylla-Migrator migrate the table schema as well, or do I have to create the exact schema in ScyllaDB, and then it can just migrate the data? … Language: en Canonical URL: https://forum.scylladb.com/t/migrating-from-dynamodb-to-scylladb/173 ## Headings Structure: H1: Migrating from DynamoDB to ScyllaDB H3: Related topics ## Main Content: H1: Migrating from DynamoDB to ScyllaDB H3: Related topics Agenda- Migrate DynamoDB tables to ScyllaDB (Schema as well as data) Does the Scylla-Migrator migrate the table schema as well, or do I have to create the exact schema in ScyllaDB, and then it can just migrate the data? *The question was asked on Stack Overflow by Aradhana Singh you need to create the keyspace and table in the remote cluster, migrator can map the old table to the new table layout if needed. The reason why we don’t migrate schema automagically is that most of the times you want to have it different (e.g., different compactions strategy or new columns), and then having this as a manual step makes sure you can review your schema before using it. That said I think it makes sense to ask for a special flag that will just migrate old schema for you to new cluster - Issues · scylladb/scylla-migrator · GitHub - can you file it there? *The answer was provided on Stack Overflow by Lubos --- ### Page: https://forum.scylladb.com/t/cassandra-vs-scylladb-memory-usage/174 Title: Cassandra Vs ScyllaDB Memory Usage - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I am doing performance comparisons of ScyllaDB and Cassandra, specifically looking at the impact of memory. The machines I am using each have 16GB and 8 cores. Based on the docs, Cassandra will default to 4GB Xmx and use… Language: en Canonical URL: https://forum.scylladb.com/t/cassandra-vs-scylladb-memory-usage/174 ## Headings Structure: H1: Cassandra Vs ScyllaDB Memory Usage H3: Related topics ## Main Content: H1: Cassandra Vs ScyllaDB Memory Usage H3: Related topics I am doing performance comparisons of ScyllaDB and Cassandra, specifically looking at the impact of memory. The machines I am using each have 16GB and 8 cores. Based on the docs, Cassandra will default to 4GB Xmx and use the remaining 12GB as file system cache. ScyllaDB instead will use all 16GB for itself. What I’m wondering is if this is a fair comparison setup (4GB Xmx for Cassandra vs 16GB for Scylla)? I realize this is what each recommends, but would a more fair test be 8GB Xmx for Cassandra and --memory 8G for ScyllaDB? My workload is mostly write intensive, and I don’t expect file system caching to always be able to help Cassandra. It’s odd to me that ScyllaDB does not expect almost any file system caching compared to Cassandra’s huge reliance on it. *The question was asked on Stack Overflow by Riley Zimmerman Scylla will use ~1/2 of the memory for MemTable, and the other half for Key/Partition caching. If your workload is mostly write, more memory will have less of an effect on performance and should be bounded by either I/O or CPU. I would recommend reading: Learn about different I/O Access Methods and what we chose for ScyllaDB to understand the way Scylla is writing information, and ScyllaDB Workload Conditioning part one: write request rate determination To understand the way Scylla is balancing I/O workloads *The answer was provided on Stack Overflow by gutkinde Cassandra will always use all of the system memory; the heap size (-Xmx) setting just determines how much is used by the heap and how much by other memory consumers (off-heap structures and the page cache). So if you limit Scylla’s memory usage, it will be at a disadvantage compared to Cassandra. *The answer was provided on Stack Overflow by Avi Kivity --- ### Page: https://forum.scylladb.com/t/how-often-should-scylladb-nodes-be-repaired/175 Title: How often should ScyllaDB nodes be repaired? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Our organization has moved on from Cassandra to ScyllaDB recently. How often should we repair ScyllaDB nodes to maintain an equal count of rows in each node, as Cassandra’s repair frequency is recommended as 5 Days? *Th… Language: en Canonical URL: https://forum.scylladb.com/t/how-often-should-scylladb-nodes-be-repaired/175 ## Headings Structure: H1: How often should ScyllaDB nodes be repaired? H3: Related topics ## Main Content: H1: How often should ScyllaDB nodes be repaired? H3: Related topics Our organization has moved on from Cassandra to ScyllaDB recently. How often should we repair ScyllaDB nodes to maintain an equal count of rows in each node, as Cassandra’s repair frequency is recommended as 5 Days? *The question was asked on Stack Overflow by Varun Nagrare Scylla Manager automates the repair process and allows you to configure how and when repair occurs. When you create a cluster a repair task is automatically scheduled. This task is set to occur each week by default, but you can change it to another time, change its parameters or add additional repair tasks if needed. Source: Repair | ScyllaDB Docs *The answer was provided on Stack Overflow by Peter Corless --- ### Page: https://forum.scylladb.com/t/how-to-bulk-fetch-data-from-scylladb/176 Title: How to bulk fetch data from ScyllaDB? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: In our use case, we’d like to fetch data from Scylladb and put it into Elasticsearch. if we take records one by one, it takes too much time. I couldn’t find a ScyllaDB binlog. What’s the right way to do this? *The que… Language: en Canonical URL: https://forum.scylladb.com/t/how-to-bulk-fetch-data-from-scylladb/176 ## Headings Structure: H1: How to bulk fetch data from ScyllaDB? H3: Data Manipulation | ScyllaDB Docs H3: Hooking up Spark and ScyllaDB: Part 1 H3: Deep Dive into the ScyllaDB Spark Migrator H3: scylla-code-samples/spark3-scylla4-demo at master · scylladb/scylla-code-samples H3: Apache Spark support | Elasticsearch for Apache Hadoop [master] | Elastic H3: Related topics ## Main Content: H1: How to bulk fetch data from ScyllaDB? H3: Data Manipulation | ScyllaDB Docs H3: Hooking up Spark and ScyllaDB: Part 1 H3: Deep Dive into the ScyllaDB Spark Migrator H3: scylla-code-samples/spark3-scylla4-demo at master · scylladb/scylla-code-samples H3: Apache Spark support | Elasticsearch for Apache Hadoop [master] | Elastic H3: Related topics In our use case, we’d like to fetch data from Scylladb and put it into Elasticsearch. if we take records one by one, it takes too much time. I couldn’t find a ScyllaDB binlog. What’s the right way to do this? *The question was asked on Stack Overflow by tianzhenjiu You might want to look at using Change Data Capture in Scylla, then using the CDC tables to feed a Kafka topic that will populate Elasticsearch. ScyllaDB’s CDC connector for Kafka is built on Debezium. You can read more about it here. *The answer was provided on Stack Overflow by Peter Corless And if you want to read everything on top of live additions using CDC, you can just write a sample scala spark application that will just load everything needing a fulltext search from Scylla to Elastic (sample apps are on the internet or have a look at series of blogs around Scylla migrator, which explain how to properly leverage dataframes). Fwiw, Scylla supports the operator LIKE, in case a simple search will cut it for you (assuming your partitions are not huge) instead of the Lucene query language and the inverted indexes Elastic uses. ScyllaDB is an Apache Cassandra-compatible NoSQL data store that can handle 1 million transactions per second on a single server. This is part one in a series on the integration of Spark and ScyllaDB. We will cover how to run both locally, data modeling, and how the connectors work. The ScyllaDB Spark Migrator is the preferred tool to migrate data from Cassandra to ScyllaDB. Learn how it works from this deep dive, with code now on Github. master/spark3-scylla4-demo Code samples for working with ScyllaDB. Contribute to scylladb/scylla-code-samples development by creating an account on GitHub. Reference documentation of elasticsearch-hadoop Not sure how useful this will be: *The answer was provided on Stack Overflow by Lubos --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-1-1/177 Title: [RELEASE] ScyllaDB 5.1.1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.1.1, a bugfix release of the ScyllaDB 5.1 stable branch. ScyllaDB Open Source 5.1.1, like all past and future 5.x.y releases, is backward compatible and supports rolling… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-1-1/177 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.1.1 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.1.1 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.1.1, a bugfix release of the ScyllaDB 5.1 stable branch. ScyllaDB Open Source 5.1.1, like all past and future 5.x.y releases, is backward compatible and supports rolling upgrades. Users are encouraged to upgrade to 5.1.1. Issues fixed in this release: --- ### Page: https://forum.scylladb.com/t/cutting-database-costs-lessons-from-comcast-rakuten-expedia-ifood/178 Title: Cutting Database Costs: Lessons from Comcast, Rakuten, Expedia & iFood - Blog Posts - ScyllaDB Community NoSQL Forum Meta Description: [800x400-blog-cutting-database-costs] Language: en Canonical URL: https://forum.scylladb.com/t/cutting-database-costs-lessons-from-comcast-rakuten-expedia-ifood/178 ## Headings Structure: H1: Cutting Database Costs: Lessons from Comcast, Rakuten, Expedia & iFood H3: Cutting Database Costs: Lessons from Comcast, Rakuten, Expedia & iFood H3: Related topics ## Main Content: H1: Cutting Database Costs: Lessons from Comcast, Rakuten, Expedia & iFood H3: Cutting Database Costs: Lessons from Comcast, Rakuten, Expedia & iFood H3: Related topics A look at how several dev teams significantly reduced database costs while actually improving database performance. --- ### Page: https://forum.scylladb.com/t/blog-topic-suggestions/179 Title: Blog topic suggestions? - Blog Posts - ScyllaDB Community NoSQL Forum Meta Description: If you have suggestions for blog topics, please add them to this thread. Language: en Canonical URL: https://forum.scylladb.com/t/blog-topic-suggestions/179 ## Headings Structure: H1: Blog topic suggestions? H3: Related topics ## Main Content: H1: Blog topic suggestions? H3: Related topics If you have suggestions for blog topics, please add them to this thread. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-159-2022-12-18/180 Title: Last week in scylladb.git master (issue #159; 2022-12-18) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the e47794ed98…b52bd9ef6a range are covered. There were 97 non-merge commits from 19 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-159-2022-12-18/180 ## Headings Structure: H1: Last week in scylladb.git master (issue #159; 2022-12-18) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #159; 2022-12-18) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the e47794ed98…b52bd9ef6a range are covered. There were 97 non-merge commits from 19 authors in that period. Some notable commits: ScyllaDB now supports server-side DESCRIBE. This is required for the latest cqlsh, and reduces the need to update cqlsh as server features are added. However, the version number bump needed to inform cqlsh about this change was reverted as other changes related to the version number are not ready. The replacenode operation is used to replace a dead node. It has now changed to generate a new unique host ID rather than taking over the ID of the node it replaces. This makes the cluster more robust against cases where a node previously thought dead comes back alive. Building on the uniqueness of host IDs, Raft will now use the host ID rather than its own ID. This simplifies administration as there is now just one host ID type to track. A crash in the task manager was fixed. The server will warn if a peer’s address cannot be found when pinging it. The various cloud snitches, which determine the datacenter and rack association of a node, improved error checking of their communication with the cloud provider. The system.truncated table holds information about truncation times of user tables. A recent regression caused it to be unreadable by cqlsh. It is now fixed. When a table with a materialized view had a large partition, and that large partition was deleted with the USING TIMESTAMP clause, the materialized view might be only partially updated. This is now fixed. COMPACT STORAGE tables allow the user to only specify a prefix of a compound clustering key. Bugs relating to such partial keys and reversed rows were fixed. It is recommended to stay away from such obscure features. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-0-7/182 Title: [RELEASE] ScyllaDB 5.0.7 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.0.7, a bugfix release of the ScyllaDB 5.0 stable branch. ScyllaDB Open Source 5.06, like all past and future 5.x.y releases, is backward compatible and supports rolling … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-0-7/182 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.0.7 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.0.7 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.0.7, a bugfix release of the ScyllaDB 5.0 stable branch. ScyllaDB Open Source 5.06, like all past and future 5.x.y releases, is backward compatible and supports rolling upgrades. Note that the latest stable branch of ScyllaDB is 5.1, and you are encouraged to upgrade. Issues fixed in this release: --- ### Page: https://forum.scylladb.com/t/top-nosql-blogs-of-2022/184 Title: Top NoSQL Blogs of 2022 - Blog Posts - ScyllaDB Community NoSQL Forum Meta Description: [800x400-most-popular-blog-2022 (1)] Language: en Canonical URL: https://forum.scylladb.com/t/top-nosql-blogs-of-2022/184 ## Headings Structure: H1: Top NoSQL Blogs of 2022 H3: Top NoSQL Blogs of 2022 H3: Related topics ## Main Content: H1: Top NoSQL Blogs of 2022 H3: Top NoSQL Blogs of 2022 H3: Related topics A look at the top 10 NoSQL blogs written this year, plus 10 perennial favorites: Rust, Go, Kafka, Webassembly, benchmark results, and use cases from Palo Alto Neworks & Disney+ Hotstar top the list. --- ### Page: https://forum.scylladb.com/t/scylladb-best-practices-blogs-videos/185 Title: ScyllaDB Best Practices- Blogs, Videos - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: We’re collecting many of our top tips for working with ScyllaDB in this library. Browse around-- and please share any feedback you have. Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-best-practices-blogs-videos/185 ## Headings Structure: H1: ScyllaDB Best Practices- Blogs, Videos H3: Related topics ## Main Content: H1: ScyllaDB Best Practices- Blogs, Videos H3: Related topics We’re collecting many of our top tips for working with ScyllaDB in this library. Browse around-- and please share any feedback you have. These ScyllaDB University lessons are also relevant for that: --- ### Page: https://forum.scylladb.com/t/enum-in-scylladb/187 Title: Enum in ScyllaDB - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, I am learning Scylla and I have a question. How do I set up an analogue of enum? I have a status field with certain values ( WaitForStart, Started, In progress, Finished, Error, Paused, Stopped ). How will it look li… Language: en Canonical URL: https://forum.scylladb.com/t/enum-in-scylladb/187 ## Headings Structure: H1: Enum in ScyllaDB H3: Related topics ## Main Content: H1: Enum in ScyllaDB H3: Related topics Hi, I am learning Scylla and I have a question. How do I set up an analogue of enum? I have a status field with certain values ( WaitForStart, Started, In progress, Finished, Error, Paused, Stopped ). How will it look like in ScyllaDb. Thought of making a UDT, but then how to constrain values to it? Hey @Eugene.vip, AFAIK ScyllaDB (and Cassandra) don’t implement enum directly. You can achieve this behavior from the application layer. Have a look at the ScyllaDB Rust driver and the Java driver for details. Thank you for your reply. That’s what I did. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-160-2022-12-25/192 Title: Last week in scylladb.git master (issue #160; 2022-12-25) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the b52bd9ef6a…b0d95948e1 range are covered. There were 107 non-merge commits from 12 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-160-2022-12-25/192 ## Headings Structure: H1: Last week in scylladb.git master (issue #160; 2022-12-25) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #160; 2022-12-25) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the b52bd9ef6a…b0d95948e1 range are covered. There were 107 non-merge commits from 12 authors in that period. Some notable commits: ScyllaDB now supports multiple compaction groups. This is not a user-visible feature for now, but will make migration of tablets to other shards and nodes much faster. Some copies of the lists of ranges to stream were eliminated from the decommission path, reducing latency spikes. When the global index cache is disabled, a local (per query) cache was used instead. When that cache was destroyed, a stall could result, generating a latency spike. This is now fixed. Compaction manager generally reacts to events to initiate compactions, but also has an hourly timer in case an event was missed (and for tombstone compaction, which isn’t triggered by an event). This timer is now less susceptible to stalls. Repair tried to trigger off-strategy compaction even for a table that was dropped during repair, failing the entire repair. It ignores the dropped table now. The docker image is now more robust against different network conditions, which could cause it to fail to launch. Single-partition reads could, in conditions involving multi-page queries, a completely empty page, and a partition or range tombstone covering the first row, terminate paging prematurely, leading to incorrect results. This is now fixed. Documentation of the embedded tools (scylla sstable etc) is moved to the documentation tree, making it accessible online. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/scylladb-blobs-internal-treatment-in-comparison-to-ascii-and-text/193 Title: ScyllaDB Blobs internal treatment in comparison to ASCII and Text - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: How are blobs treated internally in comparison to Ascii or Text? Storage wise it should be better, but does ScyllaDB compute its own hashes for every primary key column? I want to know the details. *originally asked on … Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-blobs-internal-treatment-in-comparison-to-ascii-and-text/193 ## Headings Structure: H1: ScyllaDB Blobs internal treatment in comparison to ASCII and Text H3: Related topics ## Main Content: H1: ScyllaDB Blobs internal treatment in comparison to ASCII and Text H3: Related topics How are blobs treated internally in comparison to Ascii or Text? Storage wise it should be better, but does ScyllaDB compute its own hashes for every primary key column? I want to know the details. *originally asked on ScyllaDB’s community slack channel There is not much difference internally, actually. From ScyllaDB’s POV all values are binary blobs with different serialize/deserialize/compare routines attached to them. Scylla only computes a hash for the partition key as a whole. Individual columns are not hashed. Compare wise, blobs, as well as text and ASCII are compared lexicographically. This should yield identical ordering for all three choices. Overall I think blob is the best choice simply because it is the most compact form storage-wise and has no other disadvantages over text/ASCII for this use-case that I can think of. --- ### Page: https://forum.scylladb.com/t/get-the-approximate-number-of-rows-of-a-table/196 Title: Get the approximate number of rows of a table? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, is there a (cheap + fast) way to get the approximate number of rows of a table? I know there’s e.g. ‘Number of partitions (estimate)’ of tablestats. I’d like to be able to be able to query via CQL, as an light-weig… Language: en Canonical URL: https://forum.scylladb.com/t/get-the-approximate-number-of-rows-of-a-table/196 ## Headings Structure: H1: Get the approximate number of rows of a table? H3: Related topics ## Main Content: H1: Get the approximate number of rows of a table? H3: Related topics Hi, is there a (cheap + fast) way to get the approximate number of rows of a table? I know there’s e.g. ‘Number of partitions (estimate)’ of tablestats. I’d like to be able to be able to query via CQL, as an light-weight alternative to select count(*) from my_table. In general there is no good way to do this in a way similar to compaction statistics. While we can count the rows in an sstables, those rows could overlap the rows in another sstable (so we’d count them twice), or could overlap a tombstone in another sstable (and so should not be counted at all). Starting with ScyllaDB 5.1, SELECT COUNT(*) FROM tab is automatically parallelized across all nodes and shards. In conjunction with Consistency Level LOCAL_ONE, this is much faster that before, but still requires significant CPU and I/O resources. Understood. Thanks for the reply anyway! --- ### Page: https://forum.scylladb.com/t/nodes-reshaping-data-after-compaction-and-not-coming-back-online-reshaping-progress-indication/197 Title: Nodes reshaping data after compaction and not coming back online, reshaping progress indication - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hey all. Yesterday I forced a major compaction on our tables. Unfortunately, two of our nodes restarted and wouldn’t come back up online but spent their time reshaping data. Today we would like to get those nodes back on… Language: en Canonical URL: https://forum.scylladb.com/t/nodes-reshaping-data-after-compaction-and-not-coming-back-online-reshaping-progress-indication/197 ## Headings Structure: H1: Nodes reshaping data after compaction and not coming back online, reshaping progress indication H3: Related topics ## Main Content: H1: Nodes reshaping data after compaction and not coming back online, reshaping progress indication H3: Related topics Hey all. Yesterday I forced a major compaction on our tables. Unfortunately, two of our nodes restarted and wouldn’t come back up online but spent their time reshaping data. Today we would like to get those nodes back online because we’re not seeing what the progress is. We tried running nodetool stop RESHAPE, but it seems like it just stops the reshaping task to then immediately start reshaping the same files again. Is this potentially a bug? Any advice on how to get the nodes back online? Running version 5.1 and SizeTieredCompactionStrategy. *Originally asked on ScyllaDB’s community slack channel It would be best if you could hold your breath and let reshaping finish. Ok, cool, thanks. is there any way to get an indication (e.g., by looking at sizes on the filesystem) of the progress of the reshaping? If you list the sstables on disk and cross-reference them with the log, you can see estimate how much progress reshape has made by looking at how many sstables were created before reshape started vs. those who were created after (and by reshape). As for officially providing this information over the API, we’re working on a generic task manager that will provide an API to get tasks progress. We’re currently working on integrating it with repair tasks, and next in line are node operations and compaction tasks. Thanks! that’s what we ended up doing happily, the reshape has finished by now. --- ### Page: https://forum.scylladb.com/t/can-t-find-the-monitoring-dashboards/198 Title: Can’t find the Monitoring dashboards - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I deployed Scylla-Operator, and Scylla-Manager using Helm on GKE. I now have several scylla-datacenters in a cluster. I Already have Prometheus and Grafana in the cluster and successfully configured a job to scrape scyl… Language: en Canonical URL: https://forum.scylladb.com/t/can-t-find-the-monitoring-dashboards/198 ## Headings Structure: H1: Can’t find the Monitoring dashboards H3: scylla-monitoring/grafana/build at branch-4.1 · scylladb/scylla-monitoring H3: Related topics ## Main Content: H1: Can’t find the Monitoring dashboards H3: scylla-monitoring/grafana/build at branch-4.1 · scylladb/scylla-monitoring H3: Related topics I deployed Scylla-Operator, and Scylla-Manager using Helm on GKE. I now have several scylla-datacenters in a cluster. I Already have Prometheus and Grafana in the cluster and successfully configured a job to scrape scylla-metrics. Now I want to have dashboards. Where can I find the Scylla Dashboards for Grafana? I can’t find the Scylla dashboard on Grafana Labs. Found this GitHub - scylladb/scylla-monitoring: Simple monitoring of Scylla with Grafana But here, as I understand it, I need to run it locally and manually take the dashboards from the running Grafana. Also was trying to import dashboards from here - scylla-monitoring/grafana at master · scylladb/scylla-monitoring · GitHub, but they don’t work. I get an error - “Dashboards need to have a title.” It seems like some kind of template. Is any other way to get and import dashboards to the Grafana client? *originally asked on ScyllaDB’s community slack channel Did you install Grafana yourself? It is recommended to use GitHub - scylladb/scylla-monitoring: Simple monitoring of Scylla with Grafana for monitoring. It includes Prometheus, Grafana, and more. It comes with the Scylla dashboards. You can also take the dashboards from a version, for example: branch-4.1/grafana/build Simple monitoring of Scylla with Grafana. Contribute to scylladb/scylla-monitoring development by creating an account on GitHub. --- ### Page: https://forum.scylladb.com/t/tradeoff-between-allow-filtering-and-creating-an-index/199 Title: Tradeoff between allow filtering and creating an index - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: My table has the following primary key: PRIMARY KEY((tenant_id, group_id)) If I have to query data based on tenant_id, I have two options: select * from tablename where tenant_id=‘something’ ALLOW FILTERING Pros: N… Language: en Canonical URL: https://forum.scylladb.com/t/tradeoff-between-allow-filtering-and-creating-an-index/199 ## Headings Structure: H1: Tradeoff between allow filtering and creating an index H3: Global Secondary Indexes | ScyllaDB Docs H3: Related topics ## Main Content: H1: Tradeoff between allow filtering and creating an index H3: Global Secondary Indexes | ScyllaDB Docs H3: Related topics My table has the following primary key: PRIMARY KEY((tenant_id, group_id)) If I have to query data based on tenant_id, I have two options: *Originally asked on ScyllaDB’s community slack channel The tradeoff is this: pay all costs at query time or spread the cost over all writes, each paying a small portion of it. Understood. If I create an index on “tenant_id” and the table has records let’s say, in 1 lakh per tenant. Will this work properly with the paginated response? As explained in the below post, it uses in clause internally. ScyllaDB is an Apache Cassandra-compatible NoSQL data store that can handle 1 million transactions per second on a single server. Yes, paging works with indexed queries. --- ### Page: https://forum.scylladb.com/t/migrating-legacy-monolith-app-to-scylladb-need-consulting/200 Title: Migrating Legacy monolith app to Scylladb. Need Consulting - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: hi all , I am just starting to use scylladb to optimize a legacy monolith system using scylladb + golang. I am looking for professional consulting to review our design , setup and code so that we can do it right fir… Language: en Canonical URL: https://forum.scylladb.com/t/migrating-legacy-monolith-app-to-scylladb-need-consulting/200 ## Headings Structure: H1: Migrating Legacy monolith app to Scylladb. Need Consulting H3: Related topics ## Main Content: H1: Migrating Legacy monolith app to Scylladb. Need Consulting H3: Related topics hi all , I am just starting to use scylladb to optimize a legacy monolith system using scylladb + golang. I am looking for professional consulting to review our design , setup and code so that we can do it right first time by using valuable experience you guys would have . This would be a short one-time engagement to begin with . Is there some consultant I can connect with ? Hi - my name is Guy and I’m the solution architect director at ScyllaDB. We can have a short call to discuss. please email me at guycarmin@scylladb.com --- ### Page: https://forum.scylladb.com/t/what-is-the-right-term-for-the-dynamodb-and-cassandra-data-model/202 Title: What is the right term for the DynamoDB and Cassandra data model? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: The DynamoDB Wikipedia article states that DynamoDB is a “key-value” database. However, the term “key-value” database completely overlooks a very fundamental feature of DynamoDB, that of the sort key: keys consist of two… Language: en Canonical URL: https://forum.scylladb.com/t/what-is-the-right-term-for-the-dynamodb-and-cassandra-data-model/202 ## Headings Structure: H1: What is the right term for the DynamoDB and Cassandra data model? H3: Related topics ## Main Content: H1: What is the right term for the DynamoDB and Cassandra data model? H3: Related topics The DynamoDB Wikipedia article states that DynamoDB is a “key-value” database. However, the term “key-value” database completely overlooks a very fundamental feature of DynamoDB, that of the sort key: keys consist of two parts (partition key and sort key), and items with the same partition key can be efficiently retrieved together by the sort key. Cassandra also has the exact same function of sorting items within a partition (which is called a “clustering key”), and Cassandra’s Wikipedia article uses the term wide colum store to describe this. While the term “wide column” is better than “key value,” it is still somewhat misleading as it describes the more general situation where an element may have a large number of unrelated columns, not necessarily an ordered list of separate items. So my question is, is there a more appropriate term that can describe the data model of a database like DynamoDB and Cassandra: databases that, like a key-value store, can efficiently retrieve items for individual keys, but also items sorted by the key or just a part of it (the DynamoDB sort key or the Cassandra clustering key). *Based on a question originally asked on Stack Overflow by Nadav Har’El Before the introduction of CQL, Cassandra adhered more strictly to the wide column store data model, in which there were only rows identified by a row key and containing sorted key/value columns. With the advent of CQL, rows became known as partitions, and columns could optionally be grouped into logical rows via clustering keys. Even up to Cassandra 3.0, CQL was simply an abstraction over the original thrift data model, and the concept of CQL rows within the storage engine did not exist. They were simply a sorted set of columns with a compound key consisting of the concatenated values of the clustering keys. See this article for more details. There is now native support for CQL in the storage engine, allowing CQL data models to be stored more efficiently. However, if you think of a CQL row as a logical grouping of columns within the same partition, Cassandra could still be viewed as a wide column store. Anyway, to my knowledge, there is no other established term to describe this kind of database. *Based on an answer on Stack Overflow by J.B. Langston --- ### Page: https://forum.scylladb.com/t/scylladb-innovation-awards-nominate-your-team/204 Title: ScyllaDB Innovation Awards: Nominate Your Team - Blog Posts - ScyllaDB Community NoSQL Forum Meta Description: [800x400-blog-summit-2023-innovation-award-blog] Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-innovation-awards-nominate-your-team/204 ## Headings Structure: H1: ScyllaDB Innovation Awards: Nominate Your Team H3: ScyllaDB Innovation Awards: Nominate Your Team H3: Related topics ## Main Content: H1: ScyllaDB Innovation Awards: Nominate Your Team H3: ScyllaDB Innovation Awards: Nominate Your Team H3: Related topics Get your team’s amazing achievements the recognition they deserve — tell us why you should win a ScyllaDB Innovation Award. The 2023 ScyllaDB Innovation Awards shine a spotlight on ScyllaDB users who went above and beyond to deliver exceptional... --- ### Page: https://forum.scylladb.com/t/scylladb-users-nominate-your-team-for-an-award/205 Title: ScyllaDB Users - Nominate Your Team for an Award - Announcements - ScyllaDB Community NoSQL Forum Meta Description: Get your team’s achievements the recognition they deserve — tell us why you should win a ScyllaDB Innovation Award. All users are eligible: ScyllaDB Cloud, Enterprise, and Open Source. We have 7 categories this year. W… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-users-nominate-your-team-for-an-award/205 ## Headings Structure: H1: ScyllaDB Users - Nominate Your Team for an Award H3: Related topics ## Main Content: H1: ScyllaDB Users - Nominate Your Team for an Award H3: Related topics Get your team’s achievements the recognition they deserve — tell us why you should win a ScyllaDB Innovation Award. All users are eligible: ScyllaDB Cloud, Enterprise, and Open Source. We have 7 categories this year. Winners receive a special ScyllaDB swag pack — plus recognition in a ScyllaDB Summit keynote, blog, press release, and social media posts. Just give us a paragraph or two about your accomplishments and check off which award categories you want to apply for. This blog has more details about the categories, and a link to the nomination form: ScyllaDB Innovation Awards: Nominate Your Team - ScyllaDB --- ### Page: https://forum.scylladb.com/t/package-dependencies-is-python2-7-necessary/208 Title: Package dependencies , is python2.7 necessary? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: hello, is it possible to get rid of Python 2.7 in dependencies? Cassandra4 is already fully working on Python3 in conjunction with cqlsh Language: en Canonical URL: https://forum.scylladb.com/t/package-dependencies-is-python2-7-necessary/208 ## Headings Structure: H1: Package dependencies , is python2.7 necessary? H3: Related topics ## Main Content: H1: Package dependencies , is python2.7 necessary? H4: Repackaging cqlsh H3: Related topics hello, is it possible to get rid of Python 2.7 in dependencies? Cassandra4 is already fully working on Python3 in conjunction with cqlsh Yes, it’s in the working, almost merged into Scylla core master branch. You can track it here: cqlsh is moving into it's own repository: https://github.com/scylladb/scylla-cq…lsh * add cqlsh as submodule * update scylla-java-tools to have cqlsh remove * introduced new cqlsh artifact (rpm/deb/tar) Depends: https://github.com/scylladb/scylla-tools-java/pull/316 Ref: scylladb/scylladb#11569 Hopefully it would be in before 5.2 release starts. If it would, it’s to be expected in 5.2 / 2023.1 Also it’s a dependency only on scylla-tools package Which installs cqlsh / cassandra-stress, and a few other tools, the python2 dependency is only for cqlsh. So if you don’t need the tools, one could install the specific packages without the tools apt install scylla-server scylla-jmx Thank you for the information Repackaging cqlsh by fruch · Pull Request #11937 · scylladb/scylladb · GitHub was merged. so next release 5.3, would be have newer cqlsh version bundled (that’s not depended on python2 anymore) --- ### Page: https://forum.scylladb.com/t/knowing-scylladb-limitations/213 Title: Knowing ScyllaDB Limitations - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, We are evaluating scylla for our internal filtering and sorting use case. We want to know the following: How much data scylla supports per row. I see 20KB in the pricing calculator. Is it the limit? How many index… Language: en Canonical URL: https://forum.scylladb.com/t/knowing-scylladb-limitations/213 ## Headings Structure: H1: Knowing ScyllaDB Limitations H3: Scylla Anti-Entropy | ScyllaDB Docs H3: Scylla Large Rows and Large Cells Tables | ScyllaDB Docs H3: Maximizing Scylla Performance | ScyllaDB Docs H3: System Limits | ScyllaDB Docs H3: Materialized Views | ScyllaDB Docs H3: Related topics ## Main Content: H1: Knowing ScyllaDB Limitations H3: Scylla Anti-Entropy | ScyllaDB Docs H3: Scylla Large Rows and Large Cells Tables | ScyllaDB Docs H3: Maximizing Scylla Performance | ScyllaDB Docs H3: System Limits | ScyllaDB Docs H3: Materialized Views | ScyllaDB Docs H3: Related topics We are evaluating scylla for our internal filtering and sorting use case. We want to know the following: Thanks & Regards, Abhishek Thank you for reaching out! There is no limit to the row size, but keep in mind, if the row/cell is too big might occur performance degradation and the CPU will take more time to process the data, increasing latency. No limits, but should be used carefully to avoid performance degradation, remember that index or Materialized Views are also tables and should have small partitions as well. Yes, as mentioned the Scylla will take more time to process the rows - mostly for writes, read queries should be fine since we are a column-based database. ScyllaDB is real-time with a predictable low-tail latency database and it is eventually consistent. Here you can find some links to find out more: ScyllaDB is an Apache Cassandra-compatible NoSQL data store that can handle 1 million transactions per second on a single server. ScyllaDB is an Apache Cassandra-compatible NoSQL data store that can handle 1 million transactions per second on a single server. ScyllaDB is an Apache Cassandra-compatible NoSQL data store that can handle 1 million transactions per second on a single server. ScyllaDB is an Apache Cassandra-compatible NoSQL data store that can handle 1 million transactions per second on a single server. ScyllaDB is an Apache Cassandra-compatible NoSQL data store that can handle 1 million transactions per second on a single server. --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2021-1-17/214 Title: [RELEASE] ScyllaDB Enterprise 2021.1.17 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2021.1.17, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2021.1. Note that ScyllaDB Enterprise 2022.1 LTS is the latest stable branc… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2021-1-17/214 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2021.1.17 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2021.1.17 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2021.1.17, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2021.1. Note that ScyllaDB Enterprise 2022.1 LTS is the latest stable branch and you are encouraged to upgrade in coordination with the ScyllaDB support team. The following issues are fixed in this release (with an open source reference): --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-161-2023-01-01/215 Title: Last week in scylladb.git master (issue #161; 2023-01-01) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the b0d95948e1…a7c4a129cb range are covered. There were 9 non-merge commits from 7 authors in that period.… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-161-2023-01-01/215 ## Headings Structure: H1: Last week in scylladb.git master (issue #161; 2023-01-01) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #161; 2023-01-01) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the b0d95948e1…a7c4a129cb range are covered. There were 9 non-merge commits from 7 authors in that period. Some notable commits: Off-strategy compaction is now enabled for all streaming topology operations (adding and removing nodes). Previously it was enabled only for repair-based node operations. Off-strategy compaction takes advantage of the fact that incoming sstables are non-overlapping to perform more efficient compaction that the one performed by the regular compaction strategy. The sstable row_reads metric for m-format sstables is now properly incremented, instead of showing zeroes. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/how-do-i-change-the-node-instance-size-in-a-running-cluster/220 Title: How do I change the node instance size in a running cluster? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I have a Scylla cluster running on AWS i3en.xlarge instances. How can I change the instances to a different size, for example, i3en.2xlarge? One way would be to replace each node, that is, to add and remove a node. Is… Language: en Canonical URL: https://forum.scylladb.com/t/how-do-i-change-the-node-instance-size-in-a-running-cluster/220 ## Headings Structure: H1: How do I change the node instance size in a running cluster? H3: Related topics ## Main Content: H1: How do I change the node instance size in a running cluster? H3: Related topics I have a Scylla cluster running on AWS i3en.xlarge instances. How can I change the instances to a different size, for example, i3en.2xlarge? One way would be to replace each node, that is, to add and remove a node. Is there a better way to do it? *Based on a question originally asked on Stack Overflow by SilentCanon Indeed, one way to do it would be to add a larger node, wait for the data to stream and then remove a small node. You can read more about this procedure in how to Upscale a Cluster. Another way to do this without having to change the nodes one by one would be to: Add a new DC to the existing cluster with the desired instance type. You can read more about this procedure here. It includes: Wait for the streaming to the new DC to finish Run a full cluster repair and wait for it to complete Decommission the “original” Datacenter by following the Remove a Data-Center from a Scylla Cluster procedure Notice that it’s currently not recommended to use instances of different types in a single cluster. In the future, and by using tablet partitioning, this will change. We’re working on a new partitioner alongside Raft, which would reduce the total number of tokens or partitions, while still keeping the number high to guarantee even distribution of work. It will also allow balancing the data between partitions more flexibly. These partitions are called tablets. You can read more about this in this blog post. *Based on an answer on Stack Overflow by TomerSan Clarifying question: the linked “Adding a New Data Center Into an Existing ScyllaDB Cluster” doc requires using Ec2MultiRegionSnitch or the GossipingPropertyFileSnitch. Is that a hard requirement in the context of the answer here, where a new data center is being created to replace the old? Ec2MultiRegionSnitch requires using public IP addresses, and I’m working with a cluster where public net access is disallowed for security. The EC2Snitch docs show it being used with multiple DC, so I believe it would work… but I wanted to confirm before disregarding that part of the instructions. Since locator: Do not enforce public ip address for broadcast_rpc_address · scylladb/scylladb@77b1db4 · GitHub Ec2MultiRegionSnitch no longer enforces a public IP address. You may also want to look into Ec2MultiRegionSnitch should not enforce a public RPC address · Issue #10236 · scylladb/scylladb · GitHub as it details that both Ec2Snitch and GossipingPropertyFileSnitch can still be used. --- ### Page: https://forum.scylladb.com/t/how-do-i-know-which-scylladb-version-i-m-running/221 Title: How do I know which ScyllaDB version I’m running? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: When I run nodetool version I get one result, but when I run scylla –version I get a different result: [guy@fedora mms]$ docker exec -it scylla-node1 nodetool version ReleaseVersion: 3.0.8 [guy@fedora mms]$ docker exe… Language: en Canonical URL: https://forum.scylladb.com/t/how-do-i-know-which-scylladb-version-i-m-running/221 ## Headings Structure: H1: How do I know which ScyllaDB version I’m running? H3: Related topics ## Main Content: H1: How do I know which ScyllaDB version I’m running? H3: Related topics When I run nodetool version I get one result, but when I run scylla –version I get a different result: Why is there a difference, and which one is correct? *Based on a question originally asked on Stack Overflow by LetsNoSQL You get the correct Scylla version number by using scylla –version. nodetool version is related to the Cassandra version it was derived from. In the above example, you are running Scylla version 4.5 *Based on an answer on Stack Overflow by Avi Kivity --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-0-8/223 Title: [RELEASE] ScyllaDB 5.0.8 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.0.8, a bugfix release of the ScyllaDB 5.0 stable branch. Like all past and future 5.x.y releases, it is backward compatible and supports rolling upgrades. Note that the… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-0-8/223 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.0.8 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.0.8 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.0.8, a bugfix release of the ScyllaDB 5.0 stable branch. Like all past and future 5.x.y releases, it is backward compatible and supports rolling upgrades. Note that the latest stable branch of ScyllaDB is 5.1, and you are encouraged to upgrade. Issues fixed in this release: --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-1-2/226 Title: [RELEASE] ScyllaDB 5.1.2 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.1.2, a bugfix release of the ScyllaDB 5.1 stable branch. ScyllaDB Open Source 5.1.2, like all past and future 5.x.y releases, is backward compatible and supports rolling… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-1-2/226 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.1.2 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.1.2 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.1.2, a bugfix release of the ScyllaDB 5.1 stable branch. ScyllaDB Open Source 5.1.2, like all past and future 5.x.y releases, is backward compatible and supports rolling upgrades. Users are encouraged to upgrade to 5.1.2. Issues fixed in this release: Alternator: regression tests: Stability: use-after-free in db/view row_locker #12168 Stability: [index cache] Multiple 800ms+ reactor stalls during nodetool cleanup (with index cache disabled) due to the local index cache not evicted gently when _upper_bound existed. #12271 CQL: in Scylla, returning incomplete results when using paging #12361 (introduced in 5.1.0 and 5.07). The issue manifests only in specific circumstances: when a page starts with a dead row. This page (and consequently the query) will be wrongfully terminated immediately. In 5.1 this has to be a row covered by the partition tombstone. --- ### Page: https://forum.scylladb.com/t/scylladb-rust-driver-0-7-0-released/227 Title: ScyllaDB Rust Driver 0.7.0 released! - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce ScyllaDB Rust Driver 0.7.0, an asynchronous CQL driver for Rust, optimized for Scylla, but also compatible with Apache Cassandra! Some interesting statistics: over 67k downloa… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-rust-driver-0-7-0-released/227 ## Headings Structure: H1: ScyllaDB Rust Driver 0.7.0 released! H2: Notable changes H3: Related topics ## Main Content: H1: ScyllaDB Rust Driver 0.7.0 released! H2: Notable changes H3: Related topics The ScyllaDB team is pleased to announce ScyllaDB Rust Driver 0.7.0, an asynchronous CQL driver for Rust, optimized for Scylla, but also compatible with Apache Cassandra! Some interesting statistics: New features/enhancements: CI / developer tools: Congrats to all contributors and thanks everyone for using our driver! The source code of the driver can be found here: The official crates.io registry entry is here: Thank you for your attention, please do not hesitate to contact us if you have any questions, issues, feature requests, or are simply interested in our driver! Contributors since the last release: --- ### Page: https://forum.scylladb.com/t/release-scylladb-operator-v1-8-0-rc-0/230 Title: [RELEASE] ScyllaDB Operator v1.8.0-rc.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of scylla-operator v1.8.0-rc.0 :rocket: We’ll welcome your feedback on the release candidate. Release notes are available on If you haven’t heard about the operato… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-operator-v1-8-0-rc-0/230 ## Headings Structure: H1: [RELEASE] ScyllaDB Operator v1.8.0-rc.0 H3: Release v1.8.0-rc.0 · scylladb/scylla-operator H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Operator v1.8.0-rc.0 H3: Release v1.8.0-rc.0 · scylladb/scylla-operator H3: Related topics The ScyllaDB team is pleased to announce the release of scylla-operator v1.8.0-rc.0 We’ll welcome your feedback on the release candidate. Release notes are available on Release notes for 1.8.0-rc.0 Container images docker.io/scylladb/scylla-operator:1.8.0-rc.0 Changes By Kind (since 1.7.4) Feature Expose node-exporter metrics in Scylla Pods (#952,@zimnx) Add hos... If you haven’t heard about the operator yet, here are some links to get you started: https://operator.docs.scylladb.com/ --- ### Page: https://forum.scylladb.com/t/no-way-to-downgrade-from-scylla-5-1/231 Title: No way to downgrade from Scylla 5.1? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, I tried to downgrade from Scylla 5.1 to 5.0. As there already are ME sstables written, I tried to start the scylla 5.1 nodes with “sstable_format: md”. But unfortunetaly I am getting an error when starting scylla ba… Language: en Canonical URL: https://forum.scylladb.com/t/no-way-to-downgrade-from-scylla-5-1/231 ## Headings Structure: H1: No way to downgrade from Scylla 5.1? H3: Related topics ## Main Content: H1: No way to downgrade from Scylla 5.1? H3: Related topics I tried to downgrade from Scylla 5.1 to 5.0. As there already are ME sstables written, I tried to start the scylla 5.1 nodes with “sstable_format: md”. But unfortunetaly I am getting an error when starting scylla back up: scylla[3148642]: [shard 0] init - Startup failed: std::runtime_error (Feature ‘ME_SSTABLE_FORMAT’ was previously enabled in the cluster but its support is disabled by this node. Set the corresponding configuration option to enable the support for the feature.) Should this setting simply prevent scylla from writing me-sstables, but still be able to read them? Is there any other way to downgrade back to 5.0? I’d try to fix the problem with 5.1 and stay on 5.1 if possible. Since with new sstables already in cluster you really have just 2 options I see. I guess I won’t revert back then. My expectation was that downgrading using the sstable setting should be possible. But I guess I have to remember to specify that right with the deployment to avoid sstable upgrades. By downgrading do you mean downgrading a cluster that was completely upgraded to newer version or rollbacking a node that is part of a cluster that still under upgrade procedure? If the former, downgrade is NOT a supported option at all. If the latter, it should work. However, both 5.0 and 5.1 are EOL OSS releases. --- ### Page: https://forum.scylladb.com/t/whats-the-best-fastest-way-to-figure-if-a-table-has-any-data-in-it/234 Title: What's the best/fastest way to figure if a table has any data in it? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Assuming I need to pick all the tables with data in them (at least one row) What’s the best/fastest way to figure that out ? naive approach would be getting all keyspaces and tables names SELECT keyspace_name, table_n… Language: en Canonical URL: https://forum.scylladb.com/t/whats-the-best-fastest-way-to-figure-if-a-table-has-any-data-in-it/234 ## Headings Structure: H1: What's the best/fastest way to figure if a table has any data in it? H3: Related topics ## Main Content: H1: What's the best/fastest way to figure if a table has any data in it? H3: Related topics Assuming I need to pick all the tables with data in them (at least one row) What’s the best/fastest way to figure that out ? naive approach would be getting all keyspaces and tables names SELECT keyspace_name, table_name, compaction FROM system_schema.tables and run this on each: SELECT * FROM {keyspace}.{table} LIMIT 1 it might become a bit problematic when there’s lots of tombstones involves, and the harmless looking query might become quite slow or event timeout. is there a better way getting this information ? If by “tables with data” you mean tables with at least one live row, then there’s no other way than running a query. You cannot avoid processing tombstones, as they may affect liveness of a row. You also have to reconcile writes from all replicas, so you should run the query with CL=ALL. A row may be live on one replica, but a tombstone from another replica may cover it. --- ### Page: https://forum.scylladb.com/t/reshape-during-node-restart-using-too-much-disk-space/236 Title: Reshape during node restart using too much disk space? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I noticed on my nodes while upgrading from 5.1.0 to 5.1.2 that on one node it decided to do a reshape operation (that took a couple hours). During this I happened to be just df’ing and noticed it used quite a bit of dis… Language: en Canonical URL: https://forum.scylladb.com/t/reshape-during-node-restart-using-too-much-disk-space/236 ## Headings Structure: H1: Reshape during node restart using too much disk space? H3: Related topics ## Main Content: H1: Reshape during node restart using too much disk space? H3: Related topics I noticed on my nodes while upgrading from 5.1.0 to 5.1.2 that on one node it decided to do a reshape operation (that took a couple hours). During this I happened to be just df’ing and noticed it used quite a bit of disk shape, roughly double the normal disk space: pre-reshape: /dev/nvme1n1 3661060799 1318457771 2342603028 37% /var/lib/scylla during reshape, near the end of it: /dev/nvme1n1 3661060799 2582195011 1078865788 71% /var/lib/scylla My concern is what would have happened if my node started out at over 50% disk used? Would the reshape be able to complete? I’m using only LCS compaction strategy, so I thought it would be “safe” to go over 50% disk used, but is that an incorrect assumption? That’s a problem we need to address. Brian, could you please open a github issue with all this info? Link to github issue: Reshape during node restart using too much disk space? · Issue #12495 · scylladb/scylladb · GitHub --- ### Page: https://forum.scylladb.com/t/5-factors-when-selecting-a-high-performance-low-latency-database/237 Title: 5 Factors when Selecting a High Performance, Low Latency Database - Blog Posts - ScyllaDB Community NoSQL Forum Meta Description: How to Tell When a Database is Right for Your Project When you are selecting databases for your latest use case (or replacing one that’s not meeting your current needs), the good news these days is that you have a lot… Language: en Canonical URL: https://forum.scylladb.com/t/5-factors-when-selecting-a-high-performance-low-latency-database/237 ## Headings Structure: H1: 5 Factors when Selecting a High Performance, Low Latency Database H2: How to Tell When a Database is Right for Your Project H3: Related topics ## Main Content: H1: 5 Factors when Selecting a High Performance, Low Latency Database H2: How to Tell When a Database is Right for Your Project H3: Related topics When you are selecting databases for your latest use case (or replacing one that’s not meeting your current needs), the good news these days is that you have a lot of options to choose from. Of course, that’s also the bad news. You have a lot to sort through. There are far more databases to consider and compare than ever before. In December 2012, the end of the first year DB-Engines.com first began ranking databases, they had a list of 73 systems (up significantly from the 18 they first started their list with). As of December 2022, they are just shy of 400 systems. This represents a Cambrian explosion of database technologies over the past decade. There is a vast sea of options to navigate: SQL, NoSQL, and a mix of “multi-model” databases that can be a mix of both SQL and NoSQL, or multiple data models of NoSQL (combining two or more options: document, key-value, wide column, graph and so on). Further, users should not confuse outright popularity with fitness for their use case. While network effects definitely have advantages (“Can’t go wrong with X if everyone is using it”), it can also lead to groupthink, stifling innovation and competition. In a recent webinar, my colleague Arthur Pesa and I took users through a consideration of five factors that users need to keep foremost when shortlisting and comparing databases. WATCH THE WEBINAR ON DEMAND NOW --- ### Page: https://forum.scylladb.com/t/how-to-get-which-keyspaces-or-tables-or-queries-have-non-token-aware-request/239 Title: How to get which keyspaces or tables or queries have non-token aware request - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Dear All, In scylladb monitoring stack, we saw some “Non-Token Aware Queries” there. Since there are lots of services in the system (most are token aware), we would like to know if there is any log we could enable to ge… Language: en Canonical URL: https://forum.scylladb.com/t/how-to-get-which-keyspaces-or-tables-or-queries-have-non-token-aware-request/239 ## Headings Structure: H1: How to get which keyspaces or tables or queries have non-token aware request H3: Related topics ## Main Content: H1: How to get which keyspaces or tables or queries have non-token aware request H3: Related topics In scylladb monitoring stack, we saw some “Non-Token Aware Queries” there. Since there are lots of services in the system (most are token aware), we would like to know if there is any log we could enable to get which keyspaces or tables or queries have non-token aware request? As far as I know there is no such log currently. We only have a generic counter, which is increased every time a query lands on a coordinator which is not itself a replica too. Is it possible to add this feature or log for identify service quickly? It can be added yes, but we have to carefully consider how we present this information: just logging it could flood the logs. Please open an issue with the feature request. Open new issue here: add more info about non-token aware queries · Issue #12577 · scylladb/scylladb · GitHub Since there is also mutation_data log there (lots of log when set it to trace level), so one possible way is to put this log at debug or trace level, then it will not impact performance by default. User could enable it when they need such info. Other ways are also welcome. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-162-2023-01-08/240 Title: Last week in scylladb.git master (issue #162; 2023-01-08) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the a7c4a129cb…08b3a9c786 range are covered. There were 60 non-merge commits from 13 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-162-2023-01-08/240 ## Headings Structure: H1: Last week in scylladb.git master (issue #162; 2023-01-08) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #162; 2023-01-08) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the a7c4a129cb…08b3a9c786 range are covered. There were 60 non-merge commits from 13 authors in that period. Some notable commits: Raft has graduated from being an experimental feature to a fully supported option. Currently is is disabled by default. Raft group 0 verbs are used for metadata management, and so always run on shard 0. It was assumed that registering those RPC verbs only on shard 0 would therefore be sufficient, but that’s not the case on all configurations. This was fixed by registering the verbs on all shards. A build time misconfiguration caused the CQL query parser to be compiled with a low optimization level. This is not important by itself, since generally prepared statements hide the parsing overhead, but the C++ compilation model can cause some functions in unrelated areas to also be compiled with the lower optimization level, reducing performance. This is now fixed. Alternator, ScyllaDB’s implementation of the DynamoDB API, uses the rapidjson library to parse queries presented as JSON. Due to a compile-time misconfiguration, aggressive inlining was not enabled. This is now fixed, yielding a substantial performance improvement. ScyllaDB sometimes reads ahead of the user request, in order to hide latency. In one case a read-ahead request which timed out caused errors to be emitted, even though this did not affect the query. The errors are now silenced. The scylla-api-client tool is now documented. The tool is suitable for shell automation of the REST API. Repair tasks can now be aborted via the task manager. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-1-4/241 Title: [RELEASE] ScyllaDB Enterprise 2022.1.4 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2022.1.4, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.1. Related Links Get ScyllaDB Enterprise 2022.1.4 (customers only, o… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-1-4/241 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2022.1.4 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2022.1.4 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2022.1.4, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.1. Below is a list of performance and stability improvements and bug fixes, each with an open source reference: Let’s make sure the main release notes page has the link to this (and every) release. Thanks! --- ### Page: https://forum.scylladb.com/t/scylladb-university-s-journey-featured-at-the-oeb-conference-in-berlin/243 Title: ScyllaDB University’s Journey Featured at the OEB Conference in Berlin - Blog Posts - ScyllaDB Community NoSQL Forum Meta Description: by @Guy Shtub Last month, I attended the OEB conference in Berlin, Germany. Its focus was on technology-supported learning and training. I was invited to give a talk about my experience and what I’ve learned from cre… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-university-s-journey-featured-at-the-oeb-conference-in-berlin/243 ## Headings Structure: H1: ScyllaDB University’s Journey Featured at the OEB Conference in Berlin H3: Related topics ## Main Content: H1: ScyllaDB University’s Journey Featured at the OEB Conference in Berlin H3: Related topics Last month, I attended the OEB conference in Berlin, Germany. Its focus was on technology-supported learning and training. I was invited to give a talk about my experience and what I’ve learned from creating ScyllaDB University. In this post, I’ll cover the gist of my talk. In a future post, I’ll share some of my experiences and what I learned at the conference. You can discuss this blog post and ask me questions in the community forum. © OEB Learning Technologies Europe GmbH used with permission Thanks Peter! Feel free to ask questions and discuss the blog post here. --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-1-5/246 Title: [RELEASE] ScyllaDB Enterprise 2022.1.5 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2022.1.5, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.1. Related Links Get ScyllaDB Enterprise 2022.1.5 (customers only, o… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-1-5/246 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2022.1.5 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2022.1.5 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2022.1.5, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.1. Below is a list of performance and stability improvements and bug fixes, each with an open-source reference if available: --- ### Page: https://forum.scylladb.com/t/wanted-to-upgrade-scylladb-2-1-6/248 Title: Wanted to upgrade scylladb 2.1.6 - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: HI All, My scylladb running with single dc OS Version Scylla Version CentOS Linux release 7.2.1511 (Core) 2.1.6-.20180701.7d2150a05 Wanted to upgrade /migrate to latest 5.1 community release, Please help as neith… Language: en Canonical URL: https://forum.scylladb.com/t/wanted-to-upgrade-scylladb-2-1-6/248 ## Headings Structure: H1: Wanted to upgrade scylladb 2.1.6 H3: Related topics ## Main Content: H1: Wanted to upgrade scylladb 2.1.6 H3: Related topics My scylladb running with single dc OS Version Scylla Version CentOS Linux release 7.2.1511 (Core) 2.1.6-.20180701.7d2150a05 Wanted to upgrade /migrate to latest 5.1 community release, Please help as neither i found any software as per the upgrade path mentioned in scylladb website for my running version . or by which way or i can make scylladb upto latest community release. 2.1.6-.20180701.7d2150a05 Hey @Mayankk , see the answer here first. I can think of three options to upgrade from such an old ScyllaDB version: --- ### Page: https://forum.scylladb.com/t/scylladb-summit-agenda-is-now-available/249 Title: ScyllaDB Summit - Agenda is now available - Announcements - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB Summit agenda is now posted: ScyllaDB Summit 2023 | Agenda With 2 days of 30+ sessions, ScyllaDB Summit (free + virtual) is an opportunity to quickly: Discover ScyllaDB’s latest innovations for data-inten… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-summit-agenda-is-now-available/249 ## Headings Structure: H1: ScyllaDB Summit - Agenda is now available H3: Related topics ## Main Content: H1: ScyllaDB Summit - Agenda is now available H3: Related topics The ScyllaDB Summit agenda is now posted: ScyllaDB Summit 2023 | Agenda With 2 days of 30+ sessions, ScyllaDB Summit (free + virtual) is an opportunity to quickly: We’ll also have our Solution Architects available to answer your questions in our ScyllaDB lounge. If you’re an open source user, this is a great chance to get your top questions and challenges addressed by our experts! Plus, you might end up with some cool ScyllaDB Monster swag. --- ### Page: https://forum.scylladb.com/t/java-dependencies/250 Title: Java dependencies - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello Is it possible to get away from the java dependencies (scylla-jmx) if I am not switching from Cassandra to scylla. The scylla cluster will be deployed on the new data? I understand that moving away from java will… Language: en Canonical URL: https://forum.scylladb.com/t/java-dependencies/250 ## Headings Structure: H1: Java dependencies H3: Related topics ## Main Content: H1: Java dependencies H3: Related topics Is it possible to get away from the java dependencies (scylla-jmx) if I am not switching from Cassandra to scylla. The scylla cluster will be deployed on the new data? I understand that moving away from java will take away the operability of nodetool, which runs through jmx and queries data from rest api. Instead of nodetool it will be its own application which directly requests data from the scylla rest api So far, no one volunteered to rewrite nodetool, so we’re stuck with Java. I have seen the sources of nodetool and a complete rewrite will take a long time. Right now I have a limited set of commands (drain,repair,rebuild,removenode,status,version) with few flags. (multidatacenter is not implemented yet) The question is to get rid of java and the question is: can I build scylla cluster from src packages without scylla-jmx dependency? I just want the developers to make sure that scylla can work without java and that the internals of the c++ code don’t overlap with java I don’t need the full functionality of nodetool, for my purposes and database monitoring, scylla-exporter and scylla-api will do P.S. I am rewriting nodetool in Golang , I can share the code later if you are interested Currently, the build system is integrated. It’s an optional dependency however. We’d be thrilled to drop the Java dependency, so please contribute the code if you get it to a working state. We’d need a full replacement, not just for some operations, to accept it however. --- ### Page: https://forum.scylladb.com/t/golang-scylladb-batch-delete-sample-code/252 Title: Golang ScyllaDB Batch delete sample code - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: We are using following git model fo scyllaDB transaction via golang in production github.com/scylladb/gocqlx/v2 We want to implement now batch delete operation in scyllaDB using golang. I have search in google as well … Language: en Canonical URL: https://forum.scylladb.com/t/golang-scylladb-batch-delete-sample-code/252 ## Headings Structure: H1: Golang ScyllaDB Batch delete sample code H3: Related topics ## Main Content: H1: Golang ScyllaDB Batch delete sample code H3: Related topics We are using following git model fo scyllaDB transaction via golang in production github.com/scylladb/gocqlx/v2 We want to implement now batch delete operation in scyllaDB using golang. I have search in google as well the git repo which I have mention above. I am not getting any sample how to do. Can sombody help us with sample code how to do batch delete using golang Requirenent. Could you describe in more detail what is the schema (columns, which columns are partition/clustering keys) of the table? What is the delete pattern: do you want to delete whole partitions, some specific rows? --- ### Page: https://forum.scylladb.com/t/live-in-person-training-at-scylladb-summit-2023-february-16/254 Title: Live In-person Training at ScyllaDB Summit 2023 - February 16 - University and Training - ScyllaDB Community NoSQL Forum Meta Description: Join us for the ScyllaDB Summit 2023 Training Day. It’s a half day of training in the San Francisco Bay Area with drinks, food, and fun, taking place on the 16th of February. We’ll offer two training sessions, which you… Language: en Canonical URL: https://forum.scylladb.com/t/live-in-person-training-at-scylladb-summit-2023-february-16/254 ## Headings Structure: H1: Live In-person Training at ScyllaDB Summit 2023 - February 16 H3: Related topics ## Main Content: H1: Live In-person Training at ScyllaDB Summit 2023 - February 16 H3: Related topics Join us for the ScyllaDB Summit 2023 Training Day. It’s a half day of training in the San Francisco Bay Area with drinks, food, and fun, taking place on the 16th of February. We’ll offer two training sessions, which you can take together or separately: ScyllaDB Core Concepts – ScyllaDB architecture, key components, and data modeling best practices. Power User Best Practices – Deep dives into optimizing performance, troubleshooting, and adopting the latest features. It’s also a chance to network and connect with the ScyllaDB community. Hope to see you there! --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-163-2023-01-15/255 Title: Last week in scylladb.git master (issue #163; 2023-01-15) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 08b3a9c786…abc43f97c9 range are covered. There were 88 non-merge commits from 18 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-163-2023-01-15/255 ## Headings Structure: H1: Last week in scylladb.git master (issue #163; 2023-01-15) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #163; 2023-01-15) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 08b3a9c786…abc43f97c9 range are covered. There were 88 non-merge commits from 18 authors in that period. Some notable commits: The experimental WASM user defined function (UDF) implementation has been switched to Rust (UDFs can be written in any language WASM supports, not just Rust). The new implementation is avoids stalls for long-running UDFs and shares memory with the rest of the database. ScyllaDB tracks transient memory used by queries. However, it did not track decompressed memory buffers, which could lead to running out of memory in some complicated queries. This is now fixed. A Seastar update fixed an ambiguity in the http parser, which caused Alternator requests to be unnecessarily slow. The Azure snitch, used to derive datacenter and rack information from instance metadata, now handles regions which have a single availability zone. The replica-side read metrics, which have been incorrect for some time, have been revamped. A very rare bug involving reads from memtables and multi-version concurrency control has been fixed. The CQL binary protocol versions 1 and 2 are no longer supported. Version 3 and above have been supported for 9 years, so it’s unlikely to be in real use. You can check for version 1 and 2 in the system.clients virtual table. The sstable tools gained Lua scripting. This is an expert feature intended for offline analysis of sstables. User defined aggregates (UDAs) are now correctly persisted and survive a cold start. UDAs are an experimental feature. There is now documentation about how NULL is treated in ScyllaDB. The schema of the system.raft_config table, used for storing Raft cluster membership data, has been streamlined. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/how-does-the-number-of-levels-affect-leveled-compaction-strategy-lcs/256 Title: How does the number of levels affect Leveled Compaction Strategy (LCS)? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I understand how LCS works in databases like ScyllaDB, RocksDB, Cassandra, and so on. I also know that different algorithms use different numbers of levels for compaction. How does the number of levels affect the compa… Language: en Canonical URL: https://forum.scylladb.com/t/how-does-the-number-of-levels-affect-leveled-compaction-strategy-lcs/256 ## Headings Structure: H1: How does the number of levels affect Leveled Compaction Strategy (LCS)? H3: Related topics ## Main Content: H1: How does the number of levels affect Leveled Compaction Strategy (LCS)? H3: Related topics I understand how LCS works in databases like ScyllaDB, RocksDB, Cassandra, and so on. I also know that different algorithms use different numbers of levels for compaction. How does the number of levels affect the compaction? Is it possible to use just two levels? *Based on a question originally asked on Stack Overflow by Bishnu LCS solves the space amplification problem of Size Tiered Compaction Strategy (STCS). It also reduces read amplification (the average number of disk reads required per read request). LCS divides the small SSTables (“fragments”) into levels: Level 0 (L0) includes the new SSTables recently flushed from the Memtables. As their number grows (and reads become slower), the goal is to move the SSTables from this level to the next levels (L1, L2, L3, and so on). Each of the other levels is a single run of exponentially increasing size: L1 is a run of 10 SSTables, L2 is a run of 100 SSTables, L3 is a run of 1000 SSTables, and so on. A factor of 10 is the default setting in both Scylla and Apache Cassandra. While the space amplification problem is solved, or at least significantly improved, LCS exacerbates another problem, write amplification. Write amplification is the number of bytes we had to write to disk for each byte of newly flushed SSTable data. Write amplification is always greater than 1.0 because we write each data item to the commit-log and then write it again to an SSTable. Every time compaction involves this data item and copies it to a new SSTable, it’s another write. You can learn more about it here: *Based on an answer on Stack Overflow by TomerSan In ScyllaDB, Leveled compaction (LCS) works very similarly to how it works in Cassandra and Rocksdb (with some minor differences). If you want a short overview of how leveled compaction works in Scylla and why, I suggest you read my Write Amplification in Leveled Compaction blog post. Your specific question on why two levels (L0 of recently flushed SSTables, Ln of disjoint-range SSTables) are not enough - is an excellent question: The main problem is that a single flushed MemTable (SSTable in L0), containing a random collection of writes, will often intersect all of the SSTables in Ln. This means rewriting the entire database every time there’s a new MemTable flushed, and the result is a super-huge amount of write amplification, which is completely unacceptable. One way to reduce this write amplification significantly (but perhaps not enough) is to introduce a cascade of intermediate levels, L0, L1, …, Ln. The end result is that we have L(n-1), which is 1/10th (say) the size of Ln, and we merge L(n-1) - not a single SSTable - into Ln. This is the approach that leveled compaction strategy (LCS) uses in all systems you mentioned. A completely different approach could be not to merge a single SSTable into Ln, but rather try to collect a large amount of data first and only then merge it into Ln. We can’t just collect 1,000 tables in L0 because this would make reads very slow. Rather, to collect this large amount of data, one could use size-tiered compaction (STCS) inside L0. In other words, this approach is a “mix” of STCS and LCS with two “levels”: L0 uses STCS on new SSTables, and Ln contains a run of SSTables (SSTables with disjoint ranges). When L0 reaches 1/10th (say) the size of Ln, L0 is compacted into Ln. Such a mixed approach could have lower write amplification than LCS, but because most of the data is in a run in Ln, it would have the same low space and read amplifications as in LCS. ScyllaDB implements this idea with Incremental Compaction Strategy (ICS). AFAIK none of the other mentioned databases ( Cassandra or Rocksdb) support such “mixed” compaction. You can read more about ICS in the Maximizing Disk Utilization with Incremental Compaction blog post and in the Incremental Compaction 2.0: A Revolutionary Space and Write Optimized Compaction Strategy blog post. *Based on an answer on Stack Overflow by Nadav Har’El --- ### Page: https://forum.scylladb.com/t/is-there-any-scylla-operator-for-aarch64-image-available/258 Title: Is there any scylla-operator for aarch64 image available? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Scylladb already support cpu aws graviton 2. but I don’t see any scylla-operator or scylla-manager-agent image that can run on aarch64 so when I deployed it to my kubernetes cluster it failed. so can you help me how to d… Language: en Canonical URL: https://forum.scylladb.com/t/is-there-any-scylla-operator-for-aarch64-image-available/258 ## Headings Structure: H1: Is there any scylla-operator for aarch64 image available? H3: Related topics ## Main Content: H1: Is there any scylla-operator for aarch64 image available? H3: Related topics Scylladb already support cpu aws graviton 2. but I don’t see any scylla-operator or scylla-manager-agent image that can run on aarch64 so when I deployed it to my kubernetes cluster it failed. so can you help me how to deploy scylladb on aarch64 using scylla-operator. Thanks There isn’t, but it’s open source so you can always build one --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-1-3/261 Title: [RELEASE] ScyllaDB 5.1.3 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.1.3, a bugfix release of the ScyllaDB 5.1 stable branch. ScyllaDB Open Source 5.1.3, like all past and future 5.x.y releases, is backward compatible and supports rolling… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-1-3/261 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.1.3 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.1.3 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.1.3, a bugfix release of the ScyllaDB 5.1 stable branch. ScyllaDB Open Source 5.1.3, like all past and future 5.x.y releases, is backward compatible and supports rolling upgrades. Users are encouraged to upgrade to 5.1.3. Issues fixed in this release: --- ### Page: https://forum.scylladb.com/t/scylladb-codefirst/262 Title: ScyllaDB CodeFirst? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, I’m .NET dev that is trying out ScyllaDB. My previous experience was with MSSQL and EF. I used code first and DbMigrations for updating my Database from code. While testing/playing around with ScyllaDB I used to cre… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-codefirst/262 ## Headings Structure: H1: ScyllaDB CodeFirst? H3: Related topics ## Main Content: H1: ScyllaDB CodeFirst? H3: Related topics Hi, I’m .NET dev that is trying out ScyllaDB. My previous experience was with MSSQL and EF. I used code first and DbMigrations for updating my Database from code. While testing/playing around with ScyllaDB I used to create new tables with cqlsh like: CREATE TABLE heartrate_v1 ( pet_chip_id uuid, time timestamp, heart_rate int, PRIMARY KEY (pet_chip_id) ); which is database first approach. But is there anyway/tool. To work with ScyllaDB as CodeFirst. So when I make new object in C# code, it make new table in Keyspace, and if I add new prop to that object, it add new column to existing table? Hi @Bruce_Haset_Lee, I’m not familiar with CodeFirst but if you looking for ORM for C# you can try datastax C# driver [1] which ScyllaDB is fully compatible with, [1] GitHub - datastax/csharp-driver: DataStax C# Driver for Apache Cassandra --- ### Page: https://forum.scylladb.com/t/replication-strategy-change-from-simple-to-network/265 Title: Replication Strategy Change from Simple to Network - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, I have several Scylla clusters that have been running in production for multiple years. I realized last year that we still use the default SimpleStrategy for replication on all of our data. This is obviously a bad ch… Language: en Canonical URL: https://forum.scylladb.com/t/replication-strategy-change-from-simple-to-network/265 ## Headings Structure: H1: Replication Strategy Change from Simple to Network H3: Related topics ## Main Content: H1: Replication Strategy Change from Simple to Network H3: Related topics Hi, I have several Scylla clusters that have been running in production for multiple years. I realized last year that we still use the default SimpleStrategy for replication on all of our data. This is obviously a bad choice for production, and I have switched over what keyspaces I could using the recommended procedure. The problem I have however, is that our largest keyspaces are 100s of TBs, and just running some exploratory repairs gives me an estimate of this whole procedure taking ~40 hours. My same estimates for the smaller keyspaces ran over almost x2, so this could take up to 80 hours. I cannot get that long of a downtime, because it will require a full company outage. So, I’m exploring other options. The main alternatives that I’ve been able to come up are not great, so hoping that maybe there is some feature of Scylla that I’m missing that will make these easier, or some other method to accomplish this: In addition to my ideas here, I got the following from Felipe on Slack: What is the nodetool status of that problematic cluster? (you can ping it on slack to me as well (Lubos) to keep privacy) Lots of things can be done(or can be skipped) if the layout of the cluster is prepared for the switch. So Garrett has 3 racks, lots of nodes and most nodes only 15% disk free and using Scylla 4.5 with STCS as compaction strategy Garret also uses nodetool repair -pr for repairs in such case having Scylla Enterprise with ICS and Scylla Manager(SM) would mitigate the situation a bit Anyhow Felippes idea is safe. If you however feel adventurous you can go the route that you repair the cluster, then if you have SM you can resume repair after writes are paused (or run repair anew). Then after ALTER if you resume writes you should have them consistent. But reads will still be inconsistent, roughly 11% of the data will be on RF=1 until you repair it. So until this repair is done, the reads will be eventually consistent despite using QUORUM. If one can live with that for the duration of second repair, you can have write downtime ideally for single repair (or for resume from SM) Speed up of repair can be done by using NullCompactionStrategy - assuming there is enough disk space to accommodate streamed repairs and writes for the duration of repair - which with 15% free won’t be flying. Anyhow, I am curious how Garrett will decide --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-0-9/266 Title: [RELEASE] ScyllaDB 5.0.9 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.0.9, a bugfix release of the ScyllaDB 5.0 stable branch. Like all past and future 5.x.y releases, it is backward compatible and supports rolling upgrades. Note that the… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-0-9/266 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.0.9 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.0.9 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.0.9, a bugfix release of the ScyllaDB 5.0 stable branch. Like all past and future 5.x.y releases, it is backward compatible and supports rolling upgrades. Note that the latest stable branch of ScyllaDB is 5.1, and you are encouraged to upgrade. Issues fixed in this release: --- ### Page: https://forum.scylladb.com/t/scylladb-summit-2023-agenda-announced/267 Title: ScyllaDB Summit 2023 Agenda Announced - Blog Posts - ScyllaDB Community NoSQL Forum Meta Description: If you haven’t already, take a moment to register for ScyllaDB Summit 2023 right now. Because it’s an online event you won’t want to miss! REGISTER NOW FOR SCYLLADB SUMMIT [THIS IS AN EXCERPT. READ THE FULL BLOG POS… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-summit-2023-agenda-announced/267 ## Headings Structure: H1: ScyllaDB Summit 2023 Agenda Announced H3: Related topics ## Main Content: H1: ScyllaDB Summit 2023 Agenda Announced H3: Related topics If you haven’t already, take a moment to register for ScyllaDB Summit 2023 right now. Because it’s an online event you won’t want to miss! REGISTER NOW FOR SCYLLADB SUMMIT [THIS IS AN EXCERPT. READ THE FULL BLOG POST] The summit is taking place later this week! Hope to see you there, register here. --- ### Page: https://forum.scylladb.com/t/scylladb-enterprise-release-2022-2-0/269 Title: ScyllaDB Enterprise Release 2022.2.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of ScyllaDB Enterprise 2022.2, a production-ready ScyllaDB Enterprise Feature release. ScyllaDB Enterprise 2022.1 is the latest Long-Term Support (LTS) release. More… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-enterprise-release-2022-2-0/269 ## Headings Structure: H1: ScyllaDB Enterprise Release 2022.2.0 H2: New Features H3: Alternator TTL H3: Limit partition access rate H3: Load and stream H3: Materialized Views: Prune H3: Performance: Eliminate exceptions from the read and write path H3: Deprecated Features H2: Updates in this Release H3: Deployment and Packaging H3: CQL API updates H3: Stability and Performance Improvements H3: Tooling H3: Storage H3: Configuration H3: Monitoring and Tracing H3: Bug Fixes H3: Related topics ## Main Content: H1: ScyllaDB Enterprise Release 2022.2.0 H2: New Features H3: Alternator TTL H3: Limit partition access rate H3: Load and stream H3: Materialized Views: Prune H3: Performance: Eliminate exceptions from the read and write path H3: Deprecated Features H2: Updates in this Release H3: Deployment and Packaging H3: CQL API updates H3: Stability and Performance Improvements H3: Tooling H3: Storage H3: Configuration H3: Monitoring and Tracing H3: Bug Fixes H3: Related topics The ScyllaDB team is pleased to announce the release of ScyllaDB Enterprise 2022.2, a production-ready ScyllaDB Enterprise Feature release. ScyllaDB Enterprise 2022.1 is the latest Long-Term Support (LTS) release. More information on the new ScyllaDB Long-Term Support (LTS) policy is available here. The ScyllaDB Enterprise 2022.2 release is based on ScyllaDB Open Source 5.1, and introduces Partition level rate limit, Alternator TTL, and more functional, performance and stability improvements. ScyllaDB Enterprise customers are encouraged to upgrade to ScyllaDB Enterprise 2022.2, and are welcome to contact our Support Team with questions. In Scylla 5.0 we introduced Time To Live (TTL) to the Amazon DynamoDB compatible API (Alternator) as an experimental feature. In ScyllaDB Enterprise 2022.2 we promote it to production ready. As in DynamoDB, Alternator items that are set to expire at a specific time will not disappear precisely at that time but only after some delay. DynamoDB guarantees that the expiration delay will be less than 48 hours (though for small tables, the delay is often much shorter). In Alternator, the expiration delay is configurable - it defaults to 24 hours but can be set with the --alternator-ttl-period-in-seconds configuration option. More Alternator updates in this release: Alternator, ScyllaDB’s implementation of the DynamoDB API, now provides the following improvements: An improvement to BatchGetItem performance by grouping requests to the same partition.#10757 It is now possible to limit read rates and writes rates into a partition with a new WITH per_partition_rate_limit clause for the CREATE TABLE and ALTER TABLE statements. This is useful to prevent hot-partition problems when high rate reads or writes are bogus (for example, arriving from spam bots). #4703 Limits are configured separately for reads and writes. Some examples: ALTER TABLE t WITH per_partition_rate_limit = { ‘max_reads_per_second’: 100, ‘max_writes_per_second’: 200 Limit reads only, no limit for writes: ALTER TABLE t WITH per_partition_rate_limit = { ‘max_reads_per_second’: 200 This feature extends nodetool refresh to allow loading arbitrary sstables that are not owned by a particular node into the cluster. It loads the sstables from the disk, calculates the data’s owning nodes, and automatically streams the data to the owning nodes. In particular, this is useful when restoring a cluster from a backup to a new cluster with a different number of nodes. One can copy the sstables from the old cluster to the new nodes and trigger the load and stream process. This can make restores and migrations much easier: Load_and_stream option also updates the relevant Materialized Views #9205 curl -X POST "http://{local-ip}:10000/storage_service/sstables/{keyspace}?cf={table}&load_and_stream=true Note there is an open bug, #282 , for the Nodetool refresh --load-and-stream operation. Until it is fixed, use the REST API above. A new CQL extension PRUNE MATERIALIZED VIEW statement can now be used to remove inconsistent rows from materialized views. A special statement is dedicated for pruning ghost rows from materialized views. A ghost row is an inconsistency issue that manifests itself by having rows in a materialized view which do not correspond to any base table rows. Such inconsistencies should be prevented altogether and ScyllaDB strives to avoid them, but if they happen, this statement can be used to restore a materialized view to a fully consistent state without rebuilding it from scratch. When a coordinator times out, it generates an exception which is then caught in a higher layer and converted to a protocol message. Since exceptions are slow, this can make a node that experiences timeouts even slower. To prevent that, the coordinator write path and read path has been converted not to use exceptions for timeout cases, treating them as another kind of result value instead. Further work on the read path and replica reduces the timeout cost, so goodput is preserved while a node is overloaded. Improvement results below The following features are deprecated and won’t be available in the following Enterprise releases: Please contact Scylla Support for advice if you use any of these features. A list of CQL bug fixes and extensions: The LIKE operator on descending order clustering keys now works. #10183 ScyllaDB would incorrectly use an index with some IN queries, leading to incorrect results. This is now fixed. In CREATE AGGREGATE statements, the INITCOND and FINALFUNC clauses are now optional (defaulting to NULL and the identity function respectively). CREATE KEYSPACE now has a WITH STORAGE clause, allowing to customize where data is stored. For now, this is only a placeholder for future extensions. When talking to drivers using the older v3 protocol, ScyllaDB did not serialize timeout exceptions correctly, resulting in the driver complaining about protocol violations. This is now fixed. #5610 The CQL grammar was relaxed to allow bind markers in collection literals, e.g. UPDATE tab SET my_set = { ?, ‘foobar’, :variable }. ScyllaDB now validates collections for NULLs more carefully. #10580 After this change, the following query INSERT INTO ks.t (list_column) VALUES (?); And the driver sending a list with null inside as the bound value, something like [1, 2, null, 4] Would result in an invalid_request_exception instead of an ugly marshaling error. The tool can be used to list the different API functions and their parameters, and to print detailed help for each function. Then, when invoking any function, scylla-api-cli performs basic validation on the function arguments and prints the result to the standard output. Note that json results msy be pretty-printed using commonly available command line utilities. It is recommended to use scylla-api-cli for interactive usage of the REST API over plain http tools, like curl, to prevent human errors. It is now possible to limit, and control in real time, the bandwidth of streaming and compaction. These and more configuration updates below: Scylla Monitoring Stack 4.1 and later includes a dashboard for ScyllaDB Enterprise 2022.2. Below are a list of monitoring and tracing related work in this release: For a full list of fixed issues see git log and 5.1 release candidates notes. --- ### Page: https://forum.scylladb.com/t/how-the-in-query-works-internally-and-when-to-use-it/271 Title: How the IN query works internally and when to use it - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I have one question regarding “IN CLAUSE.” How does it work internally? Example: My user table has the following PRIMARY KEY: PRIMARY KEY((tenant_id, user_id)) My Requirement is Get multiple users from the users table … Language: en Canonical URL: https://forum.scylladb.com/t/how-the-in-query-works-internally-and-when-to-use-it/271 ## Headings Structure: H1: How the IN query works internally and when to use it H3: Related topics ## Main Content: H1: How the IN query works internally and when to use it H3: Related topics I have one question regarding “IN CLAUSE.” How does it work internally? Example: My user table has the following PRIMARY KEY: PRIMARY KEY((tenant_id, user_id)) My Requirement is Get multiple users from the users table using tenant_id and user_id. I can use IN QUERY, e.g.: select * from user where tenant_id=‘testtenant’ and user_id in (‘64d73074-1bd1-486e-b717-265b1d46c022’,‘64d73074-1bd1-486e-b717-265b1d46c022’); I can get individual users by using the partition key. (+Redis caching of frequent users)I want to understand how “in query” works internally. If you have any posts, that would be useful. What approach is better for fetching such data if: *Originally asked on the User Slack channel. The coordinator breaks up the query and sends individual ones for each partition key. It’s generally better to send separate queries if they have different partition keys but use IN if everything is in the same partition. So it’s more of a convenience - the coordinator handles sending queries concurrently for you. --- ### Page: https://forum.scylladb.com/t/what-is-incremental-compaction-strategy-ics-how-does-it-work-and-what-are-its-benefits/272 Title: What is Incremental Compaction Strategy (ICS), how does it work, and what are its benefits? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: When should one use ICS vs STCS? Language: en Canonical URL: https://forum.scylladb.com/t/what-is-incremental-compaction-strategy-ics-how-does-it-work-and-what-are-its-benefits/272 ## Headings Structure: H1: What is Incremental Compaction Strategy (ICS), how does it work, and what are its benefits? H3: Related topics ## Main Content: H1: What is Incremental Compaction Strategy (ICS), how does it work, and what are its benefits? H3: Related topics When should one use ICS vs STCS? ICS combines the best of Leveled Compaction Strategy (LCS) and Size-tiered Compaction Strategy (STCS). It utilizes the advantages of the two strategies to benefit from the best of both worlds. STCS needs a lot of temporary space. A rule of thumb is about 50% free disk space. This is true for ScyllaDB as well as other databases like Datastax Enterprise. LCS solves that problem, but it introduces another problem: its write amplification. ICS splits each large SSTable into an SSTable run with fragments. Each fragment is a roughly fixed-size SSTable, and it holds a unique range of keys, a portion of the whole SSTable run. So instead of writing one big SSTable, ICS writes one big SSTable run that is composed of many small SSTables that have the same size (fragments). This means that instead of having to release the space only after large SSTables are compacted, leading to a worst case of 50% space amplification, in ICS, the space can be released after each fragment is compacted. This means that the space amplification is significantly reduced. A rule of thumb is to leave about 30% free disk space. Let’s look at an example of compacting two SSTables runs holding 7GB each, using 7 x 1GB SSTables: instead of writing up to 14GB into a single output SSTable file, we’ll break the output SSTable into a run of up to 14 x 1GB fragments (fragment size is 1GB by default). This new compaction approach takes runs as input and consequently outputs a new run, which is composed of one or more fragments. Also, the compaction procedure is modified to release an input fragment as soon as all of its data is safe in a new output fragment. For example, when compacting 2 SSTables together, each 100GB in size, the worst-case temporary space requirement with STCS would be 200G. On the other hand, ICS would have a worst-case requirement of roughly 2G (with the default fragment size of 1G) for precisely the same scenario. That’s because, with the incremental compaction approach, those 2 SSTables would be actually 2 SSTable runs, each 100 fragments long, making it possible to roughly release one input SSTable fragment for each new output SSTable fragment, both of which are 1GB in size. How much space is actually saved? To calculate the worst-case space requirement for ICS, you need to multiply the maximum number of ongoing compactions by the space overhead for a single compaction job. The maximum number of ongoing compactions can be figured out by multiplying the number of shards by log4 of (disk size per shard). As an example, on a setup with ten shards and a 1TB disk, the maximum number of compactions will be 33 (10 * log4(1000/10)), which results in a worst-case space requirement of 66GB. This means that 93% of the disk space can be used, given that compaction would temporarily increase usage by 6.6% at most. Keep in mind that in practice, additional disk space should be reserved for system usage, like commitlog, for example. Because the space overhead is a logarithmic function of the disk size, if you increase the disk size, from the previous example, by, say, a factor of 10 to 10TB, the compaction will temporarily increase space usage by about 2% at most. To summarize, using ICS saves you money because you don’t have to set aside 50% of the total disk space for compaction. It provides the same low write amplification as STCS. It also has the same read amplification as STCS, even though the number of SSTables in a table is increased compared to STCS. So with ICS, you get the benefits of STCS without the cost of 50% disk space. Additional Resources: Thanks @bhalevy . I’ll emphasize that the 30% free disk space above is a rule of thumb. Theoretically, systems can be pushed to 80% (or higher) with certain parameters, but I’d recommend indeed keeping it in the 70s. Especially taking into account other things like snapshots, commit logs growing (when they are kept in the same file system as data), etc. Also, keep in mind that the above recommendation is for systems that have replaced all tables from STCS to ICS. If some tables are using STCS and some are using ICS (say the compaction strategy is being changed), the calculation becomes more complex. Also, in terms of performance, it’s good not to allow disk usage to go to the edge. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-164-2023-01-22/273 Title: Last week in scylladb.git master (issue #164; 2023-01-22) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the abc43f97c9…ebc100f74f range are covered. There were 121 non-merge commits from 20 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-164-2023-01-22/273 ## Headings Structure: H1: Last week in scylladb.git master (issue #164; 2023-01-22) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #164; 2023-01-22) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the abc43f97c9…ebc100f74f range are covered. There were 121 non-merge commits from 20 authors in that period. Some notable commits: Automatically parallelized query aggregation used a node-local clock for timeouts, rather than the standard clock. This meant that automatically parallelized queries would fail, unless the nodes were started at the same time. This is now fixed. It is now possible to replace a node by mentioning its host id, rather then its IP address. This is useful in a container environments, where IP addresses are transient. “Unset” values are an obscure prepared statement feature that allows only some columns in an UPDATE or INSERT statement to be modified. It was a source of minor bugs and inconvenience in code. The feature has been refactored so it has less impact on the code and is more robust. USING TIMESTAMP allows setting the mutation timestamp on a CQL statement level. It has a sanity check that prevents setting timestamps in the future, as these can be hard to delete, but sometimes one wishes to do so anyway. There is now a configuration option that allows disabling the feature. During startup, ScyllaDB makes sstables conform to the compaction strategy in a process called reshaping, so that future reads will perform will. It is now more careful when reshaping Leveled Compaction Strategy tables, to avoid doing unnecessary work. Lightweight transactions are now more robust when the schema is changed during a transaction. Raft group 0 (responsible for managing topology and schema) now has improved availability during removenode and decommission. Alternator, ScyllaDB’s implementation of the DynamoDB protocol, has better validation of malformed base64 encoded values. The development version number was updated to 5.3.0-dev, marking the beginning of the 5.2 release stabilization cycle. ScyllaDB carefully measures the memory consumed by queries, and tries to ensure it will not exceed available memory. However, a query’s memory can grow after it already started. If this happens to all concurrently running queries, we may run out. To prevent this, two new safeguards are added: first, when one memory threshold is passed, we pause all queries except one with the intent of completing this one query and releasing memory. If this doesn’t help and memory grows even further, we fail all other queries with the intent of letting one succeed, with the rest retried later. Alternator table name validation has been optimized. The ScyllaDB source base contains several performance microbenchmarks. These are now integrated into the main Scylla binary as subcommands, so they can be run on any machine where ScyllaDB is installed e.g. scylla perf_simple_query. CQL table columns that have the list data type aren’t allowed to contain NULLs, but in certain situations list values in CQL literals or bind variables are allowed to contain NULLs (for example, in LWT IF conditions that use the IN operator). The type system was relaxed to accept NULLs where this is allowed. Previously, these cases were handled by hard-to-maintain workarounds. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/new-kafka-and-scylladb-lab-on-scylladb-university/274 Title: New Kafka and ScyllaDB lab on ScyllaDB University - University and Training - ScyllaDB Community NoSQL Forum Meta Description: I recently published a new lab on ScyllaDB University, on using Apache Kafka with ScyllaDB. In the lab, you’ll learn how to use the Scylla CDC source connector to push change events in a ScyllaDB cluster to a Kafka serv… Language: en Canonical URL: https://forum.scylladb.com/t/new-kafka-and-scylladb-lab-on-scylladb-university/274 ## Headings Structure: H1: New Kafka and ScyllaDB lab on ScyllaDB University H3: Related topics ## Main Content: H1: New Kafka and ScyllaDB lab on ScyllaDB University H3: Related topics I recently published a new lab on ScyllaDB University, on using Apache Kafka with ScyllaDB. In the lab, you’ll learn how to use the Scylla CDC source connector to push change events in a ScyllaDB cluster to a Kafka server. The main steps in the lab are: Change Data Capture, or CDC, is a feature that allows you to query the history of recent changes made to the table. CDC allows users to build streaming data pipelines that enable real-time data processing and analysis and immediately react to data changes occurring in the database. In the lab, you’ll see this hands-on. Additional Resources: Any questions or input? --- ### Page: https://forum.scylladb.com/t/my-experience-at-the-oeb-conference-in-berlin/277 Title: My Experience at the OEB Conference in Berlin - Blog Posts - ScyllaDB Community NoSQL Forum Meta Description: I recently attended the OEB Learning Technologies conference in Berlin. I was invited to speak about ScyllaDB’s journey in developing ScyllaDB University. I had a great experience and learned quite a bit. Over 2000 pe… Language: en Canonical URL: https://forum.scylladb.com/t/my-experience-at-the-oeb-conference-in-berlin/277 ## Headings Structure: H1: My Experience at the OEB Conference in Berlin H3: Related topics ## Main Content: H1: My Experience at the OEB Conference in Berlin H3: Related topics I recently attended the OEB Learning Technologies conference in Berlin. I was invited to speak about ScyllaDB’s journey in developing ScyllaDB University. I had a great experience and learned quite a bit. Over 2000 people attended, and there were dozens of talks and an exhibition with booths. Some talks focused on Academia, while others focused on corporate learning and training. It was interesting to see what my peers are doing and to learn about emerging technologies in this space and network. Also, Berlin is a fun city and beautiful at this time of year. Check out more impressions in the full blog post. --- ### Page: https://forum.scylladb.com/t/how-to-use-change-data-capture-with-apache-kafka-and-scylladb/282 Title: How to Use Change Data Capture with Apache Kafka and ScyllaDB - Blog Posts - ScyllaDB Community NoSQL Forum Meta Description: Today’s blog is taken from a lab in ScyllaDB University. Login or register now for ScyllaDB University where you can take this lab and the entire course and get credit for it free online, plus have access to all our o… Language: en Canonical URL: https://forum.scylladb.com/t/how-to-use-change-data-capture-with-apache-kafka-and-scylladb/282 ## Headings Structure: H1: How to Use Change Data Capture with Apache Kafka and ScyllaDB H3: Related topics ## Main Content: H1: How to Use Change Data Capture with Apache Kafka and ScyllaDB H3: Related topics Today’s blog is taken from a lab in ScyllaDB University. Login or register now for ScyllaDB University where you can take this lab and the entire course and get credit for it free online, plus have access to all our other free online courseware. TAKE THE LAB IN SCYLLADB UNIVERSITY [ALSO USE THIS THREAD TO DISCUSS THE LAB] --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-1-4/284 Title: [RELEASE] ScyllaDB 5.1.4 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.1.4, a bugfix release of the ScyllaDB 5.1 stable branch. ScyllaDB Open Source 5.1.4, like all past and future 5.x.y releases, is backward compatible and supports rolling… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-1-4/284 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.1.4 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.1.4 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.1.4, a bugfix release of the ScyllaDB 5.1 stable branch. ScyllaDB Open Source 5.1.4, like all past and future 5.x.y releases, is backward compatible and supports rolling upgrades. Users are encouraged to upgrade to 5.1.4. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/scylladb-summit-for-the-scylladb-curious-serious-sea-monsters/285 Title: ScyllaDB Summit: For the ScyllaDB Curious + Serious Sea Monsters - Blog Posts - ScyllaDB Community NoSQL Forum Meta Description: [800x400-blog-summit-2023-curious-veterans] Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-summit-for-the-scylladb-curious-serious-sea-monsters/285 ## Headings Structure: H1: ScyllaDB Summit: For the ScyllaDB Curious + Serious Sea Monsters H3: ScyllaDB Summit: For the ScyllaDB Curious + Serious Sea Monsters H3: Related topics ## Main Content: H1: ScyllaDB Summit: For the ScyllaDB Curious + Serious Sea Monsters H3: ScyllaDB Summit: For the ScyllaDB Curious + Serious Sea Monsters H3: Related topics A rundown of the ScyllaDB Summit tech talks that are geared specifically toward ScyllaDB users and the “ScyllaDB curious.” --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-165-2023-01-29/286 Title: Last week in scylladb.git master (issue #165; 2023-01-29) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the ebc100f74f…5eadea301e range are covered. There were 38 non-merge commits from 12 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-165-2023-01-29/286 ## Headings Structure: H1: Last week in scylladb.git master (issue #165; 2023-01-29) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #165; 2023-01-29) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the ebc100f74f…5eadea301e range are covered. There were 38 non-merge commits from 12 authors in that period. Some notable commits: The row cache will now purge expired tombstones before populating the cache, removing the performance impact of scanning tombstones. Note that non-expired tombstones are still loaded. Consistent schema management using Raft is now enabled by default for new clusters. Upgraded clusters default to raft disabled for schema management. Alternator, ScyllaDB’s implementation of the DynamoDB API, improved its performance by compiling regular expressions used to validate information during process startup. The compaction manager reloads sstables during schema change, but it did so with quadratic complexity, causing stalls for tables that had many sstables. This is now fixed. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylla-monitoring-stack-4-2-0/287 Title: [RELEASE] Scylla Monitoring Stack 4.2.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of ScyllaDB Monitoring Stack 4.2.0 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB Enterprise and ScyllaDB Open Source, based on Prometheus … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-monitoring-stack-4-2-0/287 ## Headings Structure: H1: [RELEASE] Scylla Monitoring Stack 4.2.0 H3: Related topics ## Main Content: H1: [RELEASE] Scylla Monitoring Stack 4.2.0 H3: Related topics The ScyllaDB team is pleased to announce the release of ScyllaDB Monitoring Stack 4.2.0 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB Enterprise and ScyllaDB Open Source, based on Prometheus and Grafana. ScyllaDB Monitoring Stack 4.2.0 supports: This release brings new panels and graphs, bug fixes, stability improvements and performance enhancements. Versions updates for Scylla Monitoring Stack 4.2.0 VictoriaMetrics is a time series database that, for the most part, is compatible with Prometheus. Following user requests, it is now possible to start the monitoring stack with VictoriaMetrics instead of Prometheus by passing --victoria-metrics to start-all.sh command line flags. This feature is still experimental and may change in future releases. New Information in ScyllaDB Dashboards This means that running kill-all.sh can take longer, you can now specify how long you want to wait for Prometheus to shutdown before forcefully killing it. To add more flexibility for users who use automation, the recording rules that are needed for dashboard operation are split into a different file. This is also part of a work for easier datadog integration. Relevant metrics are marked with a label for Datadog scrapping. Similar to Scylla-Cloud users, an additional set of recording rules adds a multi-level dashboard on-prem users. --- ### Page: https://forum.scylladb.com/t/datamodel-for-scylla-cassandra-for-table-partition-key-is-not-known-beforehand-static-field/289 Title: Datamodel for Scylla/Cassandra for table partition key is not known beforehand -> static field? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I am using ScyllaDb, but I think this also applies to Cassandra since ScyllaDb is compatible with Cassandra. I have the following table (I got ~5 of this kind of tables): create table batch_job_conversation ( conve… Language: en Canonical URL: https://forum.scylladb.com/t/datamodel-for-scylla-cassandra-for-table-partition-key-is-not-known-beforehand-static-field/289 ## Headings Structure: H1: Datamodel for Scylla/Cassandra for table partition key is not known beforehand -> static field? H3: Related topics ## Main Content: H1: Datamodel for Scylla/Cassandra for table partition key is not known beforehand -> static field? H3: Related topics I am using ScyllaDb, but I think this also applies to Cassandra since ScyllaDb is compatible with Cassandra. I have the following table (I got ~5 of this kind of tables): This is used by a batch job to make sure some fields are kept in sync. In the application, a lot of concurrent writes/reads can happen. Once in a while, I will correct the values with a batch job. A lot of writes can happen to the same row, so it will overwrite the rows. A batch job currently picks up rows with this query: Then the batch job will read the data at that point and makes sure things are in sync. I think this query is bad because it stresses all the partitions and the node coordinator because it needs to visit ALL partitions. My question is if it is better for this kind of tables to have a fixed field? Something like this: And than the query would be this: For each batch job I can use a different partition key. The amount of rows in these tables will be roughly the same size (a few thousand at most). The tables will overwrite the same row probably a lot of times. Is it better to have a fixed value? Is there another way to handle this? I don’t have a logical partition key I can use. Having just a single partition in the entire table is also bad because it makes the load on the cluster uneven: only certain shards of certain nodes will have any work to do. Scanning a table with just a few thousand partitions should not be a problem. --- ### Page: https://forum.scylladb.com/t/release-scylladb-operator-1-8-0/292 Title: [RELEASE] ScyllaDB Operator 1.8.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of Scylla Operator 1.8.0. ScyllaDB Operator is an open-source project that helps ScyllaDB Open Source and ScyllaDB Enterprise users run ScyllaDB on Kubernetes. Scylla… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-operator-1-8-0/292 ## Headings Structure: H1: [RELEASE] ScyllaDB Operator 1.8.0 H1: Notable changes H1: Supported versions H1: Upgrade instructions H2: Getting started with Scylla Operator H2: Related Links H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Operator 1.8.0 H1: Notable changes H1: Supported versions H1: Upgrade instructions H2: Getting started with Scylla Operator H2: Related Links H3: Related topics The ScyllaDB team is pleased to announce the release of Scylla Operator 1.8.0. ScyllaDB Operator is an open-source project that helps ScyllaDB Open Source and ScyllaDB Enterprise users run ScyllaDB on Kubernetes. ScyllaDB Operator manages ScyllaDB clusters deployed to Kubernetes and automates tasks related to operating a ScyllaDB cluster, like installation, scale out and downscale, as well as rolling upgrades. ScyllaDB Operator 1.8.0 improves stability and brings a few features. As with all of our releases, any API changes are backwards compatible. For more changes and details check out the GitHub release notes. Upgrading from v1.7.x with kubectl apply doesn’t require any extra action, just take the manifest from v1.8.0 tag and substitute the released image. Using helm requires a mandatory manual step for every release because helm can’t handle CRDs updates. For details, see our upgrade documentation. We’ll welcome your feedback! Feel free to open an issue or reach out on the #scylla-operator channel in Scylla User Slack. ScyllaDB Operator Team --- ### Page: https://forum.scylladb.com/t/how-to-optimally-design-data-model-to-select-by-date-which-couldnt-be-pk/294 Title: How to optimally design data model to select by date, which couldn't be PK? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: My table is CREATE TABLE cache_entries( created timestamp, expired timestamp, request_key text, response_key text, method text, uri text, response_code int, PRIMARY KEY (request_key, created)) WITH CLUSTERING… Language: en Canonical URL: https://forum.scylladb.com/t/how-to-optimally-design-data-model-to-select-by-date-which-couldnt-be-pk/294 ## Headings Structure: H1: How to optimally design data model to select by date, which couldn't be PK? H3: Related topics ## Main Content: H1: How to optimally design data model to select by date, which couldn't be PK? H3: Related topics I find expired data by query SELECT ... WHERE expired < ? ALLOW FILTERING and then remove files, which are linked with found records, and drop records by PK. So I cannot use TTL feature. Is there a way to avoid using of ALLOW FILTERING in my case? Whenever you find yourself having to use ALLOW FILTERING, a possible alternative is to create a secondary index on the filtered columns. This is a trade-off: creating a secondary index will speed up the queries but it is extra load on the cluster, as the secondary index has to be kept in sync with the base table and thus be updated on each write. If you do this filtering read often, using an index might be a good choice. If this filtering read is rare, then you can just keep using filtering, nothing wrong with that. --- ### Page: https://forum.scylladb.com/t/unable-to-create-new-keyspace-or-tables-in-existing-cluster/296 Title: Unable to create new keyspace or tables in existing cluster - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: OperationTimedOut: errors={‘10.14.65.30’: ‘Client request timeout. See Session.execute_async’}, last_host=10.14.65.30 Language: en Canonical URL: https://forum.scylladb.com/t/unable-to-create-new-keyspace-or-tables-in-existing-cluster/296 ## Headings Structure: H1: Unable to create new keyspace or tables in existing cluster H3: Related topics ## Main Content: H1: Unable to create new keyspace or tables in existing cluster H3: Related topics OperationTimedOut: errors={‘10.14.65.30’: ‘Client request timeout. See Session.execute_async’}, last_host=10.14.65.30 Still didn’t get any solution how to get rid off. Is this because of schema mismatch. Please post the command you issued, the cluster configuration, and check the logs for errors or warnings. --- ### Page: https://forum.scylladb.com/t/databases-rust-webassembly-event-streaming-more-at-scylladb-summit/297 Title: Databases + Rust, Webassembly, Event Streaming & More at ScyllaDB Summit - Blog Posts - ScyllaDB Community NoSQL Forum Meta Description: [800x400-blog-summit-database-rust-webassembly] Language: en Canonical URL: https://forum.scylladb.com/t/databases-rust-webassembly-event-streaming-more-at-scylladb-summit/297 ## Headings Structure: H1: Databases + Rust, Webassembly, Event Streaming & More at ScyllaDB Summit H3: Databases + Rust, Webassembly, Event Streaming & More at ScyllaDB Summit H3: Related topics ## Main Content: H1: Databases + Rust, Webassembly, Event Streaming & More at ScyllaDB Summit H3: Databases + Rust, Webassembly, Event Streaming & More at ScyllaDB Summit H3: Related topics What's new at the intersection of databases + Rust, Webassembly, and event streaming? Discover the latest trends and innovations at ScyllaDB Summit --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-166-2023-02-05/298 Title: Last week in scylladb.git master (issue #166; 2023-02-05) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 5eadea301e…61dfc9c10f range are covered. There were 64 non-merge commits from 16 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-166-2023-02-05/298 ## Headings Structure: H1: Last week in scylladb.git master (issue #166; 2023-02-05) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #166; 2023-02-05) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 5eadea301e…61dfc9c10f range are covered. There were 64 non-merge commits from 16 authors in that period. Some notable commits: Raft topology now verifies that the gossip view of the token ring matches the raft view. Yet another locking race in the materialized view update path was fixed. A size-on-disk accounting bug in commitlog was fixed. This could lead to segment recycling being stopped indefinitely, with a large reduction in write performance. A bug which could cause crashes while reporting errors in invalid CQL statements involving field selection from a user-defined type was fixed. The compaction backlog tracker computes the amount of work remaining for compaction. It is updated when inserting sstables into the table. The efficiency of this process, for leveled compaction strategy tables, was improved. io_uring support was disabled, since it appears to cause regressions. ScyllaDB is now more careful which dropping user-defined types that are used by a user-defined function. Out-of-memory management can halt allocations from all but one query, in order to get that query to complete and release memory. However, if that query was paused, it could deadlock the system. This is now fixed. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylla-monitoring-stack-4-2-1/306 Title: [RELEASE] Scylla Monitoring Stack 4.2.1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce a patch release of ScyllaDB Monitoring Stack 4.2.1 This patch release fixes the following bugs: Rename the active and queued reads in the detailed dashboard#1876 The all_sched… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-monitoring-stack-4-2-1/306 ## Headings Structure: H1: [RELEASE] Scylla Monitoring Stack 4.2.1 H3: Related topics ## Main Content: H1: [RELEASE] Scylla Monitoring Stack 4.2.1 H3: Related topics The ScyllaDB team is pleased to announce a patch release of ScyllaDB Monitoring Stack 4.2.1 This patch release fixes the following bugs: ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB Enterprise and ScyllaDB Open Source, based on Prometheus and Grafana. ScyllaDB Monitoring Stack 4.2.1 supports: --- ### Page: https://forum.scylladb.com/t/spring-boot-scylladb-and-time-series-data-lesson/307 Title: Spring Boot, ScyllaDB, and Time Series Data Lesson - University and Training - ScyllaDB Community NoSQL Forum Meta Description: We recently published a new lesson on ScyllaDB University on using Spring Boot with ScyllaDB, and Time Series data. The lesson includes six videos and a hands-on lab you can run. Any questions? Feel free to ask questio… Language: en Canonical URL: https://forum.scylladb.com/t/spring-boot-scylladb-and-time-series-data-lesson/307 ## Headings Structure: H1: Spring Boot, ScyllaDB, and Time Series Data Lesson H3: Related topics ## Main Content: H1: Spring Boot, ScyllaDB, and Time Series Data Lesson H3: Related topics We recently published a new lesson on ScyllaDB University on using Spring Boot with ScyllaDB, and Time Series data. The lesson includes six videos and a hands-on lab you can run. Any questions? Feel free to ask questions and discuss the lesson here. --- ### Page: https://forum.scylladb.com/t/data-modelling-to-replace-redis/310 Title: Data modelling to replace Redis - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi I’m evaluating Scylla to replace Redis as key-value DB, which for our simple use case I’m quite confident that it would be a good alternative. We are more bound to read latencies than throughput, though seemingly low… Language: en Canonical URL: https://forum.scylladb.com/t/data-modelling-to-replace-redis/310 ## Headings Structure: H1: Data modelling to replace Redis H3: Related topics ## Main Content: H1: Data modelling to replace Redis H3: Related topics I’m evaluating Scylla to replace Redis as key-value DB, which for our simple use case I’m quite confident that it would be a good alternative. We are more bound to read latencies than throughput, though seemingly low double-digit is quite possible for Scylla, which would meet our SLA. I have a more obscure use case currently with Redis, and I’m wondering if the following data model would be acceptable/good/bad in the Syclla world. I am mainly used to document DBs, though sharding and ops cost of Mongo is awkward at best, and mongo’s defaults are quite loose and probably should be stricter, but then it loses its speed. Right now, we compress a JSON payload and just look it up by key. We could in Mongo, not compress it and filter/order by any property in the raw JSON payload. We do need low latencies, ideally less than 50ms, but closer to single digits would be ideal. We could layer Redis over Mongo, but that makes it more complex. What I’m wondering is can we at insertion time break some of the key JSON payload properties out into columns and continue to store the compressed JSON, but use only the properties we care to sort on/filter on? I’m sure I can, but I guess the question is, is this a bad idea and will ultimately be more work and unpredictable? In my head, it would look like this Possible Scylla model And just extend this table as and when you needed more specificity or adjust sorting etc?? This should work fine, but do note that you cannot alter primary key columns in scylladb. So if you later want to add a new clustering key component, you have to create a new table for that. This includes sorting, which is a property of the clustering key. Thanks for the reply here, yeah I was aware, though thanks for that extra info/help --- ### Page: https://forum.scylladb.com/t/the-perf-flamegraph-data-from-scylladb-is-coming-out-weird/312 Title: The perf flamegraph data from scylladb is coming out weird - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: The perf flamegraph data from scylladb is coming out weird. Currently, our scylladb is running on kubernetes, and perf was performed directly in the pod, but scylladb-related function call graphs are not visible. comma… Language: en Canonical URL: https://forum.scylladb.com/t/the-perf-flamegraph-data-from-scylladb-is-coming-out-weird/312 ## Headings Structure: H1: The perf flamegraph data from scylladb is coming out weird H3: Related topics ## Main Content: H1: The perf flamegraph data from scylladb is coming out weird H4: hansh0801/files/blob/main/perf-node6.svg H3: Related topics The perf flamegraph data from scylladb is coming out weird. Currently, our scylladb is running on kubernetes, and perf was performed directly in the pod, but scylladb-related function call graphs are not visible. commands: perf record -p $(pidof scylla) --call-graph dwarf – sleep 60 perf script | ./stackcollapse-perf.pl > out.perf-folded ./flamegraph.pl out.perf-folded > perf.svg Do you have the scylla-server-dbg package installed? IIRC it’s not installed by default. --- ### Page: https://forum.scylladb.com/t/too-many-in-flight-hints-10493465/315 Title: Too many in flight hints: 10493465 - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, I was seeing a “Too many in flight hints: 10493465” error today. The problem was a 5 minute write peak. I was wondering if it would be possible to increase max_size_of_hints_in_progress from 10MB to 1000MB and ev… Language: en Canonical URL: https://forum.scylladb.com/t/too-many-in-flight-hints-10493465/315 ## Headings Structure: H1: Too many in flight hints: 10493465 H3: Related topics ## Main Content: H1: Too many in flight hints: 10493465 H3: Related topics I was seeing a “Too many in flight hints: 10493465” error today. The problem was a 5 minute write peak. I was wondering if it would be possible to increase max_size_of_hints_in_progress from 10MB to 1000MB and even throttle hints replay. After the 5 minute peak there is plenty of time to catch up on hints replay. For me it would be no problem if hints replay takes a long time. But I would like to avoid any errors during the peak. Is there any downside to increasing max_size_of_hints_in_progress? Since hints are stored on disk, why is there even such a low hints limit? I think in-flight-hints means hints that are pending to be written to disk. If so, then these in flight hints are helt in memory. Correct? Hints are kept on disk. The obvious downside of raising max_size_of_hints_in_progress is that hints will take up more disk space. In general you should not rely on hints for consistency. It is better to run repair regularly and have hints only cover temporary hiccups. Thanks for the response. Are you sure that max_size_of_hints_in_progress applies to the hint size on disk? I think it really means the size of hints that are currently being persisted locally (according to my limited understanding of manager::end_point_hints_manager::store_hint). Please correct me if I understand this code wrong: I was getting this error on the client side of my write-operations. Since hints are not reliable anyway, I would expect them to be silently dropped, instead of making writes fail. The code throws a overloaded-exception, which is pretty hard: Can’t the hint be simply dropped if it exceeds the threshold? Indeed you are right, max_size_of_hints_in_progress seems to only account for in-memory hints, that are in the process of being written to disk. This is currently a constant in the code and changing it is not a good idea. To be honest I’m even surprised it is a constant value, instead of some percentage of memory (the usual way we express memory limits). If you are hitting this limit, you disk might have trouble keeping up with the rate of incoming hints. I was getting this error on the client side of my write-operations. Since hints are not reliable anyway, I would expect them to be silently dropped, instead of making writes fail. I agree, it doesn’t make sense. I seem to remember this being heatedly discussed in the past. I suggest opening an issue about it, and discuss it further there. Thanks, ticket created: Too many in_flight_hints should not cause overloaded-exception · Issue #13383 · scylladb/scylladb · GitHub Lets hope for a heated discussion --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-0-10/318 Title: [RELEASE] ScyllaDB 5.0.10 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.0.10, a bugfix release of the ScyllaDB 5.0 stable branch. Like all past and future 5.x.y releases, it is backward compatible and supports rolling upgrades. Note that th… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-0-10/318 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.0.10 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.0.10 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.0.10, a bugfix release of the ScyllaDB 5.0 stable branch. Like all past and future 5.x.y releases, it is backward compatible and supports rolling upgrades. Note that the latest stable branch of ScyllaDB is 5.1, and you are encouraged to upgrade. Issues fixed in this release: --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-167-2023-02-12/319 Title: Last week in scylladb.git master (issue #167; 2023-02-12) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 61dfc9c10f…ca4db9bb72 range are covered. There were 129 non-merge commits from 18 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-167-2023-02-12/319 ## Headings Structure: H1: Last week in scylladb.git master (issue #167; 2023-02-12) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #167; 2023-02-12) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 61dfc9c10f…ca4db9bb72 range are covered. There were 129 non-merge commits from 18 authors in that period. Some notable commits: The row cache and memtables now hold rows and range tombstones in a unified data structure, rather than in separate data structures. This solves performance problems (throughput and latency) when a large partition had many range tombstones. A recently-introduced bug, where a prepared statement with named bind variables was executed without providing a value for one of the variables was fixed. Repair-based node operations was made the default. However, after problems surfaced in continuous integration, it was reverted again. The port option in SSTableLoader was fixed. The check for whether a user defined function is in use by a user defined aggregate before dropping it has been improved. A crash duing cql3 aggregation, for the case the query returned no results, has been fixed. ScyllaDB will now treat running out of disk quota (EDQUOT) in the same way it treats running out of disk space (ENOSPC). An edge case where a vnode token boundary coincided with a range scan boundary, but inclusiveness/exclusiveness of the token (< vs <=) did not agree, has been fixed. Some minor bugs in the cql transport server error handling have been corrected. Repair will now ignore local keyspaces. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-1-5/320 Title: [RELEASE] ScyllaDB 5.1.5 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.1.5, a bugfix release of the ScyllaDB 5.1 stable branch. ScyllaDB Open Source 5.1.5, like all past and future 5.x.y releases, is backward compatible and supports rolling… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-1-5/320 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.1.5 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.1.5 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.1.5, a bugfix release of the ScyllaDB 5.1 stable branch. ScyllaDB Open Source 5.1.5, like all past and future 5.x.y releases, is backward compatible and supports rolling upgrades. Users are encouraged to upgrade to 5.1.5. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/on-startup-commitlog-cannot-parse-the-version-of-the-file-commitlog-2-xxxxxxxxx-log/323 Title: On startup: commitlog - cannot parse the version of the file: Commitlog-2-xxxxxxxxx.log - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: After updating and starting scylla (5.1.5), I got this in the log a few times, only for shard 0: [shard 0] commitlog: Cannot parse the version of the file: Commitlog-2-xxxxxxxxx.log I don’t understand this, since the … Language: en Canonical URL: https://forum.scylladb.com/t/on-startup-commitlog-cannot-parse-the-version-of-the-file-commitlog-2-xxxxxxxxx-log/323 ## Headings Structure: H1: On startup: commitlog - cannot parse the version of the file: Commitlog-2-xxxxxxxxx.log H3: Related topics ## Main Content: H1: On startup: commitlog - cannot parse the version of the file: Commitlog-2-xxxxxxxxx.log H3: Related topics After updating and starting scylla (5.1.5), I got this in the log a few times, only for shard 0: [shard 0] commitlog: Cannot parse the version of the file: Commitlog-2-xxxxxxxxx.log I don’t understand this, since the commitlog folder was empty as I drained the node before the update. What’s the reason of this? If this happens before “init - loading system_schema sstables” then the warning is benign and means that the schema commit log is ignoring files which belong to the regular commitlog. We’re planning to fix that warning so that it doesn’t confuse users. See Multiple "commitlog - cannot parse the version" warnings during boot · Issue #11867 · scylladb/scylladb · GitHub. The is already a PR which addresses that. As for why are there commit log files despite draining the node, those commitlog segments were most likely created on that very same boot because of writes to system tables which happen before the commitlog replay. --- ### Page: https://forum.scylladb.com/t/what-are-sstables/324 Title: What are SSTables? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: What are SSTables and how are they used? Language: en Canonical URL: https://forum.scylladb.com/t/what-are-sstables/324 ## Headings Structure: H1: What are SSTables? H3: Related topics ## Main Content: H1: What are SSTables? H3: Related topics What are SSTables and how are they used? A Sorted Strings Table (SSTable) is a persistent file format of key-value string pairs that are sorted by the keys. They are used by Apache Cassandra, ScyllaDB, Bigtable, and some other NoSQL databases. SSTables are immutable, meaning they are never modified. SSTables are a main building block for the way ScyllaDB writes data, which follows the well-known Log Structured Merge (LSM) design described by Patrick O’Neil. Writes in ScyllaDB are very fast and are immediately available for reads. Data is written to a memory table (MemTable), and when that becomes too big, it is flushed to a new file. This file is sorted to make it easy to search and later merge. This is why the tables are known as Sorted String Tables or SSTables. SSTables are merged into new SSTables or deleted in a process called Compaction. --- ### Page: https://forum.scylladb.com/t/how-do-i-list-all-the-keyspaces-in-a-cluster/325 Title: How do I list all the keyspaces in a cluster? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I previously created a keyspace and some tables in a cluster, but forgot what the keyspace name is. How can I see a list of all existing keyspaces in my cluster? Language: en Canonical URL: https://forum.scylladb.com/t/how-do-i-list-all-the-keyspaces-in-a-cluster/325 ## Headings Structure: H1: How do I list all the keyspaces in a cluster? H3: Related topics ## Main Content: H1: How do I list all the keyspaces in a cluster? H3: Related topics I previously created a keyspace and some tables in a cluster, but forgot what the keyspace name is. How can I see a list of all existing keyspaces in my cluster? To list all the keyspaces in a cluster from a CQL Shell use: DESCRIBE can also be used. The above works in the same way for ScyllaDB and for Cassandra. You can read more about it in the Docs. --- ### Page: https://forum.scylladb.com/t/scylladb-summit-day-1-nosql-at-scale-with-less/328 Title: ScyllaDB Summit Day 1: NoSQL at Scale…with Less - Blog Posts - ScyllaDB Community NoSQL Forum Meta Description: [800x400-blog-summit23-speaker-promo] Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-summit-day-1-nosql-at-scale-with-less/328 ## Headings Structure: H1: ScyllaDB Summit Day 1: NoSQL at Scale…with Less H3: ScyllaDB Summit Day 1: NoSQL at Scale…with Less H3: Related topics ## Main Content: H1: ScyllaDB Summit Day 1: NoSQL at Scale…with Less H3: ScyllaDB Summit Day 1: NoSQL at Scale…with Less H3: Related topics Database experts from Discord, Strava, ScyllaDB & more shared top trends & strategies for working with data-intensive applications. --- ### Page: https://forum.scylladb.com/t/release-scylla-5-2-rc1/330 Title: [RELEASE] Scylla 5.2 RC1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce ScyllaDB Open Source 5.2 RC1, the first Release Candidate for the ScyllaDB Open Source 5.2 minor release. ScyllaDB 5.2 introduces Raft-based Strongly Consistent Schema Management… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-5-2-rc1/330 ## Headings Structure: H1: [RELEASE] Scylla 5.2 RC1 H2: Moving to Raft based cluster management H2: New Features H3: Strongly Consistent Schema Management H3: Alternator TTL H3: Large Collection Detection H3: Automating away the gc_grace_seconds parameter H3: Materialized Views: Synchronous Mode H3: Empty replica pages H3: Secondary index on collection columns H2: Improvements H3: CQL API H3: Amazon DynamoDB Compatible API (Alternator) H3: Correctness H3: Performance and stability H3: Operations H3: Deployment and install H3: Tools H3: Configuration H3: Deprecated and removed features H3: Build H3: Monitoring and tracing H3: Related topics ## Main Content: H1: [RELEASE] Scylla 5.2 RC1 H2: Moving to Raft based cluster management H2: New Features H3: Strongly Consistent Schema Management H3: Alternator TTL H3: Large Collection Detection H3: Automating away the gc_grace_seconds parameter H3: Materialized Views: Synchronous Mode H3: Empty replica pages H3: Secondary index on collection columns H2: Improvements H3: CQL API H3: Amazon DynamoDB Compatible API (Alternator) H3: Correctness H3: Performance and stability H3: Operations H3: Deployment and install H3: Tools H3: Configuration H3: Deprecated and removed features H3: Build H3: Monitoring and tracing H3: Related topics The ScyllaDB team is pleased to announce ScyllaDB Open Source 5.2 RC1, the first Release Candidate for the ScyllaDB Open Source 5.2 minor release. ScyllaDB 5.2 introduces Raft-based Strongly Consistent Schema Management, Alternator TTL, and many more improvements and bug fixes. Find the ScyllaDB Open Source 5.2 repository for your Linux distribution here. ScyllaDB 5.2 RC1 Docker is also available. We encourage you to run ScyllaDB 5.2 release candidates on your test environments; this will help ensure that an upgrade to ScyllaDB 5.2 General Availability will proceed smoothly with your workload. Use the release candidate with caution; RC1 is not production-ready yet. You can help stabilize ScyllaDB Open Source 5.2 by reporting bugs here. Only the last two minor releases of the ScyllaDB Open Source project are supported. Once ScyllaDB Open Source 5.2 is officially released, only ScyllaDB Open Source 5.2 and ScyllaDB 5.1 will be supported, and ScyllaDB 5.0 will be retired. Consistent Schema Management is the first Raft based feature in ScyllaDB, and ScyllaDB 5.2 is the first release to enable Raft by default. Starting from ScyllaDB 5.2, all new databases will be created with Raft enabled by default. Upgrading from 5.1 will only use Raft if you explicitly enable it (see upgrade to 5.2 docs). As soon as all nodes in the cluster opt-in to using Raft, the cluster will automatically migrate those subsystems to using Raft, and you should validate it is the case. Once Raft is enabled, every cluster-level operation, like updating schema, adding and removing nodes, and adding and removing data centers - requires a quorum to be executed. For example, in the following use cases, the cluster does not have a quorum and will not allow updating the schema: This is different from the behavior of a ScyllaDB cluster with Raft disabled. Nodes might be unavailable due to network issues, node issues, or other reasons. To reduce the chance of quorum loss, it is recommended to have 3 or more nodes per DC, and 3 or more DCs, for a multi-DCs cluster. To recover from a quorum loss, the best is to revive the failed nodes or fix the network partitioning. If this is impossible, see Raft manual recovery procedure. More on handling failures in Raft here. Schema management operations are DDL operations that modify the schema, like CREATE, ALTER, or DROP for KEYSPACE, TABLE, INDEX, UDT, MV, etc. Unstable schema management has been a problem in all Apache Cassandra and ScyllaDB versions. The root cause is the unsafe propagation of schema updates over gossip, as concurrent schema updates can lead to schema collisions. Once Raft is enabled, all schema management operations are serialized by the Raft consensus algorithm. Additional Raft related updates: In ScyllaDB 5.0 we introduced Time To Live (TTL) to DynamoDB compatible API (Alternator) as an experimental feature. In ScyllaDB 5.2 we promote it to production ready. #12037 #11737 Like in DynamoDB, Alternator items that are set to expire at a specific time will not disappear precisely at that time but only after some delay. DynamoDB guarantees that the expiration delay will be less than 48 hours (though for small tables, the delay is often much shorter). In Alternator, the expiration delay is configurable - it defaults to 24 hours but can be set with the --alternator-ttl-period-in-seconds configuration option. ScyllaDB records large partitions, large rows, and large cells in system tables so that the primary key can be used to deal with them. It additionally records collections with large numbers of elements, since these can cause degraded performance. The warning threshold is configurable: compaction_collection_elements_count_warning_threshold - how many elements are considered a “large” collection (default is 10,000 elements). The information about large collections is stored in the large_cells table, with a new collection_elements column that contains the number of elements of the large collection. Large_cells table retention is 30 days. #11449 Example of a large collection below: There is now optional automatic management of tombstone garbage collection, replacing gc_grace_seconds. This drops tombstones more frequently if repairs are made on time, and prevents data resurrection if repairs are delayed beyond gc_grace_seconds. Tombstones older than the most recent repair will be eligible for purging, and newer ones will be kept. The feature is disabled by default and needs to be enabled via ALTER TABLE. cqlsh> ALTER TABLE ks.cf WITH tombstone_gc = {‘mode’:‘repair’}; There is now a synchronous mode for materialized views. In ordinary, asynchronous materialized views the operation returns before the view is updated. In synchronous materialized view the operation does not return until the view is updated. This enhances consistency but reduces availability as in some situations all nodes might be required to be functional. Synchronous Mode reference in Scylla Docs Before this release, the paging code requires that pages have at least one row before filtering. This can cause an unbounded amount of work if there is a long sequence of tombstones in a partition or token range, leading to timeouts. ScyllaDB will now send empty pages to the client, allowing progress to be made before a timeout. This prevents analytics workloads from failing when processing long sequences of tombstones. #7689, #3914, #7933 Secondary indexes can now index collection columns. Individual keys and values within maps, sets, and lists can be indexed. Fixes #2962, #8745, #10707 Alternator, ScyllaDB’s implementation of the DynamoDB API, now provides the following improvements: TRUNCATE statements are usually much slower than other statements. TRUNCATE statements now support the WITH TIMEOUT clause to help deal with that. #11408 ScyllaDB uses an interval map data structure from the Boost library to quickly locate sstables needed to service a read. Due to the way we interface with the library, updating the interval map was unnecessarily slow. This is now fixed. #11669 Materialized view building will now ignore partitions no longer owned by the node after a topology change. When a materialized view processes updates to the base table, it locks the partition and clustering key. In some rare cases involving one of the locks timing out but the other not, this can cause a crash. This was fixed by acquiring the locks sequentially. #12168 The experimental WASM user defined function (UDF) implementation has been switched to Rust (UDFs can be written in any language WASM supports, not just Rust). The new implementation avoids stalls for long-running UDFs and shares memory with the rest of the database. #11351 The experimental WASM user defined function (UDF) has been re-enabled for aarch64 (ARM). User defined aggregates (UDAs) are now correctly persisted and survive a cold start. UDAs are an experimental feature. #11309 A bug that caused user-defined aggregates not to be persisted correctly was fixed. #11327 A very rare bug involving reads from memtables and multi-version concurrency control has been fixed. Automatically parallelized query aggregation, introduced in 5.1, used a node-local clock for timeouts, rather than the standard clock. This meant that automatically parallelized queries would fail, unless the nodes were started at the same time. This is now fixed. #12458 During startup, ScyllaDB makes sstables conform to the compaction strategy in a process called reshaping, so that future reads will perform well. It is now more careful when reshaping Leveled Compaction Strategy tables, to avoid doing unnecessary work. #12495. Lightweight transactions are now more robust when the schema is changed during a transaction. #10770 A very rare bug involving reads from memtables and multi-version concurrency control has been fixed. Single-partition reads could, in conditions involving multi-page queries, a completely empty page, and a partition or range tombstone covering the first row, terminate paging prematurely, leading to incorrect results. This is now fixed. #12361 Everywhere replication strategy exhibited quadratic behavior since each vNode has every node in its replica set. Since the number of vNodes is proportional to the number of nodes, the amount of work grows quadratically, slowing down topology operations on large clusters. We now special-case this path to avoid the problem. #10337 When using the EverywhereReplicationStrategy, vnodes have no impact on data distribution, and so we can avoid splitting the query across vnodes. This improves performance on tables with a small amount of data. A small race between DROP TABLE and inactive reads (readers that have just completed sending a page to the coordinator and are awaiting the next page fetch) was fixed. #11264 The internal representation of a CREATE TABLE statement now includes the equivalent of a DROP TABLE. This is normally unnecessary, since CREATE TABLE checks if the table already exists and refuses to proceed, but in the case of two CREATE TABLE statements executing concurrently on two different nodes (but creating the same table), each node’s check misses the other’s concurrently created new table. If the table definitions are different, this can lead to a cluster crash due to the mixture between the two definitions being illegal. The change ensures that the two table definitions are not intermixed and just one survives. Users are still urged not to create tables with the same name concurrently until the Raft schema management transition is complete. #11396 Vestigial support for compaction groups has been merged. Compaction groups are token ranges that have dedicated sstables and memtables, and so can be compacted and moved independently of each other. Currently, exactly one compaction group per table is supported, so this doesn’t change anything, but it will be expanded in the future. The database now performs sanity tests around TimeWindowCompactionStrategy tables, to limit the number of time windows. Too many time windows can cause the system to run out of memory or open file handles. #9029 Per partition rate limit error error reporting has been corrected to report the new error code when an updated driver is available.#11517 Support for the new error code has been merged to: Off-strategy compaction is used when sstables from an external source (such as repair) needs to be reshaped before being handed off to the table’s compaction strategy. If off-strategy compaction is stopped, ScyllaDB used to just leave those sstables in their unreshaped form without compacting them again. It will now hand off such sstables directly to the table’s compaction strategy. #11543. In compaction strategies such as LeveledCompactionStrategy and ScyllaDB Enterprise’s IncrementalCompactionStrategy, sstables are sealed when they reach a certain size (160MB and 1GB respectively). They are not split in the middle of a partition, because we were not able to recover the ordering of such split sstables. This is now possible (though not yet integrated into the compaction strategies). Performance: Long-term index caching in the global cache, as introduced in 4.6, hurts the performance for workloads where accesses to the index are sparse. To mitigate this, a new configuration parameter cache_index_pages (default true) is introduced to control index caching. Setting the flag to false causes all index reads to behave like they would in BYPASS CACHE queries. Consider using false if you notice performance problems due to lowered cache hit ratio in 4.6 or 5.0. The config API can update the parameter live (without restart). #11202 Scylla will now reject a too-low bloom_fulter_fp_chance when creating (or altering) the table, rather than crash while flushing memtables. #11524. ScyllaDB represents reads using mutation fragment streams. Several minor violations of fragment stream integrity were fixed. These could result in incorrect reads during range scans. The log-structured allocator is used to manage cache and memtable memory. When memory runs out, the allocator tries to reclaim memory by evicting cache items and by defragmenting memory. If this takes too long, the allocator logs a stall report. Due to a bug, if the report threshold was set too low then the report is generated even if a stall did not happen, slowing down the system and flooding the logs. This is now fixed. #10981 The compaction manager now ignores out-of-disk-space (ENOSPC) exceptions when shutting down, so the server doesn’t crash in these scenarios. An inaccuracy in the per_partition_rate_limit read metric was corrected. Tables with the per_partition_rate_limit property can throttle read and write activity on a per-partition basis. #11651 A large schema (with thousands of tables) could cause stalls when propagated from node to node. This is now fixed. #11574 A recently introduced regression caused a crash when speculative retry was enabled. This is now fixed. #11825 segmentation fault in cases where the base table schema change while MV schema is cached #10026, #11542 A crash when the compaction manager was asked to stop multiple times (for different reasons) was fixed. A problem with RPC connections being needlessly dropped was fixed. #11780 When a query completes a page, ScyllaDB caches the query activity as an inactive read. When the client requests the next page, ScyllaDB re-activates the read and continues where it left off. A bug in this mechanism that could cause crashes has been fixed. #11923 A crash was fixed during an illegal lightweight transaction INSERT with NULL clustering key ScyllaDB caches rows and (since 4.6) index entries in a single unified cache. It was observed that in some small-partition workloads index caching causes a performance regression, so index caching is now disabled by default. It can still be enabled for workloads that benefit from it. We plan to re-enable it when the regression is fixed. #11889 The topology management code is more relaxed about unknown endpoints to prevent crashes in tests that check for edge cases. This fixes a recent regression. #11870 Hinted handoff now checks that a node exists in topology before doing anything; this helps with a recent regression due to topology refactoring. A crash while fetching repaid ids from the repair history table was fixed. #11966 Usually repair can compare and update the same shard in different nodes, for example shard 3 in one node is compared against shard 3 in another. When the number of shards in nodes is dissimilar, this doesn’t work and each shard compares against data from multiple shards in other nodes. This is now made more efficient by reducing sstable reader thrashing for this dissimilar shard count case. #12157 The algorithm for removing nodes from the token ring was corrected and made more efficient. It’s not known that this had any user impact. #12082 The CQL server will now only run requests that benefit from concurrency (e.g. QUERY and EXECUTE) in parallel. Configuration and authentication related requests will be serialized, reducing the chance for errors in those code paths. A rare bug involving an allocation failure while updating cached rows was fixed. #12068 The system.truncated table holds information about truncation times of user tables. A recent regression caused it to be unreadable by cqlsh. It is now fixed. 12239 COMPACT STORAGE tables allow the user to only specify a prefix of a compound clustering key. Bugs relating to such partial keys and reversed rows were fixed.Note that compact storage is deprecated (see section). #12180 ScyllaDB now supports multiple compaction groups 1. This is not a user-visible feature for now. Some copies of the lists of ranges to stream were eliminated from the decommission path, reducing latency spikes. #12332 When the global index cache is disabled, a local (per query) cache was used instead. When that cache was destroyed, a stall could result, generating a latency spike. This is now fixed. #12271 Compaction manager generally reacts to events to initiate compactions, but also has an hourly timer in case an event was missed (and for tombstone compaction, which isn’t triggered by an event). This timer is now less susceptible to stalls. #12390. Repair tried to trigger off-strategy compaction even for a table that was dropped during repair, failing the entire repair. It ignores the dropped table now. Off-strategy compaction is now enabled for all streaming topology operations (adding and removing nodes). Previously it was enabled only for repair-based node operations. Off-strategy compaction takes advantage of the fact that incoming sstables are non-overlapping to perform more efficient compaction that the one performed by the regular compaction strategy. ScyllaDB sometimes reads ahead of the user request, in order to hide latency. In one case a read-ahead request which timed out caused errors to be emitted, even though this did not affect the query. The errors are now silenced. #12435 ScyllaDB tracks transient memory used by queries. However, it did not track decompressed memory buffers, which could lead to running out of memory in some complicated queries. This is now fixed. “Unset” values are an obscure prepared statement feature that allows only some columns in an UPDATE or INSERT statement to be modified. It was a source of minor bugs and inconvenience in code. The feature has been refactored so it has less impact on the code and is more robust. Fix a crash in Materialized View update row locking, caused by a race condition #12632 (RC1) Fix a crash when reporting error on invalid CQL query involving field selection from a user-defined type #12739 (RC1) Scylla Monitoring Stack release 4.3 and later will support ScyllaDB 5.2. metrics related updates below: Shard Latencies are now reported as summaries. This is part of an effort to reduce the total number of generated metrics. In addition, empty histograms and summaries will not be reported. The overall result is a 5x reduction in the number of metrics #11173. This is how a summary looks like: scylla_storage_proxy_coordinator_read_latency_summary_count{scheduling_group_name=“statement”,shard=“1”} 2 scylla_storage_proxy_coordinator_read_latency_summary{quantile=“0.990000”,scheduling_group_name=“statement”,shard=“1”} 640 There is now a metric that allows observation of update progress of materialized views from staging sstables. There are now completion percentage metrics for node operations using streaming; previously the completion metrics were only available when using repair-based node operations. #11600 The sstable row_reads metric for m-format sstables is now properly incremented, instead of showing zeroes. #12406 The replica-side read metrics, which have been incorrect for some time, have been revamped. #10065 --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-168-2023-02-19/336 Title: Last week in scylladb.git master (issue #168; 2023-02-19) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the ca4db9bb72…941407b905 range are covered. There were 155 non-merge commits from 18 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-168-2023-02-19/336 ## Headings Structure: H1: Last week in scylladb.git master (issue #168; 2023-02-19) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #168; 2023-02-19) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the ca4db9bb72…941407b905 range are covered. There were 155 non-merge commits from 18 authors in that period. Some notable commits: Change Data Capture (CDC) exports updates to the database as a table containing changes. One option is to capture not only the change, but also the state of the row before it was changed. In some cases, in a lightweight transaction (LWT) change, the preimage could return the state of the row after the change instead of before the change. This is now fixed. The Prometheus node_exporter, used to observe operating system level metrics, has been upgraded to version 1.5.0. The is now more documentation on diagnostic tools. Lightweight transaction IF evaluation has been refactored to have a common code base with the rest of the system. In a few places, semantics were slightly modified. This is not expected to have any impact on production code. An issue when a user-defined function (UDF) that is used in a user-defined aggregate (UDA) is updated has been fixed. The UDA now reflects the changes in the modified UDF immediately. An issue when using the counter data type in a WebAssembly UDF has been corrected. ScyllaDB installation will now tune the OS core dump service to allow a longer time to dump cores. This is necessary since ScyllaDB allocates all memory and therefore takes a longer time to dump core if an error is encountered. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/scylladb-summit-day-2-continuing-the-high-performance-nosql-conversation/342 Title: ScyllaDB Summit Day 2: Continuing the High-Performance NoSQL Conversation - Blog Posts - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-summit-day-2-continuing-the-high-performance-nosql-conversation/342 ## Headings Structure: H1: ScyllaDB Summit Day 2: Continuing the High-Performance NoSQL Conversation H3: ScyllaDB Summit Day 2: Continuing the High-Performance NoSQL Conversation H3: Related topics ## Main Content: H1: ScyllaDB Summit Day 2: Continuing the High-Performance NoSQL Conversation H3: ScyllaDB Summit Day 2: Continuing the High-Performance NoSQL Conversation H3: Related topics ScyllaDB Summit 2023 wrapup: a recap of NoSQL and event streaming tech talks, links to on-demand videos & decks, and more opportunities to explore high-performance NoSQL. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-169-2023-02-26/343 Title: Last week in scylladb.git master (issue #169; 2023-02-26) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 941407b905…ba919aa88a range are covered. There were 66 non-merge commits from 18 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-169-2023-02-26/343 ## Headings Structure: H1: Last week in scylladb.git master (issue #169; 2023-02-26) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #169; 2023-02-26) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 941407b905…ba919aa88a range are covered. There were 66 non-merge commits from 18 authors in that period. Some notable commits: Merging schema changes received from other nodes is now faster, when there is a large number of tables. The load-and-stream feature reads user-supplied sstables and copies the contents to the correct nodes across the cluster. It is now faster. A bug where varint or bool columns could be deserialized incorrectly in rare cases is now fixed. The cassandra-stress benchmarking tool’s -log hdrfile=… option now works with Java 11. Automatically parallelized aggregation queries on columns of COUNTER type have been fixed. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-2-1/344 Title: [RELEASE] ScyllaDB Enterprise 2022.2.1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2022.2.1, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.2. Related Links Get ScyllaDB Enterprise 2022.2.1 (customers only, o… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-2-1/344 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2022.2.1 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2022.2.1 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2022.2.1, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.2. Below is a list of performance and stability improvements and bug fixes, each with an open-source reference if available: CQL: Deleting a long base partition may leave some undeleted materialized view rows #12297 CQL: scylla: types: is_tuple(): doesn’t handle reverse types. For example, a schema with reversed clustering key component; this component will be incorrectly represented in the schema CQL dump: the UDT will lose the frozen attribute. When attempting to recreate this schema based on the dump, it will fail as the only frozen UDTs are allowed in primary key components. #12576 Alternator Streaming API: unexpected ARN values by list streams paged responses. This issue may only affect users with many, more than 100, tables with streams #12601 Stability: Memtable(s) are not flushed when cleaning up a table, leaving disowned tokens in the memtable, which might be resurrected.#1239 Stability: error “reader_concurrency_semaphore - Semaphore sl:oltp_read_concurrency_sem with 0/100 count and -74538773/0 memory resources”. Root cause is reader_concurrency_semaphore_group partitioning memory is broken. In some cases, this bug may, after a few cycles of updates, lead to semaphore without any memory units unable to admit any reads. In turn this may lead to high latency and even denial of service. Stability: reader_concurrency_semaphore: inactive readers are only evicted on the admission path #11770 Stability: Enabling table encryption (see Encryption at Rest) aborts Scylla when key_provider is not specified. CQL: USING TIMESTAMP allows setting the mutation timestamp on a CQL statement level. It has a sanity check that prevents setting timestamps in the future, as these can be hard to delete, but sometimes one wishes to do so anyway. There is now a configuration option that allows disabling the feature. #12527 Stability: During rebuild on asymmetric cluster several aborts and coredump happened #11923 (introduced by the fix for #11770 above) Stability: Usually repair can compare and update the same shard in different nodes, for example shard 3 in one node is compared against shard 3 in another. When the number of shards in nodes is dissimilar, this doesn’t work and each shard compares against data from multiple shards in other nodes. This is now made more efficient by reducing sstable reader thrashing for this dissimilar shard count case. #12157 Stability: ScyllaDB tracks transient memory used by queries. However, it did not track decompressed memory buffers, which could lead to running out of memory in some complicated queries. #12448 Stability: Cached reads may temporarily miss rows under rare conditions #12451 Parallel aggregation: Timeout point sent in forward_request verb comes from seastar::lowres_clock #12458 Stability: During startup, ScyllaDB makes sstables conform to the compaction strategy reshaping, so that future reads will perform well. It is now more careful when reshaping Leveled Compaction Strategy tables, to avoid doing unnecessary work and using too much disk space #12495. Stability: prevent heap use-after-free of forward_aggregates in parallel aggregators #12528 Stability: crash in Materialized View update row locking, caused by a race condition #12632 Stability: commitlog: segment recycling breaks on segment file removal #12645 Stability: Crash when reporting error on invalid CQL query involving field selection from a user-defined type #12739 Performance: parsers are compiled without inlining, even in release mode #12463 --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-1-2023-02-24/345 Title: Last week in scylla-cluster-tests.git master (issue #1; 2023-02-24) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from last week. Commits in the 4009b65c…8b095095 range are covered. There were 22 non-merge commits from 6 authors in that pe… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-1-2023-02-24/345 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #1; 2023-02-24) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #1; 2023-02-24) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from last week. Commits in the 4009b65c…8b095095 range are covered. There were 22 non-merge commits from 6 authors in that period. Some notable commits: When sct-runner creation failed in the test pipeline we were using builder to serve it’s purpose. This was causing various issues and now we just fail the pipeline. Recently introduced ComparableScyllaVersion class for comparing scylla versions is now used in more places. Previously used parse_scylla_version function was removed. Latency with ops performance test was adjusted to agreed acceptance criteria. Some upgrade jobs are now using raft on target version. 3 new upgrade tests that will run exclusively with raft enabled. Also, we started testing upgrade procedure on ubuntu 22.04 Latency during ops report now contains operation duration in seconds in the new column. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-2-rc2/346 Title: [RELEASE] ScyllaDB 5.2 RC2 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The Scylla team is pleased to announce ScyllaDB Open Source 5.2 RC2, the second Release Candidate for the Scylla Open Source 5.2 minor release. We encourage you to run ScyllaDB 5.2 release candidates on your test enviro… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-2-rc2/346 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.2 RC2 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.2 RC2 H3: Related topics The Scylla team is pleased to announce ScyllaDB Open Source 5.2 RC2, the second Release Candidate for the Scylla Open Source 5.2 minor release. We encourage you to run ScyllaDB 5.2 release candidates on your test environments; this will help ensure that an upgrade to ScyllaDB 5.2 General Availability will proceed smoothly with your workload. Use the release candidate with caution; RC2 is not production-ready yet. You can help stabilize Scylla Open Source 5.2 by reporting bugs here. Only the last two minor releases of the ScyllaDB Open Source project are supported. Once ScyllaDB Open Source 5.2 is officially released, ScyllaDB Open Source 5.2 and 5.1 will be supported, and ScyllaDB 5.0 will be retired. For a complete description of ScyllaDB 5.2 see ScyllaDB 5.2 RC1. Get ScyllaDB Open Source 5.2 (under “More Versions” for each distro) Updates and bug fixes since 5.2 RC1 (not including tests updates) --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-1-6/347 Title: [RELEASE] ScyllaDB 5.1.6 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.1.6, a bugfix release of the ScyllaDB 5.1 stable branch. ScyllaDB Open Source 5.1.6, like all past and future 5.x.y releases, is backward compatible and supports rolling… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-1-6/347 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.1.6 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.1.6 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.1.6, a bugfix release of the ScyllaDB 5.1 stable branch. ScyllaDB Open Source 5.1.6, like all past and future 5.x.y releases, is backward compatible and supports rolling upgrades. Users are encouraged to upgrade to 5.1.6. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/basic-rate-limiter-using-scylladb-how/349 Title: Basic rate limiter using ScyllaDB. How? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I want to use Scylla to write simple rate limiter (because I already use Scylla and don’t want to add another DB just for that). Counters look like a good tool for that, however they don’t have any kind of expiration. I… Language: en Canonical URL: https://forum.scylladb.com/t/basic-rate-limiter-using-scylladb-how/349 ## Headings Structure: H1: Basic rate limiter using ScyllaDB. How? H3: Related topics ## Main Content: H1: Basic rate limiter using ScyllaDB. How? H3: Related topics I want to use Scylla to write simple rate limiter (because I already use Scylla and don’t want to add another DB just for that). Counters look like a good tool for that, however they don’t have any kind of expiration. I already solved similar problem, where I need a limit per day. I basically create table every day and use it for writes, then create new one and use that (the old is deleted via cronjob). Now I want to rate limit requests per minute, so it’s not a good solution. I can use something like that: Where slot is basically a time slot (say floor(unixtime/60)). It would be enough, but of course I need some expiration. I could probably do the same, create new table every day, delete the old one. However, this looks messy and I don’t like that at all. Are there any better solutions? The underlying problem here is that counters don’t support expiry. So you either have to switch to a data type which does support expiry (basically anything apart from counters), or you need another cron job which removes rows from this table from time-to-time. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-2-2023-03-03/351 Title: Last week in scylla-cluster-tests.git master (issue #2; 2023-03-03) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 7f6ed833…39212be6 range are covered. There were 15 non-merge commits from 7 authors in tha… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-2-2023-03-03/351 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #2; 2023-03-03) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #2; 2023-03-03) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 7f6ed833…39212be6 range are covered. There were 15 non-merge commits from 7 authors in that period. Some notable commits: As long time passed since fix of kernel bug that was causing some issues with io setup on GCE, we dropped all related workarounds and now we always run complete io-setup on GCE instances. Added nemesis pipelines for Azure so we can better test ScyllaDB on this cloud platform. Manager 3.1 now supports restore tasks added tests using a restore from backup task This commit was accompanied by bunch of other commits to support this feature by SCT and improved existing tests. Along with docker-based loaders, we started to use most recent cassandra-stress. We now append no-warmup to all cassandra-stress commands used in SCT as warmup was causing issues with timeouting duration workflows. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-170-2023-03-05/352 Title: Last week in scylladb.git master (issue #170; 2023-03-05) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the ba919aa88a…4b71f87594 range are covered. There were 96 non-merge commits from 16 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-170-2023-03-05/352 ## Headings Structure: H1: Last week in scylladb.git master (issue #170; 2023-03-05) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #170; 2023-03-05) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the ba919aa88a…4b71f87594 range are covered. There were 96 non-merge commits from 16 authors in that period. Some notable commits: There is a new metric for prepared statement cache eviction rates. A bug in LWT handling, when a schema change happens concurrently with an operation, has been fixed. A few memory leaks, exposed by the out-of-memory query killer, have been fixed. In developer and debug mores, ScyllaDB will now configure Seastar in shared library mode. This only affects developers of ScyllaDB itself, as release more still uses static libraries. The --list-tools option now correctly lists all tools invocable via the scylla binary. Repair-based node operations (RBNO) use repair as an alternative to streaming to improve reliability and correctness. They are no (again) enabled by default. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylla-monitoring-stack-4-3-0/357 Title: [RELEASE] Scylla Monitoring Stack 4.3.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of ScyllaDB Monitoring Stack 4.3.0 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB Enterprise and ScyllaDB Open Source, based on Prometheus … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-monitoring-stack-4-3-0/357 ## Headings Structure: H1: [RELEASE] Scylla Monitoring Stack 4.3.0 H3: Related topics ## Main Content: H1: [RELEASE] Scylla Monitoring Stack 4.3.0 H3: Related topics The ScyllaDB team is pleased to announce the release of ScyllaDB Monitoring Stack 4.3.0 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB Enterprise and ScyllaDB Open Source, based on Prometheus and Grafana. ScyllaDB Monitoring Stack 4.3.0 supports: This release focuses on adaptation for the coming ScyllaDB open source 5.2 metrics changes. It is advised to upgrade the monitoring stack before upgrading ScyllaDB. Working towards a better Datadog integration, the label used for scraping was modified and by default, there will be no scrapping of the per-shard metrics. The level label was deprecated, and will be removed in future versions. Follow the Datadog integration guide and download the updated config and dashboard. Versions updates for Scylla Monitoring Stack 4.3.0 New Information in ScyllaDB Dashboards Add a CPU Starvation graph to the per-scheduling group section --- ### Page: https://forum.scylladb.com/t/release-scylladb-cloud-6-march-2023/358 Title: [RELEASE] ScyllaDB Cloud - 6 March 2023 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: We are happy to announce a new update to the ScyllaDB Cloud service. The update includes new beta releases of ScyllaDB Cloud Serverless and the Admin API and many other improvements. ScyllaDB Cloud Serverless ScyllaDB… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-cloud-6-march-2023/358 ## Headings Structure: H1: [RELEASE] ScyllaDB Cloud - 6 March 2023 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Cloud - 6 March 2023 H3: Related topics We are happy to announce a new update to the ScyllaDB Cloud service. The update includes new beta releases of ScyllaDB Cloud Serverless and the Admin API and many other improvements. ScyllaDB Cloud Serverless ScyllaDB Cloud now offers a Free Trial serverless cluster for 48 hours. Under the hood, Serverless is built on Kubernetes and Scylla Operator, among other technologies. This is the first general available step toward ScyllaDB Cloud Serverless as presented in ScyllaDB Summit. Learn more about Serverless in this blog. ScyllaDB Cloud Admin API The ScyllaDB Cloud Admin API is now available as a beta release. The API exposes operations like: ScyllaDB Cloud Terraform Provider Based on the Admin API (above), the Terraform provider makes it easier to control ScyllaDB Cloud as part of a larger infrastructure. Full documentation is available in the Terraform Registry. The project is open source under the following address GitHub - scylladb/terraform-provider-scylladbcloud: Terraform provider for ScyllaDB Cloud.. We welcome users to report feedback and we also welcome contributions - especially examples, new use cases and documentation! New Components Releases ScyllaDB now uses the following components: New clusters use the latest releases by default. Running clusters will be upgraded gradually. You can set your preference for your cluster Maintenance Window time and day. us-south1 (Dallas) was added to the supported GCP regions. ScyllaDB Cloud currently supports the following regions: Let us know if your favorite region is not included yet! --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-2-2/360 Title: [RELEASE] ScyllaDB Enterprise 2022.2.2 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2022.2.2, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.2. Related Links Get ScyllaDB Enterprise 2022.2.2 (customers only, o… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-2-2/360 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2022.2.2 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2022.2.2 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2022.2.2, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.2. A list of bug fixes below, each with an open-source reference if available: --- ### Page: https://forum.scylladb.com/t/scylladb-cloud-goes-serverless/367 Title: ScyllaDB Cloud Goes Serverless - Blog Posts - ScyllaDB Community NoSQL Forum Meta Description: [800x400-blog-future-of-serverless-scylladb] Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-cloud-goes-serverless/367 ## Headings Structure: H1: ScyllaDB Cloud Goes Serverless H3: ScyllaDB Cloud Goes Serverless H3: Related topics ## Main Content: H1: ScyllaDB Cloud Goes Serverless H3: ScyllaDB Cloud Goes Serverless H3: Related topics ScyllaDB Cloud is moving to a serverless implementation. You can trial it today with our monstrously fast and scalable NoSQL database. --- ### Page: https://forum.scylladb.com/t/whats-new-with-scylladb-march-2023/368 Title: What's New with ScyllaDB - March 2023 - Announcements - ScyllaDB Community NoSQL Forum Meta Description: BLOGS AND ARTICLES How Discord Stores Trillions of Messages Bo Ingram, senior software engineer at Discord, explains how they’ve changed from MongoDB to Cassandra to now ScyllaDB to continue supporting trillions of mes… Language: en Canonical URL: https://forum.scylladb.com/t/whats-new-with-scylladb-march-2023/368 ## Headings Structure: H1: What's New with ScyllaDB - March 2023 H3: Related topics ## Main Content: H1: What's New with ScyllaDB - March 2023 H3: Related topics How Discord Stores Trillions of Messages Bo Ingram, senior software engineer at Discord, explains how they’ve changed from MongoDB to Cassandra to now ScyllaDB to continue supporting trillions of messages. Read more: How Discord Stores Trillions of Messages How Strava’s NoSQL Move Keeps Athletes Moving How Strava uses ScyllaDB to speed up activity tracking for 100M+ athletes – and why they moved from Apache Cassandra. Read more: How Strava’s NoSQL Move Keeps Athletes Moving - ScyllaDB ScyllaDB Cloud Goes Serverless ScyllaDB Serverless is our most elastic and dynamic deployment model. Learn more about what’s available and what’s next. Read more: ScyllaDB Cloud Goes Serverless - ScyllaDB ScyllaDB Summit Day 1: NoSQL at Scale…with Less At Day 1 of ScyllaDB Summit, database experts from Discord, Strava, ScyllaDB & more shared top trends & strategies for working with data-intensive applications. Read more: ScyllaDB Summit Day 1: NoSQL at Scale…with Less - ScyllaDB ScyllaDB Summit Day 2: Continuing the High-Performance NoSQL Conversation ScyllaDB Summit 2023 wrapup: a recap of NoSQL and event streaming tech talks, links to on-demand videos & decks, and more opportunities to explore high-performance NoSQL. Read more: ScyllaDB Summit Day 2: Continuing the High-Performance NoSQL Conversation - ScyllaDB ScyllaDB Innovation Awards Honor Impressive NoSQL + Rust Low-Latency Achievements Announcing this year’s ScyllaDB Innovation Award winners ShareChat, Discord, Numberly, Jasper Visser, Digital Turbine, and Junglee Games. Read more: ScyllaDB Innovation Awards Honor Impressive NoSQL + Rust Low-Latency Achievements - ScyllaDB How IOTA Uses Distributed Ledgers and ScyllaDB for Supply Chain Digitization Explore how the IOTA Foundation is tackling supply chain digitization challenges in East Africa, including the role of open source distributed ledgers and ScyllaDB NoSQL. Read more: How IOTA Uses Distributed Ledgers and ScyllaDB for Supply Chain Digitization - ScyllaDB ScyllaDB University LIVE Spring March 21 | 8AM-12PM PT | 11AM-3PM ET | 15:00-19:00 GMT Join us for a half-day of live virtual training led by our top engineers and architects. We will have two parallel tracks – a beginners track focused on ScyllaDB Essentials and an advanced track covering Advanced Topics and Integrations. Register here: ScyllaDB University LIVE ScyllaDB Virtual Workshop March 23, 2023 | 12pm GMT | 8am ET | 5:30pm IST Ready to try out ScyllaDB and want to make sure you’re “doing it right?” Learn what ScyllaDB is all about, the core concepts you need to know, and a step-by-step demonstration of how to get started. Register here: Scylla Virtual Workshop Webinar: 4 Ways Development Teams Cut Costs with ScyllaDB March 29 | 10am PT | 1pm ET | 5pm GMT Join Tzach Livyatan, VP of Product at ScyllaDB, as he shares four ways that teams commonly cut database costs by rethinking their database strategy. Register here: 4 Ways Development Teams Cut Costs with ScyllaDB NEW SCYLLADB RELEASES AND UPDATES --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-171-2023-03-12/371 Title: Last week in scylladb.git master (issue #171; 2023-03-12) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 4b71f87594…e7250e5a3f range are covered. There were 82 non-merge commits from 14 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-171-2023-03-12/371 ## Headings Structure: H1: Last week in scylladb.git master (issue #171; 2023-03-12) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #171; 2023-03-12) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 4b71f87594…e7250e5a3f range are covered. There were 82 non-merge commits from 14 authors in that period. Some notable commits: A crash while clearing paused reads during shutdown (for example, queries that are paused due to reaching the page size limit) was fixed. When definining a user-defined function, one has to specify the language in the CREATE FUNCTION statement. The language name for WebAssembly functions was renamed from “xwasm” to “wasm” in preparation for moving it out of experimental status. CQL transport metrics were refined, and new metrics were added so one can measure request and response bandwidth, for each opcode type. The code for computing proximity (whether nodes are on the same or different rack and datacenter) was optimized. An edge case in repair-based node operations abort process was tightened. It is now possible to specify the minimum replication factor for new keyspaces via a new configuration item. This matches the same functionality in Cassandra. The CQL USING TTL clause allows on to specify an INSERT or UPDATE’s time-to-live property, after which the cells are automatically deleted. TTL 0 was misinterpreted as the default TTL (which happens to be unlimited, usually) rather than an explicitly unlimited TTL. This is now fixed. Documentation related to enterprise features was removed from the open-source documentation. It can be found in the enterprise-specific documentation. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/march-21-scylladb-university-live/372 Title: March 21 - ScyllaDB University LIVE - Announcements - ScyllaDB Community NoSQL Forum Meta Description: Our next LIVE training event is taking place next week, on March 21. It’s a half day of free virtual training by our top experts and engineers. The sessions will be interactive and not available on demand. More info he… Language: en Canonical URL: https://forum.scylladb.com/t/march-21-scylladb-university-live/372 ## Headings Structure: H1: March 21 - ScyllaDB University LIVE H3: Related topics ## Main Content: H1: March 21 - ScyllaDB University LIVE H3: Related topics Our next LIVE training event is taking place next week, on March 21. It’s a half day of free virtual training by our top experts and engineers. The sessions will be interactive and not available on demand. More info here. Hope to see you there. --- ### Page: https://forum.scylladb.com/t/virtual-meetup-high-performance-low-latency-database-architecture/373 Title: Virtual Meetup - High-Performance, Low-Latency Database Architecture - Announcements - ScyllaDB Community NoSQL Forum Meta Description: On Tuesday, March 14, I’ll speak about High-Performance, Low-Latency Database Architecture in a (virtual) meetup. More info here, I hope to see you there! Language: en Canonical URL: https://forum.scylladb.com/t/virtual-meetup-high-performance-low-latency-database-architecture/373 ## Headings Structure: H1: Virtual Meetup - High-Performance, Low-Latency Database Architecture H3: Related topics ## Main Content: H1: Virtual Meetup - High-Performance, Low-Latency Database Architecture H3: Related topics On Tuesday, March 14, I’ll speak about High-Performance, Low-Latency Database Architecture in a (virtual) meetup. More info here, I hope to see you there! --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2021-1-18/374 Title: [RELEASE] ScyllaDB Enterprise 2021.1.18 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2021.1.18, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2021.1. ScyllaDB Enterprise 2022.1 LTS is the latest Long-Term Support rele… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2021-1-18/374 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2021.1.18 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2021.1.18 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2021.1.18, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2021.1. ScyllaDB Enterprise 2022.1 LTS is the latest Long-Term Support release, and 2022.2 is the newest feature release. You are encouraged to upgrade in coordination with the ScyllaDB support team. The following issues are fixed in this release (with an open source reference): create table t (p int, c1 int, c2 int, primary key(p, c1, c2)); create index i1 on t(c1); insert into t(p, c1, c2) values (1, 11, 21); insert into t(p, c1, c2) values (2, 12, 22); select c1 from t where (c1,c2)=(11,21) allow filtering; --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-3-2023-03-10/375 Title: Last week in scylla-cluster-tests.git master (issue #3; 2023-03-10) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 7d691c9d…3866b20d range are covered. There were 25 non-merge commits from 12 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-3-2023-03-10/375 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #3; 2023-03-10) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #3; 2023-03-10) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 7d691c9d…3866b20d range are covered. There were 25 non-merge commits from 12 authors in that period. Some notable commits: Performance tests could not be run on specific ScyllaDB packages (e.g. dev builds). Now we can specify it in pipeline. We use Centos OS for Monitor in all manager jobs New tests related to a backup restore were added, e.g. one that restores during enospc where we run a backup task when node fills out whole disk space and verify restore task failed. scylla-bench ‘v0.1.16’ is used by default now containing improvements in retries mechanism. Fixed severe bug in the way we handle scylla_yaml where before the fix we dropped a key specified in AMI’s predefined scylla.yaml if it was not listed in ScyllaYaml class. Now, key is never dropped and we get warning message about missing it in ScyllaYaml class. Migrated Argus client to the REST Client which is stateless, sends only the necessary data, simplify the use of the API and is less sensitive to Argus schema changes. Switched default monitoring to branch 4.3 See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/how-do-i-list-all-tables-in-a-cluster/376 Title: How do I list all tables in a cluster? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I have a cluster with many tables. Is there a way to list all the tables in the cluster? Language: en Canonical URL: https://forum.scylladb.com/t/how-do-i-list-all-tables-in-a-cluster/376 ## Headings Structure: H1: How do I list all tables in a cluster? H3: Related topics ## Main Content: H1: How do I list all tables in a cluster? H3: Related topics I have a cluster with many tables. Is there a way to list all the tables in the cluster? If you want to see the tables of a specific keyspace, run the following from the CQL Shell: To list all the tables in all keyspaces, run: Without choosing a keyspace first. Notice that this will also list system tables; --- ### Page: https://forum.scylladb.com/t/inside-scylladb-university-live-q-a-with-guy-shtub/377 Title: Inside ScyllaDB University LIVE: Q & A with Guy Shtub - Blog Posts - ScyllaDB Community NoSQL Forum Meta Description: [800x400-blog-scylladb-u-q-a-guy-shtub] Language: en Canonical URL: https://forum.scylladb.com/t/inside-scylladb-university-live-q-a-with-guy-shtub/377 ## Headings Structure: H1: Inside ScyllaDB University LIVE: Q & A with Guy Shtub H3: Inside ScyllaDB University LIVE: Q & A with Guy Shtub H3: Related topics ## Main Content: H1: Inside ScyllaDB University LIVE: Q & A with Guy Shtub H3: Inside ScyllaDB University LIVE: Q & A with Guy Shtub H3: Related topics Get the inside scoop on ScyllaDB University LIVE: a fast and focused way to learn strategies for getting the most out of ScyllaDB NoSQL. We offer sessions for new users as well as ScyllaDB power users. Thanks Cynthia! Any questions or comments about the blog post or the LIVE event? Feel free to discuss it here. --- ### Page: https://forum.scylladb.com/t/using-read-and-write-consistency-level-of-quorum-how-many-nodes-can-be-down/380 Title: Using Read and Write Consistency Level of Quorum - how many nodes can be down? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I have a 5 nodes cluster and use a Replication Factor of 3. I’m using the Read and Write Consistency Level of Quorum. How many nodes can be down before Writes or Reads requests start to fail? Language: en Canonical URL: https://forum.scylladb.com/t/using-read-and-write-consistency-level-of-quorum-how-many-nodes-can-be-down/380 ## Headings Structure: H1: Using Read and Write Consistency Level of Quorum - how many nodes can be down? H3: Related topics ## Main Content: H1: Using Read and Write Consistency Level of Quorum - how many nodes can be down? H3: Related topics I have a 5 nodes cluster and use a Replication Factor of 3. I’m using the Read and Write Consistency Level of Quorum. How many nodes can be down before Writes or Reads requests start to fail? You can check that as well as other scenarios specific to your use case using the Consistency Level Calculator. If you want a better understanding of Availability and what happens in different scenarios the High Availability lesson on ScyllaDB University is a good place to start. --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-1-7/384 Title: [RELEASE] ScyllaDB 5.1.7 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.1.7, a bugfix release of the ScyllaDB 5.1 stable branch. ScyllaDB Open Source 5.1.7, like all past and future 5.x.y releases, is backward compatible and supports rolling… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-1-7/384 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.1.7 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.1.7 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.1.7, a bugfix release of the ScyllaDB 5.1 stable branch. ScyllaDB Open Source 5.1.7, like all past and future 5.x.y releases, is backward compatible and supports rolling upgrades. Users are encouraged to upgrade to 5.1.7. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-4-2023-03-17/397 Title: Last week in scylla-cluster-tests.git master (issue #4; 2023-03-17) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 0b3b0a4d…50931210 range are covered . There were 18 non-merge commits from 8 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-4-2023-03-17/397 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #4; 2023-03-17) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #4; 2023-03-17) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 0b3b0a4d…50931210 range are covered . There were 18 non-merge commits from 8 authors in that period. Some notable commits: We expand testing ScyllaDB on Azure platform. Made a step towards testing in different regions by implementing listing Azure ScyllaDB images with hydra tool. Added various fixes to a way we run SLA nemeses in non-SLA specific tests. We aim to cover SLA in more tests variations. Create index Nemesis was added to cover the scenario, where a user decides to create index after some time, when table is full of data. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-172-2023-03-19/398 Title: Last week in scylladb.git master (issue #172; 2023-03-19) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the e7250e5a3f…aad2afd417 range are covered. There were 140 non-merge commits from 18 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-172-2023-03-19/398 ## Headings Structure: H1: Last week in scylladb.git master (issue #172; 2023-03-19) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #172; 2023-03-19) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the e7250e5a3f…aad2afd417 range are covered. There were 140 non-merge commits from 18 authors in that period. Some notable commits: User-defined functions can now have [permission control] The CQL shell, cqlsh, has been separated into its own repository. As part of that change, cqlsh is now compatible with Python 3. (Merge 'Allow setting permissions for user-defined functions' from Woj… · scylladb/scylladb@843a5df · GitHub) via CQL GRANT statements. The C-style cast syntax ((type) expression) can now be applied to bind variables ((type) ? or (type) :var) to explicitly specify the type of bind variables. When a WebAssembly user-defined function is compiled, it is now compiled in a separate thread in order to avoid stalling the reactor and inducing high latency. The WebAssembly documentation is now user-visible, as part of making it ready for production. The --max-io-requests option, which has been obsolete for quite some time, was removed. The reader concurrency semaphore is responsible for managing concurrency for read queries, balancing memory and CPU use with enough concurrency to keep the disks busy. One case that was not handled well is if a read was blocked due to the system running low on memory, and subsequently made idle (as no one is waiting for its results). This edge case has been fixed. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/about-the-kubernetes-operator-category/400 Title: About the Kubernetes Operator category - Kubernetes Operator - ScyllaDB Community NoSQL Forum Meta Description: For all the SyllaDB Kubernetes Operator related topics. Scylla Operator is an open-source project that gives ScyllaDB Open Source and ScyllaDB Enterprise users an easy way to run and manage ScyllaDB via Kubernetes. The … Language: en Canonical URL: https://forum.scylladb.com/t/about-the-kubernetes-operator-category/400 ## Headings Structure: H1: About the Kubernetes Operator category H3: Related topics ## Main Content: H1: About the Kubernetes Operator category H3: Related topics For all the SyllaDB Kubernetes Operator related topics. Scylla Operator is an open-source project that gives ScyllaDB Open Source and ScyllaDB Enterprise users an easy way to run and manage ScyllaDB via Kubernetes. The ScyllaDB Operator automates the NoSQL cluster deployment process and tasks related to operating a ScyllaDB cluster, such as scaling, backup, auto-healing, rolling configuration changes, upgrades, and more. --- ### Page: https://forum.scylladb.com/t/scylla-operator-and-scylla-manager-oss-node-limitations/410 Title: Scylla-operator and scylla-manager OSS node limitations - Kubernetes Operator - ScyllaDB Community NoSQL Forum Meta Description: Hello! I am looking to use the scylla operator to automate the deployment of scylla clusters in our kubernetes environment for dev/test purposes. Based on some test deployments I’ve done, it looks like a scylla-manager d… Language: en Canonical URL: https://forum.scylladb.com/t/scylla-operator-and-scylla-manager-oss-node-limitations/410 ## Headings Structure: H1: Scylla-operator and scylla-manager OSS node limitations H3: Related topics ## Main Content: H1: Scylla-operator and scylla-manager OSS node limitations H3: Related topics Hello! I am looking to use the scylla operator to automate the deployment of scylla clusters in our kubernetes environment for dev/test purposes. Based on some test deployments I’ve done, it looks like a scylla-manager deployment is required to perform repairs on clusters – the operator does not appear to allow external access to either the API or the JMX ports (please correct me if I missed some documentation somewhere). According to the pricing page (Pricing information for ScyllaDB Enterprise - ScyllaDB) automated backups and restores are limited to 5 nodes for the open-source version of Scylla. Does this limitation apply to scylla clusters deployed by scylla-operator? Also how is this limitation enforced for Open-source based clusters? Scylla manager free usage itself is limit for 5 nodes. Repair or backups aren’t limited but they do require API or JMX access, if it’s not available it’s probably trivial to allow those ports It does not look like the API port is exposed from either the pods or the services that the operator creates, so I would not be able to access the API externally even if I set the api to listen on 0.0.0.0 form the scylla.yaml config. Is JMX controlled from the scylla.yaml file or the cluster resource config? That port is exposed from both pod and service but nodetool only appears to work from within each pod. I haven’t found any concrete configuration instructions for the JMX stuff, as I understand it runs as a separate service from the main scylla process. Please point me to any documentation I may have missed. the operator does not appear to allow external access to either the API or the JMX ports Indeed, ScyllaDB API port is not exposed by default, external communication with it is through token based authentication via Scylla Manager Agent which exposes a port and forwards request to ScyllaDB API. I’m not sure about licensing of Agent, whether the same limitation applies. Right now Agent is always deployed alongside ScyllaDB, and there’s no switch to change it. Is JMX controlled from the scylla.yaml file or the cluster resource config? I don’t have any experience with regards to JMX configuration, nor I don’t know any documentation resources, but you may want to check out the content of /etc/scylla/cassandra/* inside ScyllaDB image. To change any of this file you can mount additional volume using scyllacluster.spec.datacenter.rack[].volume and scyllacluster.spec.datacenter.rack[].volumeMounts. Also how is this limitation enforced for Open-source based clusters? It’s not enforced anyhow. --- ### Page: https://forum.scylladb.com/t/release-scylladb-rust-udf-helper-library-0-1-0/413 Title: [RELEASE] ScyllaDB Rust UDF Helper Library 0.1.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: We are excited to announce the first release of Rust helper library for Scylla UDFs! This can be used for writing pure Rust functions and using them as Scylla User-Defined Functions (UDFs). Features The main feature of… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-rust-udf-helper-library-0-1-0/413 ## Headings Structure: H1: [RELEASE] ScyllaDB Rust UDF Helper Library 0.1.0 H2: Features H2: Instructions H2: Stability H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Rust UDF Helper Library 0.1.0 H2: Features H2: Instructions H2: Stability H3: Related topics We are excited to announce the first release of Rust helper library for Scylla UDFs! This can be used for writing pure Rust functions and using them as Scylla User-Defined Functions (UDFs). The crate offers a comprehensive README file that guides users on the usage of the crate. The README includes sections on Preparation of the WASM UDFs, their usage in CQL Statements, and instructions on Contributing and Testing of the crate. The guide is intended to provide a smooth and easy process for users to get started with the crate. The repository also includes multiple examples displaying all features of the crate. The crate is officially supported and ready to use, but it is important to note that UDFs are still an experimental feature in ScyllaDB. As a result, the crate’s API is subject to change. We encourage users to provide feedback, bug reports, and pull requests to help improve the crate’s functionality and stability. --- ### Page: https://forum.scylladb.com/t/c-cassandra-driver-navigate-to-specific-page/414 Title: C# Cassandra driver : Navigate to specific page - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I’m trying to build a OData api on top of scylla using the c# Cassandra library. Do you have any recommendation to “navigate” between pages like getting the next / previous page or pointing directly to a specific page ? … Language: en Canonical URL: https://forum.scylladb.com/t/c-cassandra-driver-navigate-to-specific-page/414 ## Headings Structure: H1: C# Cassandra driver : Navigate to specific page H3: Related topics ## Main Content: H1: C# Cassandra driver : Navigate to specific page H4: Backward paging in cassandra c# driver H3: Related topics I’m trying to build a OData api on top of scylla using the c# Cassandra library. Do you have any recommendation to “navigate” between pages like getting the next / previous page or pointing directly to a specific page ? I already implement the pagination with the PagingState but this allow me only to move to the next page @Guy I saw your comment into the Slack channel telling us to ping you to get an answer, hope it doesn’t bother you Have a good day One suggestion is to “cache” previous PaginingState between calls And a bit more elaborate explanation on how to implement next/previous queries, by using revered queries: https://community.ibm.com/community/user/supplychain/blogs/tanvi-kakodkar1/2020/01/24/paging-filtering-apache-cassandra I didn’t tested if mixing pagingState with reverse queries, it might also do the trick. --- ### Page: https://forum.scylladb.com/t/release-scylladb-rust-driver-0-8-0/415 Title: [RELEASE] ScyllaDB Rust Driver 0.8.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce ScyllaDB Rust Driver 0.8.0, an asynchronous CQL driver for Rust, optimized for Scylla, but also compatible with Apache Cassandra! Some interesting statistics: over 145k downlo… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-rust-driver-0-8-0/415 ## Headings Structure: H1: [RELEASE] ScyllaDB Rust Driver 0.8.0 H2: Notable changes H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Rust Driver 0.8.0 H2: Notable changes H3: Related topics The ScyllaDB team is pleased to announce ScyllaDB Rust Driver 0.8.0, an asynchronous CQL driver for Rust, optimized for Scylla, but also compatible with Apache Cassandra! Some interesting statistics: New features / enhancements: Performance improvements: CI / developer tool improvements: Congrats to all contributors and thanks everyone for using our driver! The source code of the driver can be found here: The official crates.io registry entry is here: Thank you for your attention, please do not hesitate to contact us if you have any questions, issues, feature requests, or are simply interested in our driver! Contributors since the last release: --- ### Page: https://forum.scylladb.com/t/about-the-release-notes-category/418 Title: About the Release Notes category - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: Here you can find the release notes for ScyllaDB products. ScyllaDB supports the current release and one release back for ScyllaDB Enterprise and ScyllaDB Open Source. Users are encouraged to update to the latest releas… Language: en Canonical URL: https://forum.scylladb.com/t/about-the-release-notes-category/418 ## Headings Structure: H1: About the Release Notes category H3: Related topics ## Main Content: H1: About the Release Notes category H3: Related topics Here you can find the release notes for ScyllaDB products. ScyllaDB supports the current release and one release back for ScyllaDB Enterprise and ScyllaDB Open Source. Users are encouraged to update to the latest release. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-5-2023-03-24/421 Title: Last week in scylla-cluster-tests.git master (issue #5; 2023-03-24) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 2294a254…e4b37c86 range are covered. There were 15 non-merge commits from 7 authors in tha… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-5-2023-03-24/421 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #5; 2023-03-24) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #5; 2023-03-24) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 2294a254…e4b37c86 range are covered. There were 15 non-merge commits from 7 authors in that period. Some notable commits: Because of the big size, sct debug logs are now split to smaller parts to speed up downloading. We need to support upgrade from both last LTS and last short term versions to current LTS, additional versions tests were added. Since we are seeing cases the nodetool repair command doesn’t end, added timeout for the repair command See you in the next issue of last week in scylla-cluster-tests.git master! @soyacz let’s keep sending this also in email. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-173-2023-03-26/422 Title: Last week in scylladb.git master (issue #173; 2023-03-26) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the aad2afd417…46e6c639d9 range are covered. There were 155 non-merge commits from 16 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-173-2023-03-26/422 ## Headings Structure: H1: Last week in scylladb.git master (issue #173; 2023-03-26) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #173; 2023-03-26) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the aad2afd417…46e6c639d9 range are covered. There were 155 non-merge commits from 16 authors in that period. Some notable commits: Topology changes are now coordinated via Raft, when the consistent cluster management option is enabled. WebAssembly just-in-time compilation was recently moved to a separate thread. The thread now blocks all signals, preventing incorrect signal handling when it was randomly picked to handle a signal. User-defined functions recently gained permission support. We now additionally check that a function is not a builtin before applying permissions. The reader concurrency semaphore controls how many reads execute in parallel. Its diagnostics, printed to the system log, have been improved. In a similar vein, the reader concurrency semaphore now has more tracepoints, useful with CQL tracing. The CQL transport server (port 9042) recently gained per-opcode bandwidth statistics. They are now measured per service level as well. Some time ago, we reduced repair contention with itself when nodes had different shard counts. This turned out to cause deadlocks, so the fix was reverted. Materialized view updates are performed asynchronously relative to updating the main table, so failures there are not visible as an UPDATE or INSERT failure. Instead, errors are logged. Those errors are rate-limited now to avoid flooding the logs. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/scylladb-labs-training-event-april-23/423 Title: ScyllaDB Labs Training Event, April 23 - Announcements - ScyllaDB Community NoSQL Forum Meta Description: On the 19th of April, we’re hosting ScyllaDB Labs, an online (free) training event. The focus will be on getting started with ScyllaDB, and the event will include hands-on labs as well as live talks. More information h… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-labs-training-event-april-23/423 ## Headings Structure: H1: ScyllaDB Labs Training Event, April 23 H3: Related topics ## Main Content: H1: ScyllaDB Labs Training Event, April 23 H3: Related topics On the 19th of April, we’re hosting ScyllaDB Labs, an online (free) training event. The focus will be on getting started with ScyllaDB, and the event will include hands-on labs as well as live talks. More information here. Save your spot today. Hope to see you there! --- ### Page: https://forum.scylladb.com/t/how-to-fix-data-inconsistent-between-secondary-index-or-materialized-view-and-base-table/424 Title: How to fix data inconsistent between secondary index(or materialized view) and base table? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: In some failure scenarios, the data update for the base table may not be synchronized to the secondary index in time. If the failure time is long, this part of the update for the secondary index may even be lost, so how … Language: en Canonical URL: https://forum.scylladb.com/t/how-to-fix-data-inconsistent-between-secondary-index-or-materialized-view-and-base-table/424 ## Headings Structure: H1: How to fix data inconsistent between secondary index(or materialized view) and base table? H3: Related topics ## Main Content: H1: How to fix data inconsistent between secondary index(or materialized view) and base table? H3: Related topics In some failure scenarios, the data update for the base table may not be synchronized to the secondary index in time. If the failure time is long, this part of the update for the secondary index may even be lost, so how should we fix this part? What about inconsistent data? Using nodetool repair to repair the keyspace, will the base table and secondary indexes(or materialized views) be repaired at the same time? Repairing the keyspace will indeed repair both base tables and views/indexes and this might be enough to fix all inconsistencies. Note that this repair will not fix inconsistencies where a view/index update was completely lost and this is not present on any of the view/index replicas. To fix these kind of inconsistencies we would need a base-table ↔ view/index repair, which goes over each base table row and ensures the corresponding view/index row is correct. We did have proposals on how to fix it but nothing materialized out of it yet. Currently the only option for fixing such inconsistencies is to drop the view/index and re-create it from scratch. --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-2-rc3/427 Title: [RELEASE] ScyllaDB 5.2 RC3 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The Scylla team is pleased to announce ScyllaDB Open Source 5.2 RC3, a Release Candidate for ScyllaDB Open Source 5.2 minor release. We encourage you to run ScyllaDB 5.2 release candidates on your test environments; thi… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-2-rc3/427 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.2 RC3 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.2 RC3 H3: Related topics The Scylla team is pleased to announce ScyllaDB Open Source 5.2 RC3, a Release Candidate for ScyllaDB Open Source 5.2 minor release. We encourage you to run ScyllaDB 5.2 release candidates on your test environments; this will help ensure that an upgrade to ScyllaDB 5.2 General Availability will proceed smoothly with your workload. Use the release candidate with caution; RC3 is not production-ready yet. You can help stabilize ScyllaDB Open Source by reporting bugs here. Only the last two minor releases of the ScyllaDB Open Source project are supported. Once ScyllaDB Open Source 5.2 is officially released, ScyllaDB 5.2 and 5.1 will be supported, and ScyllaDB 5.0 will be retired. For a complete description of ScyllaDB 5.2 see ScyllaDB 5.2 RC1. ScyllaDB 5.2 RC1 , RC2 Get ScyllaDB Open Source 5.2 (under “More Versions” for each distro) Updates and bug fixes since 5.2 RC2 (not including tests updates): --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-6-2023-03-31/428 Title: Last week in scylla-cluster-tests.git master (issue #6; 2023-03-31) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 1bca3457…e56c761b range are covered. There were 28 non-merge commits from 9 authors in tha… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-6-2023-03-31/428 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #6; 2023-03-31) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #6; 2023-03-31) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 1bca3457…e56c761b range are covered. There were 28 non-merge commits from 9 authors in that period. Some notable commits: Adapted upgrade tests to remove scylla-cqlsh when downgrading, since cqlsh is using the same files as scylla-tools was using before (fixing the issue). New tests around ldap permissions changes where we add a new user with ldap permissions, and remove them during the test. Some refactor of csrangehistogram with better public interfaces and improved performance. Allow to change the number of stress threads in performance tests, so we can reuse them in serverless clusters testing. We fixed the issue preventing us from testing with cosistent_cluster_management set to false in ScyllaDB 5.2 due faulty scylla.yaml defaults management. More tests related Manager backup/restore: added a single node 5 TB backup and restore test which found some issues with this size, so we lowered the data size to 4 TB. Also Added a test that restores multiple snapshots where we restore schema and data of several previously created snapshots. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/rust-in-the-real-world-super-fast-data-ingestion-using-scylladb/429 Title: Rust in the Real World: Super Fast Data Ingestion Using ScyllaDB - Blog Posts - ScyllaDB Community NoSQL Forum Meta Description: [800x400-blog-rust-in-the-real-world] Language: en Canonical URL: https://forum.scylladb.com/t/rust-in-the-real-world-super-fast-data-ingestion-using-scylladb/429 ## Headings Structure: H1: Rust in the Real World: Super Fast Data Ingestion Using ScyllaDB H3: Rust in the Real World: Super Fast Data Ingestion Using ScyllaDB H3: Related topics ## Main Content: H1: Rust in the Real World: Super Fast Data Ingestion Using ScyllaDB H3: Rust in the Real World: Super Fast Data Ingestion Using ScyllaDB H3: Related topics A detailed tutorial on how to use Rust to build a fast data ingestion API that reads data from a data lake in S3 and stores it into ScyllaDB. Rust Tokio library is used to allow asynchronous computing using many threads to speed up the ingestion... --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-2-3/430 Title: [RELEASE] ScyllaDB Enterprise 2022.2.3 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2022.2.3, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.2. Related Links Get ScyllaDB Enterprise 2022.2.3 (customers only, o… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-2-3/430 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2022.2.3 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2022.2.3 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2022.2.3, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.2. A list of bug fixes below, each with an open-source reference if available: --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-0-11/431 Title: [RELEASE] ScyllaDB 5.0.11 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: Related links: ScyllaDB Open Source 5.0 Get ScyllaDB Open Source - AWS AMI, GCP Docker, binary packages and unified installer ScyllaDB Web Installer for Linux (all releases) Upgrade from ScyllaDB Open Source 5.x.y to… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-0-11/431 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.0.11 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.0.11 H3: Related topics Issues fixed in this release: --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-2-rc4/432 Title: [RELEASE] ScyllaDB 5.2 RC4 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The Scylla team is pleased to announce ScyllaDB Open Source 5.2 RC4, a Release Candidate for Scylla Open Source 5.2 minor release. We encourage you to run ScyllaDB 5.2 release candidates on your test environments; this … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-2-rc4/432 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.2 RC4 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.2 RC4 H3: Related topics The Scylla team is pleased to announce ScyllaDB Open Source 5.2 RC4, a Release Candidate for Scylla Open Source 5.2 minor release. We encourage you to run ScyllaDB 5.2 release candidates on your test environments; this will help ensure that an upgrade to ScyllaDB 5.2 General Availability will proceed smoothly with your workload. Use the release candidate with caution; RC4 is not production-ready yet. You can help stabilize Scylla Open Source 5.2 by reporting bugs here. Only the last two minor releases of the ScyllaDB Open Source project are supported. Once Scylla Open Source 5.2 is officially released, ScyllaDB Open Source 5.2 and 5.1 will be supported, and ScyllaDB 5.0 will be retired. For a complete description of ScyllaDB 5.2 see https://forum.scylladb.com/t/release-scylla-5-2-rc1/330. ScyllaDB 5.2 RC1 , RC2, RC3 Get ScyllaDB Open Source 5.2 (under “More Versions” for each distro) Updates and bug fixes since 5.2 RC3 (not including tests, docs updates) --- ### Page: https://forum.scylladb.com/t/increase-reading-and-writing-speed-with-openresty/434 Title: Increase reading and writing speed with Openresty - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, I’m using scylladb with Openresty in front to make calls. Reading and writing are really slow… wondering if you have any configuration advice for openresty or scyllabdb to increase speed. The server with scylladb ha… Language: en Canonical URL: https://forum.scylladb.com/t/increase-reading-and-writing-speed-with-openresty/434 ## Headings Structure: H1: Increase reading and writing speed with Openresty H3: Related topics ## Main Content: H1: Increase reading and writing speed with Openresty H3: Related topics Hi, I’m using scylladb with Openresty in front to make calls. Reading and writing are really slow… wondering if you have any configuration advice for openresty or scyllabdb to increase speed. The server with scylladb has 4vcpu and 8gb ram instead that of openresty 3vcpu and 4gb ram. Thanks for your help. This question is really hard to answer without more information. How is scylla deployed? Does it have its own node, or does it share it with Openresty? The slowness can be caused by either ScyllaDB or Openresty. The first step in diagnosing performance issues for ScyllaDB is monitoring. Please setup monitoring and check the latencies reported by it. --- ### Page: https://forum.scylladb.com/t/upcoming-change-to-your-scylladb-cloud-user-account/439 Title: Upcoming change to your ScyllaDB Cloud user account - Announcements - ScyllaDB Community NoSQL Forum Meta Description: At ScyllaDB Cloud, we are in the process of enhancing self-service user management capabilities for all of our customers. This includes the ability for Admins to manage account users, invite users to their account, assig… Language: en Canonical URL: https://forum.scylladb.com/t/upcoming-change-to-your-scylladb-cloud-user-account/439 ## Headings Structure: H1: Upcoming change to your ScyllaDB Cloud user account H3: Related topics ## Main Content: H1: Upcoming change to your ScyllaDB Cloud user account H3: Related topics At ScyllaDB Cloud, we are in the process of enhancing self-service user management capabilities for all of our customers. This includes the ability for Admins to manage account users, invite users to their account, assign roles to users as well as additional user and account management related features. In order to achieve the above, we are making changes to our user management infrastructure. Starting in this month, enhancements will be gradually deployed on the ScyllaDB Cloud app. We are planning on rolling out additional capabilities (such as multi-factor authentication enforcement, social login, single sign-on and much more) in future increments. What you need to know: If you are experiencing issues with your account after the password reset procedure, contact us via Slack or email at cloud-support@scylladb.com. Thanks, ScyllaDB Cloud Team --- ### Page: https://forum.scylladb.com/t/mongodb-vs-postgres-vs-scylladb-tractian-s-benchmarking-and-migration/441 Title: MongoDB vs Postgres vs ScyllaDB: Tractian’s Benchmarking and Migration - Blog Posts - ScyllaDB Community NoSQL Forum Meta Description: [800x400-blog-tractian-db-comparison] Language: en Canonical URL: https://forum.scylladb.com/t/mongodb-vs-postgres-vs-scylladb-tractian-s-benchmarking-and-migration/441 ## Headings Structure: H1: MongoDB vs Postgres vs ScyllaDB: Tractian’s Benchmarking and Migration H3: MongoDB vs Postgres vs ScyllaDB: Tractian’s Benchmarking and Migration H3: Related topics ## Main Content: H1: MongoDB vs Postgres vs ScyllaDB: Tractian’s Benchmarking and Migration H3: MongoDB vs Postgres vs ScyllaDB: Tractian’s Benchmarking and Migration H3: Related topics TRACTIAN shares their comparison of ScyllaDB vs MongoDB and PostgreSQL, then provides an overview of their MongoDB to ScyllaDB migration process, challenges & results. --- ### Page: https://forum.scylladb.com/t/how-to-remove-fields-from-an-udt/442 Title: How to remove fields from an UDT - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi everyone! Is there any way that I can remove fields from an UDT? Using the superheroes example from the docs, is there any way to remove the phone field from the address type? I understand that there is no ALTER TY… Language: en Canonical URL: https://forum.scylladb.com/t/how-to-remove-fields-from-an-udt/442 ## Headings Structure: H1: How to remove fields from an UDT H3: Related topics ## Main Content: H1: How to remove fields from an UDT H3: Related topics Is there any way that I can remove fields from an UDT? Using the superheroes example from the docs, is there any way to remove the phone field from the address type? I understand that there is no ALTER TYPE address DROP phone and I have also tried creating a new type without the phone field and running a alter table ... alter type but Scylla reports the types as incompatible. In my case, I have an UDT with lots of fields and a large portion of them are not used anymore, so I’d like to just remove them. Is this possible in any way? We don’t support removing fields or altering fields to incompatible types because it would open pandora’s box of misery. In effect, we only allow backward compatible changes to UDT, such that reading old values of the same type, with the new definition can work. I think the best course of action for you is to add a new field to your table, with a new UDT type and write new values to this new column. On the read side, you can check for the new column and fall back to the old one if that is missing. You can also run a script which scans the entire table and re-writes values from the old column, into the new one, allowing you to drop the old column. This might save you some space. Unused columns in an UDT value do take up a little bit of space. Got it. Thanks for the quick answer! --- ### Page: https://forum.scylladb.com/t/install-my-schema-after-deploying-in-kubernetes/444 Title: Install my schema after deploying in kubernetes - Kubernetes Operator - ScyllaDB Community NoSQL Forum Meta Description: Hi all! How do I install my database schema after deploying in kubernetes? The scheme is described in the cql files Language: en Canonical URL: https://forum.scylladb.com/t/install-my-schema-after-deploying-in-kubernetes/444 ## Headings Structure: H1: Install my schema after deploying in kubernetes H3: ScyllaDB Drivers H3: Related topics ## Main Content: H1: Install my schema after deploying in kubernetes H3: ScyllaDB Drivers H3: Related topics Hi all! How do I install my database schema after deploying in kubernetes? The scheme is described in the cql files I am not a kubernetes expert, but I think you can use the same way you would use in any other cluser, using cqlsh. Or do you want to automate the loading of the schema, as part of deploying the cluster? To define a schema you use the same way you would use in any other deployment, via cqlsh. In Kubernetes, cqlsh is available in Scylla image, you can execute it inside one of your Scylla Pods: Thank you for your reply. That’s right, I want to automate the loading of the schema, as part of deploying the cluster Thank you. I’ll try that option. Depends on how your automation is built, and in what language you might want to use on of the drivers for doing CQL operations ScyllaDB’s shard-aware drivers provide superior NoSQL database performance. ScyllaDB is also compatible with Apache Cassandra CQL drivers and DynamoDB SDKs. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-7-2023-04-08/445 Title: Last week in scylla-cluster-tests.git master (issue #7; 2023-04-08) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the a3173984…4048674e range are covered. There were 27 non-merge commits from 8 authors in tha… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-7-2023-04-08/445 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #7; 2023-04-08) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #7; 2023-04-08) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the a3173984…4048674e range are covered. There were 27 non-merge commits from 8 authors in that period. Some notable commits: In order to verify correctness of gc-modes added new test for switching tombstone-gc modes in scaling and high stress. Added new basic longevity for MV running in synchronous mode. It creates different materialized views. Part of MVs are running in sychronous mode, others in asynchronous mode. Updates/inserts/deletes into MVs perform in parallel with different nemeses, disruptive and non-disruptive In case of newly added Scylla node failing on boot fail fast whole test, in that case test anyway was doomed to fail. Speeding up issues investigation. New test options for Scylla on K8s allowing us to set cpu and memory limits for Scylla pods (k8s_scylla_cpu_limit and k8s_scylla_memory_limit). When defined, will be applied to each Scylla pod. Also, these options are suitable for multitenant configuration. Added 3 jenkins pipelines for longevity azure: longevity-1tb-5days-azure , longevity-200gb-48h-azure and longevity-large-partition-4days-azure . We did some performance tests for Scylla serverless with 1 and 2 vcpus and these are the configs. When testing, various operations take more time when node performance is lower or data load is higher. We aim to dynamically calculate timeout based on node performance and load. To enable it we started to measure how much time operations take along with node load and basic performance metrics. When we gather enough data, we will be able to calculate sensible timeouts dynamically. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/last-fortnight-in-scylladb-git-master-issue-174-2023-04-09/446 Title: Last fortnight in scylladb.git master (issue #174; 2023-04-09) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last two weeks (previous week’s report was not sent due to illness). Commits in the 46e6c639d9…c65bd01174 range are covered. The… Language: en Canonical URL: https://forum.scylladb.com/t/last-fortnight-in-scylladb-git-master-issue-174-2023-04-09/446 ## Headings Structure: H1: Last fortnight in scylladb.git master (issue #174; 2023-04-09) H3: Related topics ## Main Content: H1: Last fortnight in scylladb.git master (issue #174; 2023-04-09) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last two weeks (previous week’s report was not sent due to illness). Commits in the 46e6c639d9…c65bd01174 range are covered. There were 228 non-merge commits from 19 authors in that period. Some notable commits: WebAssembly functions are translated to machine code in a separate thread to reduce impact on the running system. The thread now runs with lower priority to reduce impact further. A scan on a disjoint token range (e.g. (1…100), (200…300)) could have resulted in incorrect results. It’s not possible for a user query to specify a disjoint token range; we are checking if internal sources could generate such queries. The bug itself is fixed. A large query is broken up into separate pages. When tracing, each page gets its own trace session. An optimization in ScyllaDB means a following page can reuse state from the previous page. The trace will now link to the previous session when this happens, improving visibility. ScyllaDB uses a separate commitlog for tables holding the schema, so that schema changes do not suffer high latency under heavy write loads. This separate commitlog will now be used for all raft-managed system tables to guarantee atomicity. The JMX support application, used to support nodetool, now runs under Java 11. An incorrect read result under complex conditions involving static columns was fixed. It is being investigated whether this has user impact. Query timeouts in configuration (e.g. read_request_timeout_in_ms) can now be hot-reloaded using SIGHUP. The separate commitlog for schema is now stored in a separate directory. Minor bugs in compaction manager, that could cause a compaction not to be triggered, were fixed. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-1-8/447 Title: [RELEASE] ScyllaDB 5.1.8 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.1.8, a bugfix release of the ScyllaDB 5.1 stable branch. ScyllaDB Open Source 5.1.8, like all past and future 5.x.y releases, is backward compatible and supports rolling… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-1-8/447 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.1.8 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.1.8 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.1.8, a bugfix release of the ScyllaDB 5.1 stable branch. ScyllaDB Open Source 5.1.8, like all past and future 5.x.y releases, is backward compatible and supports rolling upgrades. Users are encouraged to upgrade to 5.1.8. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/how-do-we-create-global-secondary-index-on-set-types/451 Title: How do we create global secondary index on set types? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi good people, I was trying to create secondary index on a column of type “set” but not sure if its possible. for example, I have a table as follows: CREATE KEYSPACE company WITH replication = {'class': 'NetworkTopol… Language: en Canonical URL: https://forum.scylladb.com/t/how-do-we-create-global-secondary-index-on-set-types/451 ## Headings Structure: H1: How do we create global secondary index on set types? H3: Related topics ## Main Content: H1: How do we create global secondary index on set types? H3: Related topics I was trying to create secondary index on a column of type “set” but not sure if its possible. for example, I have a table as follows: here, I want to create secondary index / materialized views on subordinate_ids . please suggest a way on how this can be achieved. Scylla does support global secondary indexes on collection columns, starting with ScyllaDB 5.2. This is currently in the process of being released, currently at RC3. I suggest taking it for a spin trying out this new feature. Thanks for the reply @Botond_Denes … could you please share any document regarding creating secondary indexes with collection type? See our general documentation on secondary indexes here: Global Secondary Indexes | ScyllaDB Docs and scylladb/secondary_index.md at master · scylladb/scylladb · GitHub. For a column of set type, the create statement looks something like this: worked like a charm! thanks @Botond_Denes I had to get rid of . in index name for it to work though… Thanks, I edited my answer, can you please check if it is correct now? works as it is now, thanks @Botond_Denes --- ### Page: https://forum.scylladb.com/t/free-hands-on-nosql-training-in-asia-friendly-time-zones/452 Title: Free Hands-On NoSQL Training in Asia-Friendly Time Zones - Blog Posts - ScyllaDB Community NoSQL Forum Meta Description: [800x400-blog-freehands-nosql-training] Language: en Canonical URL: https://forum.scylladb.com/t/free-hands-on-nosql-training-in-asia-friendly-time-zones/452 ## Headings Structure: H1: Free Hands-On NoSQL Training in Asia-Friendly Time Zones H3: Free Hands-On NoSQL Training in Asia-Friendly Time Zones H3: Related topics ## Main Content: H1: Free Hands-On NoSQL Training in Asia-Friendly Time Zones H3: Free Hands-On NoSQL Training in Asia-Friendly Time Zones H3: Related topics Introducing a new (hands-on + free) "Intro to NoSQL" live training event designed specifically for the tech community across India, Singapore, Malaysia, Australia, and neighboring areas. This is taking place tomorrow, you can still register here. Hope to see you there! Thanks to all the participants! If you registered and didn’t get a link to the on-demand material, send me a PM. This thread is a good place for questions that you came up with after the event and any input/feedback you have. --- ### Page: https://forum.scylladb.com/t/readtimeout-issue/454 Title: ReadTimeOut issue - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: ReadTimeout: Error from server: code=1200 [Coordinator node timed out waiting for replica nodes’ responses] message=“Operation timed out for guitardb.active_line_status_details - received only 0 responses from 1 CL=ONE.”… Language: en Canonical URL: https://forum.scylladb.com/t/readtimeout-issue/454 ## Headings Structure: H1: ReadTimeOut issue H3: Related topics ## Main Content: H1: ReadTimeOut issue H3: Related topics ReadTimeout: Error from server: code=1200 [Coordinator node timed out waiting for replica nodes’ responses] message=“Operation timed out for guitardb.active_line_status_details - received only 0 responses from 1 CL=ONE.” info={‘received_responses’: 0, ‘required_responses’: 1, ‘consistency’: ‘ONE’} how can I resolve this? This is not enough information to go on. Please check the logs and setup monitoring. Also, please post cluster topology. I see you also posted on slack, where you mention this is a full scan query. Those do often time out on larger tables. I will paste here the answer from slack for posterity: Well you are clearly running a full table scan, so that can easily timeout indeed. This is a bad query, so it is no surprise it is timing out. This is how you should be doing full table scans Efficient full table scans with ScyllaDB 1.6 - ScyllaDB An alternative is to raise the timeouts. You can do that with per query timeout, see https://docs.scylladb.com/stable/cql/cql-extensions.html#using-timeout. This allows raising the timeout for just for a single query, without affecting others. Note that you might also have to adjust your driver’s timeouts as well. select sum(unit_price) AS line_total from active_line_status_new WHERE order_createts>=‘2022-10-01 00:00:00’ AND order_createts<=‘2022-10-27 23:59:59’ ALLOW FILTERING; here I have created secondary index on order_createts column.Still Im facing this issue. The fact that you had to add ALLOW FILTERING signals that the query is still a full scan behind the scenes and faces the same challenges as a full scan. SELECT queries that have a constrain on a column value like col >= x AND col <= y are always filtering queries, unless col is a clustering key component (and x and y are key prefixes). Creating a secondary index on order_states will create a material view, where order_states is the partition key. A SELECT query then still has to filter, because partitions are not ordered according to their values, but according to their token. Hi @Botond_Denes! Though Im using this query select sum(unit_price) AS line_total from active_line_status_new WHERE order_createts>=‘2022-10-01 00:00:00’ AND order_createts<=‘2022-10-27 23:59:59’ ALLOW FILTERING; , still im facing same issue. here index is created on order_createts column. I recommend increasing the range scan timeouts, seeing that you cannot get around a scanning query. This can be changed by the range_request_timeout_in_ms config item in scylla.yaml. Currently this requires a restart of the node to take effect. Make sure you also adjust timeouts on the client side accordingly. --- ### Page: https://forum.scylladb.com/t/how-work-scylla-with-2-nodes-and-replication-factor-2/455 Title: How work scylla with 2 nodes and replication_factor 2 - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, I had this question, I created a cluster with two nodes and the keyspace with these replicas “{‘class’: ‘NetworkTopologyStrategy’, ‘replication_factor’ : 2} AND durable_writes = true;”. Doing an ab test with 500,000 … Language: en Canonical URL: https://forum.scylladb.com/t/how-work-scylla-with-2-nodes-and-replication-factor-2/455 ## Headings Structure: H1: How work scylla with 2 nodes and replication_factor 2 H3: Related topics ## Main Content: H1: How work scylla with 2 nodes and replication_factor 2 H3: Related topics Hi, I had this question, I created a cluster with two nodes and the keyspace with these replicas “{‘class’: ‘NetworkTopologyStrategy’, ‘replication_factor’ : 2} AND durable_writes = true;”. Doing an ab test with 500,000 connections, the ab test doesn’t give me any errors but if I do a count inside the table I find less and less data, almost half. Could anyone give me an explanation? What do you mean by “ab” test? Do you have another database with the same dataset and you execute reads against both of them, comparing results? Which version of ScyllaDB is this? ab stands for apache-bench which is a tool for testing concurrent users and connections. The version of scylla is 5.1 What do you mean by “less and less”? Does a count query return different results on each invocation? Could you please try writing a script which does a full scan of the table and manually counts the rows? See if that has a correct result? I insert 500,000 rows with the test but doing a count on the db gives me much less, as in the screenshots Could you please try writing a script which does a full scan of the table and manually counts the rows? See if that has a correct result? Can you please try the experiment I asked about above, and report the result here? Curious, I wonder if it was inserting unique rows or duplicate rows, and what was the schema. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-8-2023-04-15/456 Title: Last week in scylla-cluster-tests.git master (issue #8; 2023-04-15) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 54425f02…ff477556 range are covered. There were 18 non-merge commits from 6 authors in tha… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-8-2023-04-15/456 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #8; 2023-04-15) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #8; 2023-04-15) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 54425f02…ff477556 range are covered. There were 18 non-merge commits from 6 authors in that period. Some notable commits: There is no need to install python2 or python3 for cqlsh anymore as it supports relocatable python3, so we stopped installing it. Added example jupyter notebooks that allow quick and easy way to interactively play with SCT cluster and try new code in SCT without a need to restart test. Since we bumped JRE used by scylla-jmx from openjdk-8 to openjdk-11, use openjdk-11 for testing offline installation. Added support to multi az in AWS to enable coverage of test cases in multi-rack deployments. Adapted few of them to use multi az. Adopted SCT to comply with the new procedure of ‘Handling Cluster Membership Change Failures’. Now we compare members of token ring with group0 and remove garbage member with removenode operation. We also started to check group0 token ring consistency after each nemesis. Add-remove MV nemesis was added to cover the case of adding and removing materialized view when one node is down. Run “kill_test” method only once, so we avoid endless failure loop - hopefully fixing important issue of not stopping the test properly. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/i-want-to-connect-scylladb-on-golang-and-i-tried-copying-the-code-from-connect-and-pasting-it-doesnt-work/458 Title: I want to connect scylladb on golang and i tried copying the code from connect and pasting it doesn't work - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I want to connect scylladb on golang and i tried copying the code from connect and pasting it doesn’t work. Language: en Canonical URL: https://forum.scylladb.com/t/i-want-to-connect-scylladb-on-golang-and-i-tried-copying-the-code-from-connect-and-pasting-it-doesnt-work/458 ## Headings Structure: H1: I want to connect scylladb on golang and i tried copying the code from connect and pasting it doesn't work H3: GitHub - scylladb/gocql: Package gocql implements a fast and robust ScyllaDB... H3: Related topics ## Main Content: H1: I want to connect scylladb on golang and i tried copying the code from connect and pasting it doesn't work H3: GitHub - scylladb/gocql: Package gocql implements a fast and robust ScyllaDB... H3: Related topics I want to connect scylladb on golang and i tried copying the code from connect and pasting it doesn’t work. Please share the code you used and the error you received. You can see a running example of using golang with ScyllaDB on ScyllaDB University. Here is the first less (part one out of three). This is the code I came up with and I tried it and it doesn’t work because it doesn’t find it. github.com/gocql/gocql/scyllacloud const ( connectionBundlePath = “./build.yaml” ) func main() { cluster, err := scyllacloud.NewCloudCluster(connectionBundlePath) if err != nil { log.Fatalf(“Failed to create cloud cluster config: %s”, err) } cluster.PoolConfig.HostSelectionPolicy = gocql.DCAwareRoundRobinPolicy(“us-east-1”) Are you connecting to ScyllaCloud Serverless cluster (free trial)? If yes, you probably don’t have a fork replacement directive in your go.mod. Package gocql implements a fast and robust ScyllaDB client for the Go programming language. - GitHub - scylladb/gocql: Package gocql implements a fast and robust ScyllaDB client for the Go programm... Add the following line to your project go.mod file. Then make sure the path to the bundle is correct, and your code should run fine. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-175-2023-04-16/459 Title: Last week in scylladb.git master (issue #175; 2023-04-16) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the c65bd01174…c501163f95 range are covered. There were 134 non-merge commits from 11 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-175-2023-04-16/459 ## Headings Structure: H1: Last week in scylladb.git master (issue #175; 2023-04-16) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #175; 2023-04-16) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the c65bd01174…c501163f95 range are covered. There were 134 non-merge commits from 11 authors in that period. Some notable commits: It is now possible to store sstables for selected keyspaces directly on S3 or a compatible object storage system. Note that local storage is still required for commitlog, hints, and system tables. This can reduce storage costs where low latency is not required. The in-memory footprint of sstable summaries has been reduced. This is especially noticeable with very small partitions. In alternator (ScyllaDB’s implementation of the DynamoDB API), a bug in concurrent modification of table tags has been fixed. ScyllaDB can now relabel metrics according to user-provided configuration. This can be used together with Prometheus to reduce the number of metrics reported. When resharding or staging sstables, ScyllaDB will also perform a cleanup operation, removing data that’s not needed by the node since it was migrated away to other nodes. Running major compaction will no longer increase the compaction scheduling group shares to 200. This is not necessary since major compaction runs in the maintenance/streaming group. An inactive reader is a query that was paused (often because the client is busy consuming the previous page). A bug caused inactive readers to be evicted needlessly, which would cause extra work when the next page is fetched to re-create the query. It is now fixed. The load-and-stream operation reads user-supplied sstables and streams them to the cluster. It now avoids loading the bloom filter, saving memory. Error messages for incorrect usage of the CQL TOKEN() function have been improved. The scylla sstable tool now has more ways to obtain the schema. The code base was migrated away from the standard library’s regular expression implementation to the one provided by boost. The standard library implementation was proven several times to be slow (causing stalls) and to consume too much stack space, especially on ARM. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-2-4/460 Title: [RELEASE] ScyllaDB Enterprise 2022.2.4 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2022.2.4, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.2. Related Links Get ScyllaDB Enterprise 2022.2.4 (customers only, or… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-2-4/460 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2022.2.4 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2022.2.4 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2022.2.4, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.2. A list of bug fixes below, each with an open-source reference if available: --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-0-12/461 Title: [RELEASE] ScyllaDB 5.0.12 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.0.12, a bugfix release of the ScyllaDB 5.0 stable branch. Like all past and future 5.x.y releases, it is backward compatible and supports rolling upgrades. Note that th… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-0-12/461 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.0.12 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.0.12 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.0.12, a bugfix release of the ScyllaDB 5.0 stable branch. Like all past and future 5.x.y releases, it is backward compatible and supports rolling upgrades. Note that the latest stable branch of ScyllaDB is 5.1, and you are encouraged to upgrade. Issues fixed in this release: --- ### Page: https://forum.scylladb.com/t/scylladb-user-talk-takeaways-ceo-dor-laor-s-perspective/465 Title: ScyllaDB User Talk Takeaways: CEO Dor Laor’s Perspective - Blog Posts - ScyllaDB Community NoSQL Forum Meta Description: [800x400-blog-summit-2023-user-talk-takeaway] Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-user-talk-takeaways-ceo-dor-laor-s-perspective/465 ## Headings Structure: H1: ScyllaDB User Talk Takeaways: CEO Dor Laor’s Perspective H3: ScyllaDB User Talk Takeaways: CEO Dor Laor’s Perspective H3: Related topics ## Main Content: H1: ScyllaDB User Talk Takeaways: CEO Dor Laor’s Perspective H3: ScyllaDB User Talk Takeaways: CEO Dor Laor’s Perspective H3: Related topics A rundown of the impressive NoSQL achievements of 11 teams, spanning stories of extreme scale/| latency/throughput, many database migrations and innovative approaches to Rust & event streaming/queing --- ### Page: https://forum.scylladb.com/t/release-scylladb-rust-driver-0-8-1/466 Title: [RELEASE] ScyllaDB Rust Driver 0.8.1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: Bugfixes: Token awareness now works correctly with execute_iter. (#700) A deadlock that occurs immediately when trying to send a query with latency aware policy set is now fixed. (#697) The documentation used to… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-rust-driver-0-8-1/466 ## Headings Structure: H1: [RELEASE] ScyllaDB Rust Driver 0.8.1 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Rust Driver 0.8.1 H3: Related topics Token awareness now works correctly with execute_iter. (#700) A deadlock that occurs immediately when trying to send a query with latency aware policy set is now fixed. (#697) The documentation used to accidentally strip rust attributes from code examples, which is now fixed. (#684) If the database responds with an error during connection handshake, it will now be properly propagated to users. (#686) The set_retry_policy/get_retry_policy methods were brought back (they were removed previously with the introduction of execution profiles). (#707) The default load balancing policy now supports rack-aware load balancing. (#666) Serialization/deserialization is now possible for array types. (#693) Various extensions to the execution profile API needed for the cpp-rust-driver project were added. (#690) The documentation now contains an example that shows how to connect to a serverless cluster. (#685) --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-9-2023-04-22/467 Title: Last week in scylla-cluster-tests.git master (issue #9; 2023-04-22) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 32c239b8…0a7c775c range are covered. There were 36 non-merge commits from 9 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-9-2023-04-22/467 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #9; 2023-04-22) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #9; 2023-04-22) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 32c239b8…0a7c775c range are covered. There were 36 non-merge commits from 9 authors in that period. Some notable commits: Rocky Linux 9 was recently released, thus we added rocky9 support and new tests and pipelines for testing it. Because we use ‘scylla-machine-image’ docker image based on Scylla 4.1, after dropping 4.1 repos, we lost capability of testing opn EKS. This was fixed by disabling scylla package repos. Default load balancing policy in scylla-tools-java is DCAwareRoundRobinPolicy. For multi DC tests now we add datacenter to c-s command, so it will send requests to all DCs evenly. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-2-5/468 Title: [RELEASE] ScyllaDB Enterprise 2022.2.5 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2022.2.5, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.2. Related Links Get ScyllaDB Enterprise 2022.2.5 (customers only, o… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-2-5/468 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2022.2.5 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2022.2.5 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2022.2.5, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.2. One bug was fixed in this release: --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-176-2023-04-23/469 Title: Last week in scylladb.git master (issue #176; 2023-04-23) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the c501163f95…bd0b299322 range are covered. There were 106 non-merge commits from 11 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-176-2023-04-23/469 ## Headings Structure: H1: Last week in scylladb.git master (issue #176; 2023-04-23) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #176; 2023-04-23) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the c501163f95…bd0b299322 range are covered. There were 106 non-merge commits from 11 authors in that period. Some notable commits: A recent regression that caused SUM() of FLOAT or DOUBLE data types to return an error if the input contained both +Inf and -Inf has been fixed. The jmx process provides compatibility with Cassandra jmx and an interface to nodetool. It now requires Java 11 instead of Java 8. A bug in the nodetool command to disable autocompaction has been fixed. The sstable parser is now able to detect more types of corruption involving premature end-of-file with compressed sstables. A bug in topology management with Raft, when starting up a node, has been fixed. The old gossip-based failure detector has been removed. We now use the direct failure detector exclusively. Raft now uses microsecond resolution for its schema and topology tombstones, rather than millisecond resolution, correcting problems when rapidly deleting entries. The tombstone_gc feature, which allows tombstones to be garbage-collected as soon as repair completes, has been marked as ready for general use and is no longer experimental. When Raft-based schema and topology management is in use, it will also manage the Change Data Capture (CDC) generation table. This increases reliability of this operation. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/scylladb-cloud-monitoring-code-error-04002/470 Title: ScyllaDB Cloud Monitoring - Code error - "04002" - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: When I try to access the ScyllaDB Cloud Monitoring Dashboard, I get the following code error: 04002 (see below). How can I solve this? Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-cloud-monitoring-code-error-04002/470 ## Headings Structure: H1: ScyllaDB Cloud Monitoring - Code error - "04002" H3: Related topics ## Main Content: H1: ScyllaDB Cloud Monitoring - Code error - "04002" H3: Related topics When I try to access the ScyllaDB Cloud Monitoring Dashboard, I get the following code error: 04002 (see below). How can I solve this? To solve the issue, clear the browser cached data for ScyllaDB Cloud. On Chrome: In the search field type “scylla” Click on the trash can icon next to lines with Scylla Cloud and choose “clear” Close all the tabs with Scylla Cloud, open a new one, sign in, and browse to the monitoring page Notice that it might be different on a different browser. --- ### Page: https://forum.scylladb.com/t/cqlsh-connection-error-cqlsh-py-error-documents-is-not-a-valid-port-number/472 Title: CQLSH Connection error: cqlsh.py: error: 'Documents' is not a valid port number - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: When trying to connect to my cluster with the CQL Shell I get an error: “cqlsh.py: error: ‘Documents’ is not a valid port number” Please help. Language: en Canonical URL: https://forum.scylladb.com/t/cqlsh-connection-error-cqlsh-py-error-documents-is-not-a-valid-port-number/472 ## Headings Structure: H1: CQLSH Connection error: cqlsh.py: error: 'Documents' is not a valid port number H3: Related topics ## Main Content: H1: CQLSH Connection error: cqlsh.py: error: 'Documents' is not a valid port number H3: Related topics When trying to connect to my cluster with the CQL Shell I get an error: “cqlsh.py: error: ‘Documents’ is not a valid port number” This can happen when a password contains special characters. To solve it, enclose the password string with single quotes like so: --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2021-1-19/474 Title: [RELEASE] ScyllaDB Enterprise 2021.1.19 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2021.1.19, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2021.1. ScyllaDB Enterprise 2022.1 LTS is the latest Long-Term Support rele… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2021-1-19/474 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2021.1.19 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2021.1.19 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2021.1.19, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2021.1. ScyllaDB Enterprise 2022.1 LTS is the latest Long-Term Support release, and 2022.2 is the newest feature release. You are encouraged to upgrade in coordination with the ScyllaDB support team. The following issues are fixed in this release (with an open-source reference): --- ### Page: https://forum.scylladb.com/t/building-a-feature-store-with-scylladb-sample-app-on-github/478 Title: Building a feature store with ScyllaDB (sample app on GitHub) - Announcements - ScyllaDB Community NoSQL Forum Meta Description: We’ve recently published a GitHub repository that contains a sample machine learning feature store application using ScyllaDB. The project provides an easy first step for you to get more familiar with ScyllaDB and how y… Language: en Canonical URL: https://forum.scylladb.com/t/building-a-feature-store-with-scylladb-sample-app-on-github/478 ## Headings Structure: H1: Building a feature store with ScyllaDB (sample app on GitHub) H3: Related topics ## Main Content: H1: Building a feature store with ScyllaDB (sample app on GitHub) H3: Related topics We’ve recently published a GitHub repository that contains a sample machine learning feature store application using ScyllaDB. The project provides an easy first step for you to get more familiar with ScyllaDB and how you can incorporate ScyllaDB into your machine learning project as a feature store - more specifically as an online storage solution. The project is still in development so feel free to share your questions and suggestions below and how we should improve the tutorial. Get started on GitHub! --- ### Page: https://forum.scylladb.com/t/any-restriction-on-number-of-tables-in-a-keyspace/482 Title: Any restriction on number of tables in a keyspace? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Are there any system or operational considerations to creating a large number of tables within a key space? e.g. creating one table per user in a scenario where they may be thousands of users? Language: en Canonical URL: https://forum.scylladb.com/t/any-restriction-on-number-of-tables-in-a-keyspace/482 ## Headings Structure: H1: Any restriction on number of tables in a keyspace? H3: Related topics ## Main Content: H1: Any restriction on number of tables in a keyspace? H3: Related topics Are there any system or operational considerations to creating a large number of tables within a key space? e.g. creating one table per user in a scenario where they may be thousands of users? Having big number of tables would have an effect on topology changes, it would slow them down. It would slow down any other operation that works per table, like repairs backups. We are testing with 5000 tables in one keyspace regularly Same reasons as in cassandra: Impacts of many tables in a Cassandra data model. It’s not recommended, scylla is known to work with a few thousand tables, but it has a price. and I would recommend measuring it before hand, that you are willing to pay it What about num of keyspaces? If we use Alternator interface, it will create a separate keyspace for each table, so the number of keyspaces is equal to the number of tables. Many keyspaces cause problems of similar nature to may tables. They bloat data structures and everything involving keyspaces will take longer. Just to clarify, by many, I mean thousands. I think you should be fine with a 3 digit number amount of them. --- ### Page: https://forum.scylladb.com/t/release-scylladb-cloud-24-april-2023/483 Title: [RELEASE] ScyllaDB Cloud - 24 April 2023 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: We are happy to announce a new update to the ScyllaDB Cloud service. This release contains new self-service user management capabilities. It includes the ability for Admins to manage account users, invite users to their… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-cloud-24-april-2023/483 ## Headings Structure: H1: [RELEASE] ScyllaDB Cloud - 24 April 2023 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Cloud - 24 April 2023 H3: Related topics We are happy to announce a new update to the ScyllaDB Cloud service. This release contains new self-service user management capabilities. It includes the ability for Admins to manage account users, invite users to their account, assign roles to users as well as additional user and account management related features. Upon your next App login, you will receive a one-time password reset email (as per our recent announcement). As soon as you reset your password you will be able to enter your existing account and will have access to the new user management capabilities. The new user management features can be found by clicking the “Settings” button from the bottom left area of the ScyllaDB Cloud app as part of the sidebar. In addition, if you are a member of multiple ScyllaDB Cloud accounts, you will be able to switch between such accounts via the dropdown located at the top of the sidebar. For more information, please visit our docs here. --- ### Page: https://forum.scylladb.com/t/alternator-api-connect-to-scylla-using-golang/484 Title: Alternator API - connect to Scylla using Golang - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I am trying to write a client in Golang using AWS SDK to connect to Scylla through Alternator interface. I was able to do that in python using this script: import boto3 dynamodb = boto3.resource('dynamodb', endpoint_url… Language: en Canonical URL: https://forum.scylladb.com/t/alternator-api-connect-to-scylla-using-golang/484 ## Headings Structure: H1: Alternator API - connect to Scylla using Golang H3: Related topics ## Main Content: H1: Alternator API - connect to Scylla using Golang H3: Related topics I am trying to write a client in Golang using AWS SDK to connect to Scylla through Alternator interface. I was able to do that in python using this script: As previously answered via Slack, you can find a nice example on how to integrate the AWS SDK for Go with Alternator under alternator-load-balancing/try.go at master · scylladb/alternator-load-balancing · GitHub The example in question makes use of our Alternator Load Balancing library to ensure requests are load balanced automatically on the client side. Should you not be using it, just remove the references to the library and proceed as usual. --- ### Page: https://forum.scylladb.com/t/scylladb-consistency-level-client-degradation/486 Title: ScyllaDB Consistency level client degradation - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, we faced with problem of degradation of consistency level under the stress tests. Problem: an error “Cassandra.InvalidQueryException: SERIAL is not supported as conditional update commit consistency. Use ANY if you … Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-consistency-level-client-degradation/486 ## Headings Structure: H1: ScyllaDB Consistency level client degradation H3: Related topics ## Main Content: H1: ScyllaDB Consistency level client degradation H3: Related topics Hi, we faced with problem of degradation of consistency level under the stress tests. Problem: an error “Cassandra.InvalidQueryException: SERIAL is not supported as conditional update commit consistency. Use ANY if you mean “make sure it is accepted but I don’t care how many replicas commit it for non-SERIAL reads”” was occured, but there is no setup of CL = Serial in our client code. Cluster configuration: 6 nodes, RF = 3, CL = LocalQuorum DefaultRetryPolicyWithConfiguredWriteTimeoutAttempts code: Driver: cassandra c# driver v3.18.0 Problem details: After 2 minutes of stress test we found errors like: Known Issue: all nodes except 1 was unavailable during network connectivity errors: Could you help with understanding why scylla/cassandra driver decided to change consistency level to serial in that situation? Please try explicitly specifying ConsistencyLevel in QueryOptions. Somehow they inherit a wrong default for you - either from the environment, or it’s a “feature” of Cassandra C# driver. Cassandra doesn’t complain about misuse of consistency levels, so often such mistakes go unnoticed when working with it. LEARN consistency is meant here. It is not the same as Paxos commit consistency. it has to be one of the QUORUM, LOCAL_QUORUM, ANY, etc. It’s how many acks we wait before we return to user at Paxos LEARN round. The learn round does not impact much, the commit happens in the previous round, so this consistency level doesn’t impact durability. Thanks for your answer! I explicitly pass execution profile name into the batch operations via mapper (stress test above check performance of the write operations) like that: and there are many log records with correct retries: but after few minutes of test an error with CL Serial was occured. The local quorum is necessary because of the requirement of strong consistency (not eventually consistency). After successful write operation I should be able to read data with 100% confidence that I’ll receive the actual last version of data. Please note that you need to read with SERIAL to guarantee true read linearizability. I think the next step is to trace tcpdump. It can be a driver issue. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-10-2023-04-28/488 Title: Last week in scylla-cluster-tests.git master (issue #10; 2023-04-28) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 6fcdaa4d…ab7a253b range are covered. There were 17 non-merge commits from 6 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-10-2023-04-28/488 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #10; 2023-04-28) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #10; 2023-04-28) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 6fcdaa4d…ab7a253b range are covered. There were 17 non-merge commits from 6 authors in that period. Some notable commits: We fixed workload in multi-dc tests by creating loader in each region, where due to DCAwareRoundRobinPolicy we were sending all requests to one DC only. Since recently we enabled print_kernel_callstack by default, to prevent errors, now we set kernel.perf_event_paranoid=0 during k8s setup. Provision step in jenkins now supports multi-az (AWS), so we provision instances more efficiently. New 3-day longevity test with 1TB of data using STCS was added. Improved haproxy ingress logs collection and now we have logs even when it was restarted many times during the test. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/different-cpu-consumption-by-scylla-threads-with-different-linux-kernels-after-the-nodetool-drain-command/494 Title: Different cpu consumption by scylla threads with different linux kernels after the "nodetool drain" command - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi All, The question is about different cpu consumption by scylla threads with different linux kernels after the nodetool drain command. All the results are from a single node system, but behavior of multi-node systems… Language: en Canonical URL: https://forum.scylladb.com/t/different-cpu-consumption-by-scylla-threads-with-different-linux-kernels-after-the-nodetool-drain-command/494 ## Headings Structure: H1: Different cpu consumption by scylla threads with different linux kernels after the "nodetool drain" command H3: Related topics ## Main Content: H1: Different cpu consumption by scylla threads with different linux kernels after the "nodetool drain" command H3: Related topics The question is about different cpu consumption by scylla threads with different linux kernels after the nodetool drain command. All the results are from a single node system, but behavior of multi-node systems is nearly the same. scylladb 5.1.5 Open Source. We suspect, that different kernel functions are called depending on an OS kernel version at least, and this explains different behavior. The details are below. Questions: Is this known behavior? Does it work as designed? top [-1] -H -n1 -b -p $(pidof scylla) Linux kernels 3.x / 4.x Ubuntu 18.04, Centos 7/8, RHEL 8.1 Linux kernels 5.x Ubuntu 20.04, Centos 7 (5.x kernel is installed manually) On distros with the top -1 option available we see that first 1 or 2 threads are 100% busy. The situation is slightly different in a multi-node environment: On the drained node the reactor-1 thread consumes 100% cpu as well (2 theads are 100% busy in this case). But not on other nodes, where the main scylla process consumes 100 of cpu only. strace -p $(pidof scylla) -c Linux kernels 3.x / 4.x Ubuntu 18.04, Centos 7/8, RHEL 8.1 Linux kernels 5.x Ubuntu 20.04, Centos 7 (5.x kernel is installed manually) The answer is in the strace output. On the older kernels ScyllaDB falls back to use epoll for polling the kernel for I/O. This is not as efficient and involves the DB busy-polling the kernel, hence the 100% CPU usage, even when idle. On newer kernels, where it is available, ScyllaDB will use the AIO kernel interface to poll for I/O completion, which is much more efficient. Botond, thanks for your answer. Are these statements correct? Yes, your assessment is correct. To provide some more detail on the kernel check: Seastar, the application framework on top of which ScyllaDB is implemented, has a various reactor backend implementations. It chooses the best one, based on what the kernel it is running on supports. Currently, the following backends exist (in order of preference): Seastar will try to create the best one it can, finally falling back to epoll which should be supported on any kernel currently still supported on any distro. Note that there is also a command-line flag (--reactor-backend) which allows you to select the desired backend (out of thoose supported on your kernel). You can look at the available options (scylla --help-seastar) to see what is available. Botond, thanks a lot for your detailed explanation! Hello, we have found that reactor_backend_epoll::kernel_submit_work() calls epoll_wait with zero timeout and it leads the problem with 100% cpu consumption. Could you please explain us the idea of using epoll_wait with zero timeout here? Is it possible to release cpu in some cases? The event loop of seastar (the reactor) must not be blocked, it is always looking for work (polling). Events that unblock a currently blocked fiber can come from multiple sources. Therefore, the event loop must not block waiting for any individual source. With the AIO backend, the seastar reactor has a sleep mode, which it can use when it has nothing to do (waiting for events). I do not know why this is not available with the epoll backend and why it has to resort to busy polling. Thank you for the quick answer. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-177-2023-04-30/497 Title: Last week in scylladb.git master (issue #177; 2023-04-30) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the bd0b299322…a93e5698b0 range are covered. There were 124 non-merge commits from 19 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-177-2023-04-30/497 ## Headings Structure: H1: Last week in scylladb.git master (issue #177; 2023-04-30) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #177; 2023-04-30) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the bd0b299322…a93e5698b0 range are covered. There were 124 non-merge commits from 19 authors in that period. Some notable commits: ScyllaDB currently decides on replica sets for each partition using vnodes. Vnodes can be agreed on by the cluster using eventual consistency, but the trade-off is that the replication decisions are very simple. We now have initial support for a new replication algorithm, tablets. Tablets allow each table to make different choices about how it is replicated and for those choices to change dynamically as the table grows and shrinks. Tablets require strong consistency from the control plane; this is provided by Raft. Tablets are experimental and currently lack many basic features. ScyllaDB avoids allocating large contiguous memory buffers, as these stress the memory allocator. Instead, ScyllaDB uses fragmented buffers which are easier to allocate. However, most compression libraries do not work with fragmented buffers, so large linear buffers are sometimes necessary. ScyllaDB had a mechanism in place to reuse such buffers in order to avoid allocating them for every request, and here it is tightened so that reallocations are even less common. In addition its usage is corrected in sstables. After a replacenode operation, if the new node had the same IP address as the node it was replacing, the IP address was not moved from pending state to normal state. This is now fixed. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/how-alternator-creates-mv-with-2-columns-not-present-in-pk-of-original-table/499 Title: How alternator creates MV with 2 columns not present in PK of original table? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I want to utilize single table design in Scylla without resorting to DynamoDB. I suspected it does something non trivial to support Dynamo GSI Hash and Sort keys so I created a sample table through alternator and when I … Language: en Canonical URL: https://forum.scylladb.com/t/how-alternator-creates-mv-with-2-columns-not-present-in-pk-of-original-table/499 ## Headings Structure: H1: How alternator creates MV with 2 columns not present in PK of original table? H3: Related topics ## Main Content: H1: How alternator creates MV with 2 columns not present in PK of original table? H3: Related topics I want to utilize single table design in Scylla without resorting to DynamoDB. I suspected it does something non trivial to support Dynamo GSI Hash and Sort keys so I created a sample table through alternator and when I ran DESC on the keyspace I see the impossible: Evidently I get error trying to replicate this schema manually since only 1 extra column not present in PK is allowed. How is this possible and what should I do to be able to create alternator-like secondary hash and sort keys? As you noticed, in CQL (both Cassandra and Scylla), when you create a materialized view you are limited to adding at most one regular column from the base table to the view’s key. You are not allowed to add two. The reason for this limitation appears sporadically on documents on the web, but isn’t documented well enough and already three years ago I opened an issue Document better the restriction of a view to have only one new key column · Issue #6714 · scylladb/scylladb · GitHub to document it better (but unfortunately never got around to it). The reason why adding two base attributes to the view key was not allowed wasn’t some laziness or oversight, but rather a genuine problem that can happen if the user can manipulate the “liveness” properties of these two attributes - their timestamp and ttl - separately - and whether we can handle in such case the liveness of the view row. So, you may be asking now, why did we allow doing this seemingly-dangerous thing in Alternator? The thinking was (and issue #6714 was created to explain this in more detail) that because Alternator does not have TTLs, and doesn’t let you control timestamps (writes always use the current timestamps, later writes always have a later timestamp), it is not vulnerable to this problem. So we allow this case in Alternator, which is lucky because we need this feature for DynamoDB compatibility. In the aforementioned issue, I also suggested that we should “discuss whether it might be possible for a CQL user who promises to adhere to certain limitations on the TTLs or timestamps, that are similar to Alternator’s, could also be allowed to use multiple regular base columns for the view key.”. But unfortunately, this doesn’t currently exist - the code doesn’t allow you to create such table via CQL. By the way, Cassandra also doesn’t have this feature, and a request to add it has been open for 8 years now - [CASSANDRA-9928] Add Support for multiple non-primary key columns in Materialized View primary keys - ASF JIRA. Thanks a ton! I stumbled upon this exact issue and even ended up down to reading the same unresolved issues with Cassandra. I understand the problem is genuine and requires user to restrict the usage of some advanced features of Scylla. Still, it would be great to eventually expose an option to allow this from CQL and remember not to touch that features. A better option would be to enforce that on DDL level so the table knows it has multiple column MVs and disables conflicting subset of CQL. --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-2-rc5/501 Title: [RELEASE] ScyllaDB 5.2 RC5 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The Scylla team is pleased to announce ScyllaDB Open Source 5.2 RC5, a Release Candidate for Scylla Open Source 5.2 minor release. We encourage you to run ScyllaDB 5.2 release candidates on your test environments; this … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-2-rc5/501 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.2 RC5 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.2 RC5 H3: Related topics The Scylla team is pleased to announce ScyllaDB Open Source 5.2 RC5, a Release Candidate for Scylla Open Source 5.2 minor release. We encourage you to run ScyllaDB 5.2 release candidates on your test environments; this will help ensure that an upgrade to ScyllaDB 5.2 General Availability will proceed smoothly with your workload. Use the release candidate with caution; RC5 is not production-ready yet. You can help stabilize Scylla Open Source 5.2 by reporting bugs here. Only the last two minor releases of the ScyllaDB Open Source project are supported. Once Scylla Open Source 5.2 is officially released, ScyllaDB Open Source 5.2 and 5.1 will be supported, and ScyllaDB 5.0 will be retired. For a complete description of ScyllaDB 5.2 see https://forum.scylladb.com/t/release-scylla-5-2-rc1/330. ScyllaDB 5.2 RC1 , RC2, RC3, RC4 Get ScyllaDB Open Source 5.2 (under “More Versions” for each distro) Updates and bug fixes since 5.2 RC4 (not including tests, docs updates) Stability: Adding nodes to a large cluster (90+ nodes) may cause existing nodes to crash. The root cause is quadratic behavior in get_address_ranges function #12724 Stability: a rare issue when updating the schema more than once in once per millisecond #13594 Stability: Compaction manager “periodic reevaluation” is one-off. The Compaction manager intended to periodically reevaluate compaction needs for each registered table. But it’s not working as intended. The reevaluation is one-off. This means that compaction was not kicking in later for a table, with low to none write activity, that had expired data 1 hour from now. #13430 Stability: Internal error in a COUNT request with empty IN. The query “select count(*) from {table1} where p in ()” should result in the count 0, because the empty p in () matches no row. However, what we get in Scylla now is an internal error. #12475 Stability: a very rare failure in a mutation validation failure test, when Materialized View building stops just as it reaches the end of the base table, deciding it needs to stop and also reset the reader to the start token at the same time while the compaction state also thinking the last read partition will be resumed. #12629 Tools: total disk space used metric incorrectly tells the amount of disk space ever used, which is wrong. It should tell the size of all SSTable being used plus the ones waiting to be deleted. Live disk space used shouldn’t account for the ones waiting to be deleted, and live SSTable Count shouldn’t account SSTable waiting to be deleted. #12717 Stability: reader_concurrency_semaphore: inactive reader eviction is too aggressive, leading to Reads timing out, especially in range scans. #11803 Stability: scylla hangs on shutdown in test_total_space_limit_of_commitlog dtest #12810 Stability: Segmentation fault upon wrong named bind markers passed by driver (identify in Rust drivers tests) #12727 Performance: SSTable set is left uncompacted post off-strategy compaction completion, which may lead to suboptimal read and space amplification until the next compaction #13429 CQL: TTL unexpected behavior when setting to 0 on a table with default_time_to_live #6447 CQL: UDA keep using old UDF even after the UDF is replaced #12709 UDF and UDA are experimental in Scylla 5.2 Stability: a rare crash due to null pointer dereference: clear_gently of disengaged unique_ptr dereferences nullptr #13636 --- ### Page: https://forum.scylladb.com/t/nosql-walkthrough-using-spring-boot-time-series-data-with-scylladb/504 Title: NoSQL Walkthrough: Using Spring Boot & Time Series Data with ScyllaDB - Blog Posts - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/nosql-walkthrough-using-spring-boot-time-series-data-with-scylladb/504 ## Headings Structure: H1: NoSQL Walkthrough: Using Spring Boot & Time Series Data with ScyllaDB H3: Tutorial: Spring Boot & Time Series Data in ScyllaDB H3: Related topics ## Main Content: H1: NoSQL Walkthrough: Using Spring Boot & Time Series Data with ScyllaDB H3: Tutorial: Spring Boot & Time Series Data in ScyllaDB H3: Related topics Learn how to use Spring Boot apps with ScyllaDB for time series data, taking advantage of shard-aware drivers and prepared statements. --- ### Page: https://forum.scylladb.com/t/cdc-all-columns-in-delete-operations/510 Title: CDC - All columns in delete operations - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hey guys, We are using CDC on some scylladb tables to capture the operations that take place. A doubt arose regarding the log of delete operations, is it possible to configure so that delete operations fill all columns… Language: en Canonical URL: https://forum.scylladb.com/t/cdc-all-columns-in-delete-operations/510 ## Headings Structure: H1: CDC - All columns in delete operations H3: Preimages and postimages | ScyllaDB Docs H3: Related topics ## Main Content: H1: CDC - All columns in delete operations H3: Preimages and postimages | ScyllaDB Docs H3: Related topics We are using CDC on some scylladb tables to capture the operations that take place. A doubt arose regarding the log of delete operations, is it possible to configure so that delete operations fill all columns with data from before deletion? Currently only primary key columns are populated with information in the CDC log table The docs says that you can configure CDC to receive the previous values like: CREATE TABLE ks.t (pk int, ck int, v1 int, v2 map, PRIMARY KEY (pk, ck)) WITH cdc = {‘enabled’: true, ‘preimage’: ‘full’}; However: "Preimage rows are only created for rows modified using inserts, updates, or row deletes. They are not created for range deletes or partition deletes. " ScyllaDB is an Apache Cassandra-compatible NoSQL data store that can handle 1 million transactions per second on a single server. as Dor said - you can try enabling preimages, which will show you the state of the row from before the deletion (but preimages don’t appear for range/partition deletions). Note however that preimages/postimages have downsides which you should consider before enabling them: If you use only delta rows (which appear when you enable CDC without any additional options such as ‘preimage’), you won’t get previous data; you’ll only get the applied difference (“delta”) which, in case of deletions, is described by the primary key and cdc$operation column with the appropriate value. Thank you @Dor_Laor and @kbr !! --- ### Page: https://forum.scylladb.com/t/release-scylladb-cloud-2-may-2023/511 Title: [RELEASE] ScyllaDB Cloud - 2 May 2023 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: We are happy to announce a new update to the ScyllaDB Cloud service. New features and enhancements: A new Serverless free trial providing you with free credit hours which you can spend across multiple Serverless data… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-cloud-2-may-2023/511 ## Headings Structure: H1: [RELEASE] ScyllaDB Cloud - 2 May 2023 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Cloud - 2 May 2023 H3: Related topics We are happy to announce a new update to the ScyllaDB Cloud service. New features and enhancements: A new Serverless free trial providing you with free credit hours which you can spend across multiple Serverless database clusters. The trial is designed to help you get familiar with our drivers, code samples, and developer resources for building low-latency distributed applications. We’ve added support for i4i and i3en AWS instances in the ap-south-1 (Mumbai) region. --- ### Page: https://forum.scylladb.com/t/top-mistakes-with-scylladb-intro-infrastructure/512 Title: Top Mistakes with ScyllaDB: Intro & Infrastructure - Blog Posts - ScyllaDB Community NoSQL Forum Meta Description: Top Mistakes with ScyllaDB: Intro & Infrastructure - ScyllaDB A new blog series on how to avoid common mistakes with ScyllaDB. First up: selecting a suitable infrastructure. Language: en Canonical URL: https://forum.scylladb.com/t/top-mistakes-with-scylladb-intro-infrastructure/512 ## Headings Structure: H1: Top Mistakes with ScyllaDB: Intro & Infrastructure H3: Related topics ## Main Content: H1: Top Mistakes with ScyllaDB: Intro & Infrastructure H3: Related topics Top Mistakes with ScyllaDB: Intro & Infrastructure - ScyllaDB A new blog series on how to avoid common mistakes with ScyllaDB. First up: selecting a suitable infrastructure. --- ### Page: https://forum.scylladb.com/t/scylla-manager-agent-cannot-connect-to-scylla-api-in-kubernetes/514 Title: Scylla manager agent cannot connect to Scylla API in kubernetes - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: We have scylla running in a cluster deployed using helmfile in kubernetes. Last Friday, we upgraded our scylla pods by bumping their CPU. Since then one of the pods has been stuck in CrashLoopBackOff with the following … Language: en Canonical URL: https://forum.scylladb.com/t/scylla-manager-agent-cannot-connect-to-scylla-api-in-kubernetes/514 ## Headings Structure: H1: Scylla manager agent cannot connect to Scylla API in kubernetes H3: Related topics ## Main Content: H1: Scylla manager agent cannot connect to Scylla API in kubernetes H3: Related topics We have scylla running in a cluster deployed using helmfile in kubernetes. Last Friday, we upgraded our scylla pods by bumping their CPU. Since then one of the pods has been stuck in CrashLoopBackOff with the following error in scylla-manager-agent: Today another one of our pods fell over with the same error logs in the manager agent. The scylla logs also show: What can we do to get these pods to connect to the Scylla Server? I also see these errors in the scylla container: We found out that on affected pods, scylla API Is not running at all: Does anybody know why this would be happening? We haven’t found any error logs indicating why this would be happening other than: We figured out the issue. We over-provisioned CPU beyond the limitations of the underlying node. Kindly share detail config of CPU parameter? thank you in advance! --- ### Page: https://forum.scylladb.com/t/release-scylla-5-2-0-release-part-1/519 Title: [RELEASE] Scylla 5.2.0 Release - part 1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of ScyllaDB Open Source 5.2.0, a production-ready release of our open-source NoSQL database. ScyllaDB 5.2 introduces Raft-based Strongly Consistent Schema Management,… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-5-2-0-release-part-1/519 ## Headings Structure: H1: [RELEASE] Scylla 5.2.0 Release - part 1 H2: Moving to Raft based cluster management H2: New Features H3: Strongly Consistent Schema Management H3: Alternator TTL H3: Large Collection Detection H3: Automating away the gc_grace_seconds parameter H3: Materialized Views: Synchronous Mode H3: Empty replica pages H3: Secondary index on collection columns H3: Related topics ## Main Content: H1: [RELEASE] Scylla 5.2.0 Release - part 1 H2: Moving to Raft based cluster management H2: New Features H3: Strongly Consistent Schema Management H3: Alternator TTL H3: Large Collection Detection H3: Automating away the gc_grace_seconds parameter H3: Materialized Views: Synchronous Mode H3: Empty replica pages H3: Secondary index on collection columns H3: Related topics The ScyllaDB team is pleased to announce the release of ScyllaDB Open Source 5.2.0, a production-ready release of our open-source NoSQL database. ScyllaDB 5.2 introduces Raft-based Strongly Consistent Schema Management, Alternator TTL, and many more improvements and bug fixes. Only the last two minor releases of the ScyllaDB Open Source project are supported. Once ScyllaDB Open Source 5.2 is officially released, only ScyllaDB Open Source 5.2 and ScyllaDB 5.1 will be supported, and ScyllaDB 5.0 will be retired. Get ScyllaDB Open Source 5.2 as binary packages (RPM/DEB), AWS AMI, GCP Image and Docker image Upgrade from ScyllaDB 5.1 to ScyllaDB 5.2 Consistent Schema Management is the first Raft based feature in ScyllaDB, and ScyllaDB 5.2 is the first release to enable Raft by default. #12572 Starting from ScyllaDB 5.2, all new databases will be created with Raft enabled by default. Upgrading from 5.1 will only use Raft if you explicitly enable it (see upgrade to 5.2 docs). As soon as all nodes in the cluster opt-in to using Raft, the cluster will automatically migrate those subsystems to using Raft, and you should validate it is the case. Once Raft is enabled, updating schema requires a quorum to be executed. For example, in the following use cases, the cluster does not have a quorum and will not allow updating the schema: This is different from the behavior of a ScyllaDB cluster with Raft disabled. Nodes might be unavailable due to network issues, node issues, or other reasons. To reduce the chance of quorum loss, it is recommended to have 3 or more nodes per DC, and 3 or more DCs, for a multi-DCs cluster. To recover from a quorum loss, reviving the failed nodes or fixing the network partitioning is best. If this is impossible, see Raft manual recovery procedure. More on handling failures in Raft here. Schema management operations are DDL operations that modify the schema, like CREATE, ALTER, or DROP for KEYSPACE, TABLE, INDEX, UDT, MV, etc. Unstable schema management has been a problem in past ScyllaDB releases. The root cause is the unsafe propagation of schema updates over gossip, as concurrent schema updates can lead to schema collisions. Once Raft is enabled, all schema management operations are serialized by the Raft consensus algorithm. Additional Raft related updates: In ScyllaDB 5.0 we introduced Time To Live (TTL) to DynamoDB compatible API (Alternator) as an experimental feature. In ScyllaDB 5.2 we promote it to production ready. #12037 #11737 Like in DynamoDB, Alternator items that are set to expire at a specific time will not disappear precisely at that time but only after some delay. DynamoDB guarantees that the expiration delay will be less than 48 hours (though for small tables, the delay is often much shorter). In Alternator, the expiration delay is configurable - it defaults to 24 hours but can be set with the --alternator-ttl-period-in-seconds configuration option. ScyllaDB records large partitions, large rows, and large cells in system tables so that the primary key can be used to deal with them. It additionally records collections with large numbers of elements, since these can cause degraded performance. The warning threshold is configurable: compaction_collection_elements_count_warning_threshold - how many elements are considered a “large” collection (default is 10,000 elements). The information about large collections is stored in the large_cells table, with a new collection_elements column that contains the number of elements of the large collection. Large_cells table retention is 30 days. #11449 Example of a large collection below: There is now optional automatic management of tombstone garbage collection, replacing gc_grace_seconds. This drops tombstones more frequently if repairs are made on time, and prevents data resurrection if repairs are delayed beyond gc_grace_seconds. Tombstones older than the most recent repair will be eligible for purging, and newer ones will be kept. The feature is disabled by default and needs to be enabled via ALTER TABLE. cqlsh> ALTER TABLE ks.cf WITH tombstone_gc = {‘mode’:‘repair’}; There is now a synchronous mode for materialized views. In ordinary, asynchronous materialized views, the operation returns before the view is updated. In synchronous materialized view, the operation does not return until the view is updated - each base replica waits for a view replica. This enhances consistency but reduces availability as, in some situations, all nodes might be required to be functional. CREATE MATERIALIZED VIEW main.mv AS SELECT * FROM main.t WITH synchronous_updates = true; ALTER MATERIALIZED VIEW main.mv WITH synchronous_updates = true; Synchronous Mode reference in Scylla Docs Before this release, the paging code requires that pages have at least one row before filtering. This can cause an unbounded amount of work if there is a long sequence of tombstones in a partition or token range, leading to timeouts. ScyllaDB will now send empty pages to the client, allowing progress to be made before a timeout. This prevents analytics workloads from failing when processing long sequences of tombstones. #7689, #3914, #7933 Secondary indexes can now index collection columns. Individual keys and values within maps, sets, and lists can be indexed. Fixes #2962, #8745, #10707 Part 2 of the release notes --- ### Page: https://forum.scylladb.com/t/release-scylla-5-2-0-release-part-2/520 Title: [RELEASE] Scylla 5.2.0 Release - part 2 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: For part one Other Improvements CQL API ScyllaDB now supports server-side DESCRIBE. This is required for the latest cqlsh, and reduces the need to update cqlsh as server features are added. However, the version number … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-5-2-0-release-part-2/520 ## Headings Structure: H1: [RELEASE] Scylla 5.2.0 Release - part 2 H2: Other Improvements H3: CQL API H3: Amazon DynamoDB Compatible API (Alternator) H3: Correctness H3: Performance and stability H3: Operations H3: Deployment and install H3: Tools H3: Configuration H3: Deprecated and removed features H3: Build H3: Monitoring and tracing H3: Related topics ## Main Content: H1: [RELEASE] Scylla 5.2.0 Release - part 2 H2: Other Improvements H3: CQL API H3: Amazon DynamoDB Compatible API (Alternator) H3: Correctness H3: Performance and stability H3: Operations H3: Deployment and install H3: Tools H3: Configuration H3: Deprecated and removed features H3: Build H3: Monitoring and tracing H3: Related topics Alternator, ScyllaDB’s implementation of the DynamoDB API, now provides the following improvements: Support for the new error code has been merged to: Scylla Rust Driver support since version 0.6.0 Gocq - merged to master but it’s not a part of any release. Off-strategy compaction is used when sstables from an external source (such as repair) needs to be reshaped before being handed off to the table’s compaction strategy. If off-strategy compaction is stopped, ScyllaDB used to just leave those sstables in their unreshaped form without compacting them again. It will now hand off such sstables directly to the table’s compaction strategy. #11543. In compaction strategies such as LeveledCompactionStrategy and ScyllaDB Enterprise’s IncrementalCompactionStrategy, sstables are sealed when they reach a certain size (160MB and 1GB respectively). They are not split in the middle of a partition, because we were not able to recover the ordering of such split sstables. This is now possible (though not yet integrated into the compaction strategies). Performance: Long-term index caching in the global cache, as introduced in 4.6, hurts the performance for workloads where accesses to the index are sparse. To mitigate this, a new configuration parameter cache_index_pages (default true) is introduced to control index caching. Setting the flag to false causes all index reads to behave like they would in BYPASS CACHE queries. Consider using false if you notice performance problems due to lowered cache hit ratio in 4.6 or 5.0. The config API can update the parameter live (without restart). #11202 Scylla will now reject a too-low bloom_fulter_fp_chance when creating (or altering) the table, rather than crash while flushing memtables. #11524. ScyllaDB represents reads using mutation fragment streams. Several minor violations of fragment stream integrity were fixed. These could result in incorrect reads during range scans. The log-structured allocator is used to manage cache and memtable memory. When memory runs out, the allocator tries to reclaim memory by evicting cache items and by defragmenting memory. If this takes too long, the allocator logs a stall report. Due to a bug, if the report threshold was set too low then the report is generated even if a stall did not happen, slowing down the system and flooding the logs. This is now fixed. #10981 The compaction manager now ignores out-of-disk-space (ENOSPC) exceptions when shutting down, so the server doesn’t crash in these scenarios. An inaccuracy in the per_partition_rate_limit read metric was corrected. Tables with the per_partition_rate_limit property can throttle read and write activity on a per-partition basis. #11651 A large schema (with thousands of tables) could cause stalls when propagated from node to node. This is now fixed. #11574 A recently introduced regression caused a crash when speculative retry was enabled. This is now fixed. #11825 segmentation fault in cases where the base table schema change while MV schema is cached #10026, #11542 A crash when the compaction manager was asked to stop multiple times (for different reasons) was fixed. A problem with RPC connections being needlessly dropped was fixed. #11780 When a query completes a page, ScyllaDB caches the query activity as an inactive read. When the client requests the next page, ScyllaDB re-activates the read and continues where it left off. A bug in this mechanism that could cause crashes has been fixed. #11923 A crash was fixed during an illegal lightweight transaction INSERT with NULL clustering key ScyllaDB caches rows and (since 4.6) index entries in a single unified cache. It was observed that in some small-partition workloads index caching causes a performance regression, so index caching is now disabled by default. It can still be enabled for workloads that benefit from it. We plan to re-enable it when the regression is fixed. #11889 The topology management code is more relaxed about unknown endpoints to prevent crashes in tests that check for edge cases. This fixes a recent regression. #11870 Hinted handoff now checks that a node exists in topology before doing anything; this helps with a recent regression due to topology refactoring. A crash while fetching repaid ids from the repair history table was fixed. #11966 Usually repair can compare and update the same shard in different nodes, for example shard 3 in one node is compared against shard 3 in another. When the number of shards in nodes is dissimilar, this doesn’t work and each shard compares against data from multiple shards in other nodes. This is now made more efficient by reducing sstable reader thrashing for this dissimilar shard count case. #12157 The algorithm for removing nodes from the token ring was corrected and made more efficient. It’s not known that this had any user impact. #12082 The CQL server will now only run requests that benefit from concurrency (e.g. QUERY and EXECUTE) in parallel. Configuration and authentication related requests will be serialized, reducing the chance for errors in those code paths. A rare bug involving an allocation failure while updating cached rows was fixed. #12068 The system.truncated table holds information about truncation times of user tables. A recent regression caused it to be unreadable by cqlsh. It is now fixed. 12239 COMPACT STORAGE tables allow the user to only specify a prefix of a compound clustering key. Bugs relating to such partial keys and reversed rows were fixed.Note that compact storage is deprecated (see section). #12180 ScyllaDB now supports multiple compaction groups 1. This is not a user-visible feature for now. Some copies of the lists of ranges to stream were eliminated from the decommission path, reducing latency spikes. #12332 When the global index cache is disabled, a local (per query) cache was used instead. When that cache was destroyed, a stall could result, generating a latency spike. This is now fixed. #12271 Compaction manager generally reacts to events to initiate compactions, but also has an hourly timer in case an event was missed (and for tombstone compaction, which isn’t triggered by an event). This timer is now less susceptible to stalls. #12390. Repair tried to trigger off-strategy compaction even for a table that was dropped during repair, failing the entire repair. It ignores the dropped table now. Off-strategy compaction is now enabled for all streaming topology operations (adding and removing nodes). Previously it was enabled only for repair-based node operations. Off-strategy compaction takes advantage of the fact that incoming sstables are non-overlapping to perform more efficient compaction that the one performed by the regular compaction strategy. ScyllaDB sometimes reads ahead of the user request, in order to hide latency. In one case a read-ahead request which timed out caused errors to be emitted, even though this did not affect the query. The errors are now silenced. #12435 ScyllaDB tracks transient memory used by queries. However, it did not track decompressed memory buffers, which could lead to running out of memory in some complicated queries. This is now fixed. “Unset” values are an obscure prepared statement feature that allows only some columns in an UPDATE or INSERT statement to be modified. It was a source of minor bugs and inconvenience in code. The feature has been refactored so it has less impact on the code and is more robust. Fix a crash in Materialized View update row locking, caused by a race condition #12632 (RC1) Fix a crash when reporting error on invalid CQL query involving field selection from a user-defined type #12739 (RC1) Scylla Monitoring Stack release 4.3 and later will support ScyllaDB 5.2. metrics related updates below: --- ### Page: https://forum.scylladb.com/t/scylladb-s-path-to-strong-consistency-a-new-milestone/521 Title: ScyllaDB’s Path to Strong Consistency: A New Milestone - Blog Posts - ScyllaDB Community NoSQL Forum Meta Description: [800x400-blog-scylladbs-path-to-strong-consistency-part-1 (1)] Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-s-path-to-strong-consistency-a-new-milestone/521 ## Headings Structure: H1: ScyllaDB’s Path to Strong Consistency: A New Milestone H3: ScyllaDB’s Path to Strong Consistency: A New Milestone H3: Related topics ## Main Content: H1: ScyllaDB’s Path to Strong Consistency: A New Milestone H3: ScyllaDB’s Path to Strong Consistency: A New Milestone H3: Related topics In ScyllaDB 5.2, Raft is GA and used for the propagation of schema changes. Quickly assembling a fresh cluster, performing concurrent schema changes, updating node's IP addresses – all of this is now possible. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-178-2023-05-07/522 Title: Last week in scylladb.git master (issue #178; 2023-05-07) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the a93e5698b0…aba31ad06c range are covered. There were 142 non-merge commits from 15 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-178-2023-05-07/522 ## Headings Structure: H1: Last week in scylladb.git master (issue #178; 2023-05-07) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #178; 2023-05-07) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the a93e5698b0…aba31ad06c range are covered. There were 142 non-merge commits from 15 authors in that period. Some notable commits: The check for altering permissions of functions in the system keyspace has been tightened. Error messages involving the CQL token function have been improved. In Alternator, ScyllaDB’s implementation of the DynamoDB API, validation of decimal numbers has been improved. ScyllaDB sometimes caches a running query in order to resume it later. In order to do that, it must also store the position at which the query is at, since the cached query might get purged before it is resumed. To do that, it must scan over range tombstones to get to a non-ambiguous primary key position. A bug caused this scan to not terminate, resulting in the query running out of memory. This is now fixed. A crash during shutdown due to incorrect service ordering was fixed. The S3 object-storage driver can now sign multipart-upload requests, enabling it to work with Amazon S3, not just its clones. The sstable validator had a bug in range tombstone validation fixed. The reader_concurrency_semaphore managed memory and concurrency for read queries on replicas. Over time it has evolved, so various internal state names have grown out-of-sync with their actual meaning. The names have been adjusted to reflect their current meaning. The documentation now reflects that ScyllaDB 5.2.0 has been released and is the stable branch. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylla-manager-3-1-release/524 Title: [RELEASE] Scylla Manager 3.1 Release - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB Manager team is pleased to announce the release of Scylla Manager 3.1, a production-ready version of ScyllaDB Manager for ScyllaDB Enterprise customers and ScyllaDB Open Source users. ScyllaDB Manager is a c… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-manager-3-1-release/524 ## Headings Structure: H1: [RELEASE] Scylla Manager 3.1 Release H3: Automated cluster restore procedure: H3: Health Check service: H3: Monitoring H3: Bug Fixes: H3: Related topics ## Main Content: H1: [RELEASE] Scylla Manager 3.1 Release H3: Automated cluster restore procedure: H3: Health Check service: H3: Monitoring H3: Bug Fixes: H3: Related topics The ScyllaDB Manager team is pleased to announce the release of Scylla Manager 3.1, a production-ready version of ScyllaDB Manager for ScyllaDB Enterprise customers and ScyllaDB Open Source users. ScyllaDB Manager is a centralized cluster administration and recurrent tasks automation tool. ScyllaDB Manager 3.1 brings a new, improved backup and restore procedure, as well as other bug fixes. ScyllaDB Enterprise customers are encouraged to upgrade to ScyllaDB Manager 3.1 in coordination with the ScyllaDB support team. The new release includes upgrades of both ScyllaDB Manager Server and Agent. Scylla Manager 3.1 support the following Scylla Enterprise releases: And the following Open Source release (limited to 5 nodes see license): You can install and run Scylla Manager on Kubernetes using Scylla Operator. More here. We introduce a new way of performing the cluster restore. It replaces the old approach that required to run and monitor ansible script with a task creation that defines the backup location and the snapshot tag that end user wants to restore. New restore procedure doesn’t limit the restore to the same cluster topology. Any backups can be restored on any new cluster topologies, assuming it has enough storage, taking advantage of the ScyllaDB Load and Stream feature There is a possibility for restoring either the schema or the data of the cluster. To restore the schema, backup must contain system_schema keyspace data. To restore the data, the schema must be already defined on the destination cluster. Automated restore procedure works well with older backups made with Scylla Manager < 3.1. More information about the new restore procedure is available in the restore section of scylla manager 3.1 documentation. Configuration of health check service is simplified. It requires just one parameter now, which is the max_timeout . It defines the threshold for all ping types after which the TIMEOUT error is reported. Note, that health check reports can be observed via metric exposed by the scylla manager named scylla_manager_healthcheck_{cql | alternator | rest}l_rtt_ms . This metrics shows the response time of CQL, ALTERNATOR and REST pings. See Scylla Monitoring release 4.3.4 or later for Manager 3.1 dashboard, including a new restore status panel. The following metrics are added in Manager 3.1: Manager 3.1 fixes a problem in the backup process: If a node has been replaced and keeps the same host-id it starts enumerating SSTables from zero. Manager deduplication may assume an SSTable was already backed up and skip the latest, new SSTable of the new node. Manager 3.1 fixed this issue moving forward by adding versioning to SSTables. Additional bug fixes: Compatibility issue with ScyllaDB, introduced in ScyllaDB Open Source 5.0. Run repairs on full ranges when tables are below small table threshold or fully replicated. It was possible to fail the repair task when manager metrics weren’t initialized. --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-1-9/525 Title: [RELEASE] ScyllaDB 5.1.9 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.1.9, a bugfix release of the ScyllaDB 5.1 stable branch. ScyllaDB Open Source 5.1.9, like all past and future 5.x.y releases are backward compatible and support rolling … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-1-9/525 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.1.9 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.1.9 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.1.9, a bugfix release of the ScyllaDB 5.1 stable branch. ScyllaDB Open Source 5.1.9, like all past and future 5.x.y releases are backward compatible and support rolling upgrades. Note that the latest stable ScyllaDB Open Source release is 5.2, and you are encouraged to upgrade to it. Issue fixed in this release: CQL: TTL unexpected behavior when setting to 0 on a table with default_time_to_live #6447 CQL API: cql transport server report an unhandled error as SERVER_ERROR to client #12104 Stability: Adding nodes to a large cluster (90+ nodes) may cause existing nodes to crash. The root cause is quadratic behavior in get_address_ranges function #12724 Stability: fix a rare race condition when using LWT with CDC, that causes a test flakiness #12098 Stability: a very rare failure in a mutation validation failure test, when Materialized View building stops just as it reaches the end of the base table, deciding it needs to stop and also reset the reader to the start token at the same time while the compaction state also thinking the last read partition will be resumed. #12629 Stability: reader_concurrency_semaphore: inactive reader eviction is too aggressive, leading to Reads timing out, especially in range scans. #11803 Streaming: stale entries in system_distributed.view_build_status causes unnecessary view building during streaming #11905 Tools: total disk space used metric incorrectly tells the amount of disk space ever used, which is wrong. It should tell the size of all SSTable being used plus the ones waiting to be deleted. Live disk space used shouldn’t account for the ones waiting to be deleted, and live SSTable Count shouldn’t account SSTable waiting to be deleted. #12717 Docker: The docker image is now more robust against different network conditions, which could cause it to fail to launch. #12011 --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-1-6/527 Title: [RELEASE] ScyllaDB Enterprise 2022.1.6 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2022.1.6, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.1. Note that there is a newer Feature Enterprise release: ScyllaDB Ente… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-1-6/527 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2022.1.6 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2022.1.6 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2022.1.6, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.1. Note that there is a newer Feature Enterprise release: ScyllaDB Enterprise 2022.2. While we will continue to support 2022.1 LTS, you can get additional features with 2022.2. Below is a list of performance and stability improvements and bug fixes, each with an open-source reference if available: CQL: Deleting a long base partition may leave some undeleted materialized view rows #12297 CQL: USING TIMESTAMP allows setting the mutation timestamp on a CQL statement level. It has a sanity check that prevents setting timestamps in the future, as these can be hard to delete, but sometimes one wishes to do so anyway. There is now a configuration option that allows disabling the feature. #12527 CQL: scylla: types: is_tuple(): doesn’t handle reverse types. For example, a schema with reversed clustering key component; this component will be incorrectly represented in the schema CQL dump: the UDT will lose the frozen attribute. When attempting to recreate this schema based on the dump, it will fail as the only frozen UDTs are allowed in primary key components. #12576 Alternator Streaming API: unexpected ARN values by list streams paged responses. This issue may only affect users with many, more than 100, tables with streams #12601 Stability: Gossip now uses a better source to get live node information for truncate, reducing false failures of the TRUNCATE operation. #10296, #11928 Stability: Cached reads may temporarily miss rows under rare conditions #12451 Stability: Crash when reporting error on invalid CQL query involving field selection from a user-defined type #12739 Stability: reader_concurrency_semaphore: inactive readers are only evicted on the admission path #11770 Stability: During rebuild on asymmetric cluster several aborts and coredump happened #11923 (introduced by the fix for #11770 above) Stability: Memtable(s) are not flushed when cleaning up a table, leaving disowned tokens in the memtable, which might be resurrected.#1239 Stability: Enabling table encryption (see Encryption at Rest) aborts Scylla when key_provider is not specified. Stability: error “reader_concurrency_semaphore - Semaphore sl:oltp_read_concurrency_sem with 0/100 count and -74538773/0 memory resources”. Root cause is reader_concurrency_semaphore_group partitioning memory is broken. In some cases, this bug may, after a few cycles of updates, lead to semaphore without any memory units unable to admit any reads. In turn this may lead to high latency and even denial of service. Stability: Nodes SegFaulted in short succession after restore. The root cause was that the Workload Prioritization scheduling group was not robust enough. Stability: ScyllaDB tracks transient memory used by queries. However, it did not track decompressed memory buffers, which could lead to running out of memory in some complicated queries. #12448 Stability: crash in Materialized View update row locking, caused by a race condition #12632 Stability: preemption pitfall around waiting for readmission (found is fuzzy tests) #10187 Stability: error “reader_concurrency_semaphore - Semaphore sl:oltp_read_concurrency_sem with 0/100 count and -74538773/0 memory resources”. Root cause is reader_concurrency_semaphore_group partitioning memory is broken. In some cases, this bug may, after a few cycles of updates, lead to semaphore without any memory units unable to admit any reads. In turn this may lead to high latency and even denial of service. Performance: parsers are compiled without inlining, even in release mode #12463 --- ### Page: https://forum.scylladb.com/t/scylla-manager-backup-dry-run-fails-with-giving-up-after-2-attempts-after-30s-context-deadline-exceeded/529 Title: Scylla Manager backup dry run fails with: giving up after 2 attempts: after 30s: context deadline exceeded - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: We are setting up scylladb cluster to be backed up in AWS S3 compatible storage and setup access_key, secet, endpoint in scylla-manager-agent.yaml correctly. From each node the command also works fine: scylla-manager-ag… Language: en Canonical URL: https://forum.scylladb.com/t/scylla-manager-backup-dry-run-fails-with-giving-up-after-2-attempts-after-30s-context-deadline-exceeded/529 ## Headings Structure: H1: Scylla Manager backup dry run fails with: giving up after 2 attempts: after 30s: context deadline exceeded H1: sctool backup -c ‘Scylla Prod 01’ -L ‘s3:’ --dry-run H3: Related topics ## Main Content: H1: Scylla Manager backup dry run fails with: giving up after 2 attempts: after 30s: context deadline exceeded H1: sctool backup -c ‘Scylla Prod 01’ -L ‘s3:’ --dry-run H3: Related topics We are setting up scylladb cluster to be backed up in AWS S3 compatible storage and setup access_key, secet, endpoint in scylla-manager-agent.yaml correctly. From each node the command also works fine: scylla-manager-agent check-location --debug --location s3: But when we try and test backup from scyllamanger node it fails with: Error: get backup target: location is not accessible NOTICE: this may take a while, we are performing disk size calculations on the nodes 2a0b:b580:0:0:xx:xxxx:xx:xx: giving up after 2 attempts: after 30s: context deadline exceeded 2a0b:b580:0:0:xx:xxxx:xx:xx: giving up after 2 attempts: after 30s: context deadline exceeded … The scylladb node manager-agent has these errors: @roy.susmit Thank you for reporting this issue. Please find the scylla-manager git repository GitHub - scylladb/scylla-manager: The Scylla Manager That’s the best place to create issues related to scylla manger. I miss a bit of the information here ,like: The error message you attached: error: no put permission: context canceled is misleading in the context of the issue reported here. The real root cause is the timeout that comes from the scylla-manager. When you perform sctool backup -c ‘Scylla Prod 01’ -L ‘s3:’ --dry-run then it’s the scylla-manager that calls scylla-manager-agent REST API to check the permissions. Scylla manager calls to scylla-manager-agent REST API are timed out after 30s. Scylla-manager-agent tried to interact with S3 API to put a file into the bucket, but the call took more than 30s and was basically cancelled. scylla-manager-agent check-location --debug --location s3: calls are not timed out. They are performed directly on the manager-agent nodes without anything in between the caller and the manager-agent. Pls let us know how much time it took to check the location from the node directly. I suspect that the S3 compatible storage response time is > 30s. Hello @Karol_Kokoszka I wanted to know if you can help I am facing issue testing the dry run backup to azure blob storage the VMs are on prem so I used storage account name and key to update the yaml config file I get this error when I run the sctool backup -c mycluster -L ‘mybackupname’ --dry-run Error: create backup target: location is not accessible ** MYIP: giving up after 2 attempts: agent [HTTP 500] init location: Failed to acquire MSI token: MSI is not enabled on this VM: Get “http://169.254.169.254/metadata/identity/oauth2/token?api-version=2018-02-01&resource=https%3A%2F%2Fstorage.azure.com”: dial tcp 169.254.169.254:80: connect: no route to host - make sure the location is correct and credentials are set, to debug SSH to MYIP and run “scylla-manager-agent check-location -L MYbackupstorage --debug”** when i run this scylla-manager-agent check-location -L MYbackupstorage --debug on all of the nodes it returned fine also when I run the scylla-manager-agent check-location -L azure: on all nodes it returns fine too without any errors the debug on each node returns good ending with the deletion of the test as below on all nodes {“L”:“DEBUG”,“T”:“2023-10-17T13:13:02.875+0100”,“N”:“rclone”,“M”:“Waiting for deletions to finish”} {“L”:“DEBUG”,“T”:“2023-10-17T13:13:03.229+0100”,“N”:“rclone”,“M”:“test: Deleted”} Please how can I resolve the challenge of performing the dry run backup test I also tried to do an adhoc backup same error too Error: create backup target: location is not accessible ** MYIP: giving up after 2 attempts: agent [HTTP 500] init location: Failed to acquire MSI token: MSI is not enabled on this VM: Get “http://169.254.169.254/metadata/identity/oauth2/token?api-version=2018-02-01&resource=https%3A%2F%2Fstorage.azure.com”: dial tcp 169.254.169.254:80: connect: no route to host - make sure the location is correct and credentials are set, to debug SSH to MYIP and run “scylla-manager-agent check-location -L MYbackupstorage --debug”** Hello @Karol_Kokoszka We are also facing the similar problem while running below dry run command from scylla manager got Error: get backup target: location is not accessible , got below logs in scylla manager while at node scylla-manager-agent we got below error In cluster of 3 node, 2 are working fine and getting issue only on one node. All three nodes have same permission, security group , IAM role . scylla-manager-agent check-location --debug --location s3:scylla-manager-backup command works fine on all three nodes. Here main challenge is that we are not able to identify why this issue is coming, as scylla-manager to scylla-manager-agent is kind of blackbox for me. scylla-manager version : 3.2.3 scylla-manager-agent version : 3.2.3 Please help use fix the issue, Thanks when i run this scylla-manager-agent check-location -L MYbackupstorage --debug on all of the nodes it returned fine also when I run the scylla-manager-agent check-location -L azure: on all nodes it returns fine too without any errors Did you restart the scylla-manager-agent service after applying changes to the YAML configuration ? check-location started with scylla-manager-agent CLI is reading the configuration when it’s executed, but when you do the actual backup, it’s the scylla-manager-agent.service that validates the location. If the service is not restarted, then it keeps the old configuration in memory. @Vikram_Pratap_Singh Please double check that you restarted the scylla-manager-agent.service on all nodes after you set up the backup location. @sumohx Did you use use_msi parameter in you scylla-manager-agent.yaml file ? The error “Failed to acquire MSI token: MSI is not enabled on this VM” may appear when this option is enabled for azure storage. Hello @Karol_Kokoszka Thank you for your response. No I didn’t use the use_msi in the yaml file. Infact there is no such parameter There is no such a parameter on guide we provide, but there is the 3rd part lib RClone under the hood that supports it and the param can be provided in scylla-manager-agent.yaml. By using this parameter, user indicates that would like to use MSI authentication on Azure. Anyway, you didn’t use it, so seems that my suspicion is wrong. Is there a chance to see the scylla-manager-agent config YAML file ? @Karol_Kokoszka Hello wanted to know if you were able to take a look at the yaml file I uploaded Thanks @Karol_Kokoszka Hi @Karol_Kokoszka , Also having this same issue. Deployed scylladb, scylla-operator and scylla-managr with helm on Kubernetes and while trying to configure backup to AWS S3 compatible storage, we get the Error: get backup target: location is not accessible. access_key, secet, endpoint in scylla-manager-agent.yaml are correctly set on the pod. From each node the scylla-manager-agent check-location --debug --location s3:command also works fine: But when we try and test backup from scyllamanger node it fails with the error: Error: get backup target: location is not accessible On the side, I’ve also tried using scylla nodetool, while this works when executed inside the scylla pod, when itry to run the command nodetool -h "$SCYLLADB_HOST" -p 7199 snapshot $keyspace from the backup pod in my cluster (In the same namespace), I get the connection refused error. I’ve made the necessary authentication correction to the cassandra-env.sh file and the scylla-jmx file, But i still get the error nodetool: Failed to connect to '[hostname.domain.com:7199]- ConnectException: 'Connection refused (Connection refused)'. Kindly, help with a clear backup documentation for users of scyllaDB on kubernetes. @a3ts can you add --debug=true flag to the check-location command you execute ? It should give bit more meaningful information. I encountered the same issue and found the reason: In the documentation at Setup S3 compatible storage | ScyllaDB Docs, step 6 should be sudo systemctl restart scylla-manager-agent, not start. --- ### Page: https://forum.scylladb.com/t/scylla-alternator-interface-aws-golang-sdk/533 Title: Scylla Alternator interface + AWS Golang SDK - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I have setup a table through the Scylla alternator interface like so (through python script): table_definition = dynamodb.create_table( TableName='display_name_tagging', KeySchema=[ { 'Attrib… Language: en Canonical URL: https://forum.scylladb.com/t/scylla-alternator-interface-aws-golang-sdk/533 ## Headings Structure: H1: Scylla Alternator interface + AWS Golang SDK H3: Related topics ## Main Content: H1: Scylla Alternator interface + AWS Golang SDK H3: Related topics I have setup a table through the Scylla alternator interface like so (through python script): Now using Golang and AWS SDK I am connecting to that keyspace and trying to write a record into a table: but when I run my program, I get this error: ValidationException: Key column a_party not found Doing search online has not produced any results. Any help would be appreciated. I tried to do the same in Python (I’m not familiar with Go), and it worked just fine: The error you got is exactly the kind of error you get when the item does not have an a_party field, for example: Gives you, as expected (because the item is missing the key column a_party): An error occurred (ValidationException) when calling the PutItem operation: Key column a_party not found. So I can only guess that the item you passed to PutItem has the wrong format. I’m not familiar with the Go SDK, but did you really need to call that MarshalMap thing on the item instead of just passing “item” (in the Python SDK, for example, this conversion is done automatically for you)? Can you try just passing “item”? And if it’s necessary, did you use the correct way? Can you please print “nr” after the marshalling and show me what it looks like? Don’t see the full code but looks like a_party is a private field of Item so MarshalMap would produce empty map. Since a_party is a key you can’t put an item without it as nyh wrote above. BTW when writing new project aws advises to use golang aws sdk v2. --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-2-6/536 Title: [RELEASE] ScyllaDB Enterprise 2022.2.6 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2022.2.6, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.2. Related Links Get ScyllaDB Enterprise 2022.2.6 (customers only, or… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-2-6/536 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2022.2.6 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2022.2.6 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2022.2.6, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.2. The following issues are fixed in this release (with an open-source reference, if available): Stability: Adding nodes to a large cluster (90+ nodes) may cause existing nodes to crash. The root cause is quadratic behavior in get_address_ranges function #12724 Stability: a very rare failure in a mutation validation failure test, when Materialized View building stops just as it reaches the end of the base table, deciding it needs to stop and also reset the reader to the start token at the same time while the compaction state also thinking the last read partition will be resumed. #12629 Tools: total disk space used metric incorrectly tells the amount of disk space ever used, which is wrong. It should tell the size of all SSTable being used plus the ones waiting to be deleted. Live disk space used shouldn’t account for the ones waiting to be deleted, and live SSTable Count shouldn’t account SSTable waiting to be deleted. #12717 Stability: reader_concurrency_semaphore: inactive reader eviction is too aggressive, leading to Reads timing out, especially in range scans. #11803 --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-0-13-the-last-patch-release-of-the-5-0-branch/537 Title: [RELEASE] ScyllaDB 5.0.13 - the last patch release of the 5.0 branch - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.0.13, a bugfix release of the ScyllaDB 5.0 stable branch. Like all past and future 5.x.y releases, it is backward compatible and supports rolling upgrades. Note that th… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-0-13-the-last-patch-release-of-the-5-0-branch/537 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.0.13 - the last patch release of the 5.0 branch H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.0.13 - the last patch release of the 5.0 branch H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.0.13, a bugfix release of the ScyllaDB 5.0 stable branch. Like all past and future 5.x.y releases, it is backward compatible and supports rolling upgrades. Note that the latest stable branch of ScyllaDB is 5.2. ScyllaDB 5.0 branch is no longer supported, and this is the last patch release on this branch. Issues fixed in this release: Stability: distributed_loader should first detect the highest generation in a table before allowing new sstables to be created in it #11793. This is a root cause for: View building fail to move staging SSTables to base dir because generations are taken #11789 Stability: reader_concurrency_semaphore: inactive reader eviction is too aggressive, leading to Reads timing out, especially in range scans. #11803 Stability: a very rare failure in a mutation validation failure test, when Materialized View building stops just as it reaches the end of the base table, deciding it needs to stop and also reset the reader to the start token at the same time while the compaction state also thinking the last read partition will be resumed. #12629 Tools: total disk space used metric incorrectly tells the amount of disk space ever used, which is wrong. It should tell the size of all SSTable being used plus the ones waiting to be deleted. Live disk space used shouldn’t account for the ones waiting to be deleted, and live SSTable Count shouldn’t account SSTable waiting to be deleted. #12717 Stability: Adding nodes to a large cluster (90+ nodes) may cause existing nodes to crash. The root cause is quadratic behavior in get_address_ranges function #12724 --- ### Page: https://forum.scylladb.com/t/performance-issues-with-cluster/538 Title: Performance issues with cluster - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I’m having some performance issues with my cluster. It seems like some queries are taking way too long. How can I troubleshoot this? Language: en Canonical URL: https://forum.scylladb.com/t/performance-issues-with-cluster/538 ## Headings Structure: H1: Performance issues with cluster H3: Related topics ## Main Content: H1: Performance issues with cluster H3: Related topics I’m having some performance issues with my cluster. It seems like some queries are taking way too long. How can I troubleshoot this? One of the first things you should do is make sure you have Monitoring in place. Take a look at the Monitoring dashboards, and understand what’s happening with your cluster. Another useful feature is Tracing. This feature allows you to debug queries that perform poorly. Tracing enables analyzing of internal data flows in a cluster. It’s useful for observing the behavior of specific queries. It can help you look into network issues, slow queries, data transfers, and more. There are two types of tracing, client-side tracing and server-side tracing (probabilistic tracing). In addition, there is Slow Query Logging. To use client-side tracing: cqlsh> TRACING ON Which returns: Now Tracing is enabled Continue with normal queries. Each query would now return the result as before, but also a tracing session. This is useful if you have a query that you suspect is causing problems and want to examine it. The tracing data is stored in the system_traces keyspace, which can also be queried directly, for example: cqlsh> select * from system_traces.sessions where session_id=227aff60-4f21-11e6-8835-000000000000; cqlsh> select * from system_traces.events where session_id=227aff60-4f21-11e6-8835-000000000000; Probabilistic Tracing randomly chooses a request to be traced with some pre-defined probability. This is set, per node, using the nodetool settraceprobability command. See more about it here. To set this for an entire cluster, use the command on all nodes. For example, to trace %0.01 of all the queries in the node, use: nodetool settraceprobability 0.0001 Notice that this has an impact on performance, so use it carefully in production. Some example values for the settraceprobability command: So what number should you set? That depends on your workload and on the ops/second in your cluster. Typically you want to turn this on for a specific time window, so don’t forget to turn this off. For example, to collect information for 5 minutes only and then disable tracing (value 0 stops collecting information): Connect to the cluster and on each node run: nodetool settraceprobability 0.001; sleep 5m; nodetool settraceprobability 0 Traces are stored in the system_traces keyspace for 24 hours. The keyspace consists of two tables: So in the above example, you can retrieve the traced sessions or event data using the following statements: SELECT*FROM system_traces.sessions SELECT*FROM system_traces.events Slow Query Logging captures queries that take more time than the given threshold. This is useful if you’re not sure what’s happening in the cluster and you want to find out what’s causing performance issues. Whenever you use tracing, remember it has a performance impact. Don’t enable it by default and use it for small periods of time. More information about Tracing is available in this ScyllaB University lesson and in the Docs, including more advanced topics like Lightweight slow-queries logging mode, Large Partition Tracing, finding Hot Partitions, and more. --- ### Page: https://forum.scylladb.com/t/scylla-cdc-rust-v0-1-released/539 Title: Scylla CDC Rust v0.1 released! - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: We are excited to announce scylla-cdc-rust, a library that makes it easy to develop Rust applications that consume data from a Scylla Change Data Capture Log. Features A simple callback-based interface for consuming ch… Language: en Canonical URL: https://forum.scylladb.com/t/scylla-cdc-rust-v0-1-released/539 ## Headings Structure: H1: Scylla CDC Rust v0.1 released! H2: Features H2: Getting started H3: Related topics ## Main Content: H1: Scylla CDC Rust v0.1 released! H2: Features H2: Getting started H3: Related topics We are excited to announce scylla-cdc-rust, a library that makes it easy to develop Rust applications that consume data from a Scylla Change Data Capture Log. Check out the README for links to a tutorial and the documentation. --- ### Page: https://forum.scylladb.com/t/scylladb-5-2-0-image-for-azure/541 Title: ScyllaDB 5.2.0 Image for Azure - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: As of version ScyllaDB 5.2.0, we now produce a ScyllaDB image for Azure. We encourage Azure users to try it and provide us with feedback. Details on how to utilize this image can be found here. Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-5-2-0-image-for-azure/541 ## Headings Structure: H1: ScyllaDB 5.2.0 Image for Azure H3: Related topics ## Main Content: H1: ScyllaDB 5.2.0 Image for Azure H3: Related topics As of version ScyllaDB 5.2.0, we now produce a ScyllaDB image for Azure. We encourage Azure users to try it and provide us with feedback. Details on how to utilize this image can be found here. --- ### Page: https://forum.scylladb.com/t/last-fortnight-in-scylla-cluster-tests-git-master-issue-11-2023-05-13/543 Title: Last fortnight in scylla-cluster-tests.git master (issue #11; 2023-05-13) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last two weeks. Commits in the a8da7cf7…08f927dd range are covered. There were 35 non-merge commits from 10 authors… Language: en Canonical URL: https://forum.scylladb.com/t/last-fortnight-in-scylla-cluster-tests-git-master-issue-11-2023-05-13/543 ## Headings Structure: H1: Last fortnight in scylla-cluster-tests.git master (issue #11; 2023-05-13) H3: Related topics ## Main Content: H1: Last fortnight in scylla-cluster-tests.git master (issue #11; 2023-05-13) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last two weeks. Commits in the a8da7cf7…08f927dd range are covered. There were 35 non-merge commits from 10 authors in that period. Some notable commits: We reduced number of concurrent load threads in large partitions test to avoid IO/net limits and achieve better CPU utilization. In K8s tests for getting logs we use kubectl command, but it sometimes causes logs to be missing. We implemented logs retrieval with k8s python client with hope of better reliability. Introduced 4 chaos-mesh failures: packet loss, corruption, network delay and limit bandwidth to K8s tests. Disabled speculative retry for performance tests because it could affect on max throughput. Bumped the K8S version to 1.25. Bumped default manager version to 3.1. Since Ubuntu 18 version is EOL, we stopped testing it, for tests that used it we switched to Ubuntu 22. LogCollector now collects log of the manager’s scylla backend as sometimes is required for the investigations of manager bugs. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-179-2023-05-14/544 Title: Last week in scylladb.git master (issue #179; 2023-05-14) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the aba31ad06c…31e820e5a1 range are covered. There were 125 non-merge commits from 18 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-179-2023-05-14/544 ## Headings Structure: H1: Last week in scylladb.git master (issue #179; 2023-05-14) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #179; 2023-05-14) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the aba31ad06c…31e820e5a1 range are covered. There were 125 non-merge commits from 18 authors in that period. Some notable commits: The object storage driver now supports object timestamps, needed for tombstone garbage collection compaction. The build system now automatically compiles Rust and C++ user-defined functions into WebAssembly. Previously, the compiled wasm text was inlined into individual tests. The installer now wipes filesystem signatures from the individual disks making up a RAID array, preventing problems with reuse of disks. A corner case in replica-side concurrency control related to failed requests was fixed. The sstable validation facility now validates the sstable index file (-Index.db). The sstable validator now properly terminates its internal read-ahead, preventing a crash. Raft remote procedure call (RPC) verbs now check that the call arrived at its intended recipient and not somewhere else. During topology changes (adding and removing nodes), the system first reads from the old replica set while writing to both old and new replica sets, then switches to reading and writing from the new replica set. We are now prepared for an intermediate step (not yet used), where we write to both old and new replica set, while reading from the new replica set, to make the transition more robust. User defined function permissions are now dropped when the keyspace containing them is dropped. Permissions are now checked for a user-defined aggregate (UDA) that uses user-defined functions (UDFs). The selector path (expressions in the SELECT clause) now use non-contiguous memory. This reduces latency when selecting large blobs, as non-contiguous memory doesn’t suffer from fragmentation. Immediate mode tombstone garbage collection is a schema feature that requests tombstones to be garbage-collected immediately (without waiting for gc_grace_seconds), but it it ended up expiring TTLed data too early. This is now fixed. When a node synchronizes the schema from another node, if Raft is in use, it will issue a read barrier first to make sure it’s not missing any keyspaces. It’s now possible to disable and enable tombstone compaction on a per-node basis using a REST API endpoint. This is useful if the user knows that all DELETEs were performed with CL=ALL and so there is no risk of data resurrection. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/should-scylla-manager-be-run-per-region-or-one-manager-instance-for-a-multi-region-cluster/546 Title: Should Scylla manager be run per-region, or one manager instance for a multi region cluster? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Since the open source version is limited to 5 instances, could I in theory have 5 instances per region, and each region has a dedicated scylla manager? I suspect that this is not advised, but in theory will this work or… Language: en Canonical URL: https://forum.scylladb.com/t/should-scylla-manager-be-run-per-region-or-one-manager-instance-for-a-multi-region-cluster/546 ## Headings Structure: H1: Should Scylla manager be run per-region, or one manager instance for a multi region cluster? H3: Nodetool repair | ScyllaDB Docs H3: Related topics ## Main Content: H1: Should Scylla manager be run per-region, or one manager instance for a multi region cluster? H3: Nodetool repair | ScyllaDB Docs H3: Related topics Since the open source version is limited to 5 instances, could I in theory have 5 instances per region, and each region has a dedicated scylla manager? I suspect that this is not advised, but in theory will this work or break the cluster? I could stagger their repairs so it’s never concurrent Pining @Guy since he asked Technically it might work, but not advice. You will need to ensure Managers tasks do not overlap and are limited per DC. By default, Manager tasks, like repair, are cluster wide. How would you ensure that they are limited to a single region? Or I guess that’s a silly question, since it seems like the agents call back to a manager instead, so it would just be a matter of pointing an agent to the correct manager. And for reference, I am using a somewhat unusual use case where I want many small nodes in many regions to reduce latency for lookups, so I would be running into the limit of requiring the enterprise license to have scylla manager handle repairs automatically. All these nodes will be the minimum size since each node is not hit very frequently, but I want all data in all regions. Would there be another recommended way to run this in so many regions without doing manual repairs? I’m not aware of other tools for running recurrent repairs on ScyllaDB. Note that repair works between all nodes, including nodes from different data centers. Understood thanks! To clarify you are confirming that running a nodetool repair will repair on all nodes across the cluster in that single command? I.e. I only need to run that once for the whole cluster manually, not on each node? My understanding was that this is only per-node. nodetool repair sync between the node you are running the command from, and its replicas. It does not repair all the nodes. ScyllaDB is an Apache Cassandra-compatible NoSQL data store that can handle 1 million transactions per second on a single server. --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-2-1/552 Title: [RELEASE] ScyllaDB 5.2.1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.2.1, a bugfix release of the ScyllaDB 5.2 stable branch. ScyllaDB Open Source 5.2.1, like all past and future 5.x.y releases, is backward compatible and supports rolling… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-2-1/552 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.2.1 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.2.1 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.2.1, a bugfix release of the ScyllaDB 5.2 stable branch. ScyllaDB Open Source 5.2.1, like all past and future 5.x.y releases, is backward compatible and supports rolling upgrades. Users are encouraged to upgrade to 5.2.1. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/how-does-adding-replacing-decommissioning-node-affect-the-sstables-if-using-twcs/553 Title: How does adding/replacing/decommissioning node affect the SSTables if using TWCS - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: There are a couple of documents and articles online that suggest turning off read repair for tables using the time window compaction strategy (TWCS), as it mixes data that should belong to old buckets into new ones, whic… Language: en Canonical URL: https://forum.scylladb.com/t/how-does-adding-replacing-decommissioning-node-affect-the-sstables-if-using-twcs/553 ## Headings Structure: H1: How does adding/replacing/decommissioning node affect the SSTables if using TWCS H3: Related topics ## Main Content: H1: How does adding/replacing/decommissioning node affect the SSTables if using TWCS H3: Related topics There are a couple of documents and articles online that suggest turning off read repair for tables using the time window compaction strategy (TWCS), as it mixes data that should belong to old buckets into new ones, which leads to higher read amplification. I wonder if some common node operations will cause similar issues. Specifically, adding a node, replacing a dead node, running nodetool repair, running nodetool decommission or running nodetool rebuild. Also, has ScyllaDB implemented some optimizations regarding TWCS compared to Cassandra? @Guy could you help with this, thank you All node operations, as well as repair, will split data into buckets appropriate for the TWCS parameters, so neither streaming, nor repair should cause any mixing of data from different buckets. You can repair your TWCS table, and you can add/remove/rebuild nodes, it shouldn’t cause any mixing of old/new data. You can learn more about TWCS in this lesson. This is a hands-on lab where you can see TWCS in action. AFAIK there are no major differences between TWCS in ScyllaDB and Cassandra. --- ### Page: https://forum.scylladb.com/t/how-can-we-migrate-multiple-tables-simultaneously-with-scylla-migrator/555 Title: How can we migrate multiple tables simultaneously with scylla-migrator - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello Team, Please help, How can we migrate multiple tables from a Keyspace at the same time using the Scylla migration tool? I tried making changes to config.yml, but it did not work. error: Exception in thread “mai… Language: en Canonical URL: https://forum.scylladb.com/t/how-can-we-migrate-multiple-tables-simultaneously-with-scylla-migrator/555 ## Headings Structure: H1: How can we migrate multiple tables simultaneously with scylla-migrator H3: Related topics ## Main Content: H1: How can we migrate multiple tables simultaneously with scylla-migrator H3: Related topics Please help, How can we migrate multiple tables from a Keyspace at the same time using the Scylla migration tool? I tried making changes to config.yml, but it did not work. error: Exception in thread “main” DecodingFailure(Attempt to decode value on failed cursor, List(DownField(table), DownField(source))) Could you show us what your config.yaml looks like? Hi @Attila_Toth , here you see the config.yaml see Multiple tables cannot be migrated simultaneously · Issue #96 · scylladb/scylla-migrator · GitHub --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-1-10/556 Title: [RELEASE] ScyllaDB 5.1.10 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.1.10, a bugfix release of the ScyllaDB 5.1 stable branch. ScyllaDB Open Source 5.1.10, like all past and future 5.x.y releases are backward compatible and support rollin… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-1-10/556 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.1.10 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.1.10 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.1.10, a bugfix release of the ScyllaDB 5.1 stable branch. ScyllaDB Open Source 5.1.10, like all past and future 5.x.y releases are backward compatible and support rolling upgrades. Note that the latest stable ScyllaDB Open Source release is 5.2, and you are encouraged to upgrade to it. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/relese-scylla-monitoring-stack-4-3-1/558 Title: [RELESE] Scylla Monitoring Stack 4.3.1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce a patch release of ScyllaDB Monitoring Stack 4.3.1 Bug Fixes Gaps in recording rules graphs in clusters with many cores #1935 Language: en Canonical URL: https://forum.scylladb.com/t/relese-scylla-monitoring-stack-4-3-1/558 ## Headings Structure: H1: [RELESE] Scylla Monitoring Stack 4.3.1 H3: Related topics ## Main Content: H1: [RELESE] Scylla Monitoring Stack 4.3.1 H3: Related topics The ScyllaDB team is pleased to announce a patch release of ScyllaDB Monitoring Stack 4.3.1 --- ### Page: https://forum.scylladb.com/t/relese-scylla-monitoring-stack-4-3-2/559 Title: [RELESE] Scylla Monitoring Stack 4.3.2 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce a patch release of ScyllaDB Monitoring Stack 4.3.2 Following Grafana Security Release cve-2023-1410 This patch release update Grafana version to 9.3.11 Language: en Canonical URL: https://forum.scylladb.com/t/relese-scylla-monitoring-stack-4-3-2/559 ## Headings Structure: H1: [RELESE] Scylla Monitoring Stack 4.3.2 H3: Related topics ## Main Content: H1: [RELESE] Scylla Monitoring Stack 4.3.2 H3: Related topics The ScyllaDB team is pleased to announce a patch release of ScyllaDB Monitoring Stack 4.3.2 Following Grafana Security Release cve-2023-1410 This patch release update Grafana version to 9.3.11 --- ### Page: https://forum.scylladb.com/t/relese-scylla-monitoring-stack-4-3-4/560 Title: [RELESE] Scylla Monitoring Stack 4.3.4 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce a patch release of ScyllaDB Monitoring Stack 4.3.4 This patch release adds a dashboard for ScyllaDB Manager 3.1.x Language: en Canonical URL: https://forum.scylladb.com/t/relese-scylla-monitoring-stack-4-3-4/560 ## Headings Structure: H1: [RELESE] Scylla Monitoring Stack 4.3.4 H3: Related topics ## Main Content: H1: [RELESE] Scylla Monitoring Stack 4.3.4 H3: Related topics The ScyllaDB team is pleased to announce a patch release of ScyllaDB Monitoring Stack 4.3.4 This patch release adds a dashboard for ScyllaDB Manager 3.1.x --- ### Page: https://forum.scylladb.com/t/scylla-monitoring-advisor-inbalanced-connections/561 Title: Scylla Monitoring Advisor Inbalanced Connections - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi We’re running Scylla in Kubernetes with gocq driver, I’ve replaced the cql driver with the scylla one and so far I don’t have any cql optimizations showing in monitoring. Though I consistently see the connections ar… Language: en Canonical URL: https://forum.scylladb.com/t/scylla-monitoring-advisor-inbalanced-connections/561 ## Headings Structure: H1: Scylla Monitoring Advisor Inbalanced Connections H3: Related topics ## Main Content: H1: Scylla Monitoring Advisor Inbalanced Connections H3: Related topics We’re running Scylla in Kubernetes with gocq driver, I’ve replaced the cql driver with the scylla one and so far I don’t have any cql optimizations showing in monitoring. Though I consistently see the connections are imbalanced and at times the SSTable is also showing a warning. I posted a couple of screenshots, I see the load at least on the charts is not that even at times, but the reads and writes seem pretty even. We have 3 nodes, with an RF of 3 and a CL or quorum. Each node has its own k8 service and we have a load balancer k8 service, which balances across the 3 nodes. This is the service I used to connect our client with and use 19042 port. I’m just trying to figure out why we have these warnings and if connecting with the individual k8 service for each node would be a fix? Or what else I can do to determine the source of the warnings? Any more information needed on this? What is the exact warning you are getting? Also do you have any hot partitions? Sorry for the late reply. It’s more so the dashboard is constantly lit up with unbalanced connections No hot partitions, the only thing I notice and I assume this is a potential reason for the warnings now, though correct me if I’m wrong is: wrong shard being assigned; please check that you are not behind a NAT or AddressTranslater which changes source ports; falling back to non-shard-aware port --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-12-2023-05-20/563 Title: Last week in scylla-cluster-tests.git master (issue #12; 2023-05-20) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 7d6dba9f…6277fb67 range are covered. There were 22 non-merge commits from 7 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-12-2023-05-20/563 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #12; 2023-05-20) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #12; 2023-05-20) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 7d6dba9f…6277fb67 range are covered. There were 22 non-merge commits from 7 authors in that period. Some notable commits: Added support for dynamic local volume provisioner to k8s based tests. It allows storage volumes to be created on-demand by managing directories created on disks attached to instances. On supported filesystems, directories have quota limitations to ensure volume size limits. Currently, it’s fully supported only on EKS backend, and partially supported on local K8S backend (‘kind’). scylla perf-simple-query microbenchmark now is tested with SCT. There was an issue with gossip after restart and new ‘nodetool drain’ K8S functional test was added to reproduce the issue. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/how-does-docker-compose-example-know-where-the-other-nodes-are-it-uses-localhost/564 Title: How does docker-compose example know where the other nodes are? (it uses `localhost`) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Following the example found at scylla-code-samples/mms at master · scylladb/scylla-code-samples · GitHub I see that the configuration used within the first cluster has: # Address or interface to bind to and tell other S… Language: en Canonical URL: https://forum.scylladb.com/t/how-does-docker-compose-example-know-where-the-other-nodes-are-it-uses-localhost/564 ## Headings Structure: H1: How does docker-compose example know where the other nodes are? (it uses `localhost`) H3: Related topics ## Main Content: H1: How does docker-compose example know where the other nodes are? (it uses `localhost`) H3: Related topics Following the example found at scylla-code-samples/mms at master · scylladb/scylla-code-samples · GitHub I see that the configuration used within the first cluster has: Yet running nodetool status I see the cluster: Given that they are advertising localhost, I would assume they would not be able to connect to each other. I do see the seeds, but my assumption is that they wouldn’t be able to talk to each other like this for operations such as repairs and relaying queries. How does localhost allow them to both listen and discover? Is this just some docker networking magic? In the docker-compose of the MMS example, the nodes use a bridge network to allow the containers to communicate. The seed nodes are defined with an IP address, which allows them to discover each other. You can read more about running ScyllaDB on Docker here. I guess I’m wondering why it’s not trying to use local host despite it advertising that as default, but instead overrides and uses the docker ip. Generally, using the command line flags overwrites scylla.yaml settings. In this case, the seed nodes are provided when starting Docker, and that’s how they find each other. The listen_address setting for the node sets the address or interface to bind to and tells other Scylla nodes how to connect to the node. The listen_address is still local host though, it’s not being overridden by a flag. My questions is more how that works but it’s not a burning one. Asking others apparently Docker can resolve that to the containers ip within the network, but not sure if that’s true @danthegoodman1 it’s localhost since it’s using docker, each scylla is listen to it’s own localhost each docker instance has it’s own virtual network device when you configure each one to connect with the others it would be using the actually address, i.e. what set in the broadcast_address --- ### Page: https://forum.scylladb.com/t/issues-running-with-ipv6/565 Title: Issues running with ipv6 - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I’m trying to run scylladb on fly.io and running into issues around ipv6. I have the following custom Dockerfile: FROM scylladb/scylla:5.2 ADD ./cassandra-rackdc.properties.dc1 /etc/scylla/cassandra-rackdc.properties … Language: en Canonical URL: https://forum.scylladb.com/t/issues-running-with-ipv6/565 ## Headings Structure: H1: Issues running with ipv6 H3: Related topics ## Main Content: H1: Issues running with ipv6 H3: Related topics I’m trying to run scylladb on fly.io and running into issues around ipv6. I have the following custom Dockerfile: And I am using the following to create machines with the machines api: Finally, when I run netstat -l on the machine I see: As well as nodetool having issues: One specific line I see of concern is ERROR 2023-05-21 11:37:09,295 [shard 0] init - Startup failed: std::_Nested_exception (Couldn't resolve listen_address): std::system_error (error C-Ares:12, 32874e1dc11785.vm.scylla-db.internal: Timeout) I have added a yaml file with enable_ipv6_dns_lookup: true and ADD ./scylla.yaml /etc/scylla/scylla.yaml to my docker build but I still see this issue. This address resolves fine if I ssh into the node. I also notice in the netstat output that port 9042 is not present, nor 10000 I believe I solved the issue!!! The issue was that it tries to resolve the 32874e1dc11785.vm.scylla-db.internal,4d891407f20787.vm.scylla-db.internal records as IPv4s (A record). When I placed in the IPv6 directly, it worked just fine for one node (scylla-1). However scylla-0 goes into this loop of these logs: It seems like it gets stuck compacting (weird since there is no data on the node), while the scylla-1 node is able to get past this: This is especially strange as these 2 nodes are identical. Here is more detailed log output (tried to clean the prefixes as best I could) https://pastebin.com/raw/Se8fz13Y Seems like the nodes diverge after the gossip - No gossip backlog; proceeding line. The scylla-1 node starts, where as the scylla-0 shuts down. I see ‘–listen-address 0.0.0.0 --rpc-address 0.0.0.0’ looks like a mixture of IPv4 and IPv6 ? Seems like that was it! Just needed a second pair of eyes, tysm! --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-180-2023-05-21/566 Title: Last week in scylladb.git master (issue #180; 2023-05-21) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 31e820e5a1…3b424e391b range are covered. There were 98 non-merge commits from 18 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-180-2023-05-21/566 ## Headings Structure: H1: Last week in scylladb.git master (issue #180; 2023-05-21) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #180; 2023-05-21) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 31e820e5a1…3b424e391b range are covered. There were 98 non-merge commits from 18 authors in that period. Some notable commits: We now disallow the CREATE permission on a named user defined function (UDF). It is meaningless since a named function must already have been created. The permissions granted when a keyspace or function are created have been adjusted. Tables backed by object storage no longer use a deletion log to orchestrate sstable deletion. Since those tables have an sstable ownership table managed by Raft, they use this ownership table to orchestrate deletions. ScyllaDB has an error injection facility, used by QA to test error paths. It can now be enabled via configuration. In Alternator, ScyllaDB’s implementation of the DynamoDB API, the timeout configuration value can be hot-updated without restarting the node. The REST API that accept sstable generation numbers now use a string value, in preparation for using UUID generations. The WebAssembly runtime used to evaluate user defined functions has been updated to address vulnerabilities. An edge case when converting range tombstones to the internal format used by sstables has been corrected. Very large compactions, involving hundreds of sstables, could cause stalls when the log message announcing the compaction is printed due to quadratic complexity. This is now fixed. A significant performance regression in compactions that process a lot of tombstones has been fixed. Error messages involving CQL expressions will not be printed in a more user-friendly way. Previously they contained some debug information. The API for performing sstable cleanup will now wait for staging sstables to be cleaned up too. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/lightweight-transactions-and-atomic-batches-performance-considerations/568 Title: Lightweight Transactions and Atomic batches - performance considerations - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello Scylla experts! My team is trying out Scylla Cloud for the first time and are already impressed with the immediate performance improvements on the read path - we’re heavy users of LWTs and atomic batches in C* and … Language: en Canonical URL: https://forum.scylladb.com/t/lightweight-transactions-and-atomic-batches-performance-considerations/568 ## Headings Structure: H1: Lightweight Transactions and Atomic batches - performance considerations H3: Related topics ## Main Content: H1: Lightweight Transactions and Atomic batches - performance considerations H3: Related topics Hello Scylla experts! My team is trying out Scylla Cloud for the first time and are already impressed with the immediate performance improvements on the read path - we’re heavy users of LWTs and atomic batches in C* and we’re noticing regular timeouts when attempting large batches that all start with one CAS operation followed by 10k INSERT INTO statements for different keys (Cassandra timeout during CAS write query at consistency SERIAL). It seems like these actually happen more frequently in Scylla than when using DataStax Astra. Our use case for atomic batches is essentially to “fence” older writers; this isn’t exactly it but you can think about it as each write operation having a unique “generation” and then the first statement of the batch is UPDATE foo SET generation = 10 WHERE partitionKey = '0' IF generation < 10; to ensure that no writes with generation greater than 10 have succeeded previously. All other statements are just basic INSERT INTOs with no CAS operation. In steady state, there should only be a single writer to a partition (hence the fencing!) - which means contention should be nearly zero.Two questions: Note that I haven’t upgraded to the Scylla driver yet (I’m just keeping the old DSE driver from our code) if it’s possible that this could help I can easily do that (just involves some CI/CD toil). *transcript from a discussion on the user slack channel Hi! Upgrading to Scylla drivers would help, python and Java drivers are LWT aware and prefer sending LWT writes to the primary replica, which reduces paxos conflicts. The main performance consideration is a thing called disk write dma alignment, which comes from the underlying filesystem used. By default it’s 4k, which means every LWT related write to the commit log is at least 4k of commit log space. If you have low concurrency (i.e. few requests per shard) that could be a lot of writeamp. In other words, unless your workload is exclusively using LWT, you don’t need to worry. I don’t know if there is a better way to do barriers than the one with generation id you described. I don’t fully understand the scenario, in order to do it, I would need to see all queries to this data and all the nuances how the results are used. The way you describe it - to ensure that no writes with generation greater than 10 have succeeded previously - is not making a lot of sense to me, since just another write can take place right after the conditional statement and succeed before the results of the conditional statement are delivered to the client. I don’t know what you mean under atomic batches. We call batches that have at least one conditional statement in them a conditional batch. There is only 1 Paxos round for such batch. Hope this helps. This is super helpful, thanks! The way you describe it - to ensure that no writes with generation greater than 10 have succeeded previously - is not making a lot of sense to me, since just another write can take place right after the conditional statement By “atomic batch” I meant “conditional batch” - all writes to Scylla would go through these conditional batches and the first statement of each batch is the CAS, so I imagine only one writer can “win” unless I’m not understanding something? (e.g. imagine node A attempts to write with generation ID 10 while node B attempts to write with generation ID 9 for the same key - if the batch from A goes through first, then the batch from B should fail). We use the barrier on every write. We call batches that have at least one conditional statement in them a conditional batch. There is only 1 Paxos round for such batch. That’s good news for us! I suspect that means to answer my second question it would hypothetically be better to do fewer larger conditional batches than more smaller ones? Each batch is likely to far exceed 4K so hopefully the alignment consideration isn’t too big of a deal, but I’ll do some investigation to confirm that. Are there any specific metrics we should keep an eye on to figure out whether we’re poorly utilizing the commitlog (hitting write amplification)? I guess so long as we’re below the node’s IOPs limit we should be good from that perspective… yes, if you have a batch like: the entire batch is conditional and it goes through a single paxos round. It still means some writeamp (write amplification), since the data is written 4 times instead of two: first time into system.paxos and its commit log, second time into the base table and its commit log, instead of just the base table and its commit log. I think though that overall it’s going to be quite efficient. As to poorly using the commit log, I think if you’re writing in batches like above this is not relevant to you. What I would keep an eye on is how much space your commit log takes overall and how many availiable segments there are over time. I believe both metrics are present in our stock grafana monitor. --- ### Page: https://forum.scylladb.com/t/release-scylla-monitoring-stack-4-4-0/569 Title: [RELEASE] Scylla Monitoring Stack 4.4.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of ScyllaDB Monitoring Stack 4.4.0 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB Enterprise and ScyllaDB Open Source, based on Prometheus … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-monitoring-stack-4-4-0/569 ## Headings Structure: H1: [RELEASE] Scylla Monitoring Stack 4.4.0 H3: Related topics ## Main Content: H1: [RELEASE] Scylla Monitoring Stack 4.4.0 H3: Related topics The ScyllaDB team is pleased to announce the release of ScyllaDB Monitoring Stack 4.4.0 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB Enterprise and ScyllaDB Open Source, based on Prometheus and Grafana. ScyllaDB Monitoring Stack 4.4.0 supports: This release focuses on adaptation for ScyllaDB Open Source 5.2 and the upcoming ScyllaDB Enterprise 2023.1 metrics changes and reduced load from Prometheus servers for large clusters. It is advised to upgrade the monitoring stack before upgrading ScyllaDB. Versions updates for Scylla Monitoring Stack 4.4.0 New Information in ScyllaDB Dashboards Overview Dashboard Changes The extended use of both internal and external SERVICE LEVEL (e.g. scheduling groups) made it complicated to understand the true latency in the overview dashboard. The new filter would allow users to explicitly choose which scheduling groups are shown. Internal SERVICE LEVEL (e.g. scheduling groups) are used for internal tasks, like compaction, are typically set to run at lower priority not to interfere with user activity. They are now removed from the p99 report as they do not represent real traffic. The disk size panel was updated to make it clearer what is the used part out of the entire available disk. Following the changes in Grafana integrated Alerts, ScyllaDB Monitoring now uses Grafana’s Alert table for clearer representation. The advisor section was updated to match Grafana’s alerts system. It is now completely alert based. Detailed Dashboard Changes There is a new section under the per-scheduling group section for payload size. It reflects the network usage for different kinds of CQL messages. There are also estimation panels for read and write, they give a ballpark estimate of an average read/write message. The section helps identify issues that result from large messages. Scylla-Manager Dashboard Changes Scylla Manager 3.1 supports a new restore from backup implementation, which exposes how much data remains to complete restore, in bytes. A new panel in the Manager dashboard graphs this data. [image] As part of an effort to make the dashboard clearer, a description (a popup with an explanation) was added to the panels where it was missing. --- ### Page: https://forum.scylladb.com/t/performance-issue-monitoring-dashboard/572 Title: Performance issue? Monitoring dashboard - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: version:4.6.4 os: Ubuntu 16 cluster nodes: 6 used for loki foundition storage how about cluster performance are? is there some improvement for the cluster?thanks in advance Language: en Canonical URL: https://forum.scylladb.com/t/performance-issue-monitoring-dashboard/572 ## Headings Structure: H1: Performance issue? Monitoring dashboard H1: nodetool toppartitions H3: Related topics ## Main Content: H1: Performance issue? Monitoring dashboard H1: nodetool toppartitions H3: Related topics version:4.6.4 os: Ubuntu 16 cluster nodes: 6 used for loki foundition storage how about cluster performance are? is there some improvement for the cluster?thanks in advance I don’t understand the problem. Both you read and write latencies look really excellent. I don’t see anything wrong here. Can you please expand on what it is you would like to improve on? thanks . the reads per instance metric looks like unbalanced,how to avoid it ? use nodetool to analyze and found top 2 request nodes‘s local read count is very high,as following: but the other 4 nodes are normal: The imbalance is simply due to the low cardinality of the system_auth.* table. These tables have only a handful of entries and whichever shard owns one (or more) of these few entries, will see all of the traffic aimed at this table. You can see this from the fact that on the coordinator level, your requests are well balanced, the imbalance is entirely on the replica side. I think there is nothing to worry here. The only reason this imbalance is even visible is because your cluster is very lightly loaded. If the imbalance persists on higher loads or even becomes more pronounced, that might indicate a problem. The toppartitions result is inaccurate (look at the +/- column). You can increase the accuracy with the -s capacity parameter (default 256). Increase it and try again, until the error drops to reasonable levels. --- ### Page: https://forum.scylladb.com/t/what-s-next-on-scylladb-s-path-to-strong-consistency/573 Title: What’s Next on ScyllaDB’s Path to Strong Consistency - Blog Posts - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/what-s-next-on-scylladb-s-path-to-strong-consistency/573 ## Headings Structure: H1: What’s Next on ScyllaDB’s Path to Strong Consistency H3: What’s Next on ScyllaDB’s Path to Strong Consistency H3: Related topics ## Main Content: H1: What’s Next on ScyllaDB’s Path to Strong Consistency H3: What’s Next on ScyllaDB’s Path to Strong Consistency H3: Related topics What's next on ScyllaDB's path to strong consistency with Raft? Topology changes will be safe and fast: we won't be bound by ring delay any longer, and operator mistakes won't corrupt the cluster. This would be a great place to discuss the topic of Strong Consistency or ask Kostja (the author) questions. Did you find the post helpful? --- ### Page: https://forum.scylladb.com/t/feature-request-serverless-scylladb-cloud-on-azure/575 Title: [FEATURE REQUEST] Serverless ScyllaDB Cloud on Azure - Database Community - ScyllaDB Community NoSQL Forum Meta Description: Currently ScyllaDB Serverless Cloud supports AWS and GCP (see this blog post). It would be really great to have Serverless option available for Microsoft Azure Cloud. Currently Azure Cloud is the second most popular cl… Language: en Canonical URL: https://forum.scylladb.com/t/feature-request-serverless-scylladb-cloud-on-azure/575 ## Headings Structure: H1: [FEATURE REQUEST] Serverless ScyllaDB Cloud on Azure H3: Related topics ## Main Content: H1: [FEATURE REQUEST] Serverless ScyllaDB Cloud on Azure H3: Related topics Currently ScyllaDB Serverless Cloud supports AWS and GCP (see this blog post). It would be really great to have Serverless option available for Microsoft Azure Cloud. Currently Azure Cloud is the second most popular cloud infrastructure provider (after AWS). Thus, adding support for this Cloud provider gives a number of users the possibility to use Serverless ScyllaDB for their use cases. Hi porunov Thanks for the feedback. First, Scylla Cloud Serverless runs on AWS only, not GCP. Second, we plan to add Azure support, but there has yet to be a timeline for it. Quick question: are you primarily interested in Serverless or Azure support? In other words, would a dedicated VM on Azure be relevant for you? Hi @tzach ! Thank you for replay. We are primarily interested in Serverless deployment on Azure. We are currently using AstraDB Serverless deployment on Azure. It would be interesting to test and compare AstraDB Serverless deployment and ScyllaDB Serverless deployment and see how they differ in term of latency, storage cost, write cost, read cost, and the features available. Unfortunately, a dedicated VM on Azure isn’t an option for us at this moment. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-13-2023-05-26/577 Title: Last week in scylla-cluster-tests.git master (issue #13; 2023-05-26) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This brief report highlights some noteworthy commits to the scylla-cluster-tests.git master from the past week. We’re covering the commits in the 92568e18…c1a97e3d range. In this period, 16 non-merge commits were made b… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-13-2023-05-26/577 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #13; 2023-05-26) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #13; 2023-05-26) H3: Related topics This brief report highlights some noteworthy commits to the scylla-cluster-tests.git master from the past week. We’re covering the commits in the 92568e18…c1a97e3d range. In this period, 16 non-merge commits were made by 7 different authors. Here are a few remarkable commits: The explicit setting of the compaction strategy within individual nemesis tests has been removed. This adjustment was made to facilitate the testing of other compaction strategies, such as ICS, which is the default in enterprise releases. The system_auth table replication ought to be configured with the number of nodes and utilize the NetworkTopologyStrategy on production systems. We’ve adapted SCT to adhere to this guidance early in the node setup process. Consequently, the system_auth_rf configuration is no longer required and was removed from all configuration files. Scylla has been featuring the seedless function for quite some time now. In fact, it is the standard in Scylla Cloud. We’ve made it easy to set all DB nodes as seeds, which is now the default setting. Looking forward to bringing you more updates in the next edition of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/scylladb-labs-virtual-training-event/579 Title: ScyllaDB Labs Virtual Training Event - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: On June 27, 2023, we’ll host ScyllaDB Labs Building High-Performance Apps, a 2-hour virtual hands-on training event. It’s a great way to discover the NoSQL strategies used by top teams and apply them in a guided, suppor… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-labs-virtual-training-event/579 ## Headings Structure: H1: ScyllaDB Labs Virtual Training Event H3: Related topics ## Main Content: H1: ScyllaDB Labs Virtual Training Event H3: Related topics On June 27, 2023, we’ll host ScyllaDB Labs Building High-Performance Apps, a 2-hour virtual hands-on training event. It’s a great way to discover the NoSQL strategies used by top teams and apply them in a guided, supportive environment. Some of the topics that we will cover are: Save your spot here. Hope to see you there! This is happening next week! Participants will get a chance to run some new labs we are preparing especially for this event. You can still save your (free) spot here. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-181-2023-05-28/580 Title: Last week in scylladb.git master (issue #181; 2023-05-28) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 3b424e391b…af65d5a1e8 range are covered. There were 167 non-merge commits from 18 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-181-2023-05-28/580 ## Headings Structure: H1: Last week in scylladb.git master (issue #181; 2023-05-28) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #181; 2023-05-28) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 3b424e391b…af65d5a1e8 range are covered. There were 167 non-merge commits from 18 authors in that period. Some notable commits: The master branch has version has been bumped to 5.4.0-dev, as the 5.3 stabilization and release cycle has started. When the schema changes, the rows in the row cache have to be upgraded to the new schema. This happens on-demand as rows are hit in the cache. Until now, this happened with partition granularity - all of a partition’s rows that happened to be in cache were upgraded at the same time, causing reactor stalls and high latency when large partitions were cached. This has now been fixed, and the cache is upgraded using row granularity. Tablet metadata (stored in the system.tablets table) is now loaded after its commitlog has been replayed. The nodetool checkAndRepairCdcStreams is used to align CDC streams with the cluster topology. It now works when topology is under Raft control. Logging of node failures during repair has been improved, in order to help diagnose repair failures. We now drop per-table metrics early during teardown of a table. Previously, if a table was dropped and re-created quickly, the metrics from the old and new tables could clash, resulting in an error. Commitlog has gained its own scheduling group, to complement the already existing commitlog I/O priority class. This is in preparation for unification of CPU scheduling and I/O scheduling. The S3 client can now upload files larger than 50GB. The limit was due to multipart uploads having at most 10,000 components, and ScyllaDB using the minimum component size of 5MB in order to reduce memory footprint. There is now a dedicated performance benchmark for the S3 client. Schema pulls happen when a node receives a read or write request (as a replica) with an unknown schema; it will then ask the requesting node for an updated schema. These are now disabled when the schema is managed using Raft; instead the system will rely solely on Raft for schema distribution. The column name reported when writetime() is given a primary key column (which is illegal) is now human readable, even for humans that don’t remember the ASCII table. The NetworkTopologyStrategy replication strategy will now reject an empty value for the replication factor. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/scylla-repair-during-schema-changes/582 Title: Scylla Repair During Schema Changes - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: On the nodetool repair page there is a warning against doing basically anything to the cluster (maintenance or schema changes) during a repair: source The applications that we run can dynamically change schemas, it’s ra… Language: en Canonical URL: https://forum.scylladb.com/t/scylla-repair-during-schema-changes/582 ## Headings Structure: H1: Scylla Repair During Schema Changes H3: Related topics ## Main Content: H1: Scylla Repair During Schema Changes H3: Related topics On the nodetool repair page there is a warning against doing basically anything to the cluster (maintenance or schema changes) during a repair: source The applications that we run can dynamically change schemas, it’s rare after the initial start up (just creating their tables), but we might drop a table here or there and the apps will recreate it whenever they next run. This basically means that if I want to guarantee that there are no schema changes, I have to stop the world. This is impossible to do on the kind of schedule that repairs need to be run. My question mainly is what kind of error would repair throw if a schema changed? Is this just something that it will kill the repair and we’d need to rerun it? Or could this cause other side effects in the cluster or data that we would need to recover from? Adding or dropping tables while a repair is running should be fine. These tables may or may not be picked up by the currently running repair, depending on timing. But Scylla should be able to handle this case without problems. What is dangerous is altering existing tables during a repair. Repair might still be using the old schema and it can choke on data written after the table alter, as it will possibly have columns, repair will not be able to process (due to the old schema). What is dangerous is altering existing tables during a repair. Repair might still be using the old schema and it can choke on data written after the table alter, as it will possibly have columns, repair will not be able to process (due to the old schema). Ok, just to clarify though, this will just result in the repair failing. Then a rerun of the repair should be fine (though will need to do the full set of data again). Correct? Altering a table while repair is running can also result in a crash, or memory corruption in the worst case. This is quite unlikely but not at all impossible and we have seen it happening in the past. Ok, thanks for the information. --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-2-2/584 Title: [RELEASE] ScyllaDB 5.2.2 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.2.2, a bugfix release of the ScyllaDB 5.2 stable branch. ScyllaDB Open Source 5.2.2, like all past and future 5.x.y releases, is backward compatible and supports rolling… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-2-2/584 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.2.2 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.2.2 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.2.2, a bugfix release of the ScyllaDB 5.2 stable branch. ScyllaDB Open Source 5.2.2, like all past and future 5.x.y releases, is backward compatible and supports rolling upgrades. Users are encouraged to upgrade to 5.2.2. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-1-11/585 Title: [RELEASE] ScyllaDB 5.1.11 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.1.11, a bugfix release of the ScyllaDB 5.1 stable branch. ScyllaDB Open Source 5.1.11, like all past and future 5.x.y releases are backward compatible and support rollin… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-1-11/585 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.1.11 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.1.11 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.1.11, a bugfix release of the ScyllaDB 5.1 stable branch. ScyllaDB Open Source 5.1.11, like all past and future 5.x.y releases are backward compatible and support rolling upgrades. Note that the latest stable ScyllaDB Open Source release is 5.2, and you are encouraged to upgrade to it. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-2-7/587 Title: [RELEASE] ScyllaDB Enterprise 2022.2.7 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2022.2.7, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.2. Related Links Get ScyllaDB Enterprise 2022.2.7 (customers only, o… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-2-7/587 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2022.2.7 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2022.2.7 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2022.2.7, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.2. The following issues are fixed in this release (with an open-source reference, if available): --- ### Page: https://forum.scylladb.com/t/build-your-first-scylladb-application-new-rust-python-php-tutorials/589 Title: Build your First ScyllaDB Application: New Rust, Python & PHP Tutorials - Blog Posts - ScyllaDB Community NoSQL Forum Meta Description: [800x400-blog-build-app-petcare-iot] Language: en Canonical URL: https://forum.scylladb.com/t/build-your-first-scylladb-application-new-rust-python-php-tutorials/589 ## Headings Structure: H1: Build your First ScyllaDB Application: New Rust, Python & PHP Tutorials H3: Build your First ScyllaDB Application: New Rust, Python & PHP Tutorials H3: Related topics ## Main Content: H1: Build your First ScyllaDB Application: New Rust, Python & PHP Tutorials H3: Build your First ScyllaDB Application: New Rust, Python & PHP Tutorials H3: Related topics Get experience with ScyllaDB by creating a sample IoT application from scratch. Use new Rust, Python & PHP examples (or the previous Python, JavaScript, Java, Golang ones) -- plus a new Terraform provider. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-182-2023-06-04/590 Title: Last week in scylladb.git master (issue #182; 2023-06-04) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the af65d5a1e8…b39ca97919 range are covered. There were 59 non-merge commits from 14 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-182-2023-06-04/590 ## Headings Structure: H1: Last week in scylladb.git master (issue #182; 2023-06-04) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #182; 2023-06-04) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the af65d5a1e8…b39ca97919 range are covered. There were 59 non-merge commits from 14 authors in that period. Some notable commits: If the startup sequence is aborted by an interrupt (ctrl-C or systemd shutdown), an exception error message was shown. It is now ignored by the system and not displayed. The nodetool refresh command gained the --primary-replica-only option. During shutdown, the system will cancel pending hint writes rather than wait for their 5-minute timeout. This can prevent delays in stopping a server. In internode communications, we now avoid copies of certain heavyweight objects. A cleanup compaction is used to get rid of token ranges that are no longer owned by a node. A bug that delayed deletion of sstables being cleaned up, thus increasing the risk of running out of space, was fixed. A bootstrapping node will now wait for schema agreement before joining the cluster. This prevents conflict between the new node’s system distributed tables and the cluster’s tables. The conflict is eventually resolved, but while it exists, the cluster is under heavier load. A race condition between the startup of raft group 0 and its rpc listener was fixed. Consistent schema management using Raft is now the default not only for new clusters, but also for clusters upgrading from an older version. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-14-2023-06-04/591 Title: Last week in scylla-cluster-tests.git master (issue #14; 2023-06-04) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report highlights some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the d2a6d074…b3912362 range are covered. There were 25 non-merge commits from 8 authors during that… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-14-2023-06-04/591 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #14; 2023-06-04) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #14; 2023-06-04) H3: Related topics This short report highlights some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the d2a6d074…b3912362 range are covered. There were 25 non-merge commits from 8 authors during that period. Some notable commits include: Until now, each cleanup code made more API calls than necessary. Redundancy from the ‘clean-resources’ has been removed. This was causing issues with API call limits on the GCE side, we should no longer face this problem. Previously, if the Jenkins pipeline failed before reaching the SCT test, it was not shown in Argus. This resulted in missing some failures and not raising issues. Now, we catch failed runs earlier. The scylla-operator now supports monitoring. For some K8s tests, we install scylla-operator’s monitoring in K8S. This includes our SCT dashboard. Several commits were made to enable support for this feature, such as adding the possibility to ‘apply’ K8S files on the ‘server-side’, with the --server-side parameter for the kubectl apply operation. This allows us to apply large files that exceed the client-side limit and apply files ‘as is’ without needing to escape certain characters, which is problematic for ‘regex’ file lines. Tests running in multi-az environments generate additional network traffic costs. We have added the possibility to simulate racks by tuning the GossipingPropertyFileSnitch snitch and provisioning instances in one availability zone to avoid them. We look forward to the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/udfs-in-wasm-vs-lua/594 Title: UDFs in Wasm vs Lua - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi everyone! I’m experimenting with scylla UDFs and was wondering if there are any benchmarks on running udfs based on wasm (rust-based using the helper library) vs Lua. Is any language more “stable” in any way? It seem… Language: en Canonical URL: https://forum.scylladb.com/t/udfs-in-wasm-vs-lua/594 ## Headings Structure: H1: UDFs in Wasm vs Lua H3: Related topics ## Main Content: H1: UDFs in Wasm vs Lua H3: Related topics Hi everyone! I’m experimenting with scylla UDFs and was wondering if there are any benchmarks on running udfs based on wasm (rust-based using the helper library) vs Lua. Is any language more “stable” in any way? It seems that Lua support exists for quite a while now and wasm is more recent. I am not aware of any benchmarks we have done to compare the two. Lua UDFs indeed exist for quite a while, while WASM is a more recent addition. That said, WASM is much more actively developed as we believe it is a much more flexible and performant platform for UDFs, than Lua is. --- ### Page: https://forum.scylladb.com/t/release-scylla-monitoring-stack-4-4-1/595 Title: [RELEASE] Scylla Monitoring Stack 4.4.1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce a security patch release of ScyllaDB Monitoring Stack 4.4.1 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB Enterprise and ScyllaDB Open Source, based o… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-monitoring-stack-4-4-1/595 ## Headings Structure: H1: [RELEASE] Scylla Monitoring Stack 4.4.1 H3: Related topics ## Main Content: H1: [RELEASE] Scylla Monitoring Stack 4.4.1 H3: Related topics The ScyllaDB team is pleased to announce a security patch release of ScyllaDB Monitoring Stack 4.4.1 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB Enterprise and ScyllaDB Open Source, based on Prometheus and Grafana. ScyllaDB Monitoring Stack 4.4.1 supports: Following Grafana’s announcement cve-2023-2183 and-cve-2023-2801 this patch release updated Grafana’s version to 9.5.3. --- ### Page: https://forum.scylladb.com/t/what-is-the-maximum-number-of-records-that-a-scylla-table-can-carry/596 Title: What is the maximum number of records that a scylla table can carry? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello, guys! If I put 1 trillion records in one table, each record is a separate partition, is this a good schema design, what problems will exist? In other words, in a large-scale cluster, do we need to control the nu… Language: en Canonical URL: https://forum.scylladb.com/t/what-is-the-maximum-number-of-records-that-a-scylla-table-can-carry/596 ## Headings Structure: H1: What is the maximum number of records that a scylla table can carry? H3: Related topics ## Main Content: H1: What is the maximum number of records that a scylla table can carry? H3: Related topics If I put 1 trillion records in one table, each record is a separate partition, is this a good schema design, what problems will exist? In other words, in a large-scale cluster, do we need to control the number of partitions for a single table? When I have a lot of small partitions, does it mean that the reading efficiency will be poor, for example, there will be problems in the construction of Bloom filters. How should I tune some table-level parameters for good performance, such as compaction strategy? ScyllaDB loves many partitions, I would even say that many small partitions are the sweet spot for ScyllaDB. This is mostly because having many partitions will mean that data is well distributed among the nodes and shards in the cluster. Uneven data distribution leads to hot shards and hot nodes, which slow the cluster down. On the other hand, having partitions that are tiny can have some side effects. You may find that not as many as them fits in cache as you would expect, as there is a constant per-partition size overhead for partitions stored in memory. View update generation during repair, for tables that have materialized views or secondary indexes attached, can get quite slow with many tiny partitions. All that said, you should not see any problems with storing huge number of small partitions in ScyllaDB. Like I said above, this is the sweet spot for ScyllaDB and there is no limit (theoretical or hard) as to how many partitions you can have. Choosing the compaction strategy is more of a question of what workloads you have, than how your data is organized. as there is a constant per-partition size overhead for partitions stored in memory Do we have parameters to control this value? No, this is just the overhead of the C++ objects involved in storing and organizing partitions. Nothing we can do about that. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-15-2023-06-09/599 Title: Last week in scylla-cluster-tests.git master (issue #15; 2023-06-09) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the bf744770…405f9eb8 range are covered. There were 21 non-merge commits from 6 authors in tha… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-15-2023-06-09/599 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #15; 2023-06-09) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #15; 2023-06-09) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the bf744770…405f9eb8 range are covered. There were 21 non-merge commits from 6 authors in that period. Some notable commits: Since it is planned to drop SimpleStrategy and we want to test more production-like cases, now we use NetworkTopologyStrategy in every test scenario. Recently many changes were added to Gemini load tool and new versions were released, thus our tests upgraded Gemini version to 1.8.2. Because new docker image changed entrypoint and dir structure, some adaptations were required to the way we run it. Since we have feature we want to use that aren’t supported by libcloud for GCE, we removed all usages of it and replaced with official GCE Cloud SDK. This change affected many files and required a lot of code changes and was splitted to multiple commits: #1, #2, #3, #4. All the manager jobs will use the latest enterprise version (2022.2) instead of the latest OSS version. The scenarios that used to be dedicated to enterprise runs will now use the older enterprise version instead (2022.1). We encourage everyone to try out running SCT tests. To ease that, we fixed&refactored the readme into several parts with foucs on using hydra in the quickstart and moving the rest of the information into varius sections. Since for reporting purpose, new ScyllaDB images created by Releng are going to be created under Releng’s sub-account. SCT now supports it. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/release-scylla-operator-1-9-0-rc-0/600 Title: [RELEASE] Scylla Operator 1.9.0-rc.0 - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: The Scylla team is pleased to announce the release of scylla-operator v1.9.0-rc.0 :rocket: We’ll welcome your feedback on the release candidate. Release notes are available on If you haven’t heard about the operator … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-operator-1-9-0-rc-0/600 ## Headings Structure: H1: [RELEASE] Scylla Operator 1.9.0-rc.0 H3: Release v1.9.0-rc.0 · scylladb/scylla-operator H3: GitHub - scylladb/scylla-operator: The Kubernetes Operator for ScyllaDB - Scylla Operator H3: Related topics ## Main Content: H1: [RELEASE] Scylla Operator 1.9.0-rc.0 H3: Release v1.9.0-rc.0 · scylladb/scylla-operator H3: GitHub - scylladb/scylla-operator: The Kubernetes Operator for ScyllaDB - Scylla Operator H3: Related topics The Scylla team is pleased to announce the release of scylla-operator v1.9.0-rc.0 We’ll welcome your feedback on the release candidate. Release notes are available on Release notes for 1.9.0-rc.0 Container images docker.io/scylladb/scylla-operator:1.9.0-rc.0 Changes By Kind (since 1.8.1) Feature Automated Local Disk Setup (#1107,@zimnx) Add ScyllaDBMonitoring ... If you haven’t heard about the operator yet, here are some links to get you started: The Kubernetes Operator for ScyllaDB. Contribute to scylladb/scylla-operator development by creating an account on GitHub. https://operator.docs.scylladb.com/ --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-183-2023-06-11/603 Title: Last week in scylladb.git master (issue #183; 2023-06-11) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the b39ca97919…e464ad2568 range are covered. There were 73 non-merge commits from 20 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-183-2023-06-11/603 ## Headings Structure: H1: Last week in scylladb.git master (issue #183; 2023-06-11) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #183; 2023-06-11) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the b39ca97919…e464ad2568 range are covered. There were 73 non-merge commits from 20 authors in that period. Some notable commits: Alternator, ScyllaDB’s implementation of the DynamoDB API, now returns the full table description as a response to the DeleteTable API request. Alternator now avoids latency spikes for unrelated requests while building large responses for batch_get_item. Some queries for the internal authentication table used infinite timeouts, leading to shutdown problems. This is now fixed. The internal data dictionary could loose user defined types on ALTER KEYSPACE statements, resulting in a crash. This is now fixed. The nodetool refresh command loads foreign sstables into ScyllaDB and reshapes them for the current shard distribution. A bug could cause the clean-up after the reshape to crash. It is now fixed. ScyllaDB now uses Seastar unified scheduling, where I/O and CPU are both controlled by a single “scheduling group” concept. There should be no user-visible changes. Materialized view require the IS NOT NULL qualifier on primary key elements, but also accepted (and ignored) the qualifier on regular columns. The qualifier is now rejected when applied to regular columns. A configuration variables allows to warn about the rejected clause, emit an error and fail the request, or ignore it. Repair will now use a more accurate estimate of the partition count to create bloom filters for its sstables/. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/rust-scylladb-hands-on-dev-workshop/606 Title: Rust + ScyllaDB Hands-On Dev Workshop - Announcements - ScyllaDB Community NoSQL Forum Meta Description: Want to go hands-on with Rust and ScyllaDB? Join our upcoming (free + virtual) interactive developer workshop! Using Rust, the Tokio framework, and ScyllaDB, we’ll build a sample Rust application on our high performance … Language: en Canonical URL: https://forum.scylladb.com/t/rust-scylladb-hands-on-dev-workshop/606 ## Headings Structure: H1: Rust + ScyllaDB Hands-On Dev Workshop H3: Related topics ## Main Content: H1: Rust + ScyllaDB Hands-On Dev Workshop H3: Related topics Want to go hands-on with Rust and ScyllaDB? Join our upcoming (free + virtual) interactive developer workshop! Using Rust, the Tokio framework, and ScyllaDB, we’ll build a sample Rust application on our high performance native Rust client driver. https://lp.scylladb.com/rust-workshop-registration.html --- ### Page: https://forum.scylladb.com/t/the-data-modeling-behind-social-media-likes/607 Title: The Data Modeling Behind Social Media “Likes” - Blog Posts - ScyllaDB Community NoSQL Forum Meta Description: [800x400-blog-how-likes-are-stored-db] Language: en Canonical URL: https://forum.scylladb.com/t/the-data-modeling-behind-social-media-likes/607 ## Headings Structure: H1: The Data Modeling Behind Social Media “Likes” H3: The Data Modeling Behind Social Media “Likes” H3: Related topics ## Main Content: H1: The Data Modeling Behind Social Media “Likes” H3: The Data Modeling Behind Social Media “Likes” H3: Related topics Did you ever think about how Instagram, Twitter, Facebook or other social media platforms track who liked your posts? This post explains how to do it with ScyllaDB NoSQL. --- ### Page: https://forum.scylladb.com/t/poll-scylladb-usage-and-interest/611 Title: Poll: ScyllaDB Usage and Interest - Database Community - ScyllaDB Community NoSQL Forum Meta Description: To better serve your needs, I’m doing some research into our community’s interest in ScyllaDBThe results will be shared with the community. I’d really appreciate your input: poll poll poll poll poll poll poll poll Language: en Canonical URL: https://forum.scylladb.com/t/poll-scylladb-usage-and-interest/611 ## Headings Structure: H1: Poll: ScyllaDB Usage and Interest H3: Related topics ## Main Content: H1: Poll: ScyllaDB Usage and Interest H3: Related topics To better serve your needs, I’m doing some research into our community’s interest in ScyllaDBThe results will be shared with the community. I’d really appreciate your input: --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2021-1-20/612 Title: [RELEASE] ScyllaDB Enterprise 2021.1.20 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2021.1.20, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2021.1. ScyllaDB Enterprise 2022.1 LTS is the latest Long-Term Support rele… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2021-1-20/612 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2021.1.20 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2021.1.20 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2021.1.20, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2021.1. ScyllaDB Enterprise 2022.1 LTS is the latest Long-Term Support release, and 2022.2 is the newest feature release. You are encouraged to upgrade in coordination with the ScyllaDB support team. The following issues are fixed in this release (with an open-source reference): --- ### Page: https://forum.scylladb.com/t/a-simple-query-gives-a-syntax-error/617 Title: A simple query gives a syntax error - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I have a query that doesnt execute for some reason. Kindly take a look & let me know whats wrong query = """ use ks; CREATE TYPE themes ( id uuid PRIMARY KEY, name text ); … Language: en Canonical URL: https://forum.scylladb.com/t/a-simple-query-gives-a-syntax-error/617 ## Headings Structure: H1: A simple query gives a syntax error H3: Related topics ## Main Content: H1: A simple query gives a syntax error H3: Related topics I have a query that doesnt execute for some reason. Kindly take a look & let me know whats wrong This gives an error as follows- message: “line 3:2 : syntax error…\n”, Kindly help me figure out whats wrong I fixed the problems with your query, here is the diff: Let me explain the fixes: thankyou so much for your input sir, this helps a lot. I couldnt find the above mentioned points in the docs. Since a UDT as a whole can be used as a Primary Key along the docs I figured individual fields inside UDTs can be the same too. Also I didnt know about the frozen usage in embedded UDTs since the docs suggested they werent needed for higher versions. I really appreciate the help. Thankyou once again Also if possible id also like to ask for some advise. Do you think scylla would suit for a social network applications primary db? Since ive been advised against it. Due to scylla being good at dealing with time based data, hence it was made known to me that the data attributes are what enable high performance & not the db in itself. Say if you use snowflakes as id’s then it gives higher throughput to query against & not in other case. There also might be some query restrictions in scylla due to clustering columns operations being permitted only on things as seen above. Overall do you have any suggestions? Do you think scylla would suit for a social network applications primary db? Due to scylla being good at dealing with time based data, hence it was made known to me that the data attributes are what enable high performance & not the db in itself. I don’t really understand how the data attributes enable high performance. For sure, correct data modelling, done in a way to suit the DB which is used, goes a long way to ensure good performance. But the underlying DB also has a huge part in it. Say if you use snowflakes as id’s then it gives higher throughput to query against & not in other case. What data type your ID is and how you compute it, ScyllaDB doesn’t care. You have to choose your keys such that it best fits your use-case and query patterns. There also might be some query restrictions in scylla due to clustering columns operations being permitted only on things as seen above. Yes, ScyllaDB has a quite restrictive data-model, as it was designed around performance. Coming from an SQL DB, you might feel it is too restrictive. That said, any use-case can be made to work in it, but it does need deliberate data-modeling design, to fit the underlying model well and to extract maximum performance. Overall do you have any suggestions? Your use-case is too broad for me to come up with any specific suggestions. I recommend you check out the ScyllaDB Essentials and the ScyllaDB Data Modeling courses on ScyllaDB University. thankyour for your advise. This sure does help a lot. Really appreciate it. Thanks once again. --- ### Page: https://forum.scylladb.com/t/the-db-times-out-upon-dropping-a-keyspace/618 Title: The db times out upon dropping a keyspace - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Whenever I try to drop a keyspace, the operation rarely gets executed at the db times out. I always get this error- OperationTimedOut: errors={'172.17.0.2': 'Client request timeout. See Session.execute[_async](timeout)… Language: en Canonical URL: https://forum.scylladb.com/t/the-db-times-out-upon-dropping-a-keyspace/618 ## Headings Structure: H1: The db times out upon dropping a keyspace H3: Related topics ## Main Content: H1: The db times out upon dropping a keyspace H3: Related topics Whenever I try to drop a keyspace, the operation rarely gets executed at the db times out. I always get this error- OperationTimedOut: errors={'172.17.0.2': 'Client request timeout. See Session.execute[_async](timeout)'}, last_host=172.17.0.2 Why is this happening? It’s impossible to know without more context. Are there related messages in system log? What is the timeouts configuration? What kind of disks are you using? Is auto_snapshot enabled? Are there related messages in system log? Not sure, what that means??? What is the timeouts configuration? What kind of disks are you using? My PC, default 256 ssd 1 tb hdd for storage. Its in dev mode Is auto_snapshot enabled? Nope I couldnt find the default option true in the config file. Are there related messages in system log? Not sure, what that means??? Scylla prints all kinds of messages to a system log. I’m not sure how you installed it. Anything under /var/log/? What is the timeouts configuration? What kind of disks are you using? My PC, default 256 ssd 1 tb hdd for storage. Its in dev mode Is auto_snapshot enabled? Nope I couldnt find the default option true in the config file. Then actually it is enabled by default. This means that scylla takes a snapshot of the dropped table(s) / keyspace before deleting it to be on the safe side. It is possible that this takes too long with a magnetic drive. Try setting auto_snapshot: false in scylla.yaml and restart scylla for it to take effect. Then retry dropping the keyspace / table. Its installed via docker-compose on a linux system. Deepin to be precise. v4.something ig. Nothing relatedd to scylla under /var/logs tho at first glance, everything seems related to the internal FS. Try setting auto_snapshot: false in scylla.yaml and restart scylla for it to take effect. Then retry dropping the keyspace / table. aaah yes I saw that in the docs but I couldnt find that in my default scylla.yaml file. So I figured by default it mustve been false. Thats 1 thing ive gotta do for sure yep Thanks for the input, ill try to fix this once & for all --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-1-7/620 Title: [RELEASE] ScyllaDB Enterprise 2022.1.7 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2022.1.7, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.1. Note that there is a newer Feature Enterprise release: ScyllaDB Ente… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-1-7/620 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2022.1.7 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2022.1.7 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2022.1.7, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.1. Note that there is a newer Feature Enterprise release: ScyllaDB Enterprise 2022.2. While we will continue to support 2022.1 LTS, you can get additional features with 2022.2. Below is a list of performance and stability improvements and bug fixes, each with an open-source reference if available: --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-2-8/622 Title: [RELEASE] ScyllaDB Enterprise 2022.2.8 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2022.2.8, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.2. Related Links Get ScyllaDB Enterprise 2022.2.8 (customers only, o… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-2-8/622 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2022.2.8 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2022.2.8 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2022.2.8, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.2. The following issues are fixed in this release (with an open-source reference, if available): --- ### Page: https://forum.scylladb.com/t/scylladb-for-high-write-and-read-but-low-insert/625 Title: ScyllaDB for high write and read but low insert - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: i plan to use scylladb for multiplayer game data or for matchmaking that needs to be updated every second basically i want to use scylladb as a fast in-memory for high-update scenario, i looked into ScyllaDB because it … Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-for-high-write-and-read-but-low-insert/625 ## Headings Structure: H1: ScyllaDB for high write and read but low insert H3: Related topics ## Main Content: H1: ScyllaDB for high write and read but low insert H3: Related topics i plan to use scylladb for multiplayer game data or for matchmaking that needs to be updated every second basically i want to use scylladb as a fast in-memory for high-update scenario, i looked into ScyllaDB because it is table-based unlike redis,etcd,memcached i will have fields that could be updated every second or should i stick with redis,etcd,memcached? ScyllaDB can easily support a million writes per second on modern hardware and large nodes. The choice should be based on data size. If you have a few hundred gigabytes or more, and high traffic, then using ScyllaDB makes sense. If the row marker is alive, the row shows up when you query it, even if all its non-key columns are null. The difference between inserts and updates is that updates don’t affect the row marker, while slope unblocked inserts create an alive row marker .It is very difficult to say that which one is better but it depends on your application behaviour. however, both can give good performance for read intensive application. Cassandra can give you better availability and Mongo will give you more consistent data. --- ### Page: https://forum.scylladb.com/t/low-cost-method-for-new-startup/627 Title: Low cost method for new startup - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I have a web site AWS EC2 and I need new EC2 instance for ScyllaDB. When I deploy ScyllaDB in my own account, do I need minimum configuration details like CPU, RAM, Storage? I need to deploy ScyllaDB one node in EC2 wi… Language: en Canonical URL: https://forum.scylladb.com/t/low-cost-method-for-new-startup/627 ## Headings Structure: H1: Low cost method for new startup H3: Related topics ## Main Content: H1: Low cost method for new startup H3: Related topics I have a web site AWS EC2 and I need new EC2 instance for ScyllaDB. When I deploy ScyllaDB in my own account, do I need minimum configuration details like CPU, RAM, Storage? I need to deploy ScyllaDB one node in EC2 with low instance. This website is beginning level so I need to start from 5GB storage, 2CPU, 1GB RAM (One Node or One Instance). is this posible? You can consider t2.small instances with EBS volumes, and scale to larger instances as you grow. Eventually move to i3en.large when you pass 100GB or so. So do I need 3 number of t2.small instances ? Can’t I use 1 node (Instance) ? If you’re running ScyllaDB in production, it’s recommended to run it on at least 3 nodes. A good starting point is the Introduction lesson in the Essentials course on ScyllaDB University. Another useful resource is the Getting Started guide in the Documentation. I read instruction and getting started guide but i did not found any deployment video with ScyllaDB opensource. YouTube have TODO app project with ScyllaDB Cloud. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-184-2023-06-18/628 Title: Last week in scylladb.git master (issue #184; 2023-06-18) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the e464ad2568…b7627085cb range are covered. There were 110 non-merge commits from 15 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-184-2023-06-18/628 ## Headings Structure: H1: Last week in scylladb.git master (issue #184; 2023-06-18) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #184; 2023-06-18) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the e464ad2568…b7627085cb range are covered. There were 110 non-merge commits from 15 authors in that period. Some notable commits: In Alternator (ScyllaDB’s implementation of the DynamoDB API), validation of the table name on ordinary read/write requests is done only if the table lookup fails. This provides a small optimization. The documentation now says Ubuntu 18.04 is no longer supported, as it has reached end-of-life. The “forward” service is responsible for execution of automatically parallelized aggregation queries. It is now more careful to stop query retries if a shutdown is requested. The count(column) function is supposed to only count cells where the column is not NULL. A regression caused count(column) to behave like count(*) for collection, tuple, and user-defined column types. This is now fixed. SSTable generation numbers are integers used to give SSTables unique names. Generation numbers can now also be UUIDs, which enables placing SSTables on shared storage. When using the experimental Raft-managed topology, the cluster is able to verify that all reachable nodes are using current topology, and is able to block requests that use old topology. This lays the ground for faster and safer topology changes (addition and removal of nodes). See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-1-12/629 Title: [RELEASE] ScyllaDB 5.1.12 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.1.12, a bugfix release of the ScyllaDB 5.1 stable branch. ScyllaDB Open Source 5.1.12, like all past and future 5.x.y releases are backward compatible and support rollin… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-1-12/629 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.1.12 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.1.12 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.1.12, a bugfix release of the ScyllaDB 5.1 stable branch. ScyllaDB Open Source 5.1.12, like all past and future 5.x.y releases are backward compatible and support rolling upgrades. Note that the latest stable ScyllaDB Open Source release is 5.2, and you are encouraged to upgrade to it. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/cdc-on-system-tables/631 Title: CDC on system tables - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Looking to enable CDC on system.large_partitions to keep a tab on the large partitions. However, got the following error when enabling it on the system.large_partitions: Unauthorized: Error from server: code=2100 [Unau… Language: en Canonical URL: https://forum.scylladb.com/t/cdc-on-system-tables/631 ## Headings Structure: H1: CDC on system tables H3: Related topics ## Main Content: H1: CDC on system tables H3: Related topics Looking to enable CDC on system.large_partitions to keep a tab on the large partitions. However, got the following error when enabling it on the system.large_partitions: Unauthorized: Error from server: code=2100 [Unauthorized] message="system keyspace is not user-modifiable." Is it possible in Scylla to enable CDC on system keyspace tables? What are the alternatives to get CDC on system.large_partitions table? sorry, there is no way to enable CDC on system tables. A simple way for keeping a tab on large partitions - if you just want to know whether there’s a large partition that you should be worried about - would be to have a simple app which queries system.large_partitions periodically, e.g. once a day, and setup some alerts if the results pass certain thresholds. --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-2-3/633 Title: [RELEASE] ScyllaDB 5.2.3 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.2.3, a bugfix release of the ScyllaDB 5.2 stable branch. ScyllaDB Open Source 5.2.3, like all past and future 5.x.y releases, is backward compatible and supports rolling… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-2-3/633 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.2.3 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.2.3 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.2.3, a bugfix release of the ScyllaDB 5.2 stable branch. ScyllaDB Open Source 5.2.3, like all past and future 5.x.y releases, is backward compatible and supports rolling upgrades. Users are encouraged to upgrade to 5.2.3. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/row-invalidation-in-the-row-cache/636 Title: Row invalidation in the row cache - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello All, The question is about the nature of a row invalidation process in the row cache. We have the corresponding monitoring metric - scylla_cache_row_removals, and I’d like to understand what this number should me… Language: en Canonical URL: https://forum.scylladb.com/t/row-invalidation-in-the-row-cache/636 ## Headings Structure: H1: Row invalidation in the row cache H3: Related topics ## Main Content: H1: Row invalidation in the row cache H3: Related topics The question is about the nature of a row invalidation process in the row cache. We have the corresponding monitoring metric - scylla_cache_row_removals, and I’d like to understand what this number should mean for me. The basic questions are: scylla_cache_row_removals is incremented whenever a row is removed from the cache for reasons other than eviction. That may be a bit confusing, since “eviction” also removes the row from the cache, but it is not counted in scylla_cache_row_removals. scylla_cache_row_evictions is typically incremented when a row entry is removed due to memory pressure to free space for new memory allocations. It is also incremented when cache reader decides that a so called “dummy” row entry is redundant during scan. Those row entries hold no data but serve as anchors which represent boundaries of the clustering range used by the read. When a scan for a given clustering key range populates the cache, we insert dummy entries for the boundaries of the range in order to store the information that the whole range is continuous (populated). Because we have those “dummy” rows, which are included in the metrics, the number of row evictions or removal may be greater than what you would expect from the number of CQL rows. Each partition has at least 1 dummy row entry (last dummy). In particular, a single-row partition has 2 row entries. The last dummy is special, because it is never evicted, always removed with the partition version. scylla_cache_row_removals is incremented when a row entry is removed due to following reasons: Thanks a lot for the quick and detailed explanation! --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-2-9/638 Title: [RELEASE] ScyllaDB Enterprise 2022.2.9 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2022.2.9, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.2. Related Links Get ScyllaDB Enterprise 2022.2.9 (customers only, or… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-2-9/638 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2022.2.9 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2022.2.9 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2022.2.9, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.2. The following issue are fixed in this release (with an open-source reference, if available): Stability: All rpc::client objects have to be stopped before they are destroyed. Currently this is done in messaging_service::shutdown(). The cql_test_env does not call shutdown() currently. This can lead to use-after-free, leading to sigsegv #12244 Relax messaging_service::shutdown() #14031 --- ### Page: https://forum.scylladb.com/t/release-scylla-monitoring-stack-4-4-2/644 Title: [RELEASE] Scylla Monitoring Stack 4.4.2 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce a security patch release of ScyllaDB Monitoring Stack 4.4.2 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB Enterprise and ScyllaDB Open Source, based o… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-monitoring-stack-4-4-2/644 ## Headings Structure: H1: [RELEASE] Scylla Monitoring Stack 4.4.2 H3: Related topics ## Main Content: H1: [RELEASE] Scylla Monitoring Stack 4.4.2 H3: Related topics The ScyllaDB team is pleased to announce a security patch release of ScyllaDB Monitoring Stack 4.4.2 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB Enterprise and ScyllaDB Open Source, based on Prometheus and Grafana. ScyllaDB Monitoring Stack 4.4.2 supports: Following Grafana’s announcement cve-2023-3128 this patch release updated Grafana’s version to 9.5.5. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-185-2023-06-25/645 Title: Last week in scylladb.git master (issue #185; 2023-06-25) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the b7627085cb…be5b61b870 range are covered. There were 162 non-merge commits from 20 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-185-2023-06-25/645 ## Headings Structure: H1: Last week in scylladb.git master (issue #185; 2023-06-25) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #185; 2023-06-25) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the b7627085cb…be5b61b870 range are covered. There were 162 non-merge commits from 20 authors in that period. Some notable commits: The sylla sstable tool now supports the scrub operation, enabling offline (and off-node) scrubbing of sstables. Recently, schema changes to data in the row cache changed the upgrade granularity from partition to row, to prevent stalls when large partitions are cached. One place could use an outdated schema, which could cause a crash. This is now fixed. The S3 storage back-end now limits the number of open connections, in order to avoid S3 dropping connections on overload. An edge case where querying a datacenter that has replication factor equal to zero could lead to a crash has been fixed. The documentation URL has been changed to https://opensource.docs.scylladb.com. When performing the last-write-wins rule comparison, if the timestamp of the two versions being compared was equal, ScyllaDB first compared the cell value and then the expiration time (TTL). This is compatible with earlier versions of Cassandra. However, this could cause a NULL value to appear if the cell was overwritten with the same timestamp but a different TTL. The algorithm was changed to compare the cell value last, and check all the other metadata first, resulting in fewer surprising results. It is also compatible with current Cassandra versions. Bugs preventing a node from starting when using the new raft-based topology mechanism have been fixed. Tablets are a new, experimental replication model in ScyllaDB, contrasting with vnodes. Tablets are now restricted to a single shard, unlike vnodes which span all shards on a node. This simplifies how tablets are stored in sstables and how tablets can be migrated to other nodes. SSTable compression can be configured with a chunk size, with larger chunks trading less efficient I/O and higher latency for higher compression ratios. The chunk size is now capped at 128 kB, to avoid running out of memory. When a node is decommissioned or forcibly removed, Raft will now ban it from communicating with the cluster, to avoid a the removed node from affecting the cluster. ScyllaDB can automatically parallelize certain aggregation queries. The mechanism however had a bug when aggregating columns that had case-sensitive names. This is now fixed. A crash when DESCRIBE FUNCTION or DESCRIBE AGGREGATE were used on the wrong function type was fixed. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/scylla-db-use-case-to-replace-very-slow-rdbms-common-table-expression/647 Title: Scylla DB use case to replace very slow RDBMS Common Table Expression - Knowledge Base - ScyllaDB Community NoSQL Forum Meta Description: I currently employ a relational database CTE that iterates over parentage information to find all ancestors or all descendants of a given entity. Even on a small (150000) sample, it can take over an hour to return a resu… Language: en Canonical URL: https://forum.scylladb.com/t/scylla-db-use-case-to-replace-very-slow-rdbms-common-table-expression/647 ## Headings Structure: H1: Scylla DB use case to replace very slow RDBMS Common Table Expression H1: Ancestor query: H1: Descendant query: H1: New entity creation: H1: Overall H1: Worst case H1: How slow? H1: Number of columns H1: Too hard? H3: Related topics ## Main Content: H1: Scylla DB use case to replace very slow RDBMS Common Table Expression H1: Ancestor query: H1: Descendant query: H1: New entity creation: H1: Overall H1: Worst case H1: How slow? H1: Number of columns H1: Too hard? H3: Related topics I currently employ a relational database CTE that iterates over parentage information to find all ancestors or all descendants of a given entity. Even on a small (150000) sample, it can take over an hour to return a result. I want to query 10 to 20 million entities and return a result in less than a second. I cannot get anywhere close by iterating in response to a user query. I think Scylla may come to the rescue by storing the mass of information produced by iterating as entities are added. Then user queries don’t have to wait for the iteration. Is this possible with Scylla? I am considering the following approach: Each group of inter-related entities will have its own table (probably needing a Compound Primary Key?): If I understand it right, this table will in some cases have as many as 8 million rows. Each row will potentially have 1000s of columns, each representing an ancestor associated with a high precision measure. If it didn’t hurt performance or functionality, I would possibly add a few columns of entity metadata, but the functionality really just depends on this core. A table will need to service the following types of query: I’ll want to paginate the result set in which I hope to show every ancestor of the entity in descending order of the measure. If I added other columns about the entity, I would order the return set by one of these additional columns. Otherwise I would post process both of these SELECT results with additional metadata from the relational database. Finally, I would ideally input new entities with something like a relational database SELECT INTO statement, but it seems that can’t be done in CQL? In which case I was wondering if something along the lines of the following can be done on the database side, before running the INSERT: The idea being that each ancestor in common with the ancestors chosen would have its measure column added together. That aside, one way or another, immediate ancestors would be selected and the common measure among their ancestors would be aggregated to put together the ancestors and associated measures for the new entity. With potentially thousands of ancestors per entitiy, there is doubtlessly a way to upload new entities more efficiently, I just haven’t been able to track it down yet. I anticipate users being able to query among 10’s of millions of entities each with 1000’s of columns, and return associated measure info for either: They also need to be able to add new entities with data for thousands of columns made up by aggregating a small number (at least two) of entities. As long as results can be sorted and paged, no query should be onerous. However, the sort may occationally have to take into account up to 8 million rows in a descendants query. My initial test case will be smaller, with about 200,000 entities each having 100’s to 1000’s of columns and running on a single virtual private server. Obviously, iterating first generates a lot of data. I’m hoping ScyllaDB can handle it without keeping my users waiting long. Relational DB’s definitely cannot handle the iterate per request approach in any kind of reasonable time at any necessary scale. I think ScyllaDB should be able to handle this amount of data just fine, although thousands of columns can potentially cause problems, if this is towards the higher end (10K) or more. But, to be able to query your data fast, you either have to reshape your data or rethink your queries. Your ancestor query is ordering by a regular column, which is not supported by CQL. Queries can only be ordered by clustering columns. And this only works for single partition queries. Your descendant query is a filtering query, which means it will do a full scan on the entire table, which will not be fast. The descendent query is the basis upon which I would consider switching to wide-column store. Iterating with relational or graph databases is too slow, will it be that iterating first and storing in a wide-column store database is too slow as well? What if the number of rows in a table was constrained and the total number of tables increased, so that instead having one of the tables as large as 10 million rows, it was broken into tables no larger than about 200,000 rows? Would doing a descendants search of a smaller table be likely to return a result in 1.5 seconds or less? If not, how small (number of rows) would a table need to be to return a result in that timeframe (assuming up to 10000 columns, no row having more than about 2000 columns)? Then, in the worst case I would have some tables containing a couple hundred thousand rows in which the column queried may return a ‘measure’ value for every row. It sounds like I would need to be able to find other means to store the result (after waiting for ‘will not be fast’) so that I could sort by ‘measure’ value (‘measure’ not being a clustering column). I’m not sure how long ‘will not be fast’ might take in this worst case scenario? Then in addition, it would take time and resource to end up with a sorted list by ‘measure’ outside of ScyllaDB, since measure isn’t a clustering column. For the ancestor query, I don’t expect any one row to have much over 1 to 2 thousand columns (mostly under 500 columns), but if the total unique columns for a table are basically restricted to less than 10000, that is likely to become an issue as well. I feel like I’m running out of options if wide-column stores can’t handle this use case. Structurally, it seems like the ideal database construct for this type of problem, but then so did graph databases… Note that ScyllaDB is note really a wide-column DB anymore. You need to declare up-front all your columns that a row will have, in a schema. Then you are allowed to read/write only columns that appear in said schema, and only according to the type declared in the schema. Having many smaller tables instead of few larger ones would probably be counter-productive. Instead, you can split the entire range and query token-ranges in parallel. See Efficient full table scans with ScyllaDB 1.6 - ScyllaDB on how that works. Also look into How ScyllaDB Distributed Aggregates Reduce Query Execution Time up to 20X - ScyllaDB, it might help. Overall, the only way to find out is to set-up a test-cluster and start experimenting. Thank you Botond, you have been very insightful. In my use case, the potential columns will continuously evolve and cannot be declared up-front, though their type can. I misunderstood the compound primary key as being able to provide for dynamic columns. The compound primary key allows for many rows within a partition. The column count is fixed in the schema. This can be somewhat relaxed by using collection columns, but note that it is not recommended to use large collections (more than a few dozen elements), for performance reasons. Good to know. This whole exercise (of trying to apply wide-column db concepts to my situation) has been quite useful, I think. Pivoting out my many columns into many rows per entity using a compound primary key and clustering index may well allow for a RDBMS solution after all! --- ### Page: https://forum.scylladb.com/t/release-scylla-manager-3-1-1-release/648 Title: [RELEASE] Scylla Manager 3.1.1 Release - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces the release of ScyllaDB Manager 3.1.1 production-ready ScyllaDB Manager patch releases of the stable ScyllaDB Manager 3.1 branch. As always, ScyllaDB Manager customers and users are encourage… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-manager-3-1-1-release/648 ## Headings Structure: H1: [RELEASE] Scylla Manager 3.1.1 Release H3: Related topics ## Main Content: H1: [RELEASE] Scylla Manager 3.1.1 Release H3: Related topics The ScyllaDB team announces the release of ScyllaDB Manager 3.1.1 production-ready ScyllaDB Manager patch releases of the stable ScyllaDB Manager 3.1 branch. As always, ScyllaDB Manager customers and users are encouraged to upgrade in coordination with the ScyllaDB support team. ScyllaDB Manager is a centralized cluster administration and recurrent tasks automation tool. Note that this upgrade affects both ScyllaDB Manager Server and Manager Agent; it’s recommended to upgrade both. The release fixed the following bugs: Submit a ticket for questions or issues (ScyllaDB Enterprise users) --- ### Page: https://forum.scylladb.com/t/note-scylla-manager-3-1-0-users/649 Title: [NOTE] Scylla Manager 3.1.0 users - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Due to a bug in our release repo procedures, release candidate 3.1rc0, not 3.1.0 was tagged as the latest 3.1 patch release for a time window of a few days. As a result, users who upgrade to 3.1 might use a release cand… Language: en Canonical URL: https://forum.scylladb.com/t/note-scylla-manager-3-1-0-users/649 ## Headings Structure: H1: [NOTE] Scylla Manager 3.1.0 users H3: Related topics ## Main Content: H1: [NOTE] Scylla Manager 3.1.0 users H3: Related topics Due to a bug in our release repo procedures, release candidate 3.1rc0, not 3.1.0 was tagged as the latest 3.1 patch release for a time window of a few days. As a result, users who upgrade to 3.1 might use a release candidate instead of the official release. The result of using the release candidate is backup storage usage climbing steadily as old backups are not cleaned in time. See #3415. Backups are still valid. How to check if you are exposed: If you see something like You are using rc0. please contact support for instructions on how to upgrade to 3.1.x --- ### Page: https://forum.scylladb.com/t/scylla-in-ec2-ec2-load-high/651 Title: Scylla in EC2 EC2 load high - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I have 10 node scylla-cluster,5 of these nodes are newly joined scylla-cluster ,but 5new node High load,5 old node load is normal Language: en Canonical URL: https://forum.scylladb.com/t/scylla-in-ec2-ec2-load-high/651 ## Headings Structure: H1: Scylla in EC2 EC2 load high H3: Related topics ## Main Content: H1: Scylla in EC2 EC2 load high H3: Related topics I have 10 node scylla-cluster,5 of these nodes are newly joined scylla-cluster ,but 5new node High load,5 old node load is normal Please share more details, things like: Platform:AWS-i3.4xlarge Hardware:16C 122G Disks:raid0 3.8T Scylla-version:5.2.2 OS:centos7.5 RF:3 In the beginning have 5 node ,and than add 5new node in scylla-cluster, but only new node have high load and always play compaction very frequent. just like: avg ops 12k , max 30k @huang can you clarify what is the graph it’s operation per sec ? which clients are used ? java / python / rust ? are those scylla ones ? what are there versions ? something is unbalanced, and for understand the cause we’ll need much more information about the load, we’ll need the schema, and the few example of the typical queries being done. I think ScyllaDB University | Basic & Advanced Data Modeling goes into some details of how the schema should be built clients:JAVA scylla version:all node is 5.2.2 More details are needed to understand what’s going on, things like: which clients are used ? java / python / rust ? are those scylla ones ? what are there versions ? --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-1-13/652 Title: [RELEASE] ScyllaDB 5.1.13 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.1.13, a bugfix release of the ScyllaDB 5.1 stable branch. ScyllaDB Open Source 5.1.13, like all past and future 5.x.y releases are backward compatible and support rollin… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-1-13/652 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.1.13 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.1.13 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.1.13, a bugfix release of the ScyllaDB 5.1 stable branch. ScyllaDB Open Source 5.1.13, like all past and future 5.x.y releases are backward compatible and support rolling upgrades. Note that the latest stable ScyllaDB Open Source release is 5.2, and you are encouraged to upgrade to it. Issue fixed in this release: Stability: a lot of lsa-timing log messages during node replace cause c-s stuck and aborted. The fix update the reactor shares for default IO class from 1 to 200 #13753 Stability: ALTER KEYSPACE can break tables with UDT columns #14139 Stability: All rpc::client objects have to be stopped before they are destroyed. Currently this is done in messaging_service::shutdown(). The cql_test_env does not call shutdown() currently. This can lead to use-after-free, leading to sigsegv #12244 Stability: Level compaction is not working in cleanup #14035 (introduced in 5.1) Stability: Range-scans have a protection against using the wrong service-level to continue a suspended range-scan. This protection had a mistake, resulting in the node crashing when the protection mechanism was triggered. multishard_mutation_query: reader_context::lookup_readers() is not exception safe w.r.t. closing readers #13784 Stability: Shutting down auth service may hang #13545 --- ### Page: https://forum.scylladb.com/t/commit-log-lead-to-very-high-iops-while-using-lwt/659 Title: Commit log lead to very high iops while using LWT - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: My team is testing scylla in id mapping situation recently. And in our situation, we want to keep the first id value, so we used LWT(IF NOT EXISTS) to do insert. However, the throughput of our test is very low, just 300… Language: en Canonical URL: https://forum.scylladb.com/t/commit-log-lead-to-very-high-iops-while-using-lwt/659 ## Headings Structure: H1: Commit log lead to very high iops while using LWT H3: Related topics ## Main Content: H1: Commit log lead to very high iops while using LWT H3: Related topics My team is testing scylla in id mapping situation recently. And in our situation, we want to keep the first id value, so we used LWT(IF NOT EXISTS) to do insert. However, the throughput of our test is very low, just 3000 row/s in 2-replica keyspace, and we found that the iops have reached 10k/s in each node. I know that LWT will force to sync commit log, but it’s using too much IO. Is that normal? How could I make it better? The test cluster: 3 nodes(16c/32G) Indeed ScyllaDB needs to do a disk write for each LWT operation. Normally, ScyllaDB uses periodic sync for commitlog, meaning that at the end of each sync period, all accumulated writes are flushed to disk. This is not good enought for LWT as it need stronger guarantees on the persistence of writes, therefore writes done on behalf of LWT are immediately flushed to disk. The elevated IOPS count is therefore unavoidable with LWT. LWT performs multiple cluster write rounds - prepare, promise, accept, learn, prune - and each writes to disk. Prepare, promise and accept write in sync mode. Thank you for your reply! Is there any solution to decrease the IOPS caused by LWT? What if using condition batch? Could it be better if I decrease the paxos rounds? --- ### Page: https://forum.scylladb.com/t/live-training-event-july-25/661 Title: LIVE Training event July 25 - Announcements - ScyllaDB Community NoSQL Forum Meta Description: Our next LIVE (free, online) training event will take place on July 25. Our top engineers and architects conduct the training. It will have two parallel tracks covering ScyllaDB essentials as well as advanced topics. S… Language: en Canonical URL: https://forum.scylladb.com/t/live-training-event-july-25/661 ## Headings Structure: H1: LIVE Training event July 25 H3: Related topics ## Main Content: H1: LIVE Training event July 25 H3: Related topics Our next LIVE (free, online) training event will take place on July 25. Our top engineers and architects conduct the training. It will have two parallel tracks covering ScyllaDB essentials as well as advanced topics. Save your spot. Hope to see you there! --- ### Page: https://forum.scylladb.com/t/reading-on-a-stretched-cluster/662 Title: Reading on a stretched cluster - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: We have a test stand with stretched scylladb cluster: 10.0.0.2 and 10.0.0.5 in dc1 10.0.0.3 in dc2 RF=3 CL=2 the attacking host placed in dc1 use gocql driver use policy DCAwareRoundRobinPolicy in gocql driver (… Language: en Canonical URL: https://forum.scylladb.com/t/reading-on-a-stretched-cluster/662 ## Headings Structure: H1: Reading on a stretched cluster H3: Related topics ## Main Content: H1: Reading on a stretched cluster H3: Related topics We have a test stand with stretched scylladb cluster: the attacking host placed in dc1 what is the reason that after the time expires, the application starts reading from host 10.0.0.3 (which is obviously located further - in another data center) And after some time, the situation changes exactly the opposite. When you have a multi-DC setup, make sure the Consistency Level you use for you reads is a LOCAL_ one: LOCAL_ONE or LOCAL_QUORUM. Otherwise, the replicas will be selected from among all the nodes in the cluster, disregarding DC. See Consistency Levels | ScyllaDB Docs. --- ### Page: https://forum.scylladb.com/t/storing-backup-safely/665 Title: Storing backup safely - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, I have question how do you store a backup. Scylla manager stores and deletes backup and has to have a secret key to the Storage in its configuration. What if someone unauthorised will have access to machine with tha… Language: en Canonical URL: https://forum.scylladb.com/t/storing-backup-safely/665 ## Headings Structure: H1: Storing backup safely H3: Related topics ## Main Content: H1: Storing backup safely H3: Related topics Hi, I have question how do you store a backup. Scylla manager stores and deletes backup and has to have a secret key to the Storage in its configuration. What if someone unauthorised will have access to machine with that config and use secret key to delete the backups? What solutions do you use? What if someone unauthorised will have access to machine with that config and use secret key to delete the backups? how would someone get access to that machine? where is that machine? can you please clarify where is your configuration/setup? --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-186-2023-07-02/668 Title: Last week in scylladb.git master (issue #186; 2023-07-02) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the be5b61b870…1ab2bb69b8 range are covered. There were 96 non-merge commits from 18 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-186-2023-07-02/668 ## Headings Structure: H1: Last week in scylladb.git master (issue #186; 2023-07-02) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #186; 2023-07-02) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the be5b61b870…1ab2bb69b8 range are covered. There were 96 non-merge commits from 18 authors in that period. Some notable commits: The row cache holds frequently-read rows. When a row is written with the TTL (time-to-live) option, it is set to be automatically deleted after a certain time. If it’s in the cache, however, it will continue to occupy memory, reducing cache utilization. This is now improved, as the cache will detect and remove expired rows when they are read. Infrequently read rows will be removed from the cache using the least-recently-used mechanism. It is now possible to use TLS certificates to authenticate and authorize a user to ScyllaDB. The system can be configured to derive the user role from the client certificate and derive the permissions the user has from that role. After repair or data movement due to node additions or removals, materialized views need to be updated. This process involves reading from all sstables except those that have been streamed or repaired. This was slow, and is now optimized, speeding up repair and data movement on clusters that have materialized views. A race condition between cleanup and regular compaction has been fixed. ScyllaDB uses evictable readers in certain places to allow the system to cancel ongoing reads to reclaim memory, with the ability to resume those reads later. A bug in the mechanism caused an entire partition to be consumed unnecessarily, which could lead to out-of-memory errors. This is now fixed. While using a lightweight transaction, if inconsistent constraints were given on the clustering key, ScyllaDB would crash. This is now fixed. ScyllaDB optionally uses Raft to coordinate changes to the schema and topology. It now attempts to merge adjacent changes to reduce overhead. A GROUP BY query ought to return one row per group, except when all rows of a group are filtered out. However, ScyllaDB returned a row even for fully-filtered groups. This is now fixed, and ScyllaDB will not emit rows for filtered groups. ScyllaDB will now wait for all nodes to be healthy before attempting to bootstrap a new node. When using Raft for topology and schema changes, ScyllaDB will force the schema and topology to be transferred to new nodes. Previously, ScyllaDB relied on other mechanisms to transfer the schema. A rare stack overflow in some repair scenarios has been fixed. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/cant-find-path-to-cassandra-yaml/670 Title: Cant find path to cassandra.yaml - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I tried to reinstall scylla via docker sudo docker run --name scyllaU -d scylladb/scylla docker exec -it scyllaU nodetool status Got permission denied while trying to connect to the Docker daemon socket at unix:///var… Language: en Canonical URL: https://forum.scylladb.com/t/cant-find-path-to-cassandra-yaml/670 ## Headings Structure: H1: Cant find path to cassandra.yaml H3: Related topics ## Main Content: H1: Cant find path to cassandra.yaml H3: Related topics I tried to reinstall scylla via docker Got permission denied while trying to connect to the Docker daemon socket at unix:///var/run/docker.sock: Get http://%2Fvar%2Frun%2Fdocker.sock/v1.39/containers/scyllaU/json: dial unix /var/run/docker.sock: connect: permission denied You must set the CASSANDRA_CONF and CLASSPATH vars Idk where is it located or whats the problem in my Ubuntu 22.04 box i always experience it, not sure yet how to make it work forever, but the way to go is to run: i will try to search how to make it always be the case, as I need to change it myself on every boot back to my own user based on this comment you should run: --- ### Page: https://forum.scylladb.com/t/release-scylla-operator-1-9-0/671 Title: [RELEASE] Scylla Operator 1.9.0 - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of Scylla Operator 1.9.0. Scylla Operator is an open-source project that helps ScyllaDB Open Source and ScyllaDB Enterprise users run ScyllaDB on Kubernetes. The Scyl… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-operator-1-9-0/671 ## Headings Structure: H1: [RELEASE] Scylla Operator 1.9.0 H1: Notable changes H1: Supported versions H1: Upgrade instructions H2: Getting started with Scylla Operator H2: Related Links H3: Related topics ## Main Content: H1: [RELEASE] Scylla Operator 1.9.0 H1: Notable changes H1: Supported versions H1: Upgrade instructions H2: Getting started with Scylla Operator H2: Related Links H3: Related topics The ScyllaDB team is pleased to announce the release of Scylla Operator 1.9.0. Scylla Operator is an open-source project that helps ScyllaDB Open Source and ScyllaDB Enterprise users run ScyllaDB on Kubernetes. The Scylla Operator manages ScyllaDB clusters deployed to Kubernetes and automates tasks related to operating a ScyllaDB cluster, like installation, vertical and horizontal scaling, as well as rolling upgrades. Scylla Operator 1.9.0 improves stability and brings a few features. As with all of our releases, any API changes are backward compatible. For more changes and details check out the GitHub release notes. Upgrading from v1.8.x with kubectl apply doesn’t require any extra action, just take the manifest from v1.9.0 tag and substitute the released image. Using helm requires a mandatory manual step for every release because helm can’t handle CRDs updates. For details, see our upgrade documentation. We’ll welcome your feedback! Feel free to open an issue or reach out on the #scylla-operator channel in Scylla User Slack. --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-2-4/676 Title: [RELEASE] ScyllaDB 5.2.4 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.2.4, a bugfix release of the ScyllaDB 5.2 stable branch. ScyllaDB Open Source 5.2.4, like all past and future 5.x.y releases, is backward compatible and supports rolling… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-2-4/676 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.2.4 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.2.4 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.2.4, a bugfix release of the ScyllaDB 5.2 stable branch. ScyllaDB Open Source 5.2.4, like all past and future 5.x.y releases, is backward compatible and supports rolling upgrades. Users are encouraged to upgrade to 5.2.4. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-187-2023-07-09/677 Title: Last week in scylladb.git master (issue #187; 2023-07-09) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 1ab2bb69b8…c41f0ebd2a range are covered. There were 102 non-merge commits from 13 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-187-2023-07-09/677 ## Headings Structure: H1: Last week in scylladb.git master (issue #187; 2023-07-09) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #187; 2023-07-09) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 1ab2bb69b8…c41f0ebd2a range are covered. There were 102 non-merge commits from 13 authors in that period. Some notable commits: In gossip-managed clusters, the schema is propagated by nodes contacting each other ad-hoc. In Raft-managed clusters, the schema is centrally managed by the group 0 leader. We now disable the ad-hoc schema pull method when Raft cluster management is enabled. Raft cluster management still uses gossip to translate host IDs to IP addresses. It is now more careful not to let old IP address mappings overwrite new mappings. Raft-managed clusters run Data Definition Language (DDL) statements in a transaction. The transaction scope has been extended to also include access checking and validation, and not just the actual schema change. The experimental flag used to enable consistent topology changes has been renamed from “raft” to "consistent-topology-changes|. The failure detector detects failed nodes by pinging them. Now it does not attempt to ping itself. When topology changes, CDC streams also change. The metadata describing these streams is now committed in parts, to avoid overloading the system. In Alternator, ScyllaDB’s implementation of the DynamoDB API, a bug was fixed that could cause error handling while streaming responses to the client to crash the server. In Alternator, it’s now possible to disable the DescribeEndpoints API. This makes it possible to run the dynamodb shell against ScyllaDB. When a base table of a materialized view is updated, the affected rows are also changed in the materialized view. For DELETE statements, many rows can be affected, and so the view update code splits the work into batches. However, this split was not performed correctly when range tombstones were involved. This is now fixed. In older versions of ScyllaDB, different clauses of CQL statements were processed using different code bases. ScyllaDB is gradually moving towards a single code base for processing expressions. It is now the SELECT clause’s turn, moving us closer to the goal of a unified expression syntax. As this is an internal refactoring, there are no user visible changes, apart from some names of fields in SELECT JSON statements changing (specifically, if those fields are function evaluations). See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2021-1-21/678 Title: [RELEASE] ScyllaDB Enterprise 2021.1.21 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2021.1.21, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2021.1. ScyllaDB Enterprise 2022.1 LTS is the latest Long-Term Support rele… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2021-1-21/678 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2021.1.21 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2021.1.21 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2021.1.21, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2021.1. ScyllaDB Enterprise 2022.1 LTS is the latest Long-Term Support release, and 2022.2 is the newest feature release. You are encouraged to upgrade in coordination with the ScyllaDB support team. The following issue is fixed in this release (with an open-source reference): --- ### Page: https://forum.scylladb.com/t/i-want-to-give-unique-constraint-on-one-column-in-scylladb-table-i-am-not-able-to-find-the-solution-here-i-attach-my-query-below-thanks-in-advance/682 Title: I want to give unique constraint on one column in scyllaDb table i am not able to find the solution here i attach my query below.thanks in advance - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: CREATE TABLE IF NOT EXISTS ${KEY_SPACE}.${this.name} ( session_id text, user_id int, organization_id int, deviceData map<text,frozen<map<text,text>>>, userData map<text,frozen<map<text,text>>>, created_at TIMESTAMP… Language: en Canonical URL: https://forum.scylladb.com/t/i-want-to-give-unique-constraint-on-one-column-in-scylladb-table-i-am-not-able-to-find-the-solution-here-i-attach-my-query-below-thanks-in-advance/682 ## Headings Structure: H1: I want to give unique constraint on one column in scyllaDb table i am not able to find the solution here i attach my query below.thanks in advance H3: Related topics ## Main Content: H1: I want to give unique constraint on one column in scyllaDb table i am not able to find the solution here i attach my query below.thanks in advance H3: Related topics CREATE TABLE IF NOT EXISTS ${KEY_SPACE}.${this.name} ( session_id text, user_id int, organization_id int, deviceData map>>, userData map>>, created_at TIMESTAMP, updated_at TIMESTAMP, PRIMARY KEY ((user_id,organization_id),created_at) ) WITH CLUSTERING ORDER BY (created_at DESC); What do you mean by “unique constraint”? The UNIQUE from SQL? Note that all columns that are part of the primary key (partition or clustering keys) are unique by definition. The UNIQUE from SQL ? → yes i want to implement that Note that all columns that are part of the primary key (partition or clustering keys) are unique by definition. i tried this but it’s not working i tried this but it’s not working What do you mean not working? Partition key columns are unique across the entire table, while clustering columns are unique across the partition they are located in. if i add duplicate data then it will be stored Yes, the only way to implement what you need in ScyllaDB is to add a check on the client side. This can be potentially prone to races, so you might need to use LWT. Thanks sir, really appreciated it. is it possible to give default value to column in scyllaDb No, all columns default to null (no value). --- ### Page: https://forum.scylladb.com/t/i-want-to-setup-scylladb-on-docker-in-my-local-environment-and-then-i-want-to-connect-it-to-the-node-i-tried-but-connection-is-not-established/683 Title: I want to setup scyllaDb on docker in my local environment and then i want to connect it to the node i tried but connection is not established - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: type or paste code here Language: en Canonical URL: https://forum.scylladb.com/t/i-want-to-setup-scylladb-on-docker-in-my-local-environment-and-then-i-want-to-connect-it-to-the-node-i-tried-but-connection-is-not-established/683 ## Headings Structure: H1: I want to setup scyllaDb on docker in my local environment and then i want to connect it to the node i tried but connection is not established H3: Related topics ## Main Content: H1: I want to setup scyllaDb on docker in my local environment and then i want to connect it to the node i tried but connection is not established H3: Related topics Where did you try to connect to the node from? Note, that by default, scylla will bind to the localhost address, which is not accessible from outside the docker container. The most convenient way to achieve this is to just add --network=host to your docker run command. This is fine in development setups. A more complete solution is to specify in scylla.yaml ip addresses for listen_address and rpc_address, such that they are accessible from the client’s network and the network the other scylla nodes are in respectively (the two can be the same). See Administration Guide | ScyllaDB Docs for more details. I was able to do just that with running 3-node Scylla cluster in local docker. I also have JAVA application running in the container in the same docker that connects to Scylla. Here is how I set it up: --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-1-8/684 Title: [RELEASE] ScyllaDB Enterprise 2022.1.8 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2022.1.8, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.1. Note that there is a newer Feature Enterprise release: ScyllaDB Ente… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-1-8/684 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2022.1.8 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2022.1.8 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2022.1.8, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.1. Note that there is a newer Feature Enterprise release: ScyllaDB Enterprise 2022.2. While we will continue to support 2022.1 LTS, you can get additional features with 2022.2. Below is a list of performance and stability improvements and bug fixes, each with an open-source reference if available: Stability: Range-scans have a protection against using the wrong service-level to continue a suspended range-scan. This protection had a mistake, resulting in the node crashing when the protection mechanism was triggered. multishard_mutation_query: reader_context::lookup_readers() is not exception safe w.r.t. closing readers #13784 Performance: ScyllaDB uses an interval map data structure from the Boost library to quickly locate sstables needed to service a read. Due to the way we interface with the library, updating the interval map was unnecessarily slow. This is now fixed. #11669 Stability: All rpc::client objects have to be stopped before they are destroyed. Currently this is done in messaging_service::shutdown(). The cql_test_env does not call shutdown() currently. This can lead to use-after-free, leading to sigsegv #12244 Stability: race condition in gossip and failure detector #10547 Stability: mutation_reader_merger can overflow stack when merging many empty readers. This may happen when running a second repair right after the other. #14415 UX: Allow tombstone GC in compaction to be disabled on user request #14077 The fix adds new APIs /column_family/tombstone_gc and /storage_service/tombstone_gc, that will allow for disabling tombstone garbage collection (GC) in compaction. The table name must be in keyspace:table format --- ### Page: https://forum.scylladb.com/t/migrating-from-amazon-mysql/686 Title: Migrating from Amazon MySQL - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: We are migrating from Amazon RDS MySQL to ScyllaDB. Will ScyllaDB support AWS DMS (replication) CDC to our Amazon Redshift reporting server? Language: en Canonical URL: https://forum.scylladb.com/t/migrating-from-amazon-mysql/686 ## Headings Structure: H1: Migrating from Amazon MySQL H3: Related topics ## Main Content: H1: Migrating from Amazon MySQL H3: Related topics We are migrating from Amazon RDS MySQL to ScyllaDB. Will ScyllaDB support AWS DMS (replication) CDC to our Amazon Redshift reporting server? ScyllaDB has a Change Data Capture (CDC) implementation, which can easily be integrated with Kafka. You can process and send the information to be processed by different consumers. So although we are not integrated with DMS CDC, streaming CDC info to Redshift is not a problem. You can learn more in the CDC Lesson on ScyllaDB University. Additionally, you can see a working example in this hands-on ScyllaDB CDC Source Connector with Kafka – Lab. --- ### Page: https://forum.scylladb.com/t/scylla-manager-web-ui/687 Title: Scylla Manager Web UI - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Does Scylla Manager has web UI? Language: en Canonical URL: https://forum.scylladb.com/t/scylla-manager-web-ui/687 ## Headings Structure: H1: Scylla Manager Web UI H3: Restore | ScyllaDB Docs H3: Related topics ## Main Content: H1: Scylla Manager Web UI H3: Restore | ScyllaDB Docs H3: Related topics Does Scylla Manager has web UI? No, scylla-manager doesn’t have web ui. It’s managed with CLI sctool CLI sctool | ScyllaDB Docs Thanks, @Karol_Kokoszka I’ve additional questions around backup and restore using Scylla Manager. It’s not an option now. These backup locations are supported Backup | ScyllaDB Docs The only workaround I see is to setup backup location in Minio Setup S3 compatible storage | ScyllaDB Docs There is an option to restore the data from backup, but it has some prerequisites like: Scylla Manager does NOT validate that, so it’s user responsibility to ensure it. The only form of validation that Scylla Manager performs is checking whether restored tables are present in destination cluster, but it does not validate their columns types nor other properties. In case destination cluster is missing correct schema, it should be restored first. Otherwise, restored tables’ contents might be overwritten by the already existing ones. Note that an empty table is not necessarily truncated! ScyllaDB is an Apache Cassandra-compatible NoSQL data store that can handle 1 million transactions per second on a single server. --- ### Page: https://forum.scylladb.com/t/intsalling-sylla/689 Title: Intsalling sylla - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: ERROR: You will not be able to run Scylla on this machine because its CPU lacks the following features: pclmulqdq sse4_2 If this is a virtual machine, please update its CPU feature configuration or upgrade to a newer hy… Language: en Canonical URL: https://forum.scylladb.com/t/intsalling-sylla/689 ## Headings Structure: H1: Intsalling sylla H3: Related topics ## Main Content: H1: Intsalling sylla H3: Related topics ERROR: You will not be able to run Scylla on this machine because its CPU lacks the following features: pclmulqdq sse4_2 If this is a virtual machine, please update its CPU feature configuration or upgrade to a newer hypervisor. Where are you trying to install scylla? Looks like the machine you are trying to install on is not supported. --- ### Page: https://forum.scylladb.com/t/release-scylla-manager-3-1-2/690 Title: [RELEASE] Scylla Manager 3.1.2 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces the release of ScyllaDB Manager 3.1.2 production-ready ScyllaDB Manager patch releases of the stable ScyllaDB Manager 3.1 branch. As always, ScyllaDB Manager customers and users are encourage… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-manager-3-1-2/690 ## Headings Structure: H1: [RELEASE] Scylla Manager 3.1.2 H3: Related topics ## Main Content: H1: [RELEASE] Scylla Manager 3.1.2 H3: Related topics The ScyllaDB team announces the release of ScyllaDB Manager 3.1.2 production-ready ScyllaDB Manager patch releases of the stable ScyllaDB Manager 3.1 branch. As always, ScyllaDB Manager customers and users are encouraged to upgrade in coordination with the ScyllaDB support team. ScyllaDB Manager is a centralized cluster administration and recurrent tasks automation tool. Scylla Manager restore was failing whenever Scylla kept different values for listen_address and rpc_address scylla-manager#3411. This release fixes that issue. Submit a ticket for questions or issues (ScyllaDB Enterprise users) --- ### Page: https://forum.scylladb.com/t/creating-alternator-keyspace-using-java-driver/691 Title: Creating alternator keyspace using JAVA driver - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Right now, I am using Python’s boto3 library to create my alternator space: dynamodb = boto3.resource('dynamodb', endpoint_url=f'http://{alt_host}:8000', region_name='None', aws_access_key_id='… Language: en Canonical URL: https://forum.scylladb.com/t/creating-alternator-keyspace-using-java-driver/691 ## Headings Structure: H1: Creating alternator keyspace using JAVA driver H3: Related topics ## Main Content: H1: Creating alternator keyspace using JAVA driver H3: Related topics Right now, I am using Python’s boto3 library to create my alternator space: According to documentation Using DynamoDB API in Scylla it should support several other languages: I was not able to find any examples of this being done in other languages. Anyone can share some code samples, specifically in JAVA (as I am writing my main application in JAVA, I would like to keep it to one language) DynamoDB doesn’t have a notion of keyspaces, so it’s not possible to create one. Instead just create table, here are some samples from aws sdk (alternator is compatible with dynamodb): https://github.com/awsdocs/aws-doc-sdk-examples/tree/main/javav2/example_code/dynamodb. Additionally some example from our alternator load balancing library should demonstrate where to put unused placeholders as alternator doesn’t use AWS region configuration and endpoints: https://github.com/scylladb/alternator-load-balancing/blob/master/java/src/main/java/com/scylladb/alternator/AlternatorClient.java I ended up coming up with this code: --- ### Page: https://forum.scylladb.com/t/extracting-data-from-scylla-for-data-mining/692 Title: Extracting data from Scylla for data mining - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I would like to extract data in batches from ScyllaDB and dump them into files or an OLAP database for data mining. Do you have a suggestion? Language: en Canonical URL: https://forum.scylladb.com/t/extracting-data-from-scylla-for-data-mining/692 ## Headings Structure: H1: Extracting data from Scylla for data mining H3: Related topics ## Main Content: H1: Extracting data from Scylla for data mining H3: Related topics I would like to extract data in batches from ScyllaDB and dump them into files or an OLAP database for data mining. Do you have a suggestion? There are a number of ways to achieve this. One is using Presto (see this tech talk). Another is by using Spark with ScyllaDB. By doing so, you deploy analytics workloads on information stored in ScyllaDB. You can learn more about this and see a hands-on lab in the Using Spark with ScyllaDB lesson on ScyllaDB University. --- ### Page: https://forum.scylladb.com/t/scylladb-university-live-training-25th-of-july/695 Title: ScyllaDB University LIVE Training - 25th of July - Blog Posts - ScyllaDB Community NoSQL Forum Meta Description: ScyllaDB University LIVE is taking place on the 25th of July. It’s a half day of live (free, online) training with some of our top engineers and experts. You can read more about it in this blog post. Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-university-live-training-25th-of-july/695 ## Headings Structure: H1: ScyllaDB University LIVE Training - 25th of July H3: Related topics ## Main Content: H1: ScyllaDB University LIVE Training - 25th of July H3: Related topics ScyllaDB University LIVE is taking place on the 25th of July. It’s a half day of live (free, online) training with some of our top engineers and experts. You can read more about it in this blog post. The training event is happening this week! Save your spot here. --- ### Page: https://forum.scylladb.com/t/how-to-get-node-sharding-information-using-scylla-java-driver/696 Title: How to get Node Sharding information using Scylla Java Driver? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: In JanusGraph we are currently working on adding a new feature to allow users execute SELECT query for multiple partition keys using IN operator. Usually IN operator usage for different partitions is an anti-pattern (I … Language: en Canonical URL: https://forum.scylladb.com/t/how-to-get-node-sharding-information-using-scylla-java-driver/696 ## Headings Structure: H1: How to get Node Sharding information using Scylla Java Driver? H3: Related topics ## Main Content: H1: How to get Node Sharding information using Scylla Java Driver? H3: Related topics In JanusGraph we are currently working on adding a new feature to allow users execute SELECT query for multiple partition keys using IN operator. Usually IN operator usage for different partitions is an anti-pattern (I think that’s why there is a default protection exists in ScyllaDB max-partition-key-restrictions-per-query: Scylla FAQ | ScyllaDB Docs ). However, on JanusGraph side we want to group those partition keys based on either Token Ranges only (for Cassandra / AstraDB / Amazon Keyspaces) or based on Token Ranges plus Sharding information. We often need to fetch specific columns for many partition keys together. As for now we are executing a separate asynchronous CQL query per each partition key. Often times it leads to channel conjunctions when we need to select those columns for let’s say thousands of keys which leads to thousands of CQL queries. We found out that grouping keys by limited groups of keys (let’s say grouping 100 partition keys via IN operator) which belong to the same token range (i.e. keys which leave on the same Nodes) often results in slightly better performance on Cassandra side, as well as making some serverless deployments like AstraDB, where you pay 1 RRU for each CQL query which returns 4 KB of data way cheaper for running JanusGraph, because usually JanusGraph queries small amount of bytes which is much less than 4 KB. We were able to make the implementation (see here: Group CQL keys into a single query [cql-tests] [tp-tests] by porunov · Pull Request #3879 · JanusGraph/janusgraph · GitHub) which groups partition keys based on Token Ranges. I.e. we convert a partition key into a Token and then search for it’s TokenRange using TokenMap. Any partition keys which belong to the same TokenRange are grouped together (into small groups, so that we have a chance to touch multiple replicas, but let’s omit replicas existence for simplicity and assume each TokenRange has only 1 Node). The implementation works good for Cassandra, but we would like to have an improved implementation for ScyllaDB, because ScyllaDB routes queries not only by a RoutingToken / RoutingKey, but also by the Shard Id of that RoutingKey. Meaning that the request is routed directly to the correct CPU for processing. As we group multiple partition keys together via IN operator, we have a small challenge of: “How do we determine if two partition keys, which belong to the same TokenRange are going to be routed to the same Shard?”. One way I’m think of is simply getting all nodes for a partition key and if for two different partition keys ShardId matches for ALL the nodes - than it means that they can be grouped together. The problem is that I don’t know how exactly get ShardId if all I have is CQLSession, Node, and partition key. If anyone knows how to properly get Shard Id (or shard infomration) using Node + partition key via ScyllaDB Java Driver, it would be really great if you can share that. Any other concerns and / or suggestions are highly appreciated. In Java Driver 3.x, you can calculate the shard ID with the following code: Unfortunately, since we’ve made some Token classes private in the driver, we have to re-implement Token in the snippet. The situation is worse in Java Driver 4.x, where ShardingInfo class is not public. I created this tracking issue for possibly making those APIs public: Expose API for calculating shard ID from token · Issue #232 · scylladb/java-driver · GitHub The calculation of shard id is performed here in the driver: https://github.com/scylladb/java-driver/blob/d291df6b35f7903c0b2d935754aebcb5b35bcd81/driver-core/src/main/java/com/datastax/driver/core/ShardingInfo.java#L57-L67. Some of the values it uses (e.g. shardingIgnoreMSB) are read from the initial handshake the driver makes with ScyllaDB. Hi @piotr ! Thank you so much for chiming in! Sorry for confusion, but at JanusGraph we are using 4.x driver. So I think we will need to wait until the ticket you opened is implemented. Appreciate you quick response! --- ### Page: https://forum.scylladb.com/t/how-to-recover-data-from-a-snapshot-file/698 Title: How to recover data from a snapshot file? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: In some cases, I accidentally deleted a table in ScyllaDB. Now, how can I read the snapshot file from the snapshots directory in the ScyllaDB data directory to recover the data? The directory for snapshot files is as fo… Language: en Canonical URL: https://forum.scylladb.com/t/how-to-recover-data-from-a-snapshot-file/698 ## Headings Structure: H1: How to recover data from a snapshot file? H3: Related topics ## Main Content: H1: How to recover data from a snapshot file? H3: Related topics In some cases, I accidentally deleted a table in ScyllaDB. Now, how can I read the snapshot file from the snapshots directory in the ScyllaDB data directory to recover the data? The directory for snapshot files is as follows: To recover the data from a snapshot, copy all the sstables from the snapshot, to the upload directory in the table’s data directory, then issue nodetool refresh. If the topology has changed since the snapshot was created, you should use nodetool refresh --load-and-stream. Make sure to not copy auxiliary files, like manifest.json and schema.cql. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-188-2023-07-16/699 Title: Last week in scylladb.git master (issue #188; 2023-07-16) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the c41f0ebd2a…a93fd2b162 range are covered. There were 104 non-merge commits from 17 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-188-2023-07-16/699 ## Headings Structure: H1: Last week in scylladb.git master (issue #188; 2023-07-16) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #188; 2023-07-16) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the c41f0ebd2a…a93fd2b162 range are covered. There were 104 non-merge commits from 17 authors in that period. Some notable commits: A bug involving incorrect cross-shard access while performing a nodetool scrub command was fixed. Usually, repair reconciles a shard’s data on one node with the data on the same shard in other nodes. When the number of shards on different nodes doesn’t match, repair has to pick small ranges from all shards on the remote nodes. This adds significant overhead which is most pronounced when there is little or no data in the table. This is common in tests and slows then down, so we now have an optimization for the little-data case. A complication causing problems in bootstrap handling very recent changes in topology was fixed. A recent change extending the scope of CQL Data Definition Language (DDL) transactions caused a significant performance regression, so it was reverted. A production installation of ScyllaDB locks all memory so we don’t experience high latency due to page faults. However, this only applies from the first time the memory is accessed; the first access can still experience stalls, made larger by using transparent huge pages. To fix this, a Seastar update adds a prefault thread that attempts to access all memory ahead of the database, taking the latency hit on this new thread rather than user queries. This will be visible as increased CPU consumption during the first few seconds (up to a minute on large machines) during process start. We now tune the Linux kernel’s caching of inodes (in-memory structure representing file metadata) to favor evicting inodes quickly. This aims to reduce kernel memory fragmentation when there are large numbers of sstables, as most files comprising an sstable aren’t accessed after the process starts. When compaction completes, it reports the throughput it achieved. We now base it on the input bytes read rather than output bytes, as the latter gives incorrect results for overwrite or expiring workloads. The messaging service, responsible for inter-node communication, now initializes transport-layer security (TLS) earlier, to account for the failure detector pinging the its own node. Resharding is a process where an sstable is split into several sstables, each wholly-owned by a single shard. A recent change to integrate resharding into the task manager was found to crash the system, so it was reverted. Cleanup is a process where an sstable is rewritten to discard all partitions that no longer belong to the node (for example, after bootstrap). It has gained an optimization where we skip over the unnecessary partitions rather than reading and discarding them. Alternator, ScyllaDB’s implementation of the DynamoDB API, implemented the error path of the size() function incorrectly. This is now fixed. Recently, we started merging adjacent schema and topology changes when controlled by Raft, but this merging had a subtle bug leading to incorrect merging. This is now fixed. The system uses a reader_concurrency_semaphore to limit the number of concurrent reads, as each read can consume large amounts of memory when merging sstables. Repair has its own allocation of concurrent reads. We now limit the scope of a read more carefully, to allow new reads to issue more quickly. A recent regression involving a crash in decommission was fixed. There is now a configuration item to control the size of stream plans, as a fraction of the total number of token ranges to stream, DateTieredCompactionStrategy was removed. Users should move to TimeWindowCompactionStrategy. It is already illegal to create new DateTieredCompactionStrategy tables. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/scylladb-dbaas-update-frequency/701 Title: ScyllaDB DBaaS update frequency - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: How frequently do you update the ScyllaDB managed cloud service (and what)? Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-dbaas-update-frequency/701 ## Headings Structure: H1: ScyllaDB DBaaS update frequency H3: Related topics ## Main Content: H1: ScyllaDB DBaaS update frequency H3: Related topics How frequently do you update the ScyllaDB managed cloud service (and what)? When you run ScyllaDB Cloud, you don’t have to worry about updates as we take care of that. The Cloud services are updated quite often to include the latest ScyllaDB Enterprise version as well as the Cloud user interface and API. You can use the cloud-release tag in the forum to see all the ScyllaDB Cloud updates. In addition, we ensure customers are always running under a supported release (so they don’t have to worry about product life cycles). --- ### Page: https://forum.scylladb.com/t/considering-alternatives-to-document-based-database/702 Title: Considering alternatives to document based database - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: We have ETL pipelines that write millions of records daily from OLAP to OLTP databases using REST APIs. We’re exploring db options that would provide very high write throughput, that doesn’t degrade overtime, and ensures… Language: en Canonical URL: https://forum.scylladb.com/t/considering-alternatives-to-document-based-database/702 ## Headings Structure: H1: Considering alternatives to document based database H3: Related topics ## Main Content: H1: Considering alternatives to document based database H3: Related topics We have ETL pipelines that write millions of records daily from OLAP to OLTP databases using REST APIs. We’re exploring db options that would provide very high write throughput, that doesn’t degrade overtime, and ensures reads from our apis are not overloaded by the high write throughput. We’re exploring ScyllaDB, but our current schema is document based (denormalized) and not normalized. Thoughts? As you require high write throughput, your use case is a good fit for ScyllaDB. With regards to denormalization and data modeling in NoSQL vs. SQL in general, I recommend the Data Modeling course on ScyllaDB University and specifically the Denormalization lesson. Another feature that might be relevant for your use case is Workload Prioritization. It allows users to run OLAP and OLTP workloads on ScyllaDB without sacrificing latency or throughput. --- ### Page: https://forum.scylladb.com/t/removing-a-dead-node/703 Title: Removing a dead node - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I have a Scylla cluster running. Due to some storage issues I had to remove a node. The node is not physically alive at this moment. I tried to remove this via nodetool removenode ID But the node has been in DL state f… Language: en Canonical URL: https://forum.scylladb.com/t/removing-a-dead-node/703 ## Headings Structure: H1: Removing a dead node H3: Related topics ## Main Content: H1: Removing a dead node H3: Related topics I have a Scylla cluster running. Due to some storage issues I had to remove a node. The node is not physically alive at this moment. I tried to remove this via nodetool removenode ID But the node has been in DL state from quite sometime due to which I can not run certain operations or add new nodes as they are giving consistency errors. Is there anyway I can update the state of the cluster? If it’s possible, try to restore the node and then use nodetool decommission. Otherwise, you can use nodetool removnode. This lets the other nodes know that the node is down, and redistributes the data. You can see all the details here. Suppose you have 3 nodes. Node1, node2, node3. Due to some reasons you did node3 terminated completely. its not alive. JUST Login the node1 or node2 and check nodetool status it shows as 3 nodes and node 3 is DN state. Now you want to remove the node3 from nodetool status . use below steps. step1: capture the node3 IP and host id from nodetool status output. Now prepare below command nodetool removenode --ignore-dead-nodes Node3 IP Node3 Host id Now execute command. Once done validate nodetool status --- ### Page: https://forum.scylladb.com/t/partitions-in-scylla/705 Title: Partitions in Scylla - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I am building a chat application using Scylla, and I am creating partitions using the chat ID and a static timestamp. For example, let’s say the chat ID is equal to 1 and the timestamp is 11/2024. I am not sure how to re… Language: en Canonical URL: https://forum.scylladb.com/t/partitions-in-scylla/705 ## Headings Structure: H1: Partitions in Scylla H3: Related topics ## Main Content: H1: Partitions in Scylla H3: Related topics I am building a chat application using Scylla, and I am creating partitions using the chat ID and a static timestamp. For example, let’s say the chat ID is equal to 1 and the timestamp is 11/2024. I am not sure how to retrieve the partition that comes before this one if I want to perform a ‘load more’ action. Additionally, what should I do if the partition before that is empty and I need to retrieve the partition before it? Partitions in ScyllaDB are not ordered, so it is not possible to retrieve a partition, which is just before another partition. It can maybe be made to work by generating all possible partition-keys, but this is not efficient. I think you should instead change your schema, such that chat_id is the partition key and timestamp is the clustering key. Clustering rows are ordered and it is possible to select a range of them. --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-1-14/706 Title: [RELEASE] ScyllaDB 5.1.14 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.1.14, a bugfix release of the ScyllaDB 5.1 stable branch. ScyllaDB Open Source 5.1.14, like all past and future 5.x.y releases are backward compatible and support rollin… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-1-14/706 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.1.14 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.1.14 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.1.14, a bugfix release of the ScyllaDB 5.1 stable branch. ScyllaDB Open Source 5.1.14, like all past and future 5.x.y releases are backward compatible and support rolling upgrades. Note that the latest stable ScyllaDB Open Source release is 5.2, and you are encouraged to upgrade to it. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-2-11/707 Title: [RELEASE] ScyllaDB Enterprise 2022.2.11 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2022.2.11, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.2. Related Links Get ScyllaDB Enterprise 2022.2.11 (customers only, … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-2-11/707 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2022.2.11 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2022.2.11 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2022.2.11, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.2. The following issue are fixed in this release (with an open-source reference, if available): Stability: a rare failure in row_cache_test/test_concurrent_reads_and_eviction #12462 Stability: a lot of lsa-timing log messages during node replace cause c-s stuck and aborted. The fix update the reactor shares for default IO class from 1 to 200 #13753 Stability: All rpc::client objects have to be stopped before they are destroyed. Currently this is done in messaging_service::shutdown(). The cql_test_env does not call shutdown() currently. This can lead to use-after-free, leading to sigsegv #12244 Stability: ICS compaction is not working in cleanup #14035 (introduced in 2022.2.0) Stability: Range-scans have a protection against using the wrong service-level to continue a suspended range-scan. This protection had a mistake, resulting in the node crashing when the protection mechanism was triggered. multishard_mutation_query: reader_context::lookup_readers() is not exception safe w.r.t. closing readers #13784 Stability: bad_alloc (seastar - Failed to allocate 536870912 bytes) #13491. Root cause is a logic fault causing the reader to attempt to read all the data, consuming all memory. Can occur during sstableloader/nodetool refresh, repair or range scan. --- ### Page: https://forum.scylladb.com/t/release-scylladb-cloud-20-july-2023/710 Title: [RELEASE] ScyllaDB Cloud - 20 July 2023 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: We are happy to announce a new update to the ScyllaDB Cloud service. New features and enhancements: We’ve simplified the UI by adding a new dropdown menu in the top-right navigation area: The user’s name along with… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-cloud-20-july-2023/710 ## Headings Structure: H1: [RELEASE] ScyllaDB Cloud - 20 July 2023 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Cloud - 20 July 2023 H3: Related topics We are happy to announce a new update to the ScyllaDB Cloud service. New features and enhancements: We’ve simplified the UI by adding a new dropdown menu in the top-right navigation area: We’ve moved the Maintenance Windows tab to be under the My Clusters area. --- ### Page: https://forum.scylladb.com/t/nodetool-java-17-compatibility/711 Title: Nodetool Java 17 compatibility - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, it seems scylla-jmx is not compatible with Java 17. Are there any plans making it Java 17 compatible, or perhaps even ditch scylla-jmx for a REST-based nodetool? best regards, Christian Language: en Canonical URL: https://forum.scylladb.com/t/nodetool-java-17-compatibility/711 ## Headings Structure: H1: Nodetool Java 17 compatibility H3: Related topics ## Main Content: H1: Nodetool Java 17 compatibility H3: Related topics Hi, it seems scylla-jmx is not compatible with Java 17. Are there any plans making it Java 17 compatible, or perhaps even ditch scylla-jmx for a REST-based nodetool? best regards, Christian We do have plans to replace nodetool with a non-java one, that talks directly to the REST API. That said, we didn’t even start working on this, so I cannot say when it will be ready. I too facing the same issue. Scylla is getting failed to start with java 17. Tried with different scylla versions 5.1 and 5.2. Please find the error below. Best Regards Raghuveera --- ### Page: https://forum.scylladb.com/t/cassandra-stress-limit-items-in-the-list-data-type/712 Title: Cassandra Stress - Limit items in the List Data Type - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Using cassandra-stress to benchmark ScyllaDB 5.0. Is there any support to limit number of items in the list? CREATE TABLE keyspace1.users ( username text, user_fav_movies list, PRIMARY KEY (username) Cassandra-str… Language: en Canonical URL: https://forum.scylladb.com/t/cassandra-stress-limit-items-in-the-list-data-type/712 ## Headings Structure: H1: Cassandra Stress - Limit items in the List Data Type H3: Related topics ## Main Content: H1: Cassandra Stress - Limit items in the List Data Type H3: Related topics Using cassandra-stress to benchmark ScyllaDB 5.0. Is there any support to limit number of items in the list? CREATE TABLE keyspace1.users ( username text, user_fav_movies list, PRIMARY KEY (username) Cassandra-stress table user profile: I think this might be what you are looking for: --- ### Page: https://forum.scylladb.com/t/scylladb-stargate/714 Title: ScyllaDB + Stargate - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello ScyllaDB community, We need a document interface to access our data in ScyllaDB, but we have noticed that it is not available. We have come across the Stargate.io project, which provides a document-oriented API fo… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-stargate/714 ## Headings Structure: H1: ScyllaDB + Stargate H3: Related topics ## Main Content: H1: ScyllaDB + Stargate H3: Related topics Hello ScyllaDB community, We need a document interface to access our data in ScyllaDB, but we have noticed that it is not available. We have come across the Stargate.io project, which provides a document-oriented API for Cassandra DB. We tried to run Stargate.io with ScyllaDB since ScyllaDB claims to have API compatibility, but we were unable to do so. Does anyone have any thoughts or experience with this? Is it possible to integrate a document interface with ScyllaDB? We would greatly appreciate any insights or suggestions. Thank you in advance for your assistance. Hey Alexander, What were the issues you ran into when you tried to run Stargate.io with ScyllaDB? Please share as many details as possible, so that we can get a better idea. I am seeing similar behavior when attempting to set this up in a development environment using single node microk8s install and Helm. Result is that the coordinator fails to start with a “ServiceStartException: Unable to start persistence-cassandra-3.11” caused by the runtime exception “Unable to gossip with any peers”. Result: A runtime exception starting the stargate-coordinator: ERROR: Bundle io.stargate.db.cassandra_3_11 [1] EventDispatcher: Error during dispatch. (io.stargate.core.activator.ServiceStartException: Unable to start persistence-cassandra-3.11) io.stargate.core.activator.ServiceStartException: Unable to start persistence-cassandra-3.11 … Caused by: java.lang.RuntimeException: Unable to gossip with any peers at org.apache.cassandra.gms.Gossiper.doShadowRound(Gossiper.java:1603) … I was able to find an article that seems to indicate I may need to use ssl as the gossip protocol. I did attempt this by mounting a trust and keystore and passing the additiona java opts but same result as above. Any help appreciated as we would like to evaluate this combo in a development environment. Stargate tries to become a member of the cluster. Since ScyllaDB isn’t compatible with Cassandra at the internode protocol level, this doesn’t work. This has to be fixed in Stargate, by making it only use CQL when talking to ScyllaDB clusters. --- ### Page: https://forum.scylladb.com/t/scylladb-as-loki-backend-db/716 Title: Scylladb as Loki backend DB - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Is it possible to use ScyllaDB as backend database for Loki storage system. As per the Loki documentation it supports Apache Cassandra. But there is nothing about ScyllaDB. Does anyone know about the ScyllaDB support in … Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-as-loki-backend-db/716 ## Headings Structure: H1: Scylladb as Loki backend DB H3: Related topics ## Main Content: H1: Scylladb as Loki backend DB H3: Related topics Is it possible to use ScyllaDB as backend database for Loki storage system. As per the Loki documentation it supports Apache Cassandra. But there is nothing about ScyllaDB. Does anyone know about the ScyllaDB support in Loki. If Cassandra is supported as a backend, ScyllaDB should also work. That said, the only way to know for sure is to try. Thank you for your input. Yes. I am trying it. @Lahiru please update if you got it to work or if you ran into any issues. I know there were more people in the community that were also interested in using Loki with ScyllaDB as the backend, see here for example. --- ### Page: https://forum.scylladb.com/t/backup-and-resore/717 Title: Backup and resore - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello. What is the easiest way to backup and restore scylla db keyspace from cluster1 to cluster2? Language: en Canonical URL: https://forum.scylladb.com/t/backup-and-resore/717 ## Headings Structure: H1: Backup and resore H3: Related topics ## Main Content: H1: Backup and resore H3: Related topics Hello. What is the easiest way to backup and restore scylla db keyspace from cluster1 to cluster2? The easiest way is to use ScyllaDB Manager, see the Restore documentation. This is also covered in the Cluster Backup Using ScyllaDB Manager lesson on ScyllaDB University. Another option is to use snapshots, see the Backup and Restore documentation. --- ### Page: https://forum.scylladb.com/t/last-month-in-scylla-cluster-tests-git-master-issue-16-2023-07-21/718 Title: Last month in scylla-cluster-tests.git master (issue #16; 2023-07-21) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short(ish) report brings to light some interesting commits to scylla-cluster-tests.git master from the last ~month. Commits in the a1036e1b…9dfb2392 range are covered. There were 116 non-merge commits from 15 autho… Language: en Canonical URL: https://forum.scylladb.com/t/last-month-in-scylla-cluster-tests-git-master-issue-16-2023-07-21/718 ## Headings Structure: H1: Last month in scylla-cluster-tests.git master (issue #16; 2023-07-21) H3: Related topics ## Main Content: H1: Last month in scylla-cluster-tests.git master (issue #16; 2023-07-21) H3: Related topics This short(ish) report brings to light some interesting commits to scylla-cluster-tests.git master from the last ~month. Commits in the a1036e1b…9dfb2392 range are covered. There were 116 non-merge commits from 15 authors in that period. Some notable commits: As a first step in testing Audit functionality on larger scale we added short longevity with audit enabled by default. SCT was adapted to uuid_sstable_identifiers_enabled where sstables are identified with timeuuid instead of an integer. Check out docs: How can I connect to test machines in AWS ? Since releng is moving to create the images on another account, we have introduced the changes to list the images regardless general or releng accounts but were missing some fixes and improvements like: add account images tags, add image’s owner ID, and checking owner ID. New GH action was added to close stale issues after 2 years of inactivity and PR’s after 1 year, so we can enjoy fast and consistent pace of closing issues and PR’s. This change adds a Global Secondary Index to a 3 hour large partition test and a materialized view to the 200k pks 4days test. This change adds a random chance to use a “USING TIMESTAMP” clause instead of normal 50% of the partition in delete_by_range_using_timestamp disruption. K8S 1.27 is now supported in SCT tests Because the old ‘static’ approach for PV provisioning on GKE doesn’t work on latest K8S version, use dynamic local volume provisioner on GKE. Later static local volume provisioner was fixed using solution recommended by Google. Introduced a new perf throughput non-shard aware i4i tests K8s tests were swithed to use and test manager 3.1 New event process is presented and new context manager. Context manager allow to start/stop count events and collect some stats event process allow to filter collected events and save events to files in specified directory. Nodetool cleanup nemesis was refactored to run in parallel without specifying keyspace. In order to catch more issues with installations and upgrades that are related to instance type sizes (e.g. perftune issues related to amount of cores) we spreaded 5 common instances types in gce with higher amount of disks accordingly Now we are able to set longevity test duration from job settings. Hinted handoff feature was disabled for all tests as probably is rarely used on the field, leaving one sanity test to make sure it still works. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-189-2023-07-24/719 Title: Last week in scylladb.git master (issue #189; 2023-07-24) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the a93fd2b162…decbc841b7 range are covered. There were 144 non-merge commits from 19 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-189-2023-07-24/719 ## Headings Structure: H1: Last week in scylladb.git master (issue #189; 2023-07-24) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #189; 2023-07-24) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the a93fd2b162…decbc841b7 range are covered. There were 144 non-merge commits from 19 authors in that period. Some notable commits: There is now a REST API for configuring Prometheus metrics label rewriting. The hints synchronization point API allows an external user to wait for hints replay. Misuse of the API cookie could lead to unbounded memory usage; the cookie is now protected with a checksum. ScyllaDB uses feature flags to coordinate rolling upgrades; a feature isn’t enabled until all nodes report they support that particular feature flag. Occasionally some older feature flags are considered “always on” and aren’t negotiated. A problem with non-negotiated features and storing feature flags in Raft group 0 would have prevented upgrades, but it was fixed. Among its tasks, gossiper disseminates cluster state updates within the node. Node removal notifications were processed in the background, which could cause them to be reordered with other notifications, causing problems. This is fixed by moving processing to the foreground. A recent regression when using GROUP BY together with the ttl() and writetime() pseduo-functions was fixed. There is a new SELECT statement variant that allows seeing where the data that composes a selection comes from. Normally, cache, sstable, and memtable data are merged before output, but with this variant one can see the original source of the data. This is intended for forensics and is not a stable API. A Seastar update will reserve more memory for the operating system in situations that previously led to out-of-memory errors. These situations are ARM machines with 64kB pages (rather than the usual 4kB), and transparent hugepages enabled. As a side effect ScyllaDB will run with less memory. The compaction manager sometimes generates sstables composed of only tombstones, in order to safeguard against a crash causing data resurrection. If there isn’t a crash, these sstables can be safely deleted. However, they are sometimes picked up for compaction before they are deleted, wasting CPU cycles. They are now excluded from compaction. The cassandra-stress tool now supports the Java driver’s rack-aware policy. This can reduce cloud inter availability zone networking costs, with the downside of less even load balancing if care isn’t taken to balance the application. The CQL grammar incorrectly accepted nonsensical empty limit clauses such as SELECT * FROM tab LIMIT;. The errors were discovered later in processing, but with unhelpful error messages. They are now rejected. The CQL grammar incorrectly accepted nonsensical INSERT JSON statements such as INSERT INTO tab JSON;, causing a crash. This is now fixed. The mutation compactor now validates its input stream rather than the output stream. A mistake in function type inference, which could lead the CQL statements to claim there is ambiguity when in fact there is none, was fixed. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-cloud-24-july-2023/720 Title: [RELEASE] ScyllaDB Cloud - 24 July 2023 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: We are happy to announce a new update to the ScyllaDB Cloud service. New features and enhancements: We’ve added support for i3en AWS instances in the ap-southeast-2 (Sydney) region. Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-cloud-24-july-2023/720 ## Headings Structure: H1: [RELEASE] ScyllaDB Cloud - 24 July 2023 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Cloud - 24 July 2023 H3: Related topics We are happy to announce a new update to the ScyllaDB Cloud service. New features and enhancements: --- ### Page: https://forum.scylladb.com/t/how-scylladb-controls-data-redistribution-or-data-backfill-traffic/721 Title: How scylladb controls data redistribution or data backfill traffic - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: When we add a node to a cluster or remove a node from a cluster, data redistribution will be involved. When the disk failure of a node recovers, the data of this node will be backfilled. So when these happen, how do we c… Language: en Canonical URL: https://forum.scylladb.com/t/how-scylladb-controls-data-redistribution-or-data-backfill-traffic/721 ## Headings Structure: H1: How scylladb controls data redistribution or data backfill traffic H3: Related topics ## Main Content: H1: How scylladb controls data redistribution or data backfill traffic H3: Related topics When we add a node to a cluster or remove a node from a cluster, data redistribution will be involved. When the disk failure of a node recovers, the data of this node will be backfilled. So when these happen, how do we control the impact of traffic brought by these operations on business traffic? Will scylladb have a business traffic priority mechanism? ScyllaDB already has CPU and IO schedulers. User workloads and background maintenance work, like repair, streaming and compaction use different scheduling groups. These are isolated from each other and have different priorities: the user workload has 1000 shares and the maintenance workloads have only 200 shares. Furthermore, ScyllaDB Enterprise has Workload Prioritization, which allows for isolating different kind of user workloads from each other, assigning them different priorities (shares). --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-2-5/731 Title: [RELEASE] ScyllaDB 5.2.5 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.2.5, a bugfix release of the ScyllaDB 5.2 stable branch. ScyllaDB Open Source 5.2.5, like all past and future 5.x.y releases, is backward compatible and supports rolling… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-2-5/731 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.2.5 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.2.5 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.2.5, a bugfix release of the ScyllaDB 5.2 stable branch. ScyllaDB Open Source 5.2.5, like all past and future 5.x.y releases, is backward compatible and supports rolling upgrades. Users are encouraged to upgrade to 5.2.5. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/release-scylla-monitoring-stack-4-4-3/737 Title: [RELEASE] Scylla Monitoring Stack 4.4.3 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce a patch release of ScyllaDB Monitoring Stack 4.4.3 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB Enterprise and ScyllaDB Open Source, based on Prometh… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-monitoring-stack-4-4-3/737 ## Headings Structure: H1: [RELEASE] Scylla Monitoring Stack 4.4.3 H3: Related topics ## Main Content: H1: [RELEASE] Scylla Monitoring Stack 4.4.3 H3: Related topics The ScyllaDB team is pleased to announce a patch release of ScyllaDB Monitoring Stack 4.4.3 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB Enterprise and ScyllaDB Open Source, based on Prometheus and Grafana. ScyllaDB Monitoring Stack 4.4.3 supports: The patch release address two issues with the Prometheus recording rules and alerts: --- ### Page: https://forum.scylladb.com/t/release-scylladb-rust-driver-0-9-0/738 Title: [RELEASE] ScyllaDB Rust Driver 0.9.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce ScyllaDB Rust Driver 0.9.0, an asynchronous CQL driver for Rust, optimized for Scylla, but also compatible with Apache Cassandra! Some interesting statistics: over 403k downlo… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-rust-driver-0-9-0/738 ## Headings Structure: H1: [RELEASE] ScyllaDB Rust Driver 0.9.0 H2: Notable changes H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Rust Driver 0.9.0 H2: Notable changes H3: Related topics The ScyllaDB team is pleased to announce ScyllaDB Rust Driver 0.9.0, an asynchronous CQL driver for Rust, optimized for Scylla, but also compatible with Apache Cassandra! Some interesting statistics: API cleanups / breaking changes: New features / enhancements: CI / developer tool improvements: Congrats to all contributors and thanks everyone for using our driver! The source code of the driver can be found here: The official crates.io registry entry is here: Thank you for your attention, please do not hesitate to contact us if you have any questions, issues, feature requests, or are simply interested in our driver! Contributors since the last release: --- ### Page: https://forum.scylladb.com/t/segmentation-fault-on-shard/739 Title: Segmentation fault on shard - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, I am facing a challenge and trying to get the solution there I have set up a 3-node Scylla cluster in GCP (machine type: n1-highmem-8(8 vCPU, 52 GB), version of scylla-server: 4.5.1-0.20211024.4c0eac049, OS: Linux 8… Language: en Canonical URL: https://forum.scylladb.com/t/segmentation-fault-on-shard/739 ## Headings Structure: H1: Segmentation fault on shard H3: Related topics ## Main Content: H1: Segmentation fault on shard H3: Related topics Hi, I am facing a challenge and trying to get the solution there I have set up a 3-node Scylla cluster in GCP (machine type: n1-highmem-8(8 vCPU, 52 GB), version of scylla-server: 4.5.1-0.20211024.4c0eac049, OS: Linux 8-gcp #24~20.04.1-Ubuntu SMP Mon Sep 12 06:14:01 UTC 2022 x86_64 x86_64 x86_64 GNU/Linux). We have a keyspace, and here is the information about the keyspace: mykeyspace | True | {‘class’: ‘org.apache.cassandra.locator.SimpleStrategy’, ‘replication_factor’: ‘2’} The Scylla instance sometimes restarts, and we see this message in the logs: Jul 25 18:14:16 db-scylla1 scylla[2190565]: Segmentation fault on shard 4. Jul 25 18:14:16 db-scylla1 scylla[2190565]: Backtrace: Jul 25 18:14:16 db-scylla1 scylla[2190565]: 0x3f75e08 Jul 25 18:14:16 db-scylla1 scylla[2190565]: 0x3fa7ee6 Jul 25 18:14:16 db-scylla1 scylla[2190565]: 0x7f63611c81df Jul 25 18:14:16 db-scylla1 scylla[2190565]: 0x1c6f863 Jul 25 18:14:16 db-scylla1 scylla[2190565]: 0x217a7e2 Jul 25 18:14:16 db-scylla1 scylla[2190565]: 0x3f88c4f Jul 25 18:14:16 db-scylla1 scylla[2190565]: 0x3f89e37 Jul 25 18:14:16 db-scylla1 scylla[2190565]: 0x3fa8488 Jul 25 18:14:16 db-scylla1 scylla[2190565]: 0x3f5438a Jul 25 18:14:16 db-scylla1 scylla[2190565]: /opt/scylladb/libreloc/libpthread.so.0+0x93f8 Jul 25 18:14:16 db-scylla1 scylla[2190565]: /opt/scylladb/libreloc/libc.so.6+0x101902 Jul 25 18:19:34 db-scylla1 scylla[2212677]: Scylla version 4.5.1-0.20211024.4c0eac049 with build-id e0df888020bbfef43aee10264ff14c85581609f7 starting … Jul 25 18:19:34 db-scylla1 scylla[2212677]: command used: "/usr/bin/scylla --log-to-syslog 1 --log-to-stdout 0 --default-log-level info --network-stack posix --reserve-memory 4G --io-properties-file=> Jul 25 18:19:34 db-scylla1 scylla[2212677]: parsed command line options: [log-to-syslog: 1, log-to-stdout: 0, default-log-level: info, network-stack: posix, reserve-memory: 4G, io-properties-file: /e> /etc/scylla.d/io_properties.yaml /etc/scylla.d/cpuset.conf CPUSET="–cpuset 1-7 " ######################################## /etc/scylla.d/memory.conf MEM_CONF=“–lock-memory=1” ######################################## /etc/scylla/scylla.yaml cluster_name: ‘scylla’ num_tokens: 256 commitlog_sync: periodic commitlog_sync_period_in_ms: 10000 commitlog_segment_size_in_mb: 32 seed_provider: # Addresses of hosts that are deemed contact points. # Scylla nodes use this list of hosts to find each other and learn # the topology of the ring. You must change this if you are running # multiple nodes! - class_name: org.apache.cassandra.locator.SimpleSeedProvider parameters: # seeds is actually a comma-delimited list of addresses. # Ex: “,,” - seeds: “1.1.1.1,2.2.2.2,3.3.3.3” listen_address: 1.1.1.1 native_transport_port: 9042 native_shard_aware_transport_port: 19042 read_request_timeout_in_ms: 5000 write_request_timeout_in_ms: 2000 cas_contention_timeout_in_ms: 1000 endpoint_snitch: GossipingPropertyFileSnitch rpc_address: 1.1.1.1 rpc_port: 9160 api_port: 10000 api_address: 127.0.0.1 batch_size_warn_threshold_in_kb: 5 batch_size_fail_threshold_in_kb: 50 partitioner: org.apache.cassandra.dht.Murmur3Partitioner commitlog_total_space_in_mb: -1 murmur3_partitioner_ignore_msb_bits: 12 api_ui_dir: /opt/scylladb/swagger-ui/dist/ api_doc_dir: /opt/scylladb/api/api-doc/ ########################################### this is the result of ulimits -a core file size (blocks, -c) 0 data seg size (kbytes, -d) unlimited scheduling priority (-e) 0 file size (blocks, -f) unlimited pending signals (-i) 208819 max locked memory (kbytes, -l) 65536 max memory size (kbytes, -m) unlimited open files (-n) 1000000 pipe size (512 bytes, -p) 8 POSIX message queues (bytes, -q) 819200 real-time priority (-r) 0 stack size (kbytes, -s) 8192 cpu time (seconds, -t) unlimited max user processes (-u) 208819 virtual memory (kbytes, -v) unlimited file locks (-x) unlimited ########################################### /etc/sysctl.d/11-sysctl.conf net.ipv4.ip_forward=1 net.ipv6.conf.all.forwarding=1 net.ipv4.conf.all.accept_redirects=1 fs.file-max=2097152 net.ipv4.netfilter.ip_conntrack_max=4000000 net.netfilter.nf_conntrack_max=4000000 net.ipv4.netfilter.ip_conntrack_max=4000000 net.ipv4.tcp_window_scaling=0 net.ipv4.tcp_max_tw_buckets=10000 net.ipv4.tcp_max_syn_backlog=2048 net.core.somaxconn=128 net.core.netdev_max_backlog=1000 net.ipv4.tcp_keepalive_time=60 net.ipv4.netfilter.ip_conntrack_tcp_timeout_time_wait=5 net.ipv4.tcp_fin_timeout=10 net.ipv4.tcp_tw_reuse=1 net.ipv4.tcp_keepalive_intvl=15 net.ipv4.tcp_keepalive_probes=5 net.ipv4.tcp_sack=0 net.ipv4.tcp_timestamps=0 net.core.rmem_max=524287 net.core.wmem_max=524287 net.core.rmem_default=524287 net.core.wmem_default=524287 net.core.optmem_max=524287 ########################################### this is the command which we used to setup the scylla scylla_setup --disks /dev/sdb --nic ens4 --io-setup 1 --no-version-check --no-rsyslog-setup In the future, please report such issues on our github so they’d be better tracked: See Issues · scylladb/scylladb · GitHub I decoded the backtrace, and this seems like a known issue: LWT update with empty clustering key range causes a crash · Issue #13129 · scylladb/scylladb · GitHub that has been fixed, but the fix has not been backported to official releases yet. When it will be backported, please upgrade to the latest available release, as those contain many bug fixes over the version you’re currently using. Please note that 4.5.1 aged out and is not supported any more so the fix is bound to backported to a newer release. With that said, the root cause seems to be an invalid LWT query that results in an empty clustering key range condition. The fix will just turn the crash into a visible error, which is much better, but you should still locate to problematic query and fix it - if that’s indeed the case. The issue was due to a corrupt sstable file. Deleting the files of that sstable seems to have resolved the issue. But would it be worth handling … A segfault Core Ball will occur when a program attempts to operate on a memory location in a way that is not allowed (for example, attempts to write a read-only location would result in a segfault). Segfaults can also occur when your program runs out of stack space. Core ball is an addicting online game that you won’t be able to put down! Zigzag your way to the high score! Coreball (also named “Core Ball”) is a classic little arcade game that can be played online. The concept of this game is followed by an idea from a console game called AA Ball in 2015. The goal of Coreball is simple: you only need to get the ball into the core ball without hitting any of the other balls that are attached to it. When you have completed the current level by successfully throwing all of the balls, you will go on to the next one. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-17-2023-07-28/740 Title: Last week in scylla-cluster-tests.git master (issue #17; 2023-07-28) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the bfd83aa1…da9812ff range are covered. There were 37 non-merge commits from 12 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-17-2023-07-28/740 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #17; 2023-07-28) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #17; 2023-07-28) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the bfd83aa1…da9812ff range are covered. There were 37 non-merge commits from 12 authors in that period. Some notable commits: Google Cloud libraries were updated to enable showing labels for addresses. Also a new zone (us-east4) was added to the list of supported zones. Check out docs on how to clear monitoring docker instances created with hydra investigate show-monitor command. All tests by default use monitoring from 4.4 branch. Locators adjustment was required to fix screenshot functionality. Repairing system_auth keyspace after altering replication strategy during node setup was taking much time (especially in multi-dc tests). We speeded it up by running it in parallel on all nodes and adding -pr parameter to repair commands. New performance test: GCE throughput Generic test duration was not working in k8s multi-tenant tests. It’s fixed now. Issue with not timing out ssh command (sshlib2) when there was no output lines was fixed. It enables us to get some more data (like lsof and netstat output) when Scylla freezes during stopping/draining for further issue investigation. We are deprecating i3 tests. Most of i3 instances were replaced with i4i equivalents (i3en is still in use). No load adjustments were done yet. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/getting-started/742 Title: Getting started - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, I am following the first course in Scylla University and trying to do the hands on lab. I created my cluster in the cloud and connected to it in the web site. Following the instructions, the next step is to connect u… Language: en Canonical URL: https://forum.scylladb.com/t/getting-started/742 ## Headings Structure: H1: Getting started H3: Related topics ## Main Content: H1: Getting started H3: Related topics Hi, I am following the first course in Scylla University and trying to do the hands on lab. I created my cluster in the cloud and connected to it in the web site. Following the instructions, the next step is to connect using a command: docker run -it --rm --entrypoint cqlsh scylladb/scylla -u scylla -p *************** X.X.X.X I am assuming the * is the password so I changed it. Now what is the ip adress? In the previous instructions says " If you selected the Dedicated Cluster option, ensure your IP is added to the Allowed IPs list. This is to allow connections from the outside world to your new ScyllaDB Cloud cluster." But I never saw that option. Should I put my IP adress or the cluster adress? If it is the second, how do I know it? Thanks! Seems like you went through the Serverless Free Trial path, which doesn’t require you to add your IP to the Allowed IPs list. The Docker command you provided here is only relevant to Dedicated clusters, not for Serverless clusters. For your free Serverless trial cluster, you can go to My Cluster > select your cluster > navigate to “Connect” tab, and then follow the instructions (with a serverless cluster you will need to download a bundle which contains all of the cluster info such as it’s hostname and certificate) Let me know if you need any more clarifications. --- ### Page: https://forum.scylladb.com/t/what-s-trending-on-the-scylladb-community-forum-summmer-2023/743 Title: What’s Trending on the ScyllaDB Community Forum: Summmer 2023 - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: See the blog post here. This would be a good place to discuss and comment on the post. Language: en Canonical URL: https://forum.scylladb.com/t/what-s-trending-on-the-scylladb-community-forum-summmer-2023/743 ## Headings Structure: H1: What’s Trending on the ScyllaDB Community Forum: Summmer 2023 H3: Related topics ## Main Content: H1: What’s Trending on the ScyllaDB Community Forum: Summmer 2023 H3: Related topics See the blog post here. This would be a good place to discuss and comment on the post. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-190-2023-07-30/744 Title: Last week in scylladb.git master (issue #190; 2023-07-30) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the decbc841b7…1c3d22b717 range are covered. There were 88 non-merge commits from 13 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-190-2023-07-30/744 ## Headings Structure: H1: Last week in scylladb.git master (issue #190; 2023-07-30) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #190; 2023-07-30) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the decbc841b7…1c3d22b717 range are covered. There were 88 non-merge commits from 13 authors in that period. Some notable commits: Tablets are a new, experimental way to distribute data on the cluster. Tablets now have an automatic load balancer that detects nodes and shards that have a deficit or tablets, and migrates tablets to those nodes and shards in order to restore balance. ScyllaDB uses a separate commitlog, called the schema commitlog, for schema changes and topology operations in order to reduce the latency of these operations. The segmented size of the schema commitlog has been raised from 32MB to 128MB in order to avoid problems with large numbers of tables, as the entire schema must fit in a single segment. ScyllaDB uses objects called reader_concurrency_semaphores to limit query concurrency and to isolate different service levels. We now check if the service level changed during a query and avoid erroring out in this case. Fencing the the mechanism in which requests that were sent using an outdated view of the cluster topology are rejected , in order to avoid reading outdated data or resurrecting old data. It now applies to hints, a mechanism used to heal the cluster after a short node downtime. Recently the mechanism to update materialized views after repair was optimized. A latent use-after-free bug was discovered in the optimization, and fixed. A deadlock during shutdown in internode communication was fixed. When updating a materialized view after repair, we chunk the base table data and process each chunk individually. Chunking is based on memory consumption. However, empty partitions were not accounted for, so long runs of empty partitions could create large chunks and run the node out of memory. This is now fixed by accounting for empty partitions. ScyllaDB caches pages from the sstable primary index in order to reduce I/O. In certain cases it reads index pages ahead of the actual need to use them to reduce latency. In rare cases this caused an internal invariant to be violated, crashing the node. This is now fixed. The format of the timestamp data type is now compatible with Cassandra. ScyllaDB computes the version of the schema by hashing the mutations that describe the schema in the schema tables. This can lead to an inconsistency between nodes if tombstones are expired at different times. This is now fixed by ignoring empty partitions, making the tombstone expiration time irrelevant. Streaming and repair will now compact data before streaming it, reducing bandwidth usage if the sstables being streamed happen to contain data and tombstones that cover that data. A bug in the Seastar coroutine code, which could lead to unexpected crashes, has been fixed. The build toolchain has been updated to Fedora 38 with clang 16.0.6. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/crud-examples-using-app/745 Title: CRUD examples using app - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Do you have examples of Create/Read/Update/Delete operations on ScyllaDB using a front-end application? Language: en Canonical URL: https://forum.scylladb.com/t/crud-examples-using-app/745 ## Headings Structure: H1: CRUD examples using app H3: Related topics ## Main Content: H1: CRUD examples using app H3: Related topics Do you have examples of Create/Read/Update/Delete operations on ScyllaDB using a front-end application? Yes, we have a few examples. One is the IoT Care-pet project. It covers application creation, configuration, data modeling, and design in different languages (Go, Java, Rust, Javascript, and more). You can find the github repo for it here. ScyllaDB University has a few hands-on labs which cover CRUD operations, for example: --- ### Page: https://forum.scylladb.com/t/using-scylladb-prepared-dashboards-within-grafana-for-monitoring/746 Title: Using ScyllaDB prepared dashboards within Grafana for monitoring - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Can I utilize ScyllaDB prepared dashboards within Grafana Cloud to monitor different clusters, and if so, how should I go about it? Language: en Canonical URL: https://forum.scylladb.com/t/using-scylladb-prepared-dashboards-within-grafana-for-monitoring/746 ## Headings Structure: H1: Using ScyllaDB prepared dashboards within Grafana for monitoring H3: Related topics ## Main Content: H1: Using ScyllaDB prepared dashboards within Grafana for monitoring H3: Related topics Can I utilize ScyllaDB prepared dashboards within Grafana Cloud to monitor different clusters, and if so, how should I go about it? Yes, you can use ScyllaDB’s monitoring dashboard templates within Grafana Cloud. After manually loading the dashboards (which can be accessed from the ScyllaDB Monitoring github repo), configure Prometheus as the data source from within Grafana Cloud. --- ### Page: https://forum.scylladb.com/t/vector-search-in-scylladb/750 Title: Vector Search in Scylladb - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello all, are there any ideas or approaches to use Vector Searches in Scylladb? This should be supported with Cassandra 4.0 and I think it would be a nice feature to support similarity searches based on text embeddin… Language: en Canonical URL: https://forum.scylladb.com/t/vector-search-in-scylladb/750 ## Headings Structure: H1: Vector Search in Scylladb H3: Related topics ## Main Content: H1: Vector Search in Scylladb H3: Related topics are there any ideas or approaches to use Vector Searches in Scylladb? This should be supported with Cassandra 4.0 and I think it would be a nice feature to support similarity searches based on text embeddings. ScyllaDB doesn’t support vector searches. One possible option to include ScyllaDB in a vector search project is to use a tool like Vald that can do vector search and use ScyllaDB as backup. Hi Attila thanks for your response. I didn’t know about Vald looks very interesting I just know about the typical vector dbs like redis, pinecon, cozodb. Maybe somebody from the Scylladb Core team can give a outlook You might also find this tutorial useful. It explains the basics of Machine Learning (ML) feature stores, and it explains how ScyllaDB can be a critical part of your feature store architecture. Hi @Jeremy_Teichmann, I know it’s been a while since your original post, but I was curious—how did you end up tackling the initial need you mentioned? Would love to hear how things have been working out. Happy to connect for a quick chat if you’re up for it! Since the original post, there’s been an update on vector search in ScyllaDB. Vector search functionality is actively being worked on and is under development in ScyllaDB. The expected CQL syntax includes: Defining a vector column: Creating a vector index: Performing vector similarity queries:: With this feature, ScyllaDB will enable a wide range of use cases, including: We’ll share updates with the community as this feature becomes available. Learn more! Following up on Attila’s message, we’re excited to share that we’re finalizing the first milestone of ScyllaDB Vector Search – Public Beta. It will be available in ScyllaDB Cloud to everyone in just a few weeks. In the meantime, we’re launching an Early Access Program. If you’d like to get early access to Vector Search and help shape its roadmap, you can sign up here: ScyllaDB Vector Search Early Access To learn more about how ScyllaDB powers real-time AI, visit: Scale Real-Time AI with ScyllaDB --- ### Page: https://forum.scylladb.com/t/is-scylladb-can-be-good-alternative-to-rethinkdb/754 Title: Is ScyllaDB can be good alternative to RethinkDB - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: We are using rethinkDB for our operations to read millions of data using the rest API’s. The problem we are facing is performance, as their documentation & support are not so good so our team lead wants to migrate rethin… Language: en Canonical URL: https://forum.scylladb.com/t/is-scylladb-can-be-good-alternative-to-rethinkdb/754 ## Headings Structure: H1: Is ScyllaDB can be good alternative to RethinkDB H3: Related topics ## Main Content: H1: Is ScyllaDB can be good alternative to RethinkDB H3: Related topics We are using rethinkDB for our operations to read millions of data using the rest API’s. The problem we are facing is performance, as their documentation & support are not so good so our team lead wants to migrate rethinkdb to any other. Now I just started exploring Scylladb and want to know if this will be a good replacement to rethinkdb. If your main concern is performance then ScyllaDB is likely to be a good candidate for your project - read some benchmarks here. Keep in mind though that while RethinkDB is a document-based database, ScyllaDB is so called key-key-value store (it’s row store, but uses a partition key and clustering key). So you might need to re-think (no pun intended) your data model if you migrate to ScyllaDB. (we have a really good course on this topic) Furthermore, ScyllaDB is Cassandra compatible so you can find lots of tools and frameworks that work well with ScyllaDB. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-18-2023-08-04/756 Title: Last week in scylla-cluster-tests.git master (issue #18; 2023-08-04) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 6e3d34cd…1fcc9c99 range are covered. There were 18 non-merge commits from 8 authors in tha… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-18-2023-08-04/756 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #18; 2023-08-04) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #18; 2023-08-04) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 6e3d34cd…1fcc9c99 range are covered. There were 18 non-merge commits from 8 authors in that period. Some notable commits: 3 new nemesis were added: Full aggregation scan dispatch to other nodes is now verified using prometheus metrics instead of using tracing. SLA tests were improved by increasing load close to 100% and prolonged duration to allow to trigger more nemesis disrupt_terminate_and_replace_node operation in performance tests mixed workload is run on double-sized cluster, so we should get better results for replace node and not affect any other operations we have currently. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-1-15/757 Title: [RELEASE] ScyllaDB 5.1.15 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.1.15, a bugfix release of the ScyllaDB 5.1 stable branch. ScyllaDB Open Source 5.1.15, like all past and future 5.x.y releases are backward compatible and support rollin… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-1-15/757 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.1.15 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.1.15 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.1.15, a bugfix release of the ScyllaDB 5.1 stable branch. ScyllaDB Open Source 5.1.15, like all past and future 5.x.y releases are backward compatible and support rolling upgrades. Note that the latest stable ScyllaDB Open Source release is 5.2, and you are encouraged to upgrade to it. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-191-2023-08-05/758 Title: Last week in scylladb.git master (issue #191; 2023-08-05) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 1c3d22b717…421a5ad55c range are covered. There were 132 non-merge commits from 15 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-191-2023-08-05/758 ## Headings Structure: H1: Last week in scylladb.git master (issue #191; 2023-08-05) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #191; 2023-08-05) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 1c3d22b717…421a5ad55c range are covered. There were 132 non-merge commits from 15 authors in that period. Some notable commits: Tablets are a new, experimental way of distributing data across nodes and shards. The tablet load balancer is now able to concurrently migrate tablets to new nodes, and is able to make new decisions while tablets are being streamed. In CQL, a few functions for dealing with counter types were added. Alternator, ScyllaDB’s implementation of the DynamoDB API, now limits embedded expression length and nesting. A source of high latency in multi-partition scans was eliminated. Change Data Capture exposes multiple streams to reflect the cluster topology. It is now more careful to avoid closing and creating new streams unnecessarily. Fencing is a way to prevent a coordinator from interacting with replicas when it has an outdated view of cluster topology. This now applies to counter updates too. ScyllaDB verifies ownership and permissions for its own files. It now avoids doing this for snapshots, as they might be concurrently being deleted by an administrator or scylla-manager. Cluster features are ScyllaDB’s way of making rolling upgrades seamless - a feature isn’t enabled until all nodes support it. We now propagate cluster features via Raft rather than gossip for improved reliability. Most ScyllaDB metrics are per-shard, per-node, but not for a specific table. We now export some per-table metrics. These are exported once per node, not per shard. There is a new option to specify the number of token ranges to repair in parallel. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-1-9/760 Title: [RELEASE] ScyllaDB Enterprise 2022.1.9 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2022.1.9, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.1. Note that there is a newer Feature Enterprise release: ScyllaDB Ente… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-1-9/760 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2022.1.9 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2022.1.9 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2022.1.9, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.1. Note that there is a newer Feature Enterprise release: ScyllaDB Enterprise 2022.2. While we will continue to support 2022.1 LTS, you can get additional features with 2022.2. Below is a list of performance and stability improvements and bug fixes, each with an open-source reference if available: --- ### Page: https://forum.scylladb.com/t/maximizing-read-throughput-of-table-scan-using-shard-awareness/764 Title: Maximizing read throughput of table scan using shard-awareness - Knowledge Base - ScyllaDB Community NoSQL Forum Meta Description: I would love any feedback on this blog that I just posted. Is the general idea correct? Am I missing anything important? Am I actually misleading anyone? Thanks https://buransky.com/programming/maximizing-read-throughpu… Language: en Canonical URL: https://forum.scylladb.com/t/maximizing-read-throughput-of-table-scan-using-shard-awareness/764 ## Headings Structure: H1: Maximizing read throughput of table scan using shard-awareness H3: Related topics ## Main Content: H1: Maximizing read throughput of table scan using shard-awareness H3: Related topics I would love any feedback on this blog that I just posted. Is the general idea correct? Am I missing anything important? Am I actually misleading anyone? Thanks https://buransky.com/programming/maximizing-read-throughput-of-scylladb-table-scan-using-shard-awareness/ Ahoj Rado vyzera to dobre I would love you to explain “magic numbers” … (you start, not sure if it’s fully clear) We generally say here to have max 10-30 foreground requests per cpu, total concurrency is just a function of number of cpus per node and number of nodes multiplied to how big can be the background queue of queries to still fit your SLAs Otherwise great explanation how spark works (or should work, since it doesn’t go down to shards fully, I have a repo for shard aware spark connector, but didn’t fully test if it will get close to logic you drawn) Numberly also came close with How Numberly Replaced Kafka with a Rust-Based ScyllaDB Shard-Aware Application - ScyllaDB (and I saw few go and python ideal full scan implementation too, but without focus on per shard throttling). Shards also split and optimize access to disk using schedulers (we tried to explain them in blogs). Also routing token if you have metadata should be set dynamically by driver, why do you explicitely set, am I missing something? ( DataStax Java Driver - Load balancing ) Also spark uses shuffle to randomly shuffle tasks with ranges so they don’t get blocked on same shard, but I see you are trying to improve it - and btw. this is crucial - finding balance between too many small tasks and bigger tasks they try not to overlap is key, even for spark. This partiitoning to tasks/token ranges is something which can be optimized everywhere, spark gives nice task distribution engine, but this partiitoning is something that can be tuned in similar way as you try, so keep it up! Looking forward to next version! (or taking these ideas to spark?) --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-2-6/766 Title: [RELEASE] ScyllaDB 5.2.6 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.2.6, a bugfix release of the ScyllaDB 5.2 stable branch. ScyllaDB Open Source 5.2.6, like all past and future 5.x.y releases, is backward compatible and supports rolling… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-2-6/766 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.2.6 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.2.6 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.2.6, a bugfix release of the ScyllaDB 5.2 stable branch. ScyllaDB Open Source 5.2.6, like all past and future 5.x.y releases, is backward compatible and supports rolling upgrades. Users are encouraged to upgrade to 5.2.6. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/scylla-dtest-open-souce/768 Title: Scylla dtest open souce - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Since reading the article in Testing part 4: Distributed tests - ScyllaDB, I can not wait a minute to run some tests about scylla dtest. Would u mind sending me the open scource of scylla dtest, if convinient ? Language: en Canonical URL: https://forum.scylladb.com/t/scylla-dtest-open-souce/768 ## Headings Structure: H1: Scylla dtest open souce H3: Related topics ## Main Content: H1: Scylla dtest open souce H3: Related topics Since reading the article in Testing part 4: Distributed tests - ScyllaDB, I can not wait a minute to run some tests about scylla dtest. Would u mind sending me the open scource of scylla dtest, if convinient ? hi @qiuqiu , unfortunetly ScyllaDB’s dtests are not open source, but as we are fully compatible with Cassandra, hence you can access their dtests, and run against Scylla nodes. for that, you may need to use our CCM. Thanks a lot for your reply. What’s a pity is that ScyllaDB’s dtests are not open source. --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2021-1-22/769 Title: [RELEASE] ScyllaDB Enterprise 2021.1.22 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2021.1.22, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2021.1. ScyllaDB Enterprise 2022.1 LTS is the latest Long-Term Support rele… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2021-1-22/769 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2021.1.22 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2021.1.22 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2021.1.22, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2021.1. ScyllaDB Enterprise 2022.1 LTS is the latest Long-Term Support release, and 2022.2 is the newest feature release. You are encouraged to upgrade in coordination with the ScyllaDB support team. The following issue are fixed in this release (with an open-source reference): --- ### Page: https://forum.scylladb.com/t/scylladb-labs-free-online-training-event-19-september/771 Title: ScyllaDB Labs - free, online training event, 19 September - University and Training - ScyllaDB Community NoSQL Forum Meta Description: On the 19th of September, we’re hosting a live, online, free training event: ScyllaDB Labs Building High-Performance Apps. Save your spot here, This will be an interactive workshop. You’ll have a chance to learn by ru… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-labs-free-online-training-event-19-september/771 ## Headings Structure: H1: ScyllaDB Labs - free, online training event, 19 September H3: Related topics ## Main Content: H1: ScyllaDB Labs - free, online training event, 19 September H3: Related topics On the 19th of September, we’re hosting a live, online, free training event: ScyllaDB Labs Building High-Performance Apps. Save your spot here, This will be an interactive workshop. You’ll have a chance to learn by running some hands-on labs. Some of the topics we’ll cover include: There will also be some time for Q&A, hope to see you there! --- ### Page: https://forum.scylladb.com/t/mitigating-aws-outage/772 Title: Mitigating AWS outage - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: We are considering moving to ScyllaDB Cloud. If there is an outage in an AWS region, how do you mitigate it? Language: en Canonical URL: https://forum.scylladb.com/t/mitigating-aws-outage/772 ## Headings Structure: H1: Mitigating AWS outage H3: Related topics ## Main Content: H1: Mitigating AWS outage H3: Related topics We are considering moving to ScyllaDB Cloud. If there is an outage in an AWS region, how do you mitigate it? In a worst-case scenario of all nodes going down, a backup will be used to restore the data. In most cases, ScyllaDB’s high availability and data replication ensures that even if some nodes are down, the data remains available. Using multiple data centers, multiple regions, and multiple racks within a region increases the availability and protects against failures. You can read more about a real-world incident, a fire at a Strasbourg data center, and how kiwi.com remained available (as opposed to millions of other websites). --- ### Page: https://forum.scylladb.com/t/fips-140-2-compliance-and-openssl/773 Title: FIPS 140-2 compliance and OpenSSL - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: We are considering ScyllaDB Enterprise and want to understand your FIPS 140-2 compliance. For both system and table data, can you use different algorithms that are supported by OpenSSL in a file block encryption scheme?” … Language: en Canonical URL: https://forum.scylladb.com/t/fips-140-2-compliance-and-openssl/773 ## Headings Structure: H1: FIPS 140-2 compliance and OpenSSL H3: Related topics ## Main Content: H1: FIPS 140-2 compliance and OpenSSL H3: Related topics We are considering ScyllaDB Enterprise and want to understand your FIPS 140-2 compliance. For both system and table data, can you use different algorithms that are supported by OpenSSL in a file block encryption scheme?” Yes, OpenSSL for both encryptions. Once enabled, all communication between the client and the node is transmitted over TLS/SSL. The libraries used by Scylla for OpenSSL are FIPS 140-2 certified. See more in the documentation about client-to-node encryption and node-to-node encryption. ScyllaDB security features, including encryption, are covered in this ScyllaDB University lesson. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-19-2023-08-11/774 Title: Last week in scylla-cluster-tests.git master (issue #19; 2023-08-11) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the de8a61f7…e3649285 range are covered. There were 36 non-merge commits from 11 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-19-2023-08-11/774 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #19; 2023-08-11) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #19; 2023-08-11) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the de8a61f7…e3649285 range are covered. There were 36 non-merge commits from 11 authors in that period. Some notable commits: Added an ubuntu 20 with FIPS enabled scenario and verification if it is properly enabled Also added a 4h 100gb longevity with fips enabled which enables every kind of encryption scylla supports, while FIPS is enabled. FIPS are a set of standards for encryption algorithms (among others). Now we can filter by keyspace name in BaseCluster.get_non_system_ks_cf_list method to save time, resources and decrease code lines in cases where there are a lot of entities in the schema. New hydra image got published with an updated cryptography package (due found CVE in the old one). Update your packages in local environments. Note also autopep8 was updated and it works slightly differently now. Because of the new sstable uuid identifier feature the old sstabledump can’t currently work with those sstables. We Replaced sstabledump with scylla sstable dump-data Switched c-s commands to use LOCAL_QUORUM instead of QUORUM for multidc job. We added new nemesis that interrupts bootstrap of a new node at different stages (according to Raft Topology changed Failure instructions). Depending on which stage bootstrap was failed, node could be already added to group0, to group0 and token ring, or any other stage. This could affect on node availability in cluster. To finish test quicker after stress load ends we can now raise_exception_in_thread as a way to force kill a nemesis threads so they get stopped forcibly. This prevents also from false positive events during SLA tests which depend on the load. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-192-2023-08-13/776 Title: Last week in scylladb.git master (issue #192; 2023-08-13) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 421a5ad55c…2e2271f639 range are covered. There were 62 non-merge commits from 8 authors in that period… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-192-2023-08-13/776 ## Headings Structure: H1: Last week in scylladb.git master (issue #192; 2023-08-13) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #192; 2023-08-13) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 421a5ad55c…2e2271f639 range are covered. There were 62 non-merge commits from 8 authors in that period. Some notable commits: The experimental tablet load balancer now exports statistics on its activity. A crash in the experimental tablets feature when a table is dropped was fixed. When deleting multiple sstables at once (such as at the end of a compaction), we now avoid flushing the directory unnecessarily as we can rely on the deletion log file instead. When a node is decommissioned or removed, and Raft topology management is active, the node stops being a voter earlier in the process in order to improve availability. A SELECT statement that has the DISTINCT keyword and also GROUP BY on clustering keys is now rejected. DISTINCT implies only selecting the partition key and static rows, so grouping on the clustering keys is nonsensical. We now update the list of sstables requiring cleanup after compaction completion. This avoids a race between decommission and compaction, involving offstrategy compaction that could cause such compacted sstables not to be cleaned. The system.group0_history table now has descriptions for events. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-2-12/777 Title: [RELEASE] ScyllaDB Enterprise 2022.2.12 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2022.2.12, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.2. Related Links Get ScyllaDB Enterprise 2022.2.12 (customers only, … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-2-12/777 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2022.2.12 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2022.2.12 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2022.2.12, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.2. The following issue are fixed in this release (with an open-source reference, if available): CQL Conflict resolution: compare_cells_for_merge may wrongly prefer an expired cell over a live and expiring one #14182. Before this fix, when two cells have the same write timestamp and both are alive or expiring, we compare their value first, before checking if either of them is expiring and if both are expiring, comparing their expiration time and TTL value to determine which of them will expire later or was written later. This was based on an early version of Cassandra. However, the Cassandra implementation rightfully changed in CASSANDRA-14592, where the cell expiration is considered before the cell value. See update-ordering in Scylla docs for more. API: Allow tombstone GC in compaction to be disabled on user request #14077 The fix adds new APIs /column_family/tombstone_gc and /storage_service/tombstone_gc, that will allow for disabling tombstone garbage collection (GC) in compaction. The table name must be in keyspace:name format Setup: scylla-fstrim.timer is enabled but not started #14249 Setup: The installer now wipes filesystem signatures from the individual disks making up a RAID array, preventing problems with reuse of disks. #13737 DynamoDB API (Alternator) stability: assertion in output_stream when exception occurs during response streaming #14453. DynamoDB API (Alternator) stability: Yield while building large results in Alternator - rjson::print, executor::batch_get_item #13689 Stability: mutation_reader_merger can overflow stack when merging many empty readers. This may happen when running a second repair right after the other. #14415 Stability: compaction: excessive reallocation during input list formatting #14071. Issue is more likely with off strategy compaction. Stability: deadlock caused by view update _registration_sem and streaming reader _streaming_concurrency_sem #14676 Stability: a failure when reading metrics, caused by a rare race condition when another node is down. (seastar::metrics::double_registration (registering metrics twice for metrics: storage_proxy_coordinator_background_replica_writes_failed_remote_node)) #11017 Stability: messaging: when upgrading OSS nodes to Enterprise, service-levels are matched to the default scheduling group #13841, #12552 Stability: offstrategy compaction races with view building on staging sstables #11882. Offstrategy compaction is triggered after repair and the latter places sstables in the staging subdirectory if view building is required. Stability: partitioned_sstable_set::insert might stall when called by table::make_reader_v2_excluding_sstables. The root cause is View building from staging creates a reader from scratch for every partition, in order to calculate the diff between new staging data and data in base sstable set, and then pushes the result into the view replicas. #14244 Stability: View building crashes on large partitions with range tombstones. #14503 --- ### Page: https://forum.scylladb.com/t/is-it-possible-to-get-nodes-status-with-cqlsh/781 Title: Is it possible to get nodes status with cqlsh? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I can get list of the cluster nodes from system.peers. How i can find out status of the node ( running, down ) ? Thanks Language: en Canonical URL: https://forum.scylladb.com/t/is-it-possible-to-get-nodes-status-with-cqlsh/781 ## Headings Structure: H1: Is it possible to get nodes status with cqlsh? H3: Related topics ## Main Content: H1: Is it possible to get nodes status with cqlsh? H3: Related topics I can get list of the cluster nodes from system.peers. How i can find out status of the node ( running, down ) ? Thanks You can use the system.cluster_status virtual table. Example: select * from system.cluster_status; Thanks, but looks like i don’t have such table: Any idea how to create it? Thanks What scylla version are you on? This table was added in 4.5, it is in scylla for a long time. so looks like we have old version any other way to get cluster status for this version? Thanks You can use nodetool status. As a side-note, please consider upgrading to a supported version (5.1+). --- ### Page: https://forum.scylladb.com/t/local-read-for-token-aware-requests/783 Title: Local read for token aware requests - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Need some help in troubleshooting token-aware requests. I am using 5.2.2 ScyllaDB with 3.x java Scylla driver. The requests are reaching the right node, but the node does not read the local replica and instead requests a… Language: en Canonical URL: https://forum.scylladb.com/t/local-read-for-token-aware-requests/783 ## Headings Structure: H1: Local read for token aware requests H3: Related topics ## Main Content: H1: Local read for token aware requests H3: Related topics Need some help in troubleshooting token-aware requests. I am using 5.2.2 ScyllaDB with 3.x java Scylla driver. The requests are reaching the right node, but the node does not read the local replica and instead requests another host. I have confirmed this with the tracing enabled, the target selected is not the local node. Is there anything needed to avoid an additional hop? Is this a paged read? Paged reads will stick to the same set of replicas for the duration of the query, to be able to re-use saved readers on said readers. Doing so has more benefits than avoiding the hop. Thanks for the reply. This is not a paged read. It is a lookup for a specific partition and reads all rows in the partition. I checked the code and another thing that can force a non-local replica is cache hit-rate based load balancing. This tries to send more requests to nodes that already have warmed-up caches, slowly ramping up newly started nodes. Did you restart any of the nodes recently? Can you still observe the hop after some time has passed? The nodes have been up and running for a while. The additional hop is observed consistently. To rule out cache hit-rate-based load balancing, is there a configuration for disabling that feature? What does your query look like? Yes, it is called cache_hit_rate_read_balancing. Thanks will give it a try by disabling cache_hit_rate_read_balancing. For the query, I am retrieving the entire partition select * from tableX where partition_key = 'x' --- ### Page: https://forum.scylladb.com/t/release-scylla-operator-1-10-0-rc-0/787 Title: [RELEASE] Scylla Operator 1.10.0-rc.0 - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: The Scylla team is pleased to announce the release of scylla-operator v1.10.0-rc.0 :rocket: We’ll welcome your feedback on the release candidate. Release notes are available on https://github.com/scylladb/scylla-opera… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-operator-1-10-0-rc-0/787 ## Headings Structure: H1: [RELEASE] Scylla Operator 1.10.0-rc.0 H3: Related topics ## Main Content: H1: [RELEASE] Scylla Operator 1.10.0-rc.0 H3: Related topics The Scylla team is pleased to announce the release of scylla-operator v1.10.0-rc.0 We’ll welcome your feedback on the release candidate. Release notes are available on https://github.com/scylladb/scylla-operator/releases/tag/v1.10.0-rc.0 If you haven’t heard about the operator yet, here are some links to get you started: https://github.com/scylladb/scylla-operator/tree/master#scylla-operator https://operator.docs.scylladb.com/ --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-20-2023-08-18/790 Title: Last week in scylla-cluster-tests.git master (issue #20; 2023-08-18) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 02ddb5b7…1deb381c range are covered. There were 11 non-merge commits from 7 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-20-2023-08-18/790 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #20; 2023-08-18) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #20; 2023-08-18) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 02ddb5b7…1deb381c range are covered. There were 11 non-merge commits from 7 authors in that period. Some notable commits: Since we have lots of cases when stopping the stress command by the test is generating extra critical or error events, we now convert stress killed to warning. OEL doesn’t support i4i instances, thus we revert artifact tests to use i3 for them. Since manager 3.1 is now the latest released manager version, the manager upgrade test now starts from it. We removed repair after restore in manager test. Because in Scylla Manager 3.2, the manager executes a repair on its own, so any additional repair is unnecessary. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-193-2023-08-20/791 Title: Last week in scylladb.git master (issue #193; 2023-08-20) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 2e2271f639…7275b8967c range are covered. There were 54 non-merge commits from 9 authors in that period… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-193-2023-08-20/791 ## Headings Structure: H1: Last week in scylladb.git master (issue #193; 2023-08-20) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #193; 2023-08-20) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 2e2271f639…7275b8967c range are covered. There were 54 non-merge commits from 9 authors in that period. Some notable commits: Data definition language (DDL) statements are used to modify the schema. They are covered by a Raft transaction to ensure atomicity. The scope of the transaction has been extended to cover access checking to prevent check/use races (this change was already committed in the past but reverted due to performance regressions). Off-strategy compaction is run after repair or bootstrap on newly received sstables to reduce their count. This compaction now includes a cleanup, to avoid non-owned token ranges from sneaking into the main sstable set via this offstrategy compaction. The schema commitlog size was accidentally set to 10TB, it’s now set to a reasonable size. Since compaction tasks started to be managed by the task manager, their lifetime could be extended even after the compaction is complete. This caused the compaction input sstables to be kept on disk even after they should have been removed. They are now removed as soon as compaction is done. It is now possible to disable configuration changes via the system.config virtual table using a configuration parameter. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/scylladb-enterprise-release-2023-1-0-part-1/792 Title: ScyllaDB Enterprise Release 2023.1.0 - part 1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of ScyllaDB Enterprise 2023.1.0 LTS, a production-ready ScyllaDB Enterprise Long Term Support Major Release. With more than 5,000 commits, we’re excited to move forwar… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-enterprise-release-2023-1-0-part-1/792 ## Headings Structure: H1: ScyllaDB Enterprise Release 2023.1.0 - part 1 H2: Moving to Raft based cluster management H2: New Features H3: Strongly Consistent Schema Management H3: Distributed Aggregations H3: Large Collection Detection H3: Automating away the gc_grace_seconds parameter H3: Materialized Views: Synchronous Mode H3: Empty replica pages H3: Secondary index on collection columns H3: Azure Support H3: Alternator TTL (introduced in 2022.2) H3: Limit partition access rate (introduced in 2022.2) H3: Load and stream (introduced in 2022.2) H3: Materialized Views: Prune (introduced in 2022.2) H3: Performance: Eliminate exceptions from the read and write path (introduced in 2022.2) H3: Related topics ## Main Content: H1: ScyllaDB Enterprise Release 2023.1.0 - part 1 H2: Moving to Raft based cluster management H2: New Features H3: Strongly Consistent Schema Management H3: Distributed Aggregations H3: Large Collection Detection H3: Automating away the gc_grace_seconds parameter H3: Materialized Views: Synchronous Mode H3: Empty replica pages H3: Secondary index on collection columns H3: Azure Support H3: Alternator TTL (introduced in 2022.2) H3: Limit partition access rate (introduced in 2022.2) H3: Load and stream (introduced in 2022.2) H3: Materialized Views: Prune (introduced in 2022.2) H3: Performance: Eliminate exceptions from the read and write path (introduced in 2022.2) H3: Related topics The ScyllaDB team is pleased to announce the release of ScyllaDB Enterprise 2023.1.0 LTS, a production-ready ScyllaDB Enterprise Long Term Support Major Release. With more than 5,000 commits, we’re excited to move forward with ScyllaDB Enterprise 2023.1. With 2023.1 LTS out, ScyllaDB enterprise 2021.1 support will be ended. More information on ScyllaDB Long Term Support (LTS) policy is available here. The ScyllaDB Enterprise 2023.1 release is based on ScyllaDB Open Source 5.2, introducing a Raft-based Strongly Consistent Schema Management, Alternator TTL, introduces Partition level rate limit, distributed select count, and many more improvements and bug fixes. ScyllaDB Enterprise customers are encouraged to upgrade to ScyllaDB Enterprise 2023.1, and are welcome to contact our Support Team with questions. Consistent Schema Management is the first Raft based feature in ScyllaDB, and 2023.1 is the first Enterprise release to enable Raft by default for new deployments. #12572 Starting from ScyllaDB Enterprise 2023.1, all new databases will be created with Raft enabled by default. Upgrading from 2022.x will only use Raft if you explicitly enable it (see upgrade to 2023.1 docs). As soon as all nodes in the cluster opt-in to using Raft, the cluster will automatically migrate those subsystems to using Raft, and you should validate it is the case. Once Raft is enabled, updating schema requires a quorum, from all nodes in the cluster, to be executed. For example, in the following use cases, the cluster does not have a quorum and will not allow updating the schema: This is different from the behavior of a ScyllaDB cluster with Raft disabled. Nodes might be unavailable due to network issues, node issues, or other reasons. To reduce the chance of quorum loss, it is recommended to have 3 or more nodes per DC, and 3 or more DCs, for a multi-DCs cluster. To recover from a quorum loss, reviving the failed nodes or fixing the network partitioning is best. If this is not feasible, see Raft manual recovery procedure. More on handling failures in Raft here. Schema management operations are DDL operations that modify the schema, like CREATE, ALTER, or DROP for KEYSPACE, TABLE, INDEX, UDT, MV, etc. Unstable schema management has been a problem in past ScyllaDB releases. The root cause is the unsafe propagation of schema updates over gossip, as concurrent schema updates can lead to schema collisions. Once Raft is enabled, all schema management operations are serialized by the Raft consensus algorithm. Additional Raft related updates, with reference to open source issue when available: ScyllaDB will now automatically run aggregations statements, like SELECT COUNT(*), on all nodes and all shards in parallel, which brings a considerable speedup, even 20X in larger clusters. Distributed Aggregations supports all types of aggregations. This feature is limited to queries that do not use GROUP BY or filtering. The implementation includes a new level of coordination. A Super-Coordinator node splits aggregation queries into sub-queries, distributes them across some group of coordinators, and merges results. Like a regular coordinator, the Super-Coordinator is a per operation function. A 3 node cluster setup on powerful desktops (3x32 vCPU) Filled the cluster with ~2 * 10^8 rows using scylla-bench and run: time cqlsh --request-timeout=3600 -e “select count(*) from scylla_bench.test using timeout 1h;” Before Distributed Select: 68s After Distributed Select: 2s You can disable this feature by setting enable_parallelized_aggregation config parameter to false. ScyllaDB records large partitions, large rows, and large cells in system tables so that the primary key can be used to deal with them. It additionally records collections with large numbers of elements, since these can cause degraded performance. The warning threshold is configurable: compaction_collection_elements_count_warning_threshold - how many elements are considered a “large” collection (default is 10,000 elements). The information about large collections is stored in the large_cells table, with a new collection_elements column that contains the number of elements of the large collection. Large_cells table retention is 30 days. #11449 Example of a large collection below: There is now optional automatic management of tombstone garbage collection, replacing gc_grace_seconds. This drops tombstones more frequently if repairs are made on time, and prevents data resurrection if repairs are delayed beyond gc_grace_seconds. Tombstones older than the most recent repair will be eligible for purging, and newer ones will be kept. The feature is disabled by default and needs to be enabled via ALTER TABLE. cqlsh> ALTER TABLE ks.cf WITH tombstone_gc = {'mode':'repair'}; There is now a synchronous mode for materialized views. In ordinary, asynchronous materialized views, the operation returns before the view is updated. In synchronous materialized view, the operation does not return until the view is updated - each base replica waits for a view replica. This enhances consistency but reduces availability as, in some situations, all nodes might be required to be functional. Synchronous Mode reference in Scylla Docs Before this release, the paging code requires that pages have at least one row before filtering. This can cause an unbounded amount of work if there is a long sequence of tombstones in a partition or token range, leading to timeouts. ScyllaDB will now send empty pages to the client, allowing progress to be made before a timeout. This prevents analytics workloads from failing when processing long sequences of tombstones. #7689, #3914, #7933 Secondary indexes can now index collection columns. Individual keys and values within maps, sets, and lists can be indexed. Fixes #2962, #8745, #10707 Like in DynamoDB, Alternator items that are set to expire at a specific time will not disappear precisely at that time but only after some delay. DynamoDB guarantees that the expiration delay will be less than 48 hours (though for small tables, the delay is often much shorter). In Alternator, the expiration delay is configurable - it defaults to 24 hours but can be set with the --alternator-ttl-period-in-seconds configuration option. It is now possible to limit read rates and writes rates into a partition with a new WITH per_partition_rate_limit clause for the CREATE TABLE and ALTER TABLE statements. This is useful to prevent hot-partition problems when high rate reads or writes are bogus (for example, arriving from spam bots). #4703 Limits are configured separately for reads and writes. Some examples: Limit reads only, no limit for writes: This feature extends nodetool refresh to allow loading arbitrary sstables that do not belong to a particular node into the cluster. It loads the sstables from disk, calculates the data’s owning nodes, and automatically streams the data to the owning nodes. In particular this is useful when restoring a cluster from backup. For example, say the old cluster has 6 nodes and the new cluster has 3 nodes. One can copy the sstables from the old cluster to the new nodes and trigger the load and stream process. This can make restores and migrations much easier: Load_and_stream option also updates the relevant Materialized Views #9205 Load and stream is used as part of the new ScyllaDB Manager 3.1 Restore functionality. curl -X POST "http://{ip}:10000/storage_service/sstables/{keyspace}?cf={table}&load_and_stream=true nodetool refresh --load-and-streaml Docs: nodetool refresh A new CQL extension PRUNE MATERIALIZED VIEW statement can now be used to remove inconsistent rows from materialized views. A special statement is dedicated for pruning ghost rows from materialized views. A ghost row is an inconsistency issue which manifests itself by having rows in a materialized view which do not correspond to any base table rows. Such inconsistencies should be prevented altogether and ScyllaDB strives to avoid them, but if they happen, this statement can be used to restore a materialized view to a fully consistent state without rebuilding it from scratch. PRUNE MATERIALIZED VIEW my_view; PRUNE MATERIALIZED VIEW my_view WHERE token(v) > 7 AND token(v) < 1535250; PRUNE MATERIALIZED VIEW my_view WHERE v = 19; When a coordinator times out, it generates an exception which is then caught in a higher layer and converted to a protocol message. Since exceptions are slow, this can make a node that experiences timeouts become even slower. To prevent that, the coordinator write path and read path has been converted not to use exceptions for timeout cases, treating them as another kind of result value instead. Further work on the read path and on the replica reduces the cost of timeouts, so that goodput is preserved while a node is overloaded. --- ### Page: https://forum.scylladb.com/t/scylladb-enterprise-release-2023-1-0-part-2/793 Title: ScyllaDB Enterprise Release 2023.1.0 - part 2 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: See Part 1 More Improvements CQL API CQL Conflict resolution: compare_cells_for_merge may wrongly prefer an expired cell over a live and expiring one #14182. Before this fix, when two cells have the same write timestam… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-enterprise-release-2023-1-0-part-2/793 ## Headings Structure: H1: ScyllaDB Enterprise Release 2023.1.0 - part 2 H2: More Improvements H3: CQL API H3: Amazon DynamoDB Compatible API (Alternator) H3: Correctness H3: Performance and stability H3: Operations H3: Deployment and install H3: Security H3: Related topics ## Main Content: H1: ScyllaDB Enterprise Release 2023.1.0 - part 2 H2: More Improvements H3: CQL API H3: Amazon DynamoDB Compatible API (Alternator) H3: Correctness H3: Performance and stability H3: Operations H3: Deployment and install H3: Security H3: Related topics The fix adds new APIs /column_family/tombstone_gc and /storage_service/tombstone_gc, that will allow for disabling tombstone garbage collection (GC) in compaction. Get status: curl -s -X GET “http://127.0.0.1:10000/column_family/tombstone_gc/ks:cf” Enable GC curl -s -X POST “http://127.0.0.1:10000/column_family/tombstone_gc/ks:cf” Disable GC curl -s -X DELETE “http://127.0.0.1:10000/column_family/tombstone_gc/ks:cf” ScyllaDB now supports server-side DESCRIBE. This is required for the latest cqlsh, and reduces the need to update cqlsh as server features are added. However, the version number bump needed to inform cqlsh about this change was reverted as other changes related to the version number are not ready. Secondary indexes on static columns are now supported.#2963 The CQL binary protocol versions 1 and 2 are no longer supported. Version 3 and above have been supported for 9 years, so it’s unlikely to be in real use. You can check for version 1 and 2 in the system.clients virtual table. #10607 There is now documentation about how NULL is treated in ScyllaDB. #12494 USING TIMESTAMP allows setting the mutation timestamp on a CQL statement level. It has a sanity check that prevents setting timestamps in the future, as these can be hard to delete, but sometimes one wishes to do so anyway. There is now a configuration option that allows disabling the feature. #12527 TRUNCATE statements are usually much slower than other statements. TRUNCATE statements now support the WITH TIMEOUT clause to help deal with that. #11408 The CQL CONTAINS and CONTAINS KEY operators are used to check if a collection contains an element. CONTAINS and CONTAINS KEY have been changed to return false if it is asked whether a collection contains NULL, since no collection can contain a NULL. This is a behavior change, but is not expected to affect users since the query is not useful. #10359 In the expression element IN list, we now allow “list” to be NULL (and the expression evaluates to false). Previously, this failed with an error. A regression where CQL ignored some WHERE clause components when a multi-column restriction ((col_a, col_b) < (0, 0)) was present was fixed. #6200 #12014 In SELECT JSON statements, column names can be given aliases (just as with traditional SELECT). However, ScyllaDB ignored those aliases. It will now honor them. #8078 Evaluation of Boolean binary operators (e.g. “=”) has been refactored to use the same expression evaluation code as other expressions, paving the way for relaxation of the CQL grammar to be more similar to SQL. The experimental User-defined aggregates (UDAs) have a state type that can be initialized using the INITCOND clause. The initializer can now be a collection type. CQL: scylla: types: is_tuple(): doesn’t handle reverse types. For example, a schema with reversed clustering key component; this component will be incorrectly represented in the schema CQL dump: the UDT will lose the frozen attribute. When attempting to recreate this schema based on the dump, it will fail as the only frozen UDTs are allowed in primary key components. The LIKE operator on descending order clustering keys now works. #10183 ScyllaDB would incorrectly use an index with some IN queries, leading to incorrect results. This is now fixed. In CREATE AGGREGATE statements, the INITCOND and FINALFUNC clauses are now optional (defaulting to NULL and the identity function respectively). CREATE KEYSPACE now has a WITH STORAGE clause, allowing to customize where data is stored. For now, this is only a placeholder for future extensions. When talking to drivers using the older v3 protocol, ScyllaDB did not serialize timeout exceptions correctly, resulting in the driver complaining about protocol violations. This is now fixed. #5610 The CQL grammar was relaxed to allow bind markers in collection literals, e.g. UPDATE tab SET my_set = { ?, ‘foobar’, :variable }. ScyllaDB now validates collections for NULLs more carefully. #10580 After this change, the following query INSERT INTO ks.t (list_column) VALUES (?); And the driver sending a list with null inside as the bound value, something like [1, 2, null, 4] Would result in an invalid_request_exception instead of an ugly marshaling error. Alternator, ScyllaDB’s implementation of the DynamoDB API, now provides the following improvements: Support for the new error code has been merged to: Scylla Rust Driver support since version 0.6.0 Gocq - merged to master but it’s not a part of any release. Off-strategy compaction is used when sstables from an external source (such as repair) needs to be reshaped before being handed off to the table’s compaction strategy. If off-strategy compaction is stopped, ScyllaDB used to just leave those sstables in their unreshaped form without compacting them again. It will now hand off such sstables directly to the table’s compaction strategy. #11543. In compaction strategies such as LeveledCompactionStrategy and ScyllaDB Enterprise’s IncrementalCompactionStrategy, sstables are sealed when they reach a certain size (160MB and 1GB respectively). They are not split in the middle of a partition, because we were not able to recover the ordering of such split sstables. This is now possible (though not yet integrated into the compaction strategies). Performance: Long-term index caching in the global cache, as introduced in 4.6, hurts the performance for workloads where accesses to the index are sparse. To mitigate this, a new configuration parameter cache_index_pages (default true) is introduced to control index caching. Setting the flag to false causes all index reads to behave like they would in BYPASS CACHE queries. Consider using false if you notice performance problems due to lowered cache hit ratio in 4.6 or 5.0. The config API can update the parameter live (without restart). #11202 Scylla will now reject a too-low bloom_fulter_fp_chance when creating (or altering) the table, rather than crash while flushing memtables. #11524. ScyllaDB represents reads using mutation fragment streams. Several minor violations of fragment stream integrity were fixed. These could result in incorrect reads during range scans. The log-structured allocator is used to manage cache and memtable memory. When memory runs out, the allocator tries to reclaim memory by evicting cache items and by defragmenting memory. If this takes too long, the allocator logs a stall report. Due to a bug, if the report threshold was set too low then the report is generated even if a stall did not happen, slowing down the system and flooding the logs. This is now fixed. #10981 The compaction manager now ignores out-of-disk-space (ENOSPC) exceptions when shutting down, so the server doesn’t crash in these scenarios. An inaccuracy in the per_partition_rate_limit read metric was corrected. Tables with the per_partition_rate_limit property can throttle read and write activity on a per-partition basis. #11651 A large schema (with thousands of tables) could cause stalls when propagated from node to node. This is now fixed. #11574 A recently introduced regression caused a crash when speculative retry was enabled. This is now fixed. #11825 segmentation fault in cases where the base table schema change while MV schema is cached #10026, #11542 A crash when the compaction manager was asked to stop multiple times (for different reasons) was fixed. A problem with RPC connections being needlessly dropped was fixed. #11780 When a query completes a page, ScyllaDB caches the query activity as an inactive read. When the client requests the next page, ScyllaDB re-activates the read and continues where it left off. A bug in this mechanism that could cause crashes has been fixed. #11923 A crash was fixed during an illegal lightweight transaction INSERT with NULL clustering key ScyllaDB caches rows and (since 4.6) index entries in a single unified cache. It was observed that in some small-partition workloads index caching causes a performance regression, so index caching is now disabled by default. It can still be enabled for workloads that benefit from it. We plan to re-enable it when the regression is fixed. #11889 The topology management code is more relaxed about unknown endpoints to prevent crashes in tests that check for edge cases. This fixes a recent regression. #11870 Hinted handoff now checks that a node exists in topology before doing anything; this helps with a recent regression due to topology refactoring. A crash while fetching repaid ids from the repair history table was fixed. #11966 Usually repair can compare and update the same shard in different nodes, for example shard 3 in one node is compared against shard 3 in another. When the number of shards in nodes is dissimilar, this doesn’t work and each shard compares against data from multiple shards in other nodes. This is now made more efficient by reducing sstable reader thrashing for this dissimilar shard count case. #12157 The algorithm for removing nodes from the token ring was corrected and made more efficient. It’s not known that this had any user impact. #12082 The CQL server will now only run requests that benefit from concurrency (e.g. QUERY and EXECUTE) in parallel. Configuration and authentication related requests will be serialized, reducing the chance for errors in those code paths. A rare bug involving an allocation failure while updating cached rows was fixed. #12068 The system.truncated table holds information about truncation times of user tables. A recent regression caused it to be unreadable by cqlsh. It is now fixed. 12239 COMPACT STORAGE tables allow the user to only specify a prefix of a compound clustering key. Bugs relating to such partial keys and reversed rows were fixed.Note that compact storage is deprecated (see section). #12180 ScyllaDB now supports multiple compaction groups 1. This is not a user-visible feature for now. Some copies of the lists of ranges to stream were eliminated from the decommission path, reducing latency spikes. #12332 When the global index cache is disabled, a local (per query) cache was used instead. When that cache was destroyed, a stall could result, generating a latency spike. This is now fixed. #12271 Compaction manager generally reacts to events to initiate compactions, but also has an hourly timer in case an event was missed (and for tombstone compaction, which isn’t triggered by an event). This timer is now less susceptible to stalls. #12390. Repair tried to trigger off-strategy compaction even for a table that was dropped during repair, failing the entire repair. It ignores the dropped table now. Off-strategy compaction is now enabled for all streaming topology operations (adding and removing nodes). Previously it was enabled only for repair-based node operations. Off-strategy compaction takes advantage of the fact that incoming sstables are non-overlapping to perform more efficient compaction that the one performed by the regular compaction strategy. ScyllaDB sometimes reads ahead of the user request, in order to hide latency. In one case a read-ahead request which timed out caused errors to be emitted, even though this did not affect the query. The errors are now silenced. #12435 ScyllaDB tracks transient memory used by queries. However, it did not track decompressed memory buffers, which could lead to running out of memory in some complicated queries. This is now fixed. “Unset” values are an obscure prepared statement feature that allows only some columns in an UPDATE or INSERT statement to be modified. It was a source of minor bugs and inconvenience in code. The feature has been refactored so it has less impact on the code and is more robust. Fix a crash in Materialized View update row locking, caused by a race condition #12632 (RC1) Fix a crash when reporting error on invalid CQL query involving field selection from a user-defined type #12739 (RC1) If a file page was inserted into cache, but the insertion required a memory allocation, it could cause memory corruption. This is now fixed. File caching is part of scylla-4.6 index caching feature. #9915 A recent commitlog regression has been fixed: we might have created two commitlog segment files in parallel, confusing the queue that holds them. The problem was only present in scylla-4.6. #9896 When shutting down, compaction tasks that are sleeping while awaiting a retry are now aborted, so the shutdown is not delayed. #10112 The compaction manager’s definition of compaction backlog for size-tiered compaction strategy has changed, reducing write amplification. The compaction backlog is used to determine how much resources will be devoted for compaction. The compaction manager’s definition of compaction backlog for size-tiered compaction strategy has changed, reducing write amplification. The compaction backlog is used to determine how much resources will be devoted for compaction. See the results graphed here. An accounting bug in sstable index caching that could lead to running out of memory has been fixed. #10056 ScyllaDB will shut down gracefully (without a core dump) if it encounters a permission or space problem on startup. #9573 sstables are created in a temporary directory to avoid serializing on an XFS lock. If the sstable writing failed, this directory would be left behind. It is now deleted. #9522 When using the spark migrator, ScyllaDB might see TTL’ed cells which it thought needed repair, but could not actually find a difference in, leading to repair not resolving the difference and detecting it again on the next run. This has been fixed. #10156 Prepared batch statements are now correctly invalidated when a referenced table changes. #10129 A race in the prepared statement cache could cause invalidations (due to a schema change) to be ignored. This is now fixed. #10117 When populating the row cache with data from sstables, we will first compact the data in order to avoid populating the cache with data that will be later ignored. #3568 When reading, we now notice if data exists in sstables/cache but not in memtables, or vice-versa. In either case there is no need to perform a merge, yielding substantial performance improvements. Truncate now disables compaction temporarily on the table and its materialized views, in order to complete faster. Cleanup compactions, used after bootstrapping a new node to reclaim space in the original nodes, now have reduced write amplification. SSTable index file reads now use read-ahead. This improves performance in workloads that frequently skip over small parts of the partition (e.g. full scans with clustering key restrictions). When ScyllaDB generates a name for a secondary index, it will avoid using special characters (if they were present in the table name). #3403 Cleanup compactions, used to discard data that was moved to a new node, now have reduced write amplification on Time Window Compaction Strategy tables Reads from cache are now upgraded to use the new range tombstone representation. This completes the conversion of the read pipeline, and nets a nice performance improvement as detailed in the commit message. A concurrent DROP TABLE while streaming data to another node is now tolerated. Previously, streaming would fail, requiring a restart of the node add or decommission operation. #10395 A crash where a map subscript was NULL in certain CQL expressions was fixed. Fixes #10361 #10399 #10401 A regression that prevented partially-written sstables from being deleted was fixed. A race condition that allowed queries to be processed after a table was dropped (causing a crash) was fixed. #10450 The prepared statement cache was recently split into two sections, one for reused statements and one for single-use statements. This was done for flood protection - so that a bunch of single-use statements won’t evict reused statements from the cache. However, this created a regression when the size of the single-use section was shrunk so that it was too small for statements to promote into the reused section. This is now fixed by maintaining a minimum size for each section. #10440 Level selection for Leveled Compaction Strategy was improved, reducing write amplification. Reconciliation is the process that happens when two replicas return non-identical results for a query. Some reactor stalls were removed, reducing latency effects on concurrent queries. #2361 #10038 A crash in some cases where an sstable index cursor was at the end of the file was fixed.#10403 A compaction job that is waiting in queue is now aborted immediately, rather than waiting to start and then getting aborted. Repair-based node operations use repair to move data between nodes for bootstrap/decommission and similar operations (currently enabled by default only for replacenode). The iteration order has been changed from an outer iteration on vnodes and an inner iteration on tables to an outer iteration on tables and an inner iteration on vnodes, allowing tables to be completed earlier. This in turn allows compaction to reduce the number of sstables earlier, reducing the risk of having too many sstables open. A “promoted index” is the intra-partition index that allows seeking within a partition using the clustering key. Due to a quirk in the sstable index file format, this has to be held in memory when it is being created. As a result, huge partitions risk running out of memory. ScyllaDB will now automatically downscale the promoted index to protect itself from running out of memory. Change Data Capture (CDC) tables are no longer removed when CDC is disabled, to allow the still-queued data to be drained. #10489 Until now, a deletion (tombstone) that hit data in memtable or cache did not remove the data from memtable or cache; instead both the data and the tombstone coexisted (with the data getting removed during reads or memtable flush). This was changed to eagerly apply tombstones to memtable/cache data. This reduces write amplification for delete-intensive workloads (including the internal Raft log). #652 Recently, repair was changed to complete one table before starting the next (instead of repairing by vnodes first). We now perform off-strategy compaction immediately after a table was completed. This reshapes the sstables received from different nodes and reduces the number of sstables in the node. The Leveled Compaction Strategy was made less aggressive. #10583 Compaction now updates the compaction history table in the background, so if the compaction history table is slow, compaction throughput is not affected. Memtable flushes will now be blocked if flushes generate sstables faster than compaction can clear them. This prevents running out of memory during reads. This reduces problems with frequent schema updates, as schema updates cause memtable flushes for the schema tables. #4116 A recent regression involving a lightweight transaction conditional on a list element was fixed. #10821 A race condition between the failure detector and gossip startup was fixed. Node startup for large clusters was sped up in the common case where there are no nodes in the process of leaving the cluster. An unnecessary copy was removed from the memtable/cache read path, leading to a nice speedup. A Seastar update reduces the node-to-node RPC latency. Internal materialized view reads incorrectly read static columns, confusing the following code and causing it to crash. This is now fixed. #10851 Staging sstables are used when materialized views need to be built after token range is moved to a different node, or after repair. They are now compacted regularly, helping control read and space amplification. Dropping a keyspace while a repair is running used to kill the repair; now this is handled more gracefully. A large amount of range tombstones in the cache could cause a reactor stall; this is now fixed. Adding range tombstones to the cache or a memtable could cause quadratic complexity; this is also fixed. Under certain rare and complicated conditions, memtables could miss a range tombstone. This could result in temporary data resurrection. This is now fixed. Fixes #10913 #10830 Previously, a schema update caused a flush of all memtables containing schema information (i.e. in the system_schema keyspace). This made schema updates (e.g. ALTER TABLE) quite slow. This was because we could not ensure that commitlog replay of the schema update would come before the commitlog replay of changes that depend on it. Now, however, we have a separate commitlog domain that can be replayed before regular data mutations, and so we no longer flush schema update mutations, speeding up schema updates considerably. #10897 Gossip convergence time in large clusters has been improved by disregarding frequently changing state that is not important to cluster topology - cache hit rate and view backlog statistics. Reads and writes no longer use C++ exceptions to communicate timeouts. Since C++ exceptions are very slow, this reduces CPU consumption and allows an overloaded node to retain throughput for requests that do succeed (“goodput”). Major compaction will now happen using the maintenance scheduling group (so its CPU and I/O consumption will be limited), and regular compaction of sstables created since it was started will be allowed. This will reduce read amplification while a major compaction is ongoing. #10961 The Seastar I/O scheduler was adjusted to allow higher latency on slower disks, in order to avoid a crash on startup or just slow throughput. #10927 ScyllaDB will now clean up the table directory skeleton when a table is dropped. #10896 The repair watchdog interval has been increased, to reduce false failures on large clusters. ScyllaDB propagates cache hit rate information through gossip, to allow coordinators to send less traffic to newly started node. It will now spend less effort to do so on large clusters. #5971 Improvements to token calculation mean that large cluster bootstrap is much faster. An automatically parallelized aggregation query now limits the number of token ranges it sends to a remote node in a single request, in order to reduce large allocations. #10725 Accidentally quadratic behavior when a large number of range tombstones is present in a partition has been fixed. #11211 Row cache will miss a row if upper bound of population range is evicted and has an adjacent dummy row #11239 --- ### Page: https://forum.scylladb.com/t/scylladb-enterprise-release-2023-1-0-part-3/794 Title: ScyllaDB Enterprise Release 2023.1.0 - Part 3 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: Part 1 Part 2 Tools The sstable tools gained Lua scripting. This is an expert feature intended for offline analysis of sstables. #9679 The scylla-types tool can now compute the token and shard of a partition key,… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-enterprise-release-2023-1-0-part-3/794 ## Headings Structure: H1: ScyllaDB Enterprise Release 2023.1.0 - Part 3 H3: Tools H3: Storage H3: Configuration H3: Deprecated and removed features H3: Monitoring and tracing H2: Additional bug fixes H3: Related topics ## Main Content: H1: ScyllaDB Enterprise Release 2023.1.0 - Part 3 H3: Tools H3: Storage H3: Configuration H3: Deprecated and removed features H3: Monitoring and tracing H2: Additional bug fixes H3: Related topics The sstable tools gained Lua scripting. This is an expert feature intended for offline analysis of sstables. #9679 The scylla-types tool can now compute the token and shard of a partition key, using the tokenof and shardof subcommands. The bundled cqlsh now uses the ScyllaDB Python driver (rather than the generic Cassandra driver) and supports Scylla Cloud Serverless connection bundles. The bundled cqlsh now considers system_distributed_everywhere a system keyspace. The bundled scylla types tool can now serialize a value to the sstable binary format. The scylla-api-client tool is now documented. The tool is suitable for interactive usage as well as shell automation of the REST API. #11999. scylla-api-cli is a lightweight command line tool interfacing the ScyllaDB REST API. The tool can be used to list the different API functions and their parameters, and to print detailed help for each function. Then, when invoking any function, scylla-api-cli performs basic validation on the function arguments and prints the result to the standard output. Note that json results msy be pretty-printed using commonly available command line utilities. It is recommended to use scylla-api-cli for interactive usage of the REST API over plain http tools, like curl, to prevent human errors. The sstable utilities now emit JSON output. See example output here. There two new sstable tools, validate-checksums and decompress, allowing for more offline inspection options of sstables. The SSTableLoader code base has been updated to support “me” format sstables. The sstable parsing tools usually need the schema to interpret an sstable’s data. For the special case of system tables, the tools can now use well-known schemas. Nodetool was updated to fix IPv6 related errors (even when IPv4 is used) with update JVMs. #10442 Cassandra-derived tooling such as cqlsh and cassandra-stress was synchronized with Cassandra 3.11.3. The bundled Prometheus node_exporterm used to report OS level metrics to ScyllaDB Monitoring Stack was upgraded to version 1.3.1. Repairs that were in their preparation stage previously could not be aborted. This is now fixed. ScyllaDB documentation has been moved from the scylla-docs.git repository to scylla.git. This will allow us to provide versioned documentation. The sstable tools gained a write operation that can convert a json dump of an sstable back into an sstable. It is now possible to limit, and control in real time, the bandwidth of streaming and compaction. These and more configuration updates below: Scylla Monitoring Stack release 4.4 and later will support ScyllaDB Enterprise 2023.1 metrics related updates below: Shard Latencies are now reported as summaries. This is part of an effort to reduce the total number of generated metrics. In addition, empty histograms and summaries will not be reported. The overall result is a 5x reduction in the number of metrics #11173. This is how a summary looks like: scylla_storage_proxy_coordinator_read_latency_summary_count{scheduling_group_name="statement",shard="1"} 2 scylla_storage_proxy_coordinator_read_latency_summary{quantile="0.990000",scheduling_group_name="statement",shard="1"} 640 There is now a metric that allows observation of update progress of materialized views from staging sstables. There are now completion percentage metrics for node operations using streaming; previously the completion metrics were only available when using repair-based node operations. #11600 The sstable row_reads metric for m-format sstables is now properly incremented, instead of showing zeroes. #12406 The replica-side read metrics, which have been incorrect for some time, have been revamped. #10065 Slow query tracing only considered local times - the time from when a request first hit the replica - to determine if a request needs to be traced. This could cause some parts of slow query tracing to be missed. To fix that, slow queries on the replicas are determined using the start time on the coordinator. The system.large_partitions and similar system tables will now hold only the base name of the sstable, not the full path. This is to avoid confusion if the large partition is reported while the sstable is in one directory, but later moved to another, for example from staging to the main directory after view building is done or into the quarantine subdirectory if they are found to be inconsistent with scrub. There are now metrics showing each node’s idea of how many live nodes and how many unreachable nodes there are. This aids understanding problems where failure detection is not symmetric. #10102 The system.clients table has been virtualized. This is a refactoring with no UX impact. Aggregated queries that use an index are now properly traced. The amount of per-table metrics has been reduced by sending metric summaries instead of histograms and not sending unused metrics. The following issues have been fixed on top of what was fixed in Scylla Open Source 5.2.0, with open source reference if available. In addition, all relevant bug fixes from 2022.1.x and 2022.2.x are fixed in 2023.1.0 Stability: an extremely rare case can cause Iterator invalidation in lsa_partition_reader::reset_state(), following by process exit #14696 Stability: mutation_reader_merger can overflow stack when merging many empty readers. This may happen when running a second repair right after the other. #14415 Stability: a lot of lsa-timing log messages during node replace cause c-s stuck and aborted. The fix update the reactor shares for default IO class from 1 to 200 #13753 DynamoDB API (Alternator) stability: assertion in output_stream when exception occurs during response streaming #14453. Stability: cached_file, used by index caching, will potentially cause a crash after OOM #14814 Stability: compaction: excessive reallocation during input list formatting #14071. Issue is more likely with offstrategy compaction. Stability: deadlock caused by view update _registration_sem and streaming reader _streaming_concurrency_sem #14676 Stability: a failure when reading metrics, caused by a rare race condition when another node is down. (seastar::metrics::double_registration (registering metrics twice for metrics: storage_proxy_coordinator_background_replica_writes_failed_remote_node)) #11017 Stability: ICS compaction is not working in cleanup #14035 (introduced in 2022.2.0) Stability: messaging: when upgrading OSS nodes to Enterprise, service-levels are matched to the default scheduling group #13841, #12552 Stability: Range-scans have a protection against using the wrong service-level to continue a suspended range-scan. This protection had a mistake, resulting in the node crashing when the protection mechanism was triggered. multishard_mutation_query: reader_context::lookup_readers() is not exception safe w.r.t. closing readers #13784 Stability: partitioned_sstable_set::insert might stall when called by table::make_reader_v2_excluding_sstables. The root cause is View building from staging creates a reader from scratch for every partition, in order to calculate the diff between new staging data and data in base sstable set, and then pushes the result into the view replicas. #14244 Setup: scylla-fstrim.timer is enabled but not started #14249 Setup: The installer now wipes filesystem signatures from the individual disks making up a RAID array, preventing problems with reuse of disks. #13737 Stability: bad_alloc (seastar - Failed to allocate 536870912 bytes) #13491. Root cause is a logic fault causing the reader to attempt to read all the data, consuming all memory. Can occur during sstableloader/nodetool refresh, repair or range scan. Stability stack-use-after-return in table::make_reader_v2_excluding_staging() #14812 Stability: View building crashes on large partitions with range tombstones. #14503 DynamoDB API (Alternator) stability: Yield while building large results in Alternator - rjson::print, executor::batch_get_item #13689 Setup: fix a regression in setup, which overrides the manual update of perftune.yaml #11385 #10121 Setup: updates in perftune.py, improving performance for larger servers (32 cores and above) Stability: ‘sleep_aborted’ error during Scylla shutdown #13374 Stability: a rare failure in row_cache_test/test_concurrent_reads_and_eviction #12462 Stability: ALTER KEYSPACE can break tables with UDT columns #14139 Correctness: Decommission and removenode may lead to consistency issues if one of the nodes decides to abort during streaming #12989 UX: non informative iotune warnings in scylla_kernel_check #13373 Stability: a race condition in scylla boot, when migration_manager::sync_schema failed with seastar::rpc::closed_error causing repair to fail #12956, #12764 Stability: Node operations failures get masked by abort request failures #12798 Stability: Node operations may fail if prepare takes longer than heartbeat timeout #12969, #11011 Stability: Segmentation fault happend on alive nodes during adding new node with replace terminated one #13368 (issues introduced in 5.2) Stability: Shutting down auth service may hang #13545 Correctness: tables with the new tombstone_gc ‘immediate’ mode might delete ttl data that is not expired #13572 Stability: possible use-after-move in virtual table for secondary indexes #13396 Stability: possible use-after-move when initializing row cache with dummy entry #13400 Stability: possible use-after-move in virtual table for secondary indexes #13396 Stability: possible use-after-move when making streaming reader #13397 Stability: possible use-after-move when reading from SSTable in reverse #13394 Stability: possible use-after-move when tracking view builder progress #13395 Stability: reactor stalls in commitlog replay path due to commit log regexp processing #11710 Stability: Replication of default auth settings may fail #2852 Stability: db/view: update view generator doesn’t close staging sstable reader on exceptions #13413 Stability: direct_failure_detector::ping_with_timeout() causes exceptions to be thrown every 100ms times the number of live nodes, which spam the logs, and might slow it down #13278 Stability: on_internal_error doesn’t log an error when not aborting #13786 Packaging: RPM package dependencies issue. When installing a specific version with yum/dnf, scylla-python3 version will not match the specified version, but the latest one. #13222 Stability: bad_alloc (seastar - Failed to allocate 536870912 bytes) #13491. Root cause is a logic fault causing the reader to attempt to read all the data, consuming all memory. Can occur during sstableloader/nodetool refresh, repair or range scan. Monitoring: new metric for CQL request and response sizes #13061 Audit: do not round timestamp in the audit table Encryption at rest: rare deadlock when creating a table using encryption with replicated key provider (default) for the first time Stability: Adding nodes to a large cluster (90+ nodes) may cause existing nodes to crash. The root cause is quadratic behavior in get_address_ranges function #12724 Stability: a rare crash due to null pointer dereference: clear_gently of disengaged unique_ptr dereferences nullptr #13636 Performance: Compaction manager “periodic reevaluation” is one-off. This means that compaction was not kicking in later for a table, with low to none write activity, that had expired data 1 hour from now. #13430 Stability:: Internal error in a COUNT request with empty IN. The query “select count(*) from {table1} where p in ()” should result in the count 0, because the empty p in () matches no row. However, what we get in Scylla now is an internal error. #12475 Tools: total disk space used metric incorrectly tells the amount of disk space ever used, which is wrong. It should tell the size of all SSTable being used plus the ones waiting to be deleted. Live disk space used shouldn’t account for the ones waiting to be deleted, and live SSTable Count shouldn’t account SSTable waiting to be deleted. #12717 Stability: Bootstrap fails during replace operation while starting “off-strategy compaction”. Huge amount of “Error applying view update” errors were received #12693. The cause is commit “repair: Reduce repair reader eviction with diff shard count” introduced in 2022.2.1 Stability: CQL compression might cause reactor stalls on buffer allocation #13437 Stability: coredumps were not being generated. A fix increase systemd coredump generation timeout #5430 Performance: Fix stalls caused quadratic behavior when inserting sstables into tracker on schema change #12499 Stability: abort_source::do_request_abort(std::optionalstd::exception_ptr): Assertion ‘_subscriptions’ failed. during shutdown #12512 CQL: scylla: types: is_tuple(): doesn’t handle reverse types. For example, a schema with reversed clustering key component; this component will be incorrectly represented in the schema CQL dump: the UDT will lose the frozen attribute. When attempting to recreate this schema based on the dump, it will fail as the only frozen UDTs are allowed in primary key components. #12576 Stability: commitlog: segment recycling breaks on segment file removal #12645 Workload Prioritization improvements and bug fix: Incremental Compaction Strategy (ICS) improvements and bug fixes: --- ### Page: https://forum.scylladb.com/t/release-scylla-monitoring-stack-4-4-4/795 Title: [RELEASE] Scylla Monitoring Stack 4.4.4 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce a patch release of ScyllaDB Monitoring Stack 4.4.4 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB Enterprise and ScyllaDB Open Source, based on Prometh… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-monitoring-stack-4-4-4/795 ## Headings Structure: H1: [RELEASE] Scylla Monitoring Stack 4.4.4 H3: Related topics ## Main Content: H1: [RELEASE] Scylla Monitoring Stack 4.4.4 H3: Related topics The ScyllaDB team is pleased to announce a patch release of ScyllaDB Monitoring Stack 4.4.4 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB Enterprise and ScyllaDB Open Source, based on Prometheus and Grafana. ScyllaDB Monitoring Stack 4.4.4 supports: The patch release adds Scylla Manager 3.2.x support --- ### Page: https://forum.scylladb.com/t/best-practice-when-implementing-nearby-search-with-h3-indexing-write-more-or-read-more/797 Title: Best practice when implementing nearby search with h3 indexing (write more or read more) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi guys, I’m trying to develop a system that requires to do some geospatial search (eg, find nearby users), so my first version of model of table has a set of h3 hashes of different resolutions of h3 cell ids stored (ind… Language: en Canonical URL: https://forum.scylladb.com/t/best-practice-when-implementing-nearby-search-with-h3-indexing-write-more-or-read-more/797 ## Headings Structure: H1: Best practice when implementing nearby search with h3 indexing (write more or read more) H3: Related topics ## Main Content: H1: Best practice when implementing nearby search with h3 indexing (write more or read more) H3: Related topics Hi guys, I’m trying to develop a system that requires to do some geospatial search (eg, find nearby users), so my first version of model of table has a set of h3 hashes of different resolutions of h3 cell ids stored (indexed, with TTL). When try to query nearby users, I make multiple queries for the disc of the cell to query (will generate 5-7 queries). But I resemble reading somewhere that write queries in scylla db have less cost than read queries, so it came up to me that instead of storing only different resolutions of h3 cell of user’s current location, why not just store the whole cells in the disc and different resolutions (will increase the set size by 5-7x), but now only one query is required. My question is does the latter version better than the first one or it really depends on the frequency of user’s location update and nearby search? Thanks! You are correct that writes are cheaper in ScyllaDB than reads. What you suggest makes sense to me, but it’s hard to know without more details. I’d try to create your data model and test it to see the actual performance and see if it’s suitable for your requirements. You can use a tool like cassandra-stress to do the testing. A good resource is the Basic Data Modeling lesson on ScyllaDB University. --- ### Page: https://forum.scylladb.com/t/how-data-stores-in-hard-disk/800 Title: How data stores in hard disk? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi all , I want to know how the scylla stores data in hard disk. Are data stored on different different pages / frames of the hard disk? Is data storation a contiguous allocation process or non contiguous allocation? Language: en Canonical URL: https://forum.scylladb.com/t/how-data-stores-in-hard-disk/800 ## Headings Structure: H1: How data stores in hard disk? H3: Related topics ## Main Content: H1: How data stores in hard disk? H3: Related topics Hi all , I want to know how the scylla stores data in hard disk. Are data stored on different different pages / frames of the hard disk? Is data storation a contiguous allocation process or non contiguous allocation? In ScyllaDB (and Cassandra), data is stored in Sorted String Tables (SSTables). You can learn more about SSTables and compaction in the [Compaction Strategies].(Compaction Strategies - ScyllaDB University) ScyllaDB Universty lesson. You can read more about the Architecture of ScyllaDB and how it partitions data in the Architecture ScyllaDB University lesson. Another useful resource is this page. --- ### Page: https://forum.scylladb.com/t/error-in-setup-scylladb-runtime-error-no-nodes-present-in-the-cluster-has-this-node-finished-starting-up/802 Title: Error in setup scyllaDB ( runtime error: No nodes present in the cluster. Has this node finished starting up? ) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Full Error: root@b2-60-sgp1:/tmp/tb# nodetool status nodetool: Scylla API server HTTP GET to URL '/storage_service/ownership/' failed: runtime_exception (runtime error: No nodes present in the cluster. Has this node fin… Language: en Canonical URL: https://forum.scylladb.com/t/error-in-setup-scylladb-runtime-error-no-nodes-present-in-the-cluster-has-this-node-finished-starting-up/802 ## Headings Structure: H1: Error in setup scyllaDB ( runtime error: No nodes present in the cluster. Has this node finished starting up? ) H3: Related topics ## Main Content: H1: Error in setup scyllaDB ( runtime error: No nodes present in the cluster. Has this node finished starting up? ) H3: Related topics My Server OS : Ubuntu 20.04 Ram 32GB Core 8 Provider : OVH '/storage_service/ownership/' failed: runtime_exception (runtime error: No nodes present in the cluster. Has this node finished starting up?) Did you check if the Scylla service started? Please share scylla server + scylla-jmx output. Logs would help as well. Please check if there are firewall rules or ports blocked (if this is a multi-node cluster) Check scylla.yaml for any errors. --- ### Page: https://forum.scylladb.com/t/scylla-5-2-very-slow-startup/803 Title: Scylla 5.2 very slow startup - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I just updated a 3 node cluster from 5.1 to scylla 5.2, but it does not start up. It seems scylla just hangs after the upgrade: – A start job for unit scylla-server.service has begun execution. – The job identifier is 1… Language: en Canonical URL: https://forum.scylladb.com/t/scylla-5-2-very-slow-startup/803 ## Headings Structure: H1: Scylla 5.2 very slow startup H2: – A start job for unit scylla-server.service has begun execution. H3: Related topics ## Main Content: H1: Scylla 5.2 very slow startup H2: – A start job for unit scylla-server.service has begun execution. H3: Related topics I just updated a 3 node cluster from 5.1 to scylla 5.2, but it does not start up. It seems scylla just hangs after the upgrade: – The job identifier is 145767327. Aug 23 22:39:49 pcdev-1 scylla[2834247]: Scylla version 5.2.6-0.20230730.58acf071bf28 with build-id 17961be569f8503b27ff284a8de1e00a9d83811e starting … Aug 23 22:39:49 pcdev-1 scylla[2834247]: command used: “/usr/bin/scylla --log-to-syslog 1 --log-to-stdout 0 --default-log-level info --network-stack posix --memory 10G --reserve-memory 52G --overprovisioned --kernel-page-cache 1 --unsafe-bypass-fsync 1 --io-properties-file=/etc/scylla.d/io_properties.yaml --developer-mode=1 --cpuset 0-3 --smp 4” Aug 23 22:39:49 pcdev-1 scylla[2834247]: parsed command line options: [log-to-syslog, (positional) 1, log-to-stdout, (positional) 0, default-log-level, (positional) info, network-stack, (positional) posix, memory, (positional) 10G, reserve-memory, (positional) 52G, overprovisioned, kernel-page-cache, (positional) 1, unsafe-bypass-fsync, (positional) 1, io-properties-file: /etc/scylla.d/io_properties.yaml, developer-mode: 1, cpuset, (positional) 0-3, smp, (positional) 4] Aug 23 22:39:49 pcdev-1 scylla[2834247]: seastar - Reactor backend: linux-aio Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 0] seastar - Creation of perf_event based stall detector failed, falling back to posix timer: std::system_error (error system:13, perf_event_open() failed: Permission denied) Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 0] seastar - Created fair group io-queue-66305, capacity rate 192:50000, limit 23649164, rate 16777216 (factor 1), threshold 11227761 Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 0] seastar - IO queue uses 1.41ms latency goal for device 66305 Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 0] seastar - Created io group dev(66305), length limit 131072:65536, rate 192000:50000000 Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 0] seastar - Created io queue dev(66305) capacities: 512:11227761:16830904 1024:11270710:16884590 2048:11356610:16991964 4096:11528408:17206712 8192:11872006:17636210 16384:12559200:18495202 32768:13933590:20213190 65536:16682369:23649164 131072:22179928:X Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 0] seastar - Created fair group io-queue-0, capacity rate 2147483:2147483, limit 12582912, rate 16777216 (factor 1), threshold 2000 Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 0] seastar - IO queue uses 0.75ms latency goal for device 0 Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 0] seastar - Created io group dev(0), length limit 4194304:4194304, rate 2147483647:2147483647 Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 0] seastar - Created io queue dev(0) capacities: 512:2000:2000 1024:3000:3000 2048:5000:5000 4096:9000:9000 8192:17000:17000 16384:33000:33000 32768:65000:65000 65536:129000:129000 131072:257000:257000 Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 1] seastar - Creation of perf_event based stall detector failed, falling back to posix timer: std::system_error (error system:13, perf_event_open() failed: Permission denied) Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 3] seastar - Creation of perf_event based stall detector failed, falling back to posix timer: std::system_error (error system:13, perf_event_open() failed: Permission denied) Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 2] seastar - Creation of perf_event based stall detector failed, falling back to posix timer: std::system_error (error system:13, perf_event_open() failed: Permission denied) Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 0] seastar - updated: blocked-reactor-notify-ms=1000000 Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 1] seastar - updated: blocked-reactor-notify-ms=1000000 Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 3] seastar - updated: blocked-reactor-notify-ms=1000000 Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 2] seastar - updated: blocked-reactor-notify-ms=1000000 Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 0] init - Unknown option : max_size_of_hints_in_progress Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 0] init - installing SIGHUP handler Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 0] init - Scylla version 5.2.6-0.20230730.58acf071bf28 with build-id 17961be569f8503b27ff284a8de1e00a9d83811e starting … Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 0] init - starting prometheus API server Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 0] init - creating snitch Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 0] init - starting tokens manager Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 0] init - starting effective_replication_map factory Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 0] init - starting migration manager notifier Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 0] init - starting lifecycle notifier Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 0] init - creating tracing Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 0] init - starting API server Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 0] init - Scylla API server listening on 0.0.0.0:10000 … Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 0] service_level_controller - update_from_distributed_data: starting configuration polling loop Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 0] init - starting system keyspace Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 0] init - starting gossiper Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 0] init - seeds={192.168.178.101, 192.168.178.102, 192.168.178.103}, listen_address=192.168.178.103, broadcast_address=192.168.178.103 Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 0] init - starting Raft address map Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 0] init - starting direct failure detector pinger service Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 0] init - starting direct failure detector service Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 0] init - initializing storage service Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 0] storage_service - Started node_ops_abort_thread Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 1] storage_service - Started node_ops_abort_thread Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 2] storage_service - Started node_ops_abort_thread Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 3] storage_service - Started node_ops_abort_thread Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 0] init - starting per-shard database core Aug 23 22:39:49 pcdev-1 scylla[2834247]: [shard 0] init - creating and verifying directories CPU and disk is idling. The other two nodes are still one 5.1 and still up. Any ideas whats wrong or what I could check to identify the problem? Strangely, I let it run over night and now its started up: Whats going on those 13 minutes? The cluster does not even have much data. Do I have to worry it takes even longer for larger clusters? Between the log lines you quoted ScyllaDB just verifies that the directories for the tables exist and have the correct permissions. Do you have a slow disk and/or many keyspaces/tables? Around 2 keyspaces with 80 tables each. The system seemed to be idling to me. But maybe the disk is getting old has some issues. This should take at most a few seconds even with slow disks. Is it reproducible? If so please run iostat -x 1 in parallel and post the results. Please mark the start/end time of the event. I just saw that slow host is the one running still on Ubuntu 20, while the others have Ubuntu 22. Disks are spinning disks, so they are not fast, but still itn’t shouldn’t take so long. Here iostat output with some compactions running: btw: Disks do not seem to be able to saturate CPU during major compaction. The restart of that server is slow again. It seems to be only that server. iostat number dont look that high to me: smartmontools report the disks are fine. Honestly, I think this is a problem related to Ubuntu 20. The slow server is the the only server with ubuntu 20, the others in the cluster have Ubuntu 22.04. Sorry, I didn’t get a notification and lost track of the thread. Please share the Advanced dashboard (set to per instance) disk stats. Also make use you set --io-latency-goal-ms=100 for spinning disks. Don’t worry, I am happy about any response. What I find strange that its only the Ubuntu 20 machine having that problem. The other two nodes of that cluster are on Ubuntu 22. Can it be some issue with scylla 5.2 on Ubuntu 20.04? It seems there are no metrics for the restart-period: Its even worse (8 minutes) with the 100ms goal: ← There are a couple of commitlog errors. Is it perhaps going over Please strace -fF -ttt the process while it is happening, and attach the part where the timestamps match the area where it stalls. I don’t think the OS version has anything to do with it, but it might. You were faster. The entire cluster is quite slow now actually. That would confirm your theory that its not OS related. We will check … We ran the strace during the verifying directories for a few seconds on 5.2.7. Uploaded strace log was uploaded as: 536aa7c0-9ad3-4edf-a47b-78a2f29b5c2a In between we updated the cluster to 5.2.8, which seems to have made things much worse, to the point that the system was not usable any more. We reverted back to 5.2.7 for but will try again to confirm its really related to the version. It’s missing reactor: syscall thread: wakeup up reactor with finer granularity · scylladb/seastar@ecf98a5 · GitHub. But that’s a fairly new optimization and (a) it should work without it and (b) I don’t expect a huge speedup with it. I guess I see a problem. All access() calls come from on thread. IIRC previously we distributed distributed_loader. Please file an issue so I can tag relevant people. There’s another option - it’s not slower, but you have more files (because of code changes, or because you have more data). Can you check against a backup? The number of files in /var/lib/scylla/data is enough. The problem is on a dev system where I dont do backups. But you are indeed on to something: There was a ton of snapshots! With the snapshots removed it looks much faster now. Server 1 still is a bit slower, but I think thats really just due to the older Ubuntu 20 machine. But the strange thing is that startup times got so much worse with 5.2 compared to 5.1. And even worse, with 5.2.8 the entire cluster became unusable - not only during startup, but also after being started up. We’ll try 5.2.8 with the snapshots removed. Startup looks fine now. Maybe its still slower than 5.1, but its not noticable any more. Just upgraded to 5.2.9 and now queries are slow again. I’ll see if I find something in grafana… edit: Now even with 5.2.9 its starting to look good. Maybe the system was slow due to the repairs running. Grafana looks rather unspectacular. It looks like its populating caches. I’ll keep an eye on it… System is still behaving well. Is 5.2 maybe just honoring IO limits more strictly? Scylla scanning the whole snapshots directory was acknowledged a bug some time ago. Recently it got fixed, but it didn’t get into 5.2 (Scylla should skip mode validation of snapshot files · Issue #12010 · scylladb/scylladb · GitHub) --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-2-7/804 Title: [RELEASE] ScyllaDB 5.2.7 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.2.7, a bugfix release of the ScyllaDB 5.2 stable branch. ScyllaDB Open Source 5.2.7, like all past and future 5.x.y releases, is backward compatible and supports rolling… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-2-7/804 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.2.7 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.2.7 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.2.7, a bugfix release of the ScyllaDB 5.2 stable branch. ScyllaDB Open Source 5.2.7, like all past and future 5.x.y releases, is backward compatible and supports rolling upgrades. Users are encouraged to upgrade to 5.2.7. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-21-2023-08-25/806 Title: Last week in scylla-cluster-tests.git master (issue #21; 2023-08-25) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the c18f651c…88ce7170 range are covered. There were 26 non-merge commits from 7 authors in tha… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-21-2023-08-25/806 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #21; 2023-08-25) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #21; 2023-08-25) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the c18f651c…88ce7170 range are covered. There were 26 non-merge commits from 7 authors in that period. Some notable commits: Since manager 3.2.0 is being released, we bumped manager version to 3.2 in all tests. Added a nemesis that disables the target node’s binary protocol and gossip, initiates a major compaction, and enables the node’s gossip and binary back. This scenario was found to be used on production and required more testing. Using use_preinstalled_scylla: true as default in gce. New microbenchmark test on x86_64 (c6i.large). Recently the structure of defaults/manager_persistent_snapshots.yaml has changed drastically and mgmt_cli_test was adapted to it. During upgrade tests when one node is upgraded, sometimes we observe read from truncated table to fail with timeout (60sec). We prolonged some queries timeouts and started to measure&store actual time in ES for future analysis. Perf Emails subjects now contain the backend and node type for tests that support different backends. Find Prometheus stats to the bottom of perf emails body. Operator v1.10 runs cleanups on each of the old nodes after adding a new one (like it is recommended by the Scylla documentation). We added a new functional test to cover that functionality. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/more-information-on-the-open-source-plan/808 Title: More information on the "Open Source" plan - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I’m new to ScyllaDB; I’ve been struggling a lot to find the right database for my new online 3D videogame, but none has been good enough for me yet, now I’m considering ScyllaDB, but I have a few questions on the free pl… Language: en Canonical URL: https://forum.scylladb.com/t/more-information-on-the-open-source-plan/808 ## Headings Structure: H1: More information on the "Open Source" plan H3: Related topics ## Main Content: H1: More information on the "Open Source" plan H3: Related topics I’m new to ScyllaDB; I’ve been struggling a lot to find the right database for my new online 3D videogame, but none has been good enough for me yet, now I’m considering ScyllaDB, but I have a few questions on the free plan of ScyllaDB - Open Source. How big is the information limit for the free plan? Like, it fits 1 GB, or 5 GB, or other size? Is there any limit like “x CRUD requests/month” in the Open Source plan? Is ScyllaDB on its own even a good option for an online game? If not, do you have any alternative options? Is the free plan good enough if I’m creating a 3D multiplayer videogame, considering it has to send really a lot of requests per second to keep showing the game world to the player in nearly real-time? I’m considering a GET request every 10 milliseconds. Hi, first of all I’m not affiliated to ScyllaDB, ScyllaDB open source has no limits for node or requests, just download on nodes and start it! I don’t know if it’s a good option for a 3d multiplayer game because i don’t know about your game/usecase. It’s a simple Minecraft clone: briefly, you can make a world everybody can join, you can join a world, place and break blocks, you’re also handed out 8 unique guns you can shoot anybody with. Hi There is no limitation on ScyllaDB Open Source capacity, throughput, number, or size of nodes in a ScyllaDB Cluster. You can use ScyllaDB Open Source in production for 10s and 100s of TBs of data. ScyllaDB Enterprise has extra features, easier administration (see Scylla Manager), better performance, and commercial support. You can start with Open Source and upgrade to Enterprise at any point. BTW, I’m a huge Minecraft fan and would be happy to see your open-source alternative Tzach VP Product, ScyllaDB Thanks for the information you gave me! I will stop on ScyllaDB then; I’m planning the game as a game with no expense and no income either. I’m new to game dev and I’m just messing around with that, so it’s not guaranteed the game will turn out a masterpiece. The original game was made by a Russian individual and by the state of the game it was obvious that the game won’t meet anymore updates, and the game was taken down in around 2020-2021, so I decided to try and recreate that. I don’t want to add in too much new features, but I don’t want to miss on a lot of them either. Why stop using ScyllaDB? As I wrote, ScyllaDB Open Source is 100% production ready, with no limitations, and, IMHO, an excellent fit for open source projects, even large scale. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-194-2023-08-27/809 Title: Last week in scylladb.git master (issue #194; 2023-08-27) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 7275b8967c…9806bddf75 range are covered. There were 86 non-merge commits from 17 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-194-2023-08-27/809 ## Headings Structure: H1: Last week in scylladb.git master (issue #194; 2023-08-27) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #194; 2023-08-27) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 7275b8967c…9806bddf75 range are covered. There were 86 non-merge commits from 17 authors in that period. Some notable commits: Tablets are a new, experimental method of partitioning data among nodes, diverging from vnodes. Individual tablets are now stored in their own sstables, facilitating data movement and cleanup. A bug that caused nodes to fail to start if a tablet was migrated concurrently with its table being dropped was fixed. The system.tablets table stores the replica set of each tablet in the system. It now maintains the order of replicas, so that the notion of primary owner is maintained. The scylla.yaml configuration items are now documented in the documentation website. A recent regression causing a crash on table drop was fixed. The setup utility supported an --online-discard switch to enable/disable online discard, but it did not actually work. This is now fixed. The nodetool stop RESHAPE command is supposed to stop the reshape operation, but in fact only aborted running reshape compactions, which were promptly restarted. It now aborts the entire operation as expected. A crash during rebuild operations with experimental consistent cluster topology was fixed. ScyllaDB contains two classes of tables, system and user, and uses separate memory pools for their memtables. This avoids a deadlock when a user memtable is being flushed, and needs to allocate memtable space for a system table as part of the flush process. We now automatically designate all system tables as using the system memtable pool. Latency during repair of large numbers of small rows was improved. ScyllaDB now supports the –ignore-dead-nodes option family when experimental consistent cluster topology is enabled. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylla-operator-1-10-0/810 Title: [RELEASE] Scylla Operator 1.10.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of Scylla Operator 1.10.0. Scylla Operator is an open-source project that helps ScyllaDB Open Source and ScyllaDB Enterprise users run ScyllaDB on Kubernetes. The Scy… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-operator-1-10-0/810 ## Headings Structure: H1: [RELEASE] Scylla Operator 1.10.0 H1: Notable changes H1: Supported versions H1: Upgrade instructions H2: Getting started with Scylla Operator H2: Related Links H3: Related topics ## Main Content: H1: [RELEASE] Scylla Operator 1.10.0 H1: Notable changes H1: Supported versions H1: Upgrade instructions H2: Getting started with Scylla Operator H2: Related Links H3: Related topics The ScyllaDB team is pleased to announce the release of Scylla Operator 1.10.0. Scylla Operator is an open-source project that helps ScyllaDB Open Source and ScyllaDB Enterprise users run ScyllaDB on Kubernetes. The Scylla Operator manages ScyllaDB clusters deployed to Kubernetes and automates tasks related to operating a ScyllaDB cluster, like installation, vertical and horizontal scaling, as well as rolling upgrades. Scylla Operator 1.10.0 improves stability and brings a few features. As with all of our releases, any API changes are backward compatible. For more changes and details check out the GitHub release notes. Upgrading from v1.9.x with kubectl apply doesn’t require any extra action, just take the manifest from v1.10.0 tag and substitute the released image. Using helm requires a mandatory manual step for every release because helm can’t handle CRDs updates. For details, see our upgrade documentation. We’ll welcome your feedback! Feel free to open an issue or reach out on the #scylla-operator channel in Scylla User Slack. --- ### Page: https://forum.scylladb.com/t/will-synchronous-materialized-views-roll-back-after-a-write-failure/812 Title: Will synchronous materialized views roll back after a write failure? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: When synchronous_view_updates is set to True, it will synchronously write the base table and GSI. If writing a certain GSI fails during this period, will the other successfully written base and index tables be rolled bac… Language: en Canonical URL: https://forum.scylladb.com/t/will-synchronous-materialized-views-roll-back-after-a-write-failure/812 ## Headings Structure: H1: Will synchronous materialized views roll back after a write failure? H3: Related topics ## Main Content: H1: Will synchronous materialized views roll back after a write failure? H3: Related topics When synchronous_view_updates is set to True, it will synchronously write the base table and GSI. If writing a certain GSI fails during this period, will the other successfully written base and index tables be rolled back? It is not reflected in the introduction of document synchronous-materialized-views. I think it will. But I haven’t found the corresponding code in this PR either? Does anyone understand the working principle? Thank you. No, no previous writes will be rolled back. This is only possible with transactions, which we don’t support. --- ### Page: https://forum.scylladb.com/t/scylla-reactor-stalled/813 Title: Scylla reactor stalled - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello! I have old scylla cluster v4.3. It started from 5 nodes, then 6 and about 2 months ago i added 7th node. It worked fine untill i added 7th node. Now all nodes sometimes send error: Aug 30 14:15:04 scylla1 scylla… Language: en Canonical URL: https://forum.scylladb.com/t/scylla-reactor-stalled/813 ## Headings Structure: H1: Scylla reactor stalled H3: Related topics ## Main Content: H1: Scylla reactor stalled H3: Related topics Hello! I have old scylla cluster v4.3. It started from 5 nodes, then 6 and about 2 months ago i added 7th node. It worked fine untill i added 7th node. Now all nodes sometimes send error: Aug 30 14:15:04 scylla1 scylla[3433932]: Reactor stalled for 525 ms on shard 1. … Aug 30 14:15:04 scylla1 scylla[3433932]: /opt/scylladb/libreloc/libpthread.so.0+0x0000000000009431 Aug 30 14:15:04 scylla1 scylla[3433932]: /opt/scylladb/libreloc/libc.so.6+0x0000000000101912 or Aug 30 14:18:19 scylla1 scylla[3433932]: Reactor stalled for 526 ms on shard 3. … Aug 30 14:18:19 scylla1 scylla[3433932]: /opt/scylladb/libreloc/libpthread.so.0+0x0000000000009431 Aug 30 14:18:19 scylla1 scylla[3433932]: /opt/scylladb/libreloc/libc.so.6+0x0000000000101912 with different shards and different backtraces I’m not shure if it is related to keyspaces/tables damage, or incorrect scylla configuration or network issues or scylla bug. Tried repair tables and drain/stop new added node - it didn’t help. Please open an issue about this in the bug tracker, and be sure to also include the exact scylla version you have. --- ### Page: https://forum.scylladb.com/t/scylla-manager-task-healthcheck-cql-frequency-update/814 Title: Scylla manager task healthcheck/cql frequency update - Knowledge Base - ScyllaDB Community NoSQL Forum Meta Description: Can I update the frequency of default scylla manager task healthcheck/cql with sctool commands. Language: en Canonical URL: https://forum.scylladb.com/t/scylla-manager-task-healthcheck-cql-frequency-update/814 ## Headings Structure: H1: Scylla manager task healthcheck/cql frequency update H3: Related topics ## Main Content: H1: Scylla manager task healthcheck/cql frequency update H3: Related topics Can I update the frequency of default scylla manager task healthcheck/cql with sctool commands. Unfortunately, this is not possible with the current CLI. --- ### Page: https://forum.scylladb.com/t/release-scylla-manager-3-2-0/815 Title: [RELEASE] Scylla Manager 3.2.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces the release of ScyllaDB Manager 3.2 production-ready ScyllaDB Manager minor release of the stable ScyllaDB Manager 3.2 branch. ScyllaDB Manager is a centralized cluster administration and rec… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-manager-3-2-0/815 ## Headings Structure: H1: [RELEASE] Scylla Manager 3.2.0 H3: Restore Automation H3: Repair Improvements H3: Configuration updates H3: Monitoring H3: Known Issue H3: Related topics ## Main Content: H1: [RELEASE] Scylla Manager 3.2.0 H3: Restore Automation H4: Additional restore metrics H3: Repair Improvements H4: Sctool repair flag updates H3: Configuration updates H3: Monitoring H3: Known Issue H3: Related topics The ScyllaDB team announces the release of ScyllaDB Manager 3.2 production-ready ScyllaDB Manager minor release of the stable ScyllaDB Manager 3.2 branch. ScyllaDB Manager is a centralized cluster administration and recurrent tasks automation tool. Scylla Manager 3.2 includes performance and safety improvement to recurrent repair and automate manual parts of the restore procedure. ScyllaDB Enterprise customers are encouraged to upgrade to ScyllaDB Manager 3.2 in coordination with the ScyllaDB support team. The new release includes upgrades of both ScyllaDB Manager Server and Agent. Scylla Manager 3.2 support the following Scylla Enterprise releases: And the following Open Source release (limited to 5 nodes see license): You can install and run Scylla Manager on Kubernetes using Scylla Operator. More here. Scylla Manager Restore command allows you to restore backed-up data into a cluster. In Manager 3.1 release, one had to run the following commands manually: These steps are now automated as part of the sctool restore option. See Manager 3.2 docs here Repair is a recurrent offline task that synchronizes data across all data replicas. Scylla Manager automates the repair process and allows you to configure how and when repair occurs. When you create a cluster, a repair task is automatically scheduled. This task is set to occur each week by default, but you can change it to another time, change its parameters or add additional repair tasks if needed. This release changes the parallelism and order of the repair job for better performance and stability. The following changes have been made: The following parameter in scylla-manager.yaml Age_max is the maximum time for a backup run to be considered fresh and can be continued from the same snapshot. If exceeded, a new run with a new snapshot will be created. You can use Scylla Monitoring releases 4.4.4 and later to monitor Scylla Manager 3.2. When upgrading from Manager 3.1 to Manager 3.2, the progress percentages of a repair task created before the upgrade can show a value greater than 100% - only for the first post-upgrade run. The problem is limited to the progress metrics, not the actual repair. #3534 --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-1-16/817 Title: [RELEASE] ScyllaDB 5.1.16 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.1.16, a bugfix release of the ScyllaDB 5.1 stable branch. ScyllaDB Open Source 5.1.16, like all past and future 5.x.y releases are backward compatible and support rollin… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-1-16/817 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.1.16 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.1.16 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.1.16, a bugfix release of the ScyllaDB 5.1 stable branch. ScyllaDB Open Source 5.1.16, like all past and future 5.x.y releases are backward compatible and support rolling upgrades. Note that the latest stable ScyllaDB Open Source release is 5.2, and you are encouraged to upgrade to it. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/cluster-stuck-at-initialization-phase/819 Title: Cluster stuck at initialization phase - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello everyone, I am trying to set up a brand new 3 CPU Cluster for QA purposes. However, I am having a hard time getting up and running. I am able to start Node 1 but it seems to get stuck during the initialization pr… Language: en Canonical URL: https://forum.scylladb.com/t/cluster-stuck-at-initialization-phase/819 ## Headings Structure: H1: Cluster stuck at initialization phase H3: Related topics ## Main Content: H1: Cluster stuck at initialization phase H3: Related topics I am trying to set up a brand new 3 CPU Cluster for QA purposes. However, I am having a hard time getting up and running. I am able to start Node 1 but it seems to get stuck during the initialization procedure: As for Nodes 2 & 3, their processes exit with error code 1: I have also made UFW rules for the 11 Scylla related ports: Available at paste bin for convenience. Node 1 - Execution Log Node 2 - Execution Log Node 3 - Execution Log Thank you for your time and help. Alright, so after some very decent amount of trial and error, it seems that I got it to work. Somehow the “data” and “commitlog” directories became corrupt and were preventing Scylla from completing the intialization procedure. Hope this helps anyone who happens to come by this situation in the future. Hello, I encountered the same error as you when deploying scylladb5.2 using ubuntu 22.04, so I would like to ask, how did you solve this error? --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-22-2023-09-01/820 Title: Last week in scylla-cluster-tests.git master (issue #22; 2023-09-01) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 78a90963…6ab20635 range are covered. There were 17 non-merge commits from 9 authors in tha… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-22-2023-09-01/820 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #22; 2023-09-01) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #22; 2023-09-01) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 78a90963…6ab20635 range are covered. There were 17 non-merge commits from 9 authors in that period. Some notable commits: KMS nemesis got additional variation: key rotation. We can also enable KMS key rotation on test configuration level, test configuration is now easier to prepare. SCT uses monitoring version 4.5. SLA tests now wait and verify a service level is propagated to nodes. It was achieved by verifying the scylla ‘scylla_scheduler_group_shares’ metric on each node. We also measure its propagation time and keep it in ES for future analysis. Added nemesis filter for manager operations. To validate the manager repair operation functions as intended, added another manager repair nemesis that corrupts the data in the cluster before initiating the repair. Materialized view was created for main scylla_bench table in the large partition 4 days test. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/need-help-with-modelling-the-post-comment-architecture-to-support-complex-queries/821 Title: Need help with modelling the post/comment architecture to support complex queries - University and Training - ScyllaDB Community NoSQL Forum Meta Description: I need threaded comments in my posts but I cant seem to figure out how to model the create table. My original table- CREATE TABLE IF NOT EXISTS posts ( id string, parent_id string, comments_count bigint, shares_cou… Language: en Canonical URL: https://forum.scylladb.com/t/need-help-with-modelling-the-post-comment-architecture-to-support-complex-queries/821 ## Headings Structure: H1: Need help with modelling the post/comment architecture to support complex queries H3: Related topics ## Main Content: H1: Need help with modelling the post/comment architecture to support complex queries H3: Related topics I need threaded comments in my posts but I cant seem to figure out how to model the create table. My original table- CREATE TABLE IF NOT EXISTS posts ( id string, parent_id string, comments_count bigint, shares_count bigint, settings blob, user_id uuid, created_at timestamp, updated_at timestamp, body string, PRIMARY KEY ((user_id, parent_id), created_at) ) WITH CLUSTERING ORDER BY (created_at DESC); This was my original idea but the problem with this is I cant write a query that loads all the comments to this post since, theres no way I can pass the user_id. Any ideas? I need per user to be able to see the posts theyve posted on their timeline, & individual posts to hold all associated comments too. What is the exact query you want to run? Generally, once you know the queries it’s easier to come up with a model that works and performant I need something along the lines of- select * from posts where user_id = 3 and parent_id = :not_set; (for list of posts made from their personal account) select * from posts where parent_id = :not_set and created_at < ....; (for posts to display on homepage, this can infact use user_id if I want to show relevant posts, but ill have to pass a list of user_ids I want that are say friends of the current user so he can view their recent posts etc.) & now the most difficult problem- parent_id :not_set means that this is a post from a user. if parent_id = , then this some_id is the id of an existing post, & it means that this post is actually a comment to another post. Im trying to implement hierarchical comments using the same post model. I want to include threaded comments. But the problem with this is, theres no way for me to provide the user_id in the query that loads the comments. select * from posts where parent_id= ; where post_id is the post im trying to load the comments for So for the first query (or maybe 2nd query too) to work efficiently, I need user_id, but I cant provide user_id for the 3rd query since idk which user has commented on the post. How do we solve this? Maybe an approach you can try is adding a Materialized View with parent_id as the key? So in this case you’d have a schema like this: This would work well with the following queries: Then you’d have a MV (example): So if you need to query by the parent id you can use this view like this: Hmm, this does look great. 3 questions though- We cannot have a materialized view partition key given as a tables partition/clustering key? All PRIMARY KEY columns in the base table need to be included in the PRIMARY KEY in the materialized view Is this approach scalable for practical situations? Imagine a load equivalent to that of say- twitter Yes, materialized views are scalable, but the DB will take up more space on disk because the data gets duplicated in the MV. when 2 partition keys are mentioned, the 1st locates the node which holds the data via hash, & am I right to think that the 2nd key also does the same thing within that located node. Kind of like locating a virtual node within a node to more effectively pin point the location? In a composite partition key (two or more columns) all fields in the partition key are used to generate the hash thankyou so much for your time & help. I appreciate it. sorry, ive got a bit of a problem. i created the post table & the MV. I needed to change the post table. So once I changed the create table query, I tried to delete the original post table but I couldnt due to the MV & it seems drop isnt supported on MVs. How do I drop/delete the MV? Can’t you do drop materialized view keyspace.mv ? oh yeah this works. thanks again. I couldnt find any docs on the drop. My bad --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-195-2023-09-03/822 Title: Last week in scylladb.git master (issue #195; 2023-09-03) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 9806bddf75…cf37ab96f4 range are covered. There were 78 non-merge commits from 15 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-195-2023-09-03/822 ## Headings Structure: H1: Last week in scylladb.git master (issue #195; 2023-09-03) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #195; 2023-09-03) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 9806bddf75…cf37ab96f4 range are covered. There were 78 non-merge commits from 15 authors in that period. Some notable commits: Previously, the cache was enhanced to remove expired tombstones on read. It will now remove expired range tombstones on read as well. This prevents tombstone accumulation in cache. In preparation for streaming tablets as sstables, a replica can now capture an sstable snapshot of a tablet. Read concurrency on replicas is managed by reader_concurrency_semaphore. A deadlock while stopping it has been fixed. Gossip SYN messages, now carry the Raft cluster ID. This is used to prevent nodes from different clusters from communicating. This can happen if incorrect seed configuration was used when bootstrapping the cluster. Consistent cluster topology using Raft now supports the --ignore-dead-nodes with IP addresses. The option is now deprecated in favor of host IDs. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-1-10/823 Title: [RELEASE] ScyllaDB Enterprise 2022.1.10 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2022.1.10, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.1. Note that there is a newer Enterprise LTS release: ScyllaDB Enterpr… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-1-10/823 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2022.1.10 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2022.1.10 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2022.1.10, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.1. Note that there is a newer Enterprise LTS release: ScyllaDB Enterprise 2023.1. While we will continue to support 2022.1 LTS, you can get additional features with 2023.1. Below is a list of performance and stability improvements and bug fixes, each with an open-source reference if available: --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-2-13/824 Title: [RELEASE] ScyllaDB Enterprise 2022.2.13 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2022.2.13, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.2 Feature Release.. Note the latest ScyllaDB Enterprise release is 202… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-2-13/824 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2022.2.13 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2022.2.13 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2022.2.13, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.2 Feature Release.. Note the latest ScyllaDB Enterprise release is 2023.1 LTS, and you are encouraged to upgrade to it in coordination with the ScyllaDB Support team. The following issues are fixed in this release (with an open-source reference, if available): --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-1-17/828 Title: [RELEASE] ScyllaDB 5.1.17 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.1.17, a bugfix release of the ScyllaDB 5.1 stable branch. ScyllaDB Open Source 5.1.17, like all past and future 5.x.y releases are backward compatible and support rollin… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-1-17/828 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.1.17 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.1.17 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.1.17, a bugfix release of the ScyllaDB 5.1 stable branch. ScyllaDB Open Source 5.1.17, like all past and future 5.x.y releases are backward compatible and support rolling upgrades. Note that the latest stable ScyllaDB Open Source release is 5.2, and you are encouraged to upgrade to it. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-23-2023-09-08/830 Title: Last week in scylla-cluster-tests.git master (issue #23; 2023-09-08) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the bbe1e291…ff871707 range are covered (Thu Aug 31 09:58:09 2023 +0200 - Wed Sep 6 23:35:13 20… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-23-2023-09-08/830 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #23; 2023-09-08) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #23; 2023-09-08) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the bbe1e291…ff871707 range are covered (Thu Aug 31 09:58:09 2023 +0200 - Wed Sep 6 23:35:13 2023 +0300). There were 14 non-merge commits from 9 authors in that period. Some notable commits: Because handling of deprecated “experimental” option was stopped, we specify “experimental_features” explicitly. This should address the failures in the upgrade tests and align with the current Scylla behavior. It’s also now possible to set experimental features with the SCT_EXPERIMENTAL_FEATURES variable. Due to flakiness in performance throughput tests i4i test concurrency was adjusted by increasing the number of loaders, number of c-s commands (processes) and number of threads per CS command GCE tests now also use the same load as for i4i throughput tests. We added a t3.micro artifact test. It required to activate and handle developer mode, add additional EBS volume for data and skip verify_xfs_online_discard_enabled check. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/question-about-ttl-tombstones/831 Title: Question about TTL & tombstones - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, the documenation Time to Live (TTL) and Compaction | ScyllaDB Docs states “If the data time stamp + gc_grace_seconds is less than or equal to the current time (now), the data is thrown away and a tombstone is not cr… Language: en Canonical URL: https://forum.scylladb.com/t/question-about-ttl-tombstones/831 ## Headings Structure: H1: Question about TTL & tombstones H3: Related topics ## Main Content: H1: Question about TTL & tombstones H3: Related topics the documenation Time to Live (TTL) and Compaction | ScyllaDB Docs states “If the data time stamp + gc_grace_seconds is less than or equal to the current time (now), the data is thrown away and a tombstone is not created.”. So basically the Cassandra fix [CASSANDRA-4917] Optimize tombstone creation for ExpiringColumns - ASF JIRA is already included in Scylla? So this means I will not have any tombstones if I do the following? This way I could track my deletes and write a range tombstone if I find many of my own app-level deletes. Sorry for asking, but I was not able to find the TTL logic in the Scylla codebase. Yes, that is correct. Expired cells do not necessarily create a tombstone upon expiry, only in the case where the expiry was less than gc grace period. See this commit: https://github.com/scylladb/scylladb/commit/a6f8f4fe24 Thanks for the confirmation. Also thanks for the commit, I was already looking for the code --- ### Page: https://forum.scylladb.com/t/share-your-database-experiences-and-insights-at-scylladb-summit/832 Title: Share Your Database Experiences and Insights at ScyllaDB Summit - Announcements - ScyllaDB Community NoSQL Forum Meta Description: [1200x6128-fb-virtual-workshop-768x402] Language: en Canonical URL: https://forum.scylladb.com/t/share-your-database-experiences-and-insights-at-scylladb-summit/832 ## Headings Structure: H1: Share Your Database Experiences and Insights at ScyllaDB Summit H3: Share Your Database Experiences and Insights at ScyllaDB Summit H3: Related topics ## Main Content: H1: Share Your Database Experiences and Insights at ScyllaDB Summit H3: Share Your Database Experiences and Insights at ScyllaDB Summit H3: Related topics Share your database experiences and insights at ScyllaDB Summit 24– and join the ranks of distinguished speakers like Discord, Disney+ Hotstar, Epic Games, Palo Alto Networks, and more. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-196-2023-09-10/834 Title: Last week in scylladb.git master (issue #196; 2023-09-10) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the cf37ab96f4…0656810c28 range are covered. There were 62 non-merge commits from 11 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-196-2023-09-10/834 ## Headings Structure: H1: Last week in scylladb.git master (issue #196; 2023-09-10) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #196; 2023-09-10) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the cf37ab96f4…0656810c28 range are covered. There were 62 non-merge commits from 11 authors in that period. Some notable commits: The index cache caches the Index.db components. It was previously disabled by default due to regressions on small-partition workloads. It is now enabled by default, with its memory usage capped at 20% of cache memory. This should improve out-of-the-box large partition performance. The recently introduced stream_plan_ranges_percentage configuration item was renamed to stream_plan_ranges_fraction to conform to naming conventions. The Raft leadership monitor is now started during normal node start, not only bootstrap. There is a new guardrail on replication factor, generating a warning if the replication factor is lower than 3. The --experimental flag was removed. It was replaced some time ago with --experimental-features., which provides fine-grained control about which experimental features are enabled. A crash if the chunk_len table parameter was set to 0 was fixed. Change Data Capture (CDC) updates its view of topology from time to time. It now does so in the background, to avoid slowing down topology changes. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/how-to-roll-back-from-5-1-to-4-6/836 Title: How to roll back from 5.1 to 4.6? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: The default SStable format used in version 5.1 is me. However, 4.3, 4.4, 4.5, 4.6, and 5.0 default to md. So when I upgrade all nodes in the cluster to 5.1, wrote some data, and finally rolled back to 4.6, the following … Language: en Canonical URL: https://forum.scylladb.com/t/how-to-roll-back-from-5-1-to-4-6/836 ## Headings Structure: H1: How to roll back from 5.1 to 4.6? H3: Related topics ## Main Content: H1: How to roll back from 5.1 to 4.6? H3: Related topics The default SStable format used in version 5.1 is me. However, 4.3, 4.4, 4.5, 4.6, and 5.0 default to md. So when I upgrade all nodes in the cluster to 5.1, wrote some data, and finally rolled back to 4.6, the following error occurred: Although I have already recovered all tables in keyspace system and system_schema Finally, I tried to clean up all the me format SStable. It still reports the following error: The official document provides a rollback method from 5.1 to 5.0. How does it avoid SStable formatting errors from me to md? What is our upgrade order? From 4.6 to 5.0 and then to 5.1? Then perform a reverse rollback. Rollbacks are only supported when at least one node in the cluster is still on the old version. Once all nodes are upgraded, new features, specific to the new version are enabled, and there is no way back anymore. This is what you see: the cluster enabled the me sstable format once all nodes were upgraded to 5.1. Past versions lack this feature and therefore rollback is not possible anymore. On a related note, why do you want to roll back? 5.0 is not supported anymore. --- ### Page: https://forum.scylladb.com/t/scylladb-error-scylla-476058/838 Title: ScyllaDB error - scylla[476058] - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello All, is there anyone that has had the following errors before, I need help on resolving the issue – scylla[476058]: [shard 0] cql_server_controller - Starting listening for CQL clients on XXX.XX.XX.XXX:9042 (une… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-error-scylla-476058/838 ## Headings Structure: H1: ScyllaDB error - scylla[476058] H3: Related topics ## Main Content: H1: ScyllaDB error - scylla[476058] H3: Related topics is there anyone that has had the following errors before, I need help on resolving the issue – scylla[476058]: [shard 0] cql_server_controller - Starting listening for CQL clients on XXX.XX.XX.XXX:9042 (unencrypted, non-shard-aware) scylla[476058]: [shard 0] cql_server_controller - Starting listening for CQL clients on XXX.XX.XX.XXX:19042 (unencrypted, shard-aware) scylla[476058]: [shard 0] init - serving scylla[476058]: [shard 0] init - Scylla version 5.2.6-0.20230730.58acf071bf28 initialization completed. scylla[476058]: [shard 0] service_level_controller - update_from_distributed_data: failed to update configuration for more than 90 seconds : exceptions::read_timeout_exception (Operation timed out for system_distributed.service_levels - received only 0 responses from 1 CL=ONE.) There’s an OSS GH issue for this matter, which I recommend to follow. Engineering is still investigating --- ### Page: https://forum.scylladb.com/t/integration-with-prestodb-and-metbase-hands-on/840 Title: Integration with PrestoDB and Metbase Hands-on - Blog Posts - ScyllaDB Community NoSQL Forum Meta Description: Maxim and I wrote a blog post about integrating ScyllaDB with Presto and how it can be used for Data Analytics. It includes a hands-on code example you can run yourself. Any input or questions? This would be a good pla… Language: en Canonical URL: https://forum.scylladb.com/t/integration-with-prestodb-and-metbase-hands-on/840 ## Headings Structure: H1: Integration with PrestoDB and Metbase Hands-on H3: Related topics ## Main Content: H1: Integration with PrestoDB and Metbase Hands-on H3: Related topics Maxim and I wrote a blog post about integrating ScyllaDB with Presto and how it can be used for Data Analytics. It includes a hands-on code example you can run yourself. Any input or questions? This would be a good place to discuss this. --- ### Page: https://forum.scylladb.com/t/can-i-use-some-kind-of-client-like-sql-server-management-studio-to-visualize-the-data-in-scylladb/841 Title: Can I use some kind of client like SQL Server Management Studio to visualize the data in ScyllaDB? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I’m wondering if you can use it, and if so, how? Language: en Canonical URL: https://forum.scylladb.com/t/can-i-use-some-kind-of-client-like-sql-server-management-studio-to-visualize-the-data-in-scylladb/841 ## Headings Structure: H1: Can I use some kind of client like SQL Server Management Studio to visualize the data in ScyllaDB? H3: Related topics ## Main Content: H1: Can I use some kind of client like SQL Server Management Studio to visualize the data in ScyllaDB? H3: Related topics I’m wondering if you can use it, and if so, how? There are different third-party tools that integrate with ScyllaDB and can be used to visualize data. One of them is Hackolade. They recently showcased their tool in our data modeling masterclass. I’ve also heard of people using DBVisualizer, and there are many more out there. As a rule of thumb, if the visualization tool is compatible with Apache Cassandra, it will work with ScyllaDB. You can see a list of integrations and connectors on this documentation page. What about TablePlus? See TablePlus client support for ScyllaDB Cloud’s Database-as-a-Service (DBaaS) serverless - #3 by tzach Oh, this other post is mine: I had created it yesterday! Do you know whether DBVisualiser does support the new Serverless ScyllaDB Cloud free Cluster? @Furia I am unaware of any. The easiest path is to simply ask the developers of such tools to support ScyllaDB Cloud, which under the hood simply involves using the ScyllaDB driver and changing the connection string. DBVisualizer , and ther Hi, Guy, Does DBVisualizer support ScyllaDB Cloud’s Database-as-a-Service (DBaaS) serverless ? --- ### Page: https://forum.scylladb.com/t/where-are-scylladb-cloud-servers-located/843 Title: Where are ScyllaDB Cloud servers located? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Where are ScyllaDB Cloud servers located? Language: en Canonical URL: https://forum.scylladb.com/t/where-are-scylladb-cloud-servers-located/843 ## Headings Structure: H1: Where are ScyllaDB Cloud servers located? H3: Related topics ## Main Content: H1: Where are ScyllaDB Cloud servers located? H3: Related topics Where are ScyllaDB Cloud servers located? ScyllaDB Cloud servers can run in any combination of AWS or GCP regions. The cloud providers continuously add new regions, so if you see a region is missing from our list of supported regions, please reach out and we will make sure to add it. Also, you can learn more about launching a cluster in the Quick Start Guide. It’s also possible to start a free trial cluster and run labs from ScyllaDB University with that cluster. The Quick Wins lab is a good starting point. --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-2-8/845 Title: [RELEASE] ScyllaDB 5.2.8 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.2.8, a bugfix release of the ScyllaDB 5.2 stable branch. ScyllaDB Open Source 5.2.8, like all past and future 5.x.y releases, is backward compatible and supports rolling… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-2-8/845 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.2.8 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.2.8 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.2.8, a bugfix release of the ScyllaDB 5.2 stable branch. ScyllaDB Open Source 5.2.8, like all past and future 5.x.y releases, is backward compatible and supports rolling upgrades. Users are encouraged to upgrade to 5.2.8. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/unknown-disk-usage-of-var-lib-scylla/846 Title: Unknown disk usage of /var/lib/scylla - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: When we spin up a fresh EC2 using scylla AMI, the disk usage for /var/lib/scylla 12GB, with no data at all on that node. On a node with approx 3TB data the /var/lib/scylla usage is ~3.2TB(~200GB). scyllaadm@ip-127.0.0… Language: en Canonical URL: https://forum.scylladb.com/t/unknown-disk-usage-of-var-lib-scylla/846 ## Headings Structure: H1: Unknown disk usage of /var/lib/scylla H3: Related topics ## Main Content: H1: Unknown disk usage of /var/lib/scylla H3: Related topics When we spin up a fresh EC2 using scylla AMI, the disk usage for /var/lib/scylla 12GB, with no data at all on that node. On a node with approx 3TB data the /var/lib/scylla usage is ~3.2TB(~200GB). scyllaadm@ip-127.0.0.1:~$ df -h Filesystem Size Used Avail Use% Mounted on /dev/nvme1n1 6.9T 3.2T 3.8T 46% /var/lib/scylla scyllaadm@ip-127.0.0.1:/var/lib/scylla$ du -sh * 87G commitlog 0 coredump 2.9T data 0 hints 1.0K logs 0 saved_caches 0 view_hints What’s causing this extra disk space on the Scylla AWS AMI? There are no snapshots, the data size from nodetool status and du -sh /var/lib/scylla does not match. This behaviour is observed on the multiple nodes or even on the multiple clusters. When a new node is added to an existing cluster, it will start receiving data from other nodes right away, as part of its bootstrap. Even if the node is started up all alone, there are many internal tables created, although 12GB for these is too much. As for the other case, the difference seems to mainly come from commitlog. This contains data that is currently held up in memtables. This will be removed once those memtables are flushed. The node 1 I’ve mentioned is the single node cluster, is not part of any existing nodes. For the other case I’ve shown, even if you add commitlog size, there is still gap of 200GB. --- ### Page: https://forum.scylladb.com/t/release-scylladb-cloud-14-september-2023/847 Title: [RELEASE] ScyllaDB Cloud - 14 September 2023 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: We are happy to announce a new update to the ScyllaDB Cloud service. New features and enhancements: We’ve added support for the il-central-1 Israel (Tel Aviv) region. Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-cloud-14-september-2023/847 ## Headings Structure: H1: [RELEASE] ScyllaDB Cloud - 14 September 2023 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Cloud - 14 September 2023 H3: Related topics We are happy to announce a new update to the ScyllaDB Cloud service. New features and enhancements: --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-24-2023-09-15/849 Title: Last week in scylla-cluster-tests.git master (issue #24; 2023-09-15) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 4a1a93aa…1eb27696 range are covered. There were 16 non-merge commits from 8 authors in tha… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-24-2023-09-15/849 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #24; 2023-09-15) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #24; 2023-09-15) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 4a1a93aa…1eb27696 range are covered. There were 16 non-merge commits from 8 authors in that period. Some notable commits: SCT now supports AWS placement groups - we can create instances that are in the same rack, and thus have lower latency between them. This is useful for performance testing and is used in one of them. Also note, performance tests now use separate step for provisioning resources, like it is in typical longevities. Up until now, we were saving only the Scylla version and the git sha of the commit in ES, making chart creation less clear or logical. This is now improved by adding commit dates to order the latest entries in our charts. We added test cases for low and asymmetric loads to simulate a low load happening during repair processes, to cover potential overhead like in scylladb/scylladb#14093. New utility was created to easily create Scylla API call commands. Following issue investigation on grow and shrink cluster, we added waiting for native_transport to be available after adding a node. This is to avoid the situation where the cluster is not yet ready to serve requests, and thus the test fails. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/how-does-scylladb-compare-to-accumulo-in-terms-of-performance/852 Title: How does ScyllaDB compare to Accumulo in terms of performance? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Can you provide any details on ScyllaDB vs. Apache Accumulo performance? Language: en Canonical URL: https://forum.scylladb.com/t/how-does-scylladb-compare-to-accumulo-in-terms-of-performance/852 ## Headings Structure: H1: How does ScyllaDB compare to Accumulo in terms of performance? H3: Related topics ## Main Content: H1: How does ScyllaDB compare to Accumulo in terms of performance? H3: Related topics Can you provide any details on ScyllaDB vs. Apache Accumulo performance? Apache Accumulo is OSS based on Google Bigtable. I’m not aware of specific testing vs Accumulo. The benchmark tests vs. Bigtable will provide a good indication. Accumulo is also written in Java and has Java overhead similar to Cassandra (Garbage collection pauses, JVM tuning…). --- ### Page: https://forum.scylladb.com/t/facing-a-trouble-with-inserting-into-tables-from-kafka-using-scylla-sink-connector/853 Title: Facing a trouble with inserting into tables from kafka using scylla-sink-connector - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: So, i need to catch up messages from kafka in binary datatype (don’t know actual encoding) and insert them into table using scylla-sink-connector. The main point is that i don’t know which serialization format used (jus… Language: en Canonical URL: https://forum.scylladb.com/t/facing-a-trouble-with-inserting-into-tables-from-kafka-using-scylla-sink-connector/853 ## Headings Structure: H1: Facing a trouble with inserting into tables from kafka using scylla-sink-connector H3: Related topics ## Main Content: H1: Facing a trouble with inserting into tables from kafka using scylla-sink-connector H3: Related topics So, i need to catch up messages from kafka in binary datatype (don’t know actual encoding) and insert them into table using scylla-sink-connector. The main point is that i don’t know which serialization format used (just getting key as string and values as a blob). I’ll appreciate if someone could help me. P.S. when I try to use JsonConverter it gives an error that i should use valuetokey transformation. However, I cannot store data in json because of Object of type bytes is not JSON serializable The ScyllaDB Sink Connector accepts two data formats from kafka. They are: Without knowing how the data was serialized it may prove difficult to insert it into the table in a meaningful way. After all, the connector needs to know how to divide the data into columns. First thing you could do would be try to trace back where did this data come from and look for clues what serialization format was used. If that is not possible, you can blindly try using supported converters if they return any meaningful information. Quick way to do that would be for example using console consumer to peek what is there on your topic: Replace localhost with your host (and port). You can also set separate converter for the key and try with schemas enabled. If you are using “Confluent flavour” of Kafka, it usually contains avro console consumer which you can use to test for Avro: You may also have to include --property schema.registry.url=http://schema-registry-host:port. If any of those options does not return binary gibberish, then you have likely found converter to use. You can try also other converters than those for formats officially supported by the connector just to see what works, however connector itself probably won’t work with them. If that still leads you nowhere you can always try to transform the data on the topic. For example you could take the binary data, convert it to Base64 and keep it on another topic serialized as JSON with a single String field. Connector should handle that by copying it as a single column with text data. Other option would be converting it to Avro. If I recall correctly that would not require conversion to Base64. I don’t know if that would be good enough for your use case though. You mentioned that after trying JSON Converter you’ve encountered the error telling you to use ValueToKey transform. That’s possibly because the key is empty. If it turns out the data is readable with JSON Converter then you have to do just that, by specifying which fields to use in the key field. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-197-2023-09-18/854 Title: Last week in scylladb.git master (issue #197; 2023-09-18) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 0656810c28…4eb4ac4634 range are covered. There were 187 non-merge commits from 16 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-197-2023-09-18/854 ## Headings Structure: H1: Last week in scylladb.git master (issue #197; 2023-09-18) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #197; 2023-09-18) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 0656810c28…4eb4ac4634 range are covered. There were 187 non-merge commits from 16 authors in that period. Some notable commits: Tablets are a new, experimental way of distributing data across a cluster. We now support decommissioning with tablets enabled. First all tablets are evicted from the node, then the regular decommissioning process continues. The scylla executable can now act as nodetool by executing scylla nodetool . For now only scylla nodetool compact is implemented. The cdc_generations_v3 table stores internal information about Change Data Capture streams, when consistent cluster topology is enabled. Its schema has been changed to allow for efficiently trimming older and unneeded topology information. There is a new REST API call to recalculate schema digests. It can be useful to heal some schema disagreement problems. The bundled Prometheus node_exporter has been updated to version 1.6.1. There are new metrics exposed by the S3 object storage driver. Some internal tables were moved from the general commitlog to the private schema commitlog. As a result their memtables are flushed less often, reducing latency for topology changes. When ALTERing a table, the compaction strategy options are now validated. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/when-is-scylladb-5-3-out-it-has-already-passed-all-the-milestones-and-completed/856 Title: When is scylladb 5.3 out? it has already passed all the milestones and completed - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: was waiting for 5.3 for quite some time. when will it be out? sry, just really curious on the date of release. Language: en Canonical URL: https://forum.scylladb.com/t/when-is-scylladb-5-3-out-it-has-already-passed-all-the-milestones-and-completed/856 ## Headings Structure: H1: When is scylladb 5.3 out? it has already passed all the milestones and completed H3: Related topics ## Main Content: H1: When is scylladb 5.3 out? it has already passed all the milestones and completed H3: Related topics was waiting for 5.3 for quite some time. when will it be out? sry, just really curious on the date of release. ScyllaDB Open Source 5.3 won’t be graduated for production (5.3.0). We will soon start the process for the release of ScyllaDB Open Source 5.4. --- ### Page: https://forum.scylladb.com/t/read-concurrency-semaphore-p99-read-latency/857 Title: Read_concurrency_semaphore & p99 read latency - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Currently chasing p99 read latency in our applications that queries time-series data on period of time. Useful information : Scylla Open Source 5.1.14-0.20230716.753c9a4769be 16 core / 128 GB RAM RAID0 SSD 3 TB 3 no… Language: en Canonical URL: https://forum.scylladb.com/t/read-concurrency-semaphore-p99-read-latency/857 ## Headings Structure: H1: Read_concurrency_semaphore & p99 read latency H3: Related topics ## Main Content: H1: Read_concurrency_semaphore & p99 read latency H3: Related topics Currently chasing p99 read latency in our applications that queries time-series data on period of time. Useful information : Scylla Open Source 5.1.14-0.20230716.753c9a4769be 16 core / 128 GB RAM RAID0 SSD 3 TB 3 nodes Application (using golang gocqlx driver) query are pretty straight forward, it queries data by aggregation (always 1 row returned, using basic agg. functions) on a given time-range, from current time minus P to current time, where P is in the range 1min - 15min. (so always query most recent data) P99 read latency on the scylla dashboard is low, (1ms - 10ms). On the gocql latency report, P95 is between 1ms-10ms BUT the p99 is very high : from 700ms to 1s ! I’ve tried many things : code profile, tracing, but it seems that somehow ONE request amongst several take forever to run (we issue a batch of read every 1 second to aggregate data). I’ve noticed those exceptions in the scylla-server logs : However it’s not in direct correlation in terms of timestamp with those p99 spikes in the application i.e spikes are every 5/10 s whereas those exceptions occurs less frequently. To add more context in terms of volumetry, it’s very low, I expected scylla to handle it easily, 6k writes / s , 2k read / s It’s an isolated cluster so I can pinpoint the root cause, our prod cluster handles more req/s ( ~ 50/130k writes and 8 / 20k reads) If anyone have an idea of where to look to eliminate those p99 latency spike that would be very helpful. The window size of 1hour seems quite low. When using TWCS, ScyllaDB never compacts sstables across windows, so a small window size can lead to a huge number of sstables piling up and adversely impacting reads that have to touch multiple windows. Also, just the sheer amount of sstables will take ever larger share of memory, squeezing out cached content. What timeout for reader_concurrency_semaphore? What is result for queries in the dumping permit diagnostics? Are the queries dropped? Only recent data (max 15min in the past) gets queried here, I’ve changed the window size to 4 hours but that not to affect anything Currently inspecting slow queries logging to try to better understand what is going on Sorry but could you rephrase your question please ? Not sure to fully understand what are you asking. What timeout for reader_concurrency_semaphore? The same as the timeout of the query. This is set in the config. What is result for queries in the dumping permit diagnostics? Are the queries dropped? Timed out queries are dropped, the others proceed as usual. Timed out queries are dropped, the others proceed as usual. For example for shard with read_request_timeout_in_ms: 5000: Yes, if the reads are cache reads, then this is what happens. Once a read goes to disk, another read can get started in parallel. Timeout is 5s for both server & client. Queries does not seem to be dropped, As I explained client receives results but latency is very high (between 500ms and 900ms). I see those rogue queries in my application log, it’s only 1/40 every 1 second. The workload in the test cluster has been reduced to the strict minimum for the application so we have only between 2k/3k writes and very few reads : 200 ops/s. Problem still persists and I can find anything helpful in the slow query log. Do you do deletes? The typical explanation for such rouge reads used to be tombstones. This was fixed in 5.2, upgrading to 5.2 might solve your problem. Another typical explanation is stalls. Do you see any stalls in the logs? Data is never deleted, only TTL’ed, we insert it (never updates) in an append-only fashion. The only information that is given in the logs are those “reader_concurrency_semaphore” one. TTL’d data also generates a tombstone when deleted. How long is your TTL and gc grace period respectively? Also, how does the queries look like, especially the one which times out? You can find schema definition in the initial post, The hour field is time bucket of the datetime hour. Queries being run look all the same : Where the time range is always current_time / current_time - X sec. (so we are always requesting the same hour bucket, maybe we are in a hot-partition scenario ?) I’m testing a new version where I only request the last second, and add truncated second to a field into the cluster key so I can GROUP BY second, but it’s giving me pretty much the same results. For the record we have a distinct (exchange_code, symbol) of ~ 13k. I’m thinking about using a consistent hash value of the id field, modulo # of shard (16) and add it to the partition key for better distribution and avoid hot partition/ shard load. Hot partition is a possibility indeed. This query will read an entire partition, which should be fine on its own. You can look for hot partitions by going to the Detailed dashboard in monitoring and switching to by Instance,shard mode. If you see a particular shard/node combo being significantly more load than others, that is a good sign that there is a hot partition involved. --- ### Page: https://forum.scylladb.com/t/scylladb-vs-aerospike/858 Title: ScyllaDb vs Aerospike - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I read the whitepaper about comparing ScyllaDb vs. Aerospike ScyllaDB White Paper | ScyllaDB vs. Aerospike: A NoSQL Database Performance Comparison, and it claims to overperform Aerospike from 30 to 40 percent. And I hav… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-vs-aerospike/858 ## Headings Structure: H1: ScyllaDb vs Aerospike H3: Related topics ## Main Content: H1: ScyllaDb vs Aerospike H3: Related topics I read the whitepaper about comparing ScyllaDb vs. Aerospike ScyllaDB White Paper | ScyllaDB vs. Aerospike: A NoSQL Database Performance Comparison, and it claims to overperform Aerospike from 30 to 40 percent. And I have some questions: --- ### Page: https://forum.scylladb.com/t/scylla-support-for-object-storage-as-a-data-archival-option/860 Title: Scylla Support for Object Storage as a data Archival option - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I currently work in Fintech. Does Scylla support Object storage for data archival? Say for example, you’re a financial company which stores transactions on Scylla. With this use-case, you might want to archive financial… Language: en Canonical URL: https://forum.scylladb.com/t/scylla-support-for-object-storage-as-a-data-archival-option/860 ## Headings Structure: H1: Scylla Support for Object Storage as a data Archival option H3: Related topics ## Main Content: H1: Scylla Support for Object Storage as a data Archival option H3: Related topics I currently work in Fintech. Does Scylla support Object storage for data archival? Say for example, you’re a financial company which stores transactions on Scylla. With this use-case, you might want to archive financial transactions which are older than 6 years, for example. Is this currently possible using ScyllaDB? I’ve seen it with new players in the database field which allows companies to save on storage costs of NVMEs for data that is infrequently accessed Hi @Isaac_Kinuthia Quick answer: Yes, No,Yes --- ### Page: https://forum.scylladb.com/t/commit-log-schema-separation/861 Title: Commit Log Schema Separation - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: We recently noticed commitlog changes for new nodes, after upgrading to 5.1, and I’m not sure what to make of it. There is a new configuration force_schema_commit_log that says it will enable a separate commitlog immedi… Language: en Canonical URL: https://forum.scylladb.com/t/commit-log-schema-separation/861 ## Headings Structure: H1: Commit Log Schema Separation H3: Related topics ## Main Content: H1: Commit Log Schema Separation H3: Related topics We recently noticed commitlog changes for new nodes, after upgrading to 5.1, and I’m not sure what to make of it. There is a new configuration force_schema_commit_log that says it will enable a separate commitlog immediately when enabled, as compared to enabling it on the first reboot. This is set to true by default. I however, cannot find any documentation for what this commitlog separation on first boot is or why we need it, or even if I can enable/disable it, outside of the first boot. In the 5.1.0 release notes the main references to schema changes are Raft, which is still experimental, so I assume this isn’t Raft related, or it wouldn’t default to true. Should we have done a full cluster reboot after the 5.1 upgrade to enable this feature, and then this is all a moot point? How can I verify this is enabled on a node, besides commitlog errors on boot? You’ll have files of the form /var/lib/scylla/commilog/SchemaLog*.log. No action is required on your part. The idea behind it is to give schema and topology a commitlog that is more durable (commitlog_sync = batch instead of periodic), and to give schema and topology higher priority over your writes, so you can’t flood the database with writes and watch it malfunction as its internal metadata operations are delayed indefinitely. Ok, so we don’t even need to reboot after the cluster was upgraded to 5.1? It should already be active on all nodes due to the reboot while upgrading? I checked a few nodes that haven’t been restarted and didn’t see any of the mentioned files. Though ones that were restarted I do see. Should I just give everything a rolling reboot to make sure it’s enabled? Perhaps it does need a rolling restart, though it’s okay to delay it. --- ### Page: https://forum.scylladb.com/t/new-data-resurrection-without-cleanups/862 Title: New(?) Data Resurrection Without Cleanups - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: There was a change in the Adding a New Node Into an Existing ScyllaDB Cluster (Out Scale) documentation for 5.2, that says there is a chance of data resurrection if cleanups are not run in a timely manner. Is this relat… Language: en Canonical URL: https://forum.scylladb.com/t/new-data-resurrection-without-cleanups/862 ## Headings Structure: H1: New(?) Data Resurrection Without Cleanups H3: Related topics ## Main Content: H1: New(?) Data Resurrection Without Cleanups H3: Related topics There was a change in the Adding a New Node Into an Existing ScyllaDB Cluster (Out Scale) documentation for 5.2, that says there is a chance of data resurrection if cleanups are not run in a timely manner. Is this related to the data resurrection for not repairing often enough, or are they unrelated? My understanding in the past was that cleanups were just to recover disk space, so just trying to understand what/if anything changed with it or the full risks of not doing it between node additions and removals. The primary goal of cleanup is avoiding data resurrection. Consider a write W1 which is written to a node, N1. After a new node Nx is added, N1 no longer owns W1. If no cleanup is run, this data will stay on N1. Some time down the line W1 is deleted and the tombstone is garbage collected. Then at one point Nx is removed from the cluster and the ownership of W1 comes back to N1. Remember that this write was deleted earlier, but the tombstone was garbage collected. Since there is currently no newer entry for W1, the old value N1 has, becomes the latest value and therefore it is resurrected. To avoid this we run cleanup, which ensures that no such stale data lingers on nodes after token movement. Freeing up disk space is a secondary, albeit also important aspect. Thanks for the info Botond. I assume this has always been a possibility then, and not something new? This has some impact to our process around AWS node decommissions that we’ll have to figure out. AWS gives around 2 weeks notice that a node will be decommissioned, so we bootstrap a new node into the cluster, decommission the old one, then worry about cleanups. With this process we would need to insert the cleanups in the middle and need to get it done within the 2 week window, which could be cutting it close on some clusters (our largest has 39 nodes). Do you have any recommendation on how you would handle that situation? Decommissioning the node first would fix it, though then we have to undersize the cluster for a short window instead of oversizing it. Doing bootstrap + decomission in quick succession, then doing cleanup after should be fine. Just don’t delay running cleanup too long. Also, maybe look into replace operation. With replace, no cleanup is needed, although it has its drawback in that the cluster temporarily looses a replica, while the replace is going on, so read QUORUMs are more susceptible to failing. Interesting, I’m fairly sure that is new as well. We used to use the dead node replacement procedure, but ran into issues with data loss (on 4.X, before RBNO was enabled for replace actions). I can’t find the old docs, but I have notes saying the bootstrap/decommission was recommended at the time, so we switched to that. If the replace is considered as a valid alternative then we may switch back. Thank you. Yes, replace was made safe by using RBNO for it. This change was made in 4.6, where we enabled RBNO by default for replace. See ScyllaDB Open Source 4.6 - ScyllaDB. --- ### Page: https://forum.scylladb.com/t/release-scylladb-monitoring-stack-4-4-5/864 Title: [RELEASE] ScyllaDB Monitoring Stack 4.4.5 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce a patch release of ScyllaDB Monitoring Stack 4.4.5 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB Enterprise and ScyllaDB Open Source, based on Prometh… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-monitoring-stack-4-4-5/864 ## Headings Structure: H1: [RELEASE] ScyllaDB Monitoring Stack 4.4.5 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Monitoring Stack 4.4.5 H3: Related topics The ScyllaDB team is pleased to announce a patch release of ScyllaDB Monitoring Stack 4.4.5 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB Enterprise and ScyllaDB Open Source, based on Prometheus and Grafana. ScyllaDB Monitoring Stack 4.4.5 supports: The patch release adds ScyllaDB open source 5.4.x support and fixes the following bugs: --- ### Page: https://forum.scylladb.com/t/getting-scylla-cluster-started-with-docker-noob-questions/866 Title: Getting Scylla cluster started with docker (Noob Questions) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello, I am litterally new to Scylla, I just come from a mongodb world. I’ve read the docs and I plan to fire up Scylla and start using it for a demanding rw application. I have 2 Vps with 4 Core 24 Gb ram Each. I kin… Language: en Canonical URL: https://forum.scylladb.com/t/getting-scylla-cluster-started-with-docker-noob-questions/866 ## Headings Structure: H1: Getting Scylla cluster started with docker (Noob Questions) H3: Related topics ## Main Content: H1: Getting Scylla cluster started with docker (Noob Questions) H3: Related topics Hello, I am litterally new to Scylla, I just come from a mongodb world. I’ve read the docs and I plan to fire up Scylla and start using it for a demanding rw application. I have 2 Vps with 4 Core 24 Gb ram Each. I kind of wanted to make a Scylla cluster out of them for testing. I have the following questions: How can I make one vps node join the other in the cluster? (Basically I dont know how to make a cluster with 2 vps using docker. From what I had read, only local nodes were described in the documentation, and I dont know how to setup a connection between nodes running on separated instances) How can I enable and setup authentification between those nodes? (I want to secure everything with an user:pass) How should I connect to the database using cassandra-driver in nodejs? (This one is a noob question, but I do not understeand what code should I use to connect to my vps’s IP. I saw that the Client method takes some contact points and datacenter name, but I cannot understeand why there is no ip and even an user:password specificated. I come from a space where a connection to a db is made via ip, user, password and database name) Again sorry for those beginner questions, but I really want to use ScyllaDB for my future projects.If someone can guide me a little bit into this, I would be grateful. How can I make one vps node join the other in the cluster? (Basically I dont know how to make a cluster with 2 vps using docker. From what I had read, only local nodes were described in the documentation, and I dont know how to setup a connection between nodes running on separated instances) To setup a connection between two scylla nodes, you need to configure them as following (update /etc/scylla/scylla.yml): Also make sure you expose all the ports scylla uses. There are many-many more nuances to configuring scylla, the above is just the absolute basics, to get a cluster up so you can start experimenting. I recommend you look into courses in scylla university, in particular those around administering scylla. How can I enable and setup authentification between those nodes? (I want to secure everything with an user:pass) I’m not sure what you mean by this. Scylla nodes don’t use auth when talking between themselves. Securing a scylla cluster in this regard is usually done by using ip addresses which are only accessible from within the cluster. See $INTERNAL_IP from my reply above. If you mean client authentication, then look at Enable Authentication | ScyllaDB Docs. How should I connect to the database using cassandra-driver in nodejs? (This one is a noob question, but I do not understeand what code should I use to connect to my vps’s IP. I saw that the Client method takes some contact points and datacenter name, but I cannot understeand why there is no ip and even an user:password specificated. I come from a space where a connection to a db is made via ip, user, password and database name) I am not familiar with the nodejs driver, but in python (the driver I’m familiar with) those contact points are ip addresses in fact. Thank you very much for dedicating time into aswering my questions. I now understend how everything works. If I have further problems I will post here. Many thanks again! --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-25-2023-09-23/867 Title: Last week in scylla-cluster-tests.git master (issue #25; 2023-09-23) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the e5f2055d…730bd7a4 range are covered. There were 17 non-merge commits from 7 authors in tha… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-25-2023-09-23/867 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #25; 2023-09-23) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #25; 2023-09-23) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the e5f2055d…730bd7a4 range are covered. There were 17 non-merge commits from 7 authors in that period. Some notable commits: Scylla’s setup scripts now produce debug logs, we started collecting them into setup_scripts_errors.log. The file contains more detail of traceback with current value of variables, it helps to fix bugs which are hard to reproduce. SCT now supports using dns names in cluster. So listen, rpc and seed addresses use dns names instead of ip’s and it was enabled in one of our weekly tests. Scale test with large number of nodes has been refactored: See docs explaining how to use OKTA with created aws profile Fixed selecting random regions in pipelines instead of specifying it directly. Applied better firewall rules to close external network access for GCE. New pipelines created for EaR+KMS: 1d longevity , 6h multiDC longevity and perf CI jobs. We reverted the change related to collecting stats on truncate duration in upgrade tests because AdaptiveTimeout machinery was causing long delays in tests and failing them due to timeouts. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-198-2023-09-24/868 Title: Last week in scylladb.git master (issue #198; 2023-09-24) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 4eb4ac4634…99d83808cc range are covered. There were 124 non-merge commits from 16 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-198-2023-09-24/868 ## Headings Structure: H1: Last week in scylladb.git master (issue #198; 2023-09-24) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #198; 2023-09-24) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 4eb4ac4634…99d83808cc range are covered. There were 124 non-merge commits from 16 authors in that period. Some notable commits: A security vulnerability that allowed table permissions to be taken over by an unauthorized user has been fixed. Without Raft-managed consistent schema, each node evolves its own schema and the cluster attempts to merge each node’s schema to produce a consistent outcome. To see if a consistent outcome was achieved, each node calculates a hash of its schema and then these hashes are compared. This is a fragile computation and sometimes leads to schema disagreement even though there is no real difference between nodes. This was carried over to raft-managed consistent schema, but now when Raft is enabled the leader will decree a schema version and other nodes will learn it from the leader. This should eliminate schema disagreement problems when Raft consistent schema is enabled. Change Data Capture (CDC) funnels change information to a number of streams, which are dependent on current topology. As topology changes the streams change, and their identity is maintained in the CDC generations table. The table is now garbage-collected to remove old streams. A reactor stall when creating CDC streams for large clusters has been fixed. On startup, partially created sstables stored on object storage (e.g. S3) are now removed. The S3 driver can now use temporary credentials. The S3 driver now tracks its memory usage to avoid running out of memory. We now abort running repairs on nodetool drain commands. In consistent topology mode, the leader now prevents the previous leader from affecting the cluster before starting its own changes. Off-strategy compaction is used to reshape sstables that weren’t created from memtable flushes or regular compaction, for example repair or bootstrap. It now uses incremental compaction for run-based compaction strategies (Leveled or Incremental compaction strategies), reducing temporary storage requirements. A bug in the fromJson() CQL function when operating on NULL operands has been fixed. Compaction strategy options are now validated earlier. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylla-manager-3-2-1/870 Title: [RELEASE] Scylla Manager 3.2.1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces the release of ScyllaDB Manager 3.2.1 production-ready ScyllaDB Manager patch release of the stable ScyllaDB Manager 3.2 branch. ScyllaDB Manager is a centralized cluster administration and r… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-manager-3-2-1/870 ## Headings Structure: H1: [RELEASE] Scylla Manager 3.2.1 H3: Repair changes H3: Monitoring H3: Related topics ## Main Content: H1: [RELEASE] Scylla Manager 3.2.1 H3: Repair changes H4: Extended output of sctool cluster list H4: New ScyllaDB Manager Agent metrics H3: Monitoring H3: Related topics The ScyllaDB team announces the release of ScyllaDB Manager 3.2.1 production-ready ScyllaDB Manager patch release of the stable ScyllaDB Manager 3.2 branch. ScyllaDB Manager is a centralized cluster administration and recurrent tasks automation tool. Scylla Manager 3.2.1 includes minor changes to repair task schedule and parameters validation and allows checking if the Manager has CQL credentials to manage a cluster. ScyllaDB Enterprise customers are encouraged to upgrade to ScyllaDB Manager 3.2.1 in coordination with the ScyllaDB support team. The new release includes upgrades of both ScyllaDB Manager Server and Agent. Scylla Manager 3.2.1 support the following Scylla Enterprise releases: And the following Open Source release (limited to 5 nodes see license): You can install and run Scylla Manager on Kubernetes using Scylla Operator. More here. ScyllaDB Manager 3.2.1 contains minor repair changes: The output of this command has been extended to indicate whether the Manager has CQL credentials for a managed cluster #3531 . Manager use CQL to ensure repair stability by repairing base tables before Materialized Views and Secondary Indexes (see docs). It is recommended to set credentials via sctool cluster update command. You can use Scylla Monitoring releases 4.4.4 and later to monitor Scylla Manager 3.2.1 --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2023-1-1/871 Title: [RELEASE] ScyllaDB Enterprise 2023.1.1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2023.1.1, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2023.1 LTS Release.. This patch release enables support of KMS integration f… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2023-1-1/871 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2023.1.1 H3: Amazon KMS Integration for Encryption at Rest H3: FIPS Tolerant H3: Bug fixes H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2023.1.1 H3: Amazon KMS Integration for Encryption at Rest H3: FIPS Tolerant H3: Bug fixes H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2023.1.1, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2023.1 LTS Release.. This patch release enables support of KMS integration for Encryption at Rest, allows ScyllaDB to work with a FIPS-enabled Ubuntu, and fixes multiple minor bugs. You are encouraged to upgrade to it in coordination with the ScyllaDB Support team. Scylla Enterprise has supported Encryption at Rest (EaR) for a long time. So far, one can store the keys for EaR locally, in an encrypted table, or an external KMIP server. Release 2023.1.1 adds the ability to use Amazon KMS keys. ScyllaDB can now use Customer Managed Key (CMK), stored in KMS, to create, encrypt, and decrypt Data Keys (DEK), which are then used to encrypt and decrypt the data in storage, such as SSTables, Commit logs, Batches, and hints logs. See AWS KMS concepts, Data Keys for more information Before using KMS, you need to set KMS as a key provider and validate that ScyllaDB nodes have permission to access and use the CMK you created in KMS. Once you do that, you can use the CMK in the CRETE and ALTER TABLE commands with KmsKeyProviderFactory, as follows Where “my_key” point to a section in scylla.yaml You can also use the KMS provider to encrypt System level data. See more examples and info here. ScyllaDB Enterprise can now run on FIPS enabled Ubuntu, using libraries that were compiled with FIPS enabled, like OpenSSL, GnuTLS, and more. The following issues are fixed in this release (with an open-source reference, if available): --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-1-18/872 Title: [RELEASE] ScyllaDB 5.1.18 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.1.18, a bugfix release of the ScyllaDB 5.1 stable branch. ScyllaDB Open Source 5.1.18, like all past and future 5.x.y releases are backward compatible and support rollin… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-1-18/872 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.1.18 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.1.18 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.1.18, a bugfix release of the ScyllaDB 5.1 stable branch. ScyllaDB Open Source 5.1.18, like all past and future 5.x.y releases are backward compatible and support rolling upgrades. Note that the latest stable ScyllaDB Open Source release is 5.2, and you are encouraged to upgrade to it. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/can-i-use-nodejs-to-interact-with-scylladb-any-examples/873 Title: Can I use NodeJS to interact with ScyllaDB? Any examples? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I’m using NodeJS for my project. I’d like to know how I can use it with ScyllaDB. Language: en Canonical URL: https://forum.scylladb.com/t/can-i-use-nodejs-to-interact-with-scylladb-any-examples/873 ## Headings Structure: H1: Can I use NodeJS to interact with ScyllaDB? Any examples? H3: Related topics ## Main Content: H1: Can I use NodeJS to interact with ScyllaDB? Any examples? H3: Related topics I’m using NodeJS for my project. I’d like to know how I can use it with ScyllaDB. Yes, it’s possible to use the open-source NodeJS driver with ScyllaDB. Notice that this driver is not shard-aware (like some other drivers, for example, the Rust driver and the Go driver). Some useful resources are: You can learn more about ScyllaDB drivers in general and about share-aware drivers in the Using Scylla Drivers course on ScyllaDB University. --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2021-1-23/874 Title: [RELEASE] ScyllaDB Enterprise 2021.1.23 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2021.1.23, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2021.1. Note that ScyllaDB Enterprise 2023.1 LTS is the latest Long-Term Su… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2021-1-23/874 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2021.1.23 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2021.1.23 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2021.1.23, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2021.1. Note that ScyllaDB Enterprise 2023.1 LTS is the latest Long-Term Support release. With 2023.1 LTS out, ScyllaDB enterprise 2021.1 support will be ended. You are encouraged to upgrade in coordination with the ScyllaDB support team. The following issue are fixed in this release (with an open-source reference): --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-2-9/875 Title: [RELEASE] ScyllaDB 5.2.9 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.2.9, a bugfix release of the ScyllaDB 5.2 stable branch. ScyllaDB Open Source 5.2.9, like all past and future 5.x.y releases, is backward compatible and supports rolling… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-2-9/875 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.2.9 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.2.9 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.2.9, a bugfix release of the ScyllaDB 5.2 stable branch. ScyllaDB Open Source 5.2.9, like all past and future 5.x.y releases, is backward compatible and supports rolling upgrades. Users are encouraged to upgrade to 5.2.9. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-199-2023-10-01/878 Title: Last week in scylladb.git master (issue #199; 2023-10-01) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 99d83808cc…1640f83fdc range are covered. There were 103 non-merge commits from 18 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-199-2023-10-01/878 ## Headings Structure: H1: Last week in scylladb.git master (issue #199; 2023-10-01) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #199; 2023-10-01) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 99d83808cc…1640f83fdc range are covered. There were 103 non-merge commits from 18 authors in that period. Some notable commits: The DESCRIBE statement now includes user defined types and functions. A rare crash when a SERVICE LEVEL is dropped has been fixed. When checking the bloom filter for a partition key, we now hash the key once, rather than for every sstable being checked. The column names for SELECT CAST(b AS int) and similar expressions have been adjusted to match Cassandra. In some cases where a bind variable was used both for the partition key and to match a non-key column, ScyllaDB would not generate correct partition key routing for the driver. This is now fixed. The internal feature negotiation between joining nodes and an exising cluster has been adjusted to account for the fact that features are now stored in Raft group 0. Raft snapshot update and commit log truncation are now atomic, removing a failure case. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-cloud-2-october-2023/879 Title: [RELEASE] ScyllaDB Cloud - 2 October 2023 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: We are happy to announce a new update to the ScyllaDB Cloud service. New features and enhancements: We’ve added support for the asia-south2 (Delhi) GCP region. You can now create Personal Tokens in your ScyllaDB … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-cloud-2-october-2023/879 ## Headings Structure: H1: [RELEASE] ScyllaDB Cloud - 2 October 2023 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Cloud - 2 October 2023 H3: Related topics We are happy to announce a new update to the ScyllaDB Cloud service. New features and enhancements: We’ve added support for the asia-south2 (Delhi) GCP region. You can now create Personal Tokens in your ScyllaDB Cloud accounts, in order to access the ScyllaDB Cloud API and the ScyllaDB Cloud Terraform provider. To add a Personal Token, navigate to Settings > Personal Tokens. From there, you’ll be able to generate Personal Tokens. If you have access to multiple ScyllaDB Cloud accounts, ensure you create a new personal token in each of the relevant accounts. This feature replaces legacy API keys. Legacy API Keys will be deprecated on January 1, 2024. Users who have an active legacy API key will receive an email with steps on how to generate a Personal Token in ScyllaDB Cloud. For more information, see our API docs. --- ### Page: https://forum.scylladb.com/t/data-modeling-workshop-at-devopsdays-tel-aviv/880 Title: Data Modeling Workshop at DevopsDays Tel Aviv - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: @tzach and I will host a workshop on data modeling and high availability at the upcoming DevopsDays event in Tel Aviv. It’s on the 29th of October, hope to see you there. Language: en Canonical URL: https://forum.scylladb.com/t/data-modeling-workshop-at-devopsdays-tel-aviv/880 ## Headings Structure: H1: Data Modeling Workshop at DevopsDays Tel Aviv H3: Related topics ## Main Content: H1: Data Modeling Workshop at DevopsDays Tel Aviv H3: Related topics @tzach and I will host a workshop on data modeling and high availability at the upcoming DevopsDays event in Tel Aviv. It’s on the 29th of October, hope to see you there. --- ### Page: https://forum.scylladb.com/t/introducing-database-performance-at-scale-a-free-open-source-book/882 Title: Introducing “Database Performance at Scale”: A Free, Open Source Book - Announcements - ScyllaDB Community NoSQL Forum Meta Description: [1200x628-data-performance-book (1)] Language: en Canonical URL: https://forum.scylladb.com/t/introducing-database-performance-at-scale-a-free-open-source-book/882 ## Headings Structure: H1: Introducing “Database Performance at Scale”: A Free, Open Source Book H3: Introducing “Database Performance at Scale”: A Free, Open Source Book H3: Related topics ## Main Content: H1: Introducing “Database Performance at Scale”: A Free, Open Source Book H3: Introducing “Database Performance at Scale”: A Free, Open Source Book H3: Related topics Discover new ways to optimize database performance and avoid common mistakes that impact latency and throughput. --- ### Page: https://forum.scylladb.com/t/scylladb-is-not-starting-on-a-private-subnet-server/884 Title: ScyllaDB is not starting on a private subnet server - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I am trying to setup 2 node cluster of ScyllaDB(version - 5.2.9-0.20230920.5709d0043978) on AWS EC2 Server (Ubuntu) which is in a private subnet. Node 1 IP - 10.0.137.63 Node 2 IP - 10.0.156.228 scylla.yaml → clust… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-is-not-starting-on-a-private-subnet-server/884 ## Headings Structure: H1: ScyllaDB is not starting on a private subnet server H3: Related topics ## Main Content: H1: ScyllaDB is not starting on a private subnet server H3: Related topics I am trying to setup 2 node cluster of ScyllaDB(version - 5.2.9-0.20230920.5709d0043978) on AWS EC2 Server (Ubuntu) which is in a private subnet. Node 1 IP - 10.0.137.63 Node 2 IP - 10.0.156.228 cluster_name: ‘Test Cluster’ num_tokens: 256 commitlog_sync: periodic commitlog_sync_period_in_ms: 10000 commitlog_segment_size_in_mb: 32 schema_commitlog_segment_size_in_mb: 32 seed_provider: - class_name: org.apache.cassandra.locator.SimpleSeedProvider parameters: - seeds: “10.0.137.63” listen_address: “10.0.137.63” broadcast_address: “10.0.137.63” native_transport_port: 9042 native_shard_aware_transport_port: 19042 read_request_timeout_in_ms: 5000 write_request_timeout_in_ms: 2000 cas_contention_timeout_in_ms: 1000 endpoint_snitch: Ec2Snitch rpc_address: localhost rpc_port: 9160 api_port: 10000 api_address: “127.0.0.1” batch_size_warn_threshold_in_kb: 128 batch_size_fail_threshold_in_kb: 1024 partitioner: org.apache.cassandra.dht.Murmur3Partitioner commitlog_total_space_in_mb: -1 murmur3_partitioner_ignore_msb_bits: 12 force_schema_commit_log: true consistent_cluster_management: true api_ui_dir: /opt/scylladb/swagger-ui/dist/ api_doc_dir: /opt/scylladb/api/api-doc/ Once I start the server I am getting below status On running cmd nodetool status, getting below output Here are few thing which I have already tried: Can anyone help me what could I do to resolve it. Please let me know if any other information is required from my end. It is not clear what the problem may be as the scylla.yaml you provided shows you are trying to bootstrap the first node of the cluster (a seed of itself), whereas the systemd output indicate bootstrapping is taking a long (10 min) period, which of course explains why your nodetool output is failing. As you’re using Ec2Snitch, ensure you can reach AWS Metadata server. Then, update your scylla.yaml and properly set your rpc_address for CQL connections to listen outside of loopback only. Finally, follow the Node Cleanup Procedure and try to bootstrap it again. --- ### Page: https://forum.scylladb.com/t/how-to-integrate-apache-solr-and-apache-spark-with-scylladb/887 Title: How to integrate Apache Solr and Apache Spark with ScyllaDB? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello everyone, We are planning to migrate from DataStax Enterprise to ScyllaDB but we facing some problems because ScyllaDB doesn’t bundle with Solr, Spark, and doesn’t have SAI index including the ScyllaDB Enterprise … Language: en Canonical URL: https://forum.scylladb.com/t/how-to-integrate-apache-solr-and-apache-spark-with-scylladb/887 ## Headings Structure: H1: How to integrate Apache Solr and Apache Spark with ScyllaDB? H3: Integrate Scylla with Spark | ScyllaDB Docs H3: Scylla Integrations and Connectors | ScyllaDB Docs H3: Related topics ## Main Content: H1: How to integrate Apache Solr and Apache Spark with ScyllaDB? H3: Integrate Scylla with Spark | ScyllaDB Docs H3: Scylla Integrations and Connectors | ScyllaDB Docs H3: Related topics Hello everyone, We are planning to migrate from DataStax Enterprise to ScyllaDB but we facing some problems because ScyllaDB doesn’t bundle with Solr, Spark, and doesn’t have SAI index including the ScyllaDB Enterprise version. How to integrate Apache Solr and Apache Spark with ScyllaDB and make spark/spark-jobserver queries take advantage of indexes including Solr and Scylla indexes? Hi @asterix0108 Yes, ScyllaDB does not come with bundled 3rd parties, an does not support SAI. There are available integrations to Spark ScyllaDB is an Apache Cassandra-compatible NoSQL data store that can handle 1 million transactions per second on a single server. and other 3rd parties ScyllaDB is an Apache Cassandra-compatible NoSQL data store that can handle 1 million transactions per second on a single server. You can use these integrations to choose the best bread release of each 3rd party. Please let me know if it answers your questions. Thank you for your reply. Does Scylla have any plan to integrate 3rd parties? We are consistently adding more integrations, but we do not plan to bundle with more products in the short term. We might do it in the future if it makes sense. --- ### Page: https://forum.scylladb.com/t/are-batch-insert-update-into-different-tables-logged-safe/888 Title: Are BATCH insert/update into different tables logged + "safe"? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, if I e.g. write to 2 or 3 different tables as part of a batch, are there any benefits or guarantees compared to doing the writes separately (not batched)? The documentation mentions isolation/atomicity is only guara… Language: en Canonical URL: https://forum.scylladb.com/t/are-batch-insert-update-into-different-tables-logged-safe/888 ## Headings Structure: H1: Are BATCH insert/update into different tables logged + "safe"? H3: BATCH H3: Related topics ## Main Content: H1: Are BATCH insert/update into different tables logged + "safe"? H3: BATCH H3: Related topics Hi, if I e.g. write to 2 or 3 different tables as part of a batch, are there any benefits or guarantees compared to doing the writes separately (not batched)? The documentation mentions isolation/atomicity is only guaranteed within a single partition → of one table…? By default, Scylla uses a batch log to ensure all operations in a batch eventually complete or none will (note, however, that operations are only isolated within a single partition). The isolation/atomicity guarantees a batch statement provides is only valid within a single partition of a single table. If you batch updates to multiple partitions or multiple tables, there are no guarantees and you are better off doing separate writes. Thanks for your reply @Botond_Denes. Could you please elaborate? Does BATCH in Scylla differ from Cassandra? According to the docs, single partition batches are isolated. If multiple partitions/tables are involved atomicity should be guaranteed via the LOGGED mechanism? For multiple partition batches, logging ensures that all DML statements are applied. Either all or none of the batch operations will succeed, ensuring atomicity. Batch isolation occurs only if the batch operation is writing to a single partition. Applies multiple data modification language (DML) statements with atomicity and/or in isolation. I don’t see how any guarantees could be provided when multiple partitions are involved, beyond the built-in retry mechanism provided by the logged batch variant. ScyllaDB does not have an undo mechanism, so some writes that are part of a batch could succeed, while others not. Do not mistake batches for transactions. They are not. There are benefits to batching, when writes target a single partition, as ScyllaDB will merge these writes into a single mutation and they will be applied internally as a single write. This does have benefits and additional guarantees. But as soon as multiple partitions or multiple tables are involved, batch is nothing more than a mechanism to send multiple separate writes in a single message. Logged variants come with built-in retry, but that is pretty much it as far as I know. Performance wise, batches have no benefits, on the contrary, separate writes are preferred if you want high performance. Thanks for confirming. I always found the docs (both Cassandra and Scylla) to be quite ambiguous on the matter… Indeed, we now have tracking Docs: ScyllaDB Multi-Table BATCH behavior · Issue #15688 · scylladb/scylladb · GitHub --- ### Page: https://forum.scylladb.com/t/are-there-any-plans-for-triggers-procedures-of-any-kind-on-the-scylladb-roadmap/889 Title: Are there any plans for TRIGGERS/PROCEDURES of any kind on the ScyllaDB Roadmap - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, are there any plans for TRIGGERS of any kind on the ScyllaDB Roadmap? Also how about functions that go beyond UDF/UDA, to e.g. read/write into other tables? References Language: en Canonical URL: https://forum.scylladb.com/t/are-there-any-plans-for-triggers-procedures-of-any-kind-on-the-scylladb-roadmap/889 ## Headings Structure: H1: Are there any plans for TRIGGERS/PROCEDURES of any kind on the ScyllaDB Roadmap H3: DynamoDB Streams and AWS Lambda triggers - Amazon DynamoDB H3: CREATE TRIGGER H3: Related topics ## Main Content: H1: Are there any plans for TRIGGERS/PROCEDURES of any kind on the ScyllaDB Roadmap H3: DynamoDB Streams and AWS Lambda triggers - Amazon DynamoDB H3: CREATE TRIGGER H4: Support for triggers H3: Related topics Hi, are there any plans for TRIGGERS of any kind on the ScyllaDB Roadmap? Also how about functions that go beyond UDF/UDA, to e.g. read/write into other tables? When you need database triggers in DynamoDB, use the combined power of DynamoDB Streams and Lambda functions. Learn about creating triggers and out-of-band data aggregations to scale to new heights. CREATE TRIGGER CREATE TRIGGER — define a new trigger Synopsis CREATE [ OR REPLACE ] [ CONSTRAINT ] TRIGGER name … Hi No, there is no plan for triggers in the short-mid term. See Support CQL `CREATE TRIGGER`, `DROP TRIGGER` Similar to UDF #2204, Scylla trig…ger should **not** necessary use Java related nodetool operations: `nodetool reloadtriggers` https://issues.apache.org/jira/browse/CASSANDRA-1311 https://issues.apache.org/jira/browse/CASSANDRA-7606 https://issues.apache.org/jira/browse/CASSANDRA-4949 Note that UFA/UDA are close but not yet production-ready. Similar to the DynamoDB reference you included, one can use the ScyllaDB CDC feature to look for an event and run application-level logic. Thanks @tzach for your reply. One use case I’ve had in mind actually would have been to build sort of an outbox pattern. Since it’s not recommended/applicable to enable CDC for a high number of tables it would be interesting to create a single ‘outbox’ table with CDC enabled, and push events into this table upon writing to many other tables (via TRIGGERS). But maybe that’s actually ‘putting the cart before the horse’. I guess what I’m asking for or planning to workaround is having a DB-wide WAL?!? Would that be something worth exploring? A single outbox table is definitely the way to go - but it is not clear exactly why you would need a trigger. In fact, there are concerns to that, eg: tables can have a different partition (thus making your updates to end up on different set of replicas), or even in different DCs (should the table in keyspace live in a different keyspace), all of which could introduce an availability concern and make your trigger to fail (or worse, accumulate and over time degrade the entire system performance). What is feasible is to write to both your primary datastore and to an outbox table, and then consume the tail (with CDC or not) of the outbox table overtime. However, I see where you are coming from given Are BATCH insert/update into different tables logged + "safe"? (which is ultimately tied to scylladb/scylladb#13390), but you will ultimately need to decide whether you will batch both entries (to different tables) altogether, or concurrently write. For batching, this will incur an entry on the batchlog, and the updates will be done asynchronously. So it may happen that you consume an event from the outbox table before a record gets persisted in the main service table (though that should be rare), but it doesn’t mean that the main table record is lost, as there’s a built-in retry mechanism to ensure that all entries eventually are applied. For parallel writes, you would need to work around the old fashioned “dual write” problem, so indeed not fun. --- ### Page: https://forum.scylladb.com/t/queries-with-millions-of-records-and-pagination-in-a-distributed-database/891 Title: Queries with millions of records and pagination in a distributed database - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: In our application, we have queries that may return millions of records. Is it possible to use pagination with ScyllaDB? Also, how does paging work with distributed databases, taking into account that the data is distrib… Language: en Canonical URL: https://forum.scylladb.com/t/queries-with-millions-of-records-and-pagination-in-a-distributed-database/891 ## Headings Structure: H1: Queries with millions of records and pagination in a distributed database H3: Related topics ## Main Content: H1: Queries with millions of records and pagination in a distributed database H3: Related topics In our application, we have queries that may return millions of records. Is it possible to use pagination with ScyllaDB? Also, how does paging work with distributed databases, taking into account that the data is distributed on multiple replicas? Yes, ScyllaDB supports paging. In use cases like yours, when queries can return huge amounts of data and when the amount of data is not known in advance, it makes sense to use paging. Otherwise, different resource management issues might occur both on the client and on the database side. Generally speaking, it makes sense to use paging unless you have a good reason not to use it. Paging in the context of database queries involves transmitting query results in manageable chunks or “pages” to prevent these resource management issues. The page size can be limited by the client. Here’s how paging works: The coordinator (the node receiving a query request from the client) identifies the list of replicas (nodes with the relevant data) to fulfill the query. If using token-aware drivers, the coordinator node will also be a replica node. All read requests are sent concurrently to the selected replicas, which execute the requests and return results to the coordinator. *The coordinator merges the results from replicas and requests a page worth of data across the entire read range. The read progresses linearly across partitions in token order. Once the page is full, the page is cut, be it between or in the middle of a partition. To resume the query on the next page, the database records the query’s interruption position in a binary “paging state” cookie, which is transmitted with every page to the client and back to the database with each page request. The paging state is opaque to the client and can store other query-related states. On the next page, the coordinator checks the partition range and drops already-read partitions based on the stored paging state. To improve performance, the replica’s state is kept between pages. *The above steps 2 and 3 are a simplification. What actually happens is that there is only one data request (served by the local node if that is also a replica) and CL-1 digest requests. Only one replica will return data. All replicas participating in the requests receive the page size and are expected to process an identical amount of data. After all replicas have responded, the data and the digests are checked. If they match, the page is returned as-is. Otherwise, a read-repair is started. To address these issues, queries are made stateful by preserving the query’s state between pages using an object called the “querier.” The querier is saved in a special cache, and on the next page, it’s looked up to continue the query from where it left off. Each query is assigned a unique identifier to ensure that the correct querier is used for each query. You can learn more about paging large queries and see examples in the documentation: There is more information about ScyllaDB drivers in the Using ScyllaDB Drivers course on ScyllaDB University, including some hands-on examples for different languages. Also, see the relevant discussion about using Count Limit or Paging. --- ### Page: https://forum.scylladb.com/t/release-scylla-manager-3-2-2/892 Title: [RELEASE] Scylla Manager 3.2.2 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces the release of ScyllaDB Manager 3.2.2 production-ready ScyllaDB Manager patch release of the stable ScyllaDB Manager 3.2 branch. ScyllaDB Manager is a centralized cluster administration and r… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-manager-3-2-2/892 ## Headings Structure: H1: [RELEASE] Scylla Manager 3.2.2 H3: Related topics ## Main Content: H1: [RELEASE] Scylla Manager 3.2.2 H3: Related topics The ScyllaDB team announces the release of ScyllaDB Manager 3.2.2 production-ready ScyllaDB Manager patch release of the stable ScyllaDB Manager 3.2 branch. ScyllaDB Manager is a centralized cluster administration and recurrent tasks automation tool. Scylla Manager 3.2.2 includes one fix for a regression introduced in the 3.2.0 release, which causes Scylla Manager to use a repair optimization for small tables when repairing big tables from the keyspaces having a replication factor equal to the number of nodes. As result, p99 latency might increase during repair operation compared to the same operation with Manager 3.1. #3591 ScyllaDB Enterprise customers are encouraged to upgrade to ScyllaDB Manager 3.2.2 in coordination with the ScyllaDB support team. The new release includes upgrades of both ScyllaDB Manager Server and Agent. Scylla Manager 3.2.2 support the following Scylla Enterprise releases: And the following Open Source release (limited to 5 nodes see license): You can install and run Scylla Manager on Kubernetes using Scylla Operator. More here. --- ### Page: https://forum.scylladb.com/t/using-in-in-cql-queries-performance/893 Title: Using IN in CQL Queries, Performance - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/using-in-in-cql-queries-performance/893 ## Headings Structure: H1: Using IN in CQL Queries, Performance H3: Related topics ## Main Content: H1: Using IN in CQL Queries, Performance H3: Related topics View in #general on Slack @Hartmut: Hi, which would be the recommended/preferred query pattern or anti-pattern? (A vs. B) B) obviously isn’t shard-aware but needs to be orchestrated I guess it may depend on the actual use case, how many rows are to be fetched and so on… But still, I wonder if anyone has any experience or insights to share…? @avi: Individual queries are generally better. You’ve moving some of the coordination from the server to the client, which is more easily scaled. The single IN query cannot be made shard/token aware, so you pay with an extra hop. On the contrary, when querying a specific partition, it should be perfectly valid, correct? @avi: Yes, in this case IN is preferable --- ### Page: https://forum.scylladb.com/t/release-scylla-operator-1-11-0-rc-0/894 Title: [RELEASE] Scylla Operator 1.11.0-rc.0 - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of scylla-operator v1.11.0-rc.0 :rocket: We’ll welcome your feedback on the release candidate. Release notes are available on: https://github.com/scylladb/scylla-op… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-operator-1-11-0-rc-0/894 ## Headings Structure: H1: [RELEASE] Scylla Operator 1.11.0-rc.0 H3: Related topics ## Main Content: H1: [RELEASE] Scylla Operator 1.11.0-rc.0 H3: Related topics The ScyllaDB team is pleased to announce the release of scylla-operator v1.11.0-rc.0 We’ll welcome your feedback on the release candidate. Release notes are available on: https://github.com/scylladb/scylla-operator/releases/tag/v1.11.0-rc.0 If you haven’t heard about the operator yet, here are some links to get you started: https://github.com/scylladb/scylla-operator/tree/master#scylla-operator https://operator.docs.scylladb.com/ --- ### Page: https://forum.scylladb.com/t/release-scylladb-rust-driver-0-10-0/895 Title: [RELEASE] ScyllaDB Rust Driver 0.10.0 - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce ScyllaDB Rust Driver 0.10.0, an asynchronous CQL driver for Rust, optimized for Scylla, but also compatible with Apache Cassandra! Some interesting statistics: over 685k downl… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-rust-driver-0-10-0/895 ## Headings Structure: H1: [RELEASE] ScyllaDB Rust Driver 0.10.0 H2: Notable changes H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Rust Driver 0.10.0 H2: Notable changes H3: Related topics The ScyllaDB team is pleased to announce ScyllaDB Rust Driver 0.10.0, an asynchronous CQL driver for Rust, optimized for Scylla, but also compatible with Apache Cassandra! Some interesting statistics: API cleanups / breaking changes: New features / enhancements: CI / developer tool improvements: Congrats to all contributors and thanks everyone for using our driver! The source code of the driver can be found here: The official crates.io registry entry is here: Thank you for your attention, please do not hesitate to contact us if you have any questions, issues, feature requests, or are simply interested in our driver! Contributors since the last release: --- ### Page: https://forum.scylladb.com/t/last-2-weeks-in-scylla-cluster-tests-git-master-issue-26-2023-10-06/896 Title: Last 2 weeks in scylla-cluster-tests.git master (issue #26; 2023-10-06) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 558925b4…ecaae207 range are covered. There were 35 non-merge commits from 12 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-2-weeks-in-scylla-cluster-tests-git-master-issue-26-2023-10-06/896 ## Headings Structure: H1: Last 2 weeks in scylla-cluster-tests.git master (issue #26; 2023-10-06) H3: Related topics ## Main Content: H1: Last 2 weeks in scylla-cluster-tests.git master (issue #26; 2023-10-06) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 558925b4…ecaae207 range are covered. There were 35 non-merge commits from 12 authors in that period. Some notable commits: Performance latency test also started to use placement groups to stabilise the test results. New performance test for GrowShrinkClusterNemesis in multi-AZ environment was added Logging was improved in the upgrade test by adding multiple InfoEvent messages for all (or mostly of) the commands we run in the upgrade procedure. New performance test related to latency during upgrade was added uses 650GB dataset, and in the report each node upgrade is represented as operation cycle. Individual nemesis tests was enhanced by adding MV and large partition load Removed the compaction strategy from the cassandra-stress commands, since we’d like scylla to use its default compaction strategy and increased row size and partition count to have a more significant amount of data (a few GB). When triggering a reboot using Azure begin_restart SDK VM is not rebooted immediately, but rather scheduled for reboot in the future. This is causing timeouts and broken nemesis logic. So we switched Azure API restart to running reboot -ff to reboot immediately. Azure’s ExtensionOperations were causing problems by enabling auditd service - this was causing log flood and breaking long longevities. To tackle this issue we decided to disable azure agents. We fixed the issue with respecting disk_size option for Azure VM’s. We started to validate if sstables are truly encrypted. Because of false failures during some nemesis in ScyllaDB cloud tests we reworked SSH-based remote loggers and improved code quality. DB logs now have millisecond resolution. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-200-2023-10-08/897 Title: Last week in scylladb.git master (issue #200; 2023-10-08) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 1640f83fdc…4e6fe34501 range are covered. There were 139 non-merge commits from 14 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-200-2023-10-08/897 ## Headings Structure: H1: Last week in scylladb.git master (issue #200; 2023-10-08) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #200; 2023-10-08) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 1640f83fdc…4e6fe34501 range are covered. There were 139 non-merge commits from 14 authors in that period. Some notable commits: The native implementation of nodetool supports many more commands. Note it’s not yet part of regular nodetool invocation. Tablets are a new, experimental method of distributing data in a ScyllaDB cluster. With tablets, data movement is decoupled from node bootstrap and decommissioning. When a tablet is migrated, its sstables will now be automatically cleaned up on the source node. The log-structured allocator (LSA) will evict cache if a query fails because it needs more memory, and retry it. On the other hand, the reader concurrency semaphore will simulate an allocation failure to a query if it detects the system is under severe memory pressure. The two mechanisms work against each other, as we’ll simulate an allocation failure in order to terminate a query, but LSA will respond by retrying it. To avoid this, LSA will now detect the simulated allocation failure and let the query be terminated. A map value, when parsed from its JSON representation, did not parse the key correctly. This is now fixed. If a QUORUM (or higher) read detects a mismatch between data from different replica, it starts a process of reconciliation to bring all replicas to the same state. Previously, this did not work well when at least on replica had a large prefix of tombstones, as we would read all of them into memory. Now, we are able to incrementally process sections of the data, even if they are all tombstones. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/is-scylla-bootstrapping-inefficient/898 Title: Is scylla bootstrapping inefficient? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, bootstrapping a new node causes a lot of compactions on the new node. Is this really necessary? Couldn’t streaming be in token range order, where all ranges are written into one sstable? Also reshaping could be don… Language: en Canonical URL: https://forum.scylladb.com/t/is-scylla-bootstrapping-inefficient/898 ## Headings Structure: H1: Is scylla bootstrapping inefficient? H3: Related topics ## Main Content: H1: Is scylla bootstrapping inefficient? H3: Related topics bootstrapping a new node causes a lot of compactions on the new node. Is this really necessary? Couldn’t streaming be in token range order, where all ranges are written into one sstable? Also reshaping could be done directly when receiving, splitting data into shards. From what I see, this is not how bootstrapping works, right? First, yes it’s inefficient. Streaming is done in token range order. The problem is that each node contributes some vnodes, and those vnodes are discontiguous. Therefore there is a reshaping pass after all sstables are received, to reorganize the data into one sstable. There should be one RESHAPE compaction per shard after bootstrap, no more. With our upcoming tablets feature, there will be no reshaping or compaction needed. First, yes it’s inefficient. Streaming is done in token range order. The problem is that each node contributes some vnodes, and those vnodes are discontiguous. Therefore there is a reshaping pass after all sstables are received, to reorganize the data into one sstable. There should be one RESHAPE compaction per shard after bootstrap, no more. But currently there is much more than one reshape/compaction?! Couldn’t the bootstrapping node request token ranges in order and write those into the same sstables (one per shard)? Right now it seems to be creating many small files, I assume because its writing a separate sstable per vnode. With our upcoming tablets feature, there will be no reshaping or compaction needed. Does this mean there will be no imrpovement for the vnode setup? There are many users out there using vnodes, I think it would be good if vnode bootstrapping still works well. Or will there be a migration path to tablets? That’s what there should be. If there is more, maybe there’s a bug, or I don’t understand the process will enough. @raphaelsc should know more. We don’t request data in token order since that means only one NORMAL node will participate at a time. As it is, all NORMAL nodes (in the same rack) participate, and so they need to contribute less bandwidth. For tablets, we plan to have an automatic migration process, though that will come some time after tablets are generally available. We don’t request data in token order since that means only one NORMAL node will participate at a time. The new server could ask for multiple token ranges from different nodes and merge those into one sstable, resulting into something like this: … Problems are error handling and concurrency. I assume thats why scylla is not doing it like this. Avi, I apologize for distracting you so often! But bootstrapping really feels less optimal than in Cassandra and I’d like to understand it. It can’t ask for the ranges serially, since that’s too slow. So it asks for the range in parallel, then merges them. Maybe there’s a bug, but we need more detailed information about compaction history to see it. Just had a look at the logs from the weekend (grepped for one table for better readability), where I added a second node in that datacenter: 13:07:52 - 16:01:08 → Bootstrap/streaming (~ 3 hours) 16:05:47 - 17:35:30 → Reshape/compact (1 1/2 hours) This was bootstrapping from one other node in the datacenter. I assume streaming would have been faster if there were multiple nodes in that datacenter. I have to see if I can find some logs for that. Also this was on spinning disks. I found another log of a bootstrap into a larger cluster (adding a 5th node to a rack) and with SSDs: (I had to shorten the log a little to fit in here) 11:59:54 - 12:49:19 - Bootstrap / stream (50 minutes) 12:49:19 - 13:56:15 - Reshape / compact (1 hour 7 minutes) Adding 2nd node (previous log): 44 stream sessions / sstables Adding 5th node (in the rack): 264 stream sessions / sstables Conclusion: With a larger number of nodes, number of sstables is growing and reshape is taking longer than streaming. Please file an issue with all the information, it’s easier to track it there. Conclusion: With a larger number of nodes, number of sstables is growing and reshape is taking longer than streaming. That’s actually expected. With a small number of nodes, the scheduler on the existing node will throttle the exiting node’s streaming bandwidth to preserve queries, so the new node’s disk bandwidth will not be saturated. With a large number of nodes, each existing node needs to contribute a small amount of bandwidth, so we can saturate the new node. So streaming from a large cluster is faster than streaming from a small cluster. Reshape takes place on the new node by itself, so its performance doesn’t depend on the cluster size. I was refering to the relative time between streaming & reshaping. Why should reshape have a higher ratio when the cluster is larger? In an ideal world it should not, but I think currently its because scylla creates more&smaller sstable files when clusters are larger. btw: I am currently bootstrapping another node (again second node in DC) and its compact like crazy during streaming: I still think bootstrapping would be faster if less sstables would be created and reshape would happen during streaming Should I still open a ticket? --- ### Page: https://forum.scylladb.com/t/tableplus-client-support-for-scylladb-cloud-s-database-as-a-service-dbaas-serverless/901 Title: TablePlus client support for ScyllaDB Cloud’s Database-as-a-Service (DBaaS) serverless - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: How to connect to ScyllaDB Cloud’s Database-as-a-Service (DBaaS) serverless using a standard database client, such as TablePlus? After creating a ScyllaDB Free Cluster (only the serverless version is free), now Scylla C… Language: en Canonical URL: https://forum.scylladb.com/t/tableplus-client-support-for-scylladb-cloud-s-database-as-a-service-dbaas-serverless/901 ## Headings Structure: H1: TablePlus client support for ScyllaDB Cloud’s Database-as-a-Service (DBaaS) serverless H3: Related topics ## Main Content: H1: TablePlus client support for ScyllaDB Cloud’s Database-as-a-Service (DBaaS) serverless H3: Related topics How to connect to ScyllaDB Cloud’s Database-as-a-Service (DBaaS) serverless using a standard database client, such as TablePlus? After creating a ScyllaDB Free Cluster (only the serverless version is free), now Scylla Cloud generates one connection Bundle File, which is a YAML file with Key, Cert and “CA Cert” information. However I do not know how to extract these 3 pieces of security connection information to generate 3 specific files to load into each of Table Plus Cassandra connection page. I’m not familiar with TablePlus, but it looks like it has support for Apache Cassandra. The gap is Apache Cassandra drivers are compatible with ScyllaDB and Scylla Cloud with dedicated VMs, but not with the latest Scylla Cloud Serverless. I see two alternatives: I do not own Table Plus, but I had already asked support for them. Meanwhile, wich visual client (free or shareware) tool can connect to this new Cluster? All my students will create a cluster (that’s why it must be free)… I would advise checking CQLSh in this case. It is a CLI tool, but for ScyllaDB I think you will find it useful. ScyllaDB, like other NoSQLs, does not support complex relations between tables, so visual tools do not bring as much value as for SQL DBs. Is it easy to install in Windows? Ah, remember that “visual” query tools are much better to see results (I’m not talking about data modeling) Scylla-CQLSh is a Python base too, you can download it from scylla-cqlsh · PyPI or as Docker (If you can use Docker, you can run ScyllaDB+CQLSh as one Docker image) I think you will find CQLSh tabular output useful for most cases. I would be happy to get your feedback. I’m using ScyllaDB Serverless Cloud. And we ar not able to use Docker… Why is it so difficult to access ScyllaDB? --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-2-14/904 Title: [RELEASE] ScyllaDB Enterprise 2022.2.14 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2022.2.14, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.2 Feature Release.. This patch release enables a new configuration to … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-2-14/904 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2022.2.14 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2022.2.14 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2022.2.14, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.2 Feature Release.. This patch release enables a new configuration to control streaming bandwidth and multiple minor bug fixes. Note the latest ScyllaDB Enterprise release is 2023.1 LTS, and you are encouraged to upgrade to it in coordination with the ScyllaDB Support team. Streaming: Add stream_plan_ranges_percentage This option allows user to change the number of ranges to stream in batch per stream plan. Currently, each stream plan streams 10% of the total ranges. The default value is the same as before: 10% percentage of total ranges. #14191 The following issues are fixed in this release (with an open-source reference, if available): --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-27-2023-10-13/906 Title: Last week in scylla-cluster-tests.git master (issue #27; 2023-10-13) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 0ebb9363…dd92a435 range are covered. There were 15 non-merge commits from 4 authors in tha… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-27-2023-10-13/906 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #27; 2023-10-13) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #27; 2023-10-13) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 0ebb9363…dd92a435 range are covered. There were 15 non-merge commits from 4 authors in that period. Some notable commits: Local K8s tests got updates to the latest versions for Kind and Scylla and Scylla Manager. We reverted disabling Azure agents as we need one for running commands using Azure python SDK (used to reboot VM). To avoid log pollution issues with auditd service we disabled and masked it. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/why-is-it-so-difficult-to-find-a-good-visual-client-to-connect-to-scylladb/908 Title: Why is it so difficult to find a good visual client to connect to ScyllaDB? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, I need a good, free or trial, visual client tool to connect and make DML and DDL commands in a ScyllaDB Serverless Cloud cluster. DBever coild work, but community edition does not have a free Cassandra Driver. Tabl… Language: en Canonical URL: https://forum.scylladb.com/t/why-is-it-so-difficult-to-find-a-good-visual-client-to-connect-to-scylladb/908 ## Headings Structure: H1: Why is it so difficult to find a good visual client to connect to ScyllaDB? H3: Related topics ## Main Content: H1: Why is it so difficult to find a good visual client to connect to ScyllaDB? H3: Related topics Hi, I need a good, free or trial, visual client tool to connect and make DML and DDL commands in a ScyllaDB Serverless Cloud cluster. DBever coild work, but community edition does not have a free Cassandra Driver. TablePlus does not work with Serverless ScyllaDB DBVisualizer is hard to install a Cassandra driver and probably won’t work either. SQLSh is not visual and is hard to install. I’m a teacher and would like to use with my students a very simple and visual tool to meke commands and see results. Dozens of Students will have to install and configure the client. I do not want to make any other visual operations (such as modeling, Visual DDL,…). Hi Furia I understand the frustration, but it looks like you are asking the same question in different variants. As I wrote in one of the other threads, most ScyllaDB users use CQLSh to interact with ScyllaDB. The other variant is a question of another user… And we are not having the solution. My students use Windows computers at home and will not install several pieces of software, such as python, docker and then a command line tool… We would loose to many classes to make it work! Well, at least please provide support on how to use connect-bundle.yaml ScyllaDB Cluster connection file to get Key, Cert, Cert CA xml certificate files: --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-1-11/909 Title: [RELEASE] ScyllaDB Enterprise 2022.1.11 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2022.1.11, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.1. Note that there is a newer Enterprise LTS release: ScyllaDB Enterpr… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-1-11/909 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2022.1.11 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2022.1.11 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2022.1.11, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.1. Note that there is a newer Enterprise LTS release: ScyllaDB Enterprise 2023.1. While we will continue to support 2022.1 LTS, you can get additional features with 2023.1. Below is a list of performance and stability improvements and bug fixes, each with an open-source reference if available: --- ### Page: https://forum.scylladb.com/t/how-does-kubernetes-use-the-swap-space-on-the-host-computer/910 Title: How does Kubernetes use the swap space on the host computer? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi all! We are now using kubernetes to deploy our ScyllaDB cluster. In this case, can we use the swap space on the host computer?If so, how does Kubernetes use the swap space on the host computer? Thanks! Language: en Canonical URL: https://forum.scylladb.com/t/how-does-kubernetes-use-the-swap-space-on-the-host-computer/910 ## Headings Structure: H1: How does Kubernetes use the swap space on the host computer? H3: Related topics ## Main Content: H1: How does Kubernetes use the swap space on the host computer? H3: Related topics Hi all! We are now using kubernetes to deploy our ScyllaDB cluster. In this case, can we use the swap space on the host computer?If so, how does Kubernetes use the swap space on the host computer? Thanks! This is more of a Kubernetes question rather than a ScyllaDB one. In the past, there were no kubelet support for swapping, but this has been changing lately with KEP-2400, which is in Beta in 1.28. In general, for a ScyllaDB deployment you will want to ensure you use --lock-memory=1 in order to avoid the database process from swapping altogether. Thanks a lot! But I’m still confused that is it recommended to use swap space when we are executing bare metal deployment? You’re right. In a bare metal / VM deployment the ScyllaDB process typically runs with the --lock-memory option enabled, which causes the process to mlock() the memory area. The swap space, therefore, is left to other processes such as JMX, OS, etc. On K8s you have the flexibility to set the pod memory in your deployment, as well as limit ScyllaDB’s own memory consumption via the --memory parameter. Ideally, you will also want to --lock-memory, given that it can cause the process to swap once the aforementioned proposal lands in a K8s release. Thanks a lot for your reply. I think I have a clearer understanding of this issue now. --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2023-1-2/912 Title: [RELEASE] ScyllaDB Enterprise 2023.1.2 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2023.1.2, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2023.1 LTS Release. This patch release enables cluster-level configuration f… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2023-1-2/912 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2023.1.2 H3: Transparent Data Encryption H3: Streaming: Add stream_plan_ranges_fraction H3: Bug fixes H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2023.1.2 H3: Transparent Data Encryption H3: Streaming: Add stream_plan_ranges_fraction H3: Bug fixes H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2023.1.2, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2023.1 LTS Release. This patch release enables cluster-level configuration for Encryption at Rest, a new configuration to control streaming bandwidth, and multiple minor bug fixes. You are encouraged to upgrade to it in coordination with the ScyllaDB Support team. Scylla Enterprise has supported Encryption at Rest (EaR) for a long time. So far, one can store the keys for EaR locally, in an encrypted table, or an external KMIP server. Release 2023.1.1 added the ability to use Amazon KMS keys. Release 2023.1.2 adds Transparent Data Encryption (TDE), a way to define Encryption at Rest parameters per cluster, not only per table. This allows the system administrator to enforce encryption of all tables using the same master key, for example, from KMS, without specifying the encryption parameter per table. For example, with the following in scylla.yaml, all tables will be encrypted using encryption parameters of my-kms1 See more examples and info here. This option allows user to change the number of ranges to stream in batch per stream plan. Currently, each stream plan streams 10% of the total ranges. The default value is the same as before: 10% of total ranges. #14191 The following issues are fixed in this release (with an open-source reference, if available): --- ### Page: https://forum.scylladb.com/t/dial-tcp-127-0-0-1-connect-connection-refused/913 Title: Dial tcp 127.0.0.1:5080: connect: connection refused - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Getting below error post inallation of scylla-manager sctool version Client version: 3.2.2-0.20231002.6af4bf9b Error: Get “http://127.0.0.1:5080/api/v1/version”: dial tcp 127.0.0.1:5080: connect: connection refused Language: en Canonical URL: https://forum.scylladb.com/t/dial-tcp-127-0-0-1-connect-connection-refused/913 ## Headings Structure: H1: Dial tcp 127.0.0.1:5080: connect: connection refused H3: Install ScyllaDB Manager | ScyllaDB Docs H3: Related topics ## Main Content: H1: Dial tcp 127.0.0.1:5080: connect: connection refused H3: Install ScyllaDB Manager | ScyllaDB Docs H3: Related topics Getting below error post inallation of scylla-manager sctool version Client version: 3.2.2-0.20231002.6af4bf9b Error: Get “http://127.0.0.1:5080/api/v1/version”: dial tcp 127.0.0.1:5080: connect: connection refused @CHARAN_CHINTHA Can you post the scylla-manager logs ? It looks that the manager API server is down, but just wanted to make sure checking the logs. Please find error log scylla-manager[19421]: {“L”:“ERROR”,“T”:“2023-10-18T09:22:58.222Z”,“M”:“Bye”,“error”:“db init: gocql: unable to create session: unable to discover protocol version: Cannot achieve consistency level for cl ONE. Requires 1, alive 0”,“_trace_id”:“UWmlchK5QpqbMMsdtRdU6w”,“errorStack”:“main.glob…func2\n\tgithub.com/scylladb/scylla-manager/v3/pkg/cmd/scylla-manager/root.go:125\ngithub.com/spf13/cobra.(*Command).execute\n\tgithub.com/spf13/cobra@v1.1.1/command.go:850\ngithub.com/spf13/cobra.(*Command).ExecuteC\n\tgithub.com/spf13/cobra@v1.1.1/command.go:958\ngithub.com/spf13/cobra.(*Command).Execute\n\tgithub.com/spf13/cobra@v1.1.1/command.go:895\nmain.main\n\tgithub.com/scylladb/scylla-manager/v3/pkg/cmd/scylla-manager/main.go:12\nruntime.main\n\truntime/proc.go:250\nruntime.goexit\n\truntime/asm_amd64.s:1598\n”,“S”:“github.com/scylladb/go-log.Logger.log\n\tgithub.com/scylladb/go-log@v0.0.7/logger.go:101\ngithub.com/scylladb/go-log.Logger.Error\n\tgithub.com/scylladb/go-log@v0.0.7/logger.go:84\nmain.glob…func2.1\n\tgithub.com/scylladb/scylla-manager/v3/pkg/cmd/scylla-manager/root.go:70\nmain.glob…func2\n\tgithub.com/scylladb/scylla-manager/v3/pkg/cmd/scylla-manager/root.go:125\ngithub.com/spf13/cobra.(*Command).execute\n\tgithub.com/spf13/cobra@v1.1.1/command.go:850\ngithub.com/spf13/cobra.(*Command).ExecuteC\n\tgithub.com/spf13/cobra@v1.1.1/command.go:958\ngithub.com/spf13/cobra.(*Command).Execute\n\tgithub.com/spf13/cobra@v1.1.1/command.go:895\nmain.main\n\tgithub.com/scylladb/scylla-manager/v3/pkg/cmd/scylla-manager/main.go:12\nruntime.main\n\truntime/proc.go:250”} Oct 18 09:22:58 ip-10-21-186-170 scylla-manager[19421]: STARTUP ERROR: db init: gocql: unable to create session: unable to discover protocol version: Cannot achieve consistency level for cl ONE. Requires 1, alive 0 @CHARAN_CHINTHA Scylla-manager is using its own instance of ScyllaDB as a backend/cache. The error log you pasted shows that there is a problem in connecting to this ScyllaDB instance. Please make sure that you set it up correctly. ScyllaDB is an Apache Cassandra-compatible NoSQL data store that can handle 1 million transactions per second on a single server. Getting below error now. scylla-manager[18942]: {“L”:“INFO”,“T”:“2023-10-24T10:13:24.969Z”,“N”:“wait”,“M”:“Waiting for network connection”,“sleep”:“2s”,“error”:“dial tcp: lookup 10.21.186.153:9042: no such host; dial tcp: lookup 10.21.185.87:9042: no such host”,“_trace_id”:“2AD7MMM_THWCoTOnnVXopA”} FYI, i have enabled port and ip in firewall rules. [root@ip- ~]# sctool status Error: Get “http://ip:5080/api/v1/clusters”: dial tcp ip:5080: connect: connection refused @CHARAN_CHINTHA i got same issue for scylla db @Karol_Kokoszka i am also refer the document for scylla-manager document but not use Did you installed scylla-server and status is up and running?? And also same details you have updated scylla-manager yaml file.? --- ### Page: https://forum.scylladb.com/t/scylladb-dry-run-test-error-giving-up-after-2-attempts-agent-http-500-init-location-failed-to-acquire-msi-token-msi-is-not-enabled-on-this-vm/916 Title: ScyllaDB dry run test ERROR - giving up after 2 attempts: agent [HTTP 500] init location: Failed to acquire MSI token: MSI is not enabled on this VM - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I wanted to know if you can help I am facing issue testing the dry run backup to azure blob storage the VMs are on prem so I used storage account name and key to update the yaml config file I have 3 nodes I get this e… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-dry-run-test-error-giving-up-after-2-attempts-agent-http-500-init-location-failed-to-acquire-msi-token-msi-is-not-enabled-on-this-vm/916 ## Headings Structure: H1: ScyllaDB dry run test ERROR - giving up after 2 attempts: agent [HTTP 500] init location: Failed to acquire MSI token: MSI is not enabled on this VM H3: Related topics ## Main Content: H1: ScyllaDB dry run test ERROR - giving up after 2 attempts: agent [HTTP 500] init location: Failed to acquire MSI token: MSI is not enabled on this VM H3: Related topics I wanted to know if you can help I am facing issue testing the dry run backup to azure blob storage the VMs are on prem so I used storage account name and key to update the yaml config file I get this error when I run the sctool backup -c mycluster -L ‘mybackupname’ --dry-run Error: create backup target: location is not accessible ** MYIP: giving up after 2 attempts: agent [HTTP 500] init location: Failed to acquire MSI token: MSI is not enabled on this VM: Get “http://169.254.169.254/metadata/identity/oauth2/token?api-version=2018-02-01&resource=https%3A%2F%2Fstorage.azure.com”: dial tcp 169.254.169.254:80: connect: no route to host - make sure the location is correct and credentials are set, to debug SSH to MYIP and run “scylla-manager-agent check-location -L MYbackupstorage --debug”** when i run this scylla-manager-agent check-location -L MYbackupstorage --debug on all of the nodes it returned fine also when I run the scylla-manager-agent check-location -L azure: on all nodes it returns fine too without any errors the debug on each node returns good ending with the deletion of the test as below on all nodes {“L”:“DEBUG”,“T”:“2023-10-17T13:13:02.875+0100”,“N”:“rclone”,“M”:“Waiting for deletions to finish”} {“L”:“DEBUG”,“T”:“2023-10-17T13:13:03.229+0100”,“N”:“rclone”,“M”:“test: Deleted”} Please how can I resolve the challenge of performing the dry run backup test I also tried to do an adhoc backup same error too Error: create backup target: location is not accessible ** MYIP: giving up after 2 attempts: agent [HTTP 500] init location: Failed to acquire MSI token: MSI is not enabled on this VM: Get “http://169.254.169.254/metadata/identity/oauth2/token?api-version=2018-02-01&resource=https%3A%2F%2Fstorage.azure.com”: dial tcp 169.254.169.254:80: connect: no route to host - make sure the location is correct and credentials are set, to debug SSH to MYIP and run “scylla-manager-agent check-location -L MYbackupstorage --debug”** This may or not be related but you typically want to schedule your backup by prepending the provider beforehand, as in: That said, the error you have shown is a DNS lookup error, which makes sense, given that 169.254.169.254 is the IP of Azure’s internal metadata service which - as you are running on-prem - it won’t be able to resolve. Let’s try to use the correct backup name, and should the problem persist, raise an issue under the ScyllaDB Manager GitHub repo. Hi @felipemendes Thank you for your response. I wanted to ask like you have said, is it not possible to run the Scylla setup and the manager on the On-Prem VMs while we push the backups to the azure blob storage? How do I prepend the provider? will I add that (azure:bucket_name) to a file? I have used the correct commands and followed all guides here in and here Examples | ScyllaDB Docs I will look into the DNS an It is possible, hence why I explained in my previous post that the process shouldn’t abort upon a DNS failure, but rather likely fallback to resolving the Azure Blobstorage endpoint via normal means. How do I prepend the provider? will I add that (azure:bucket_name) to a file? In your command-line, use: sctool backup -c mycluster -L azure:mybackupname --dry-run Please follow-up with a Github issue as initially explained should this not resolve your problem. --- ### Page: https://forum.scylladb.com/t/release-scylla-manager-3-2-3/917 Title: [RELEASE] Scylla Manager 3.2.3 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces the release of ScyllaDB Manager 3.2.3 production-ready ScyllaDB Manager patch release of the stable ScyllaDB Manager 3.2 branch. ScyllaDB Manager is a centralized cluster administration and r… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-manager-3-2-3/917 ## Headings Structure: H1: [RELEASE] Scylla Manager 3.2.3 H3: Related topics ## Main Content: H1: [RELEASE] Scylla Manager 3.2.3 H3: Related topics The ScyllaDB team announces the release of ScyllaDB Manager 3.2.3 production-ready ScyllaDB Manager patch release of the stable ScyllaDB Manager 3.2 branch. ScyllaDB Manager is a centralized cluster administration and recurrent tasks automation tool. Scylla Manager 3.2.3 extends the Scylla Health Check task to include additional information into the scylla_manager_healthcheck_cql_status (#3555). This metric shows whether the manager-agent is up and running as well. Users could face a situation in which the repair progress was reported as 100%/0%, which could be interpreted as 100% success and no errors, even though a few errors appeared. We fixed that by improving the way how we round the fractions (#3547, #3581). With the 3.2.3 patch release of Scylla Manager, we are ensuring that repair intensity/parallel parameters updated on running tasks are respected on the retry (#3580). Beginning with the 3.2.3 release, we started publishing linux/arm64 docker images for scylla-manager and for scylla-manager-agent (#3278), next to linux/amd64 ones. ScyllaDB Enterprise customers are encouraged to upgrade to ScyllaDB Manager 3.2.3 in coordination with the ScyllaDB support team. The new release includes upgrades of both ScyllaDB Manager Server and Agent. Scylla Manager 3.2.2 support the following Scylla Enterprise releases: And the following Open Source release (limited to 5 nodes see license): You can install and run Scylla Manager on Kubernetes using Scylla Operator. More here. Hi, it seems like with the latest release you’ve broken the default ARCH for docker image - it is now only linux/arm64/v8, without linux/amd64. Is it supposed to be like this? It has just broken one of our setups. @bsdelnik Thank you for reporting that. This is a mistake we made. We actually should support amd64 versions. Arm64 is not ready yet. Amd64 docker image for 3.2.3 should appear soon. ScyllaDB Manager and ScyllaDB Manager Agent for open source users (up to 5 nodes) I’m new to Scylla so please be patient if this is a nonsensical question. Why is there a 5-node limitation on the ScyllaDB Manager for open source users? What are the options for that functionality for larger clusters? --- ### Page: https://forum.scylladb.com/t/release-scylladb-monitoring-stack-4-5-0/933 Title: [RELEASE] ScyllaDB Monitoring Stack 4.5.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of ScyllaDB Monitoring Stack 4.5.0 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB Enterprise and ScyllaDB Open Source, based on Prometheus … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-monitoring-stack-4-5-0/933 ## Headings Structure: H1: [RELEASE] ScyllaDB Monitoring Stack 4.5.0 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Monitoring Stack 4.5.0 H3: Related topics The ScyllaDB team is pleased to announce the release of ScyllaDB Monitoring Stack 4.5.0 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB Enterprise and ScyllaDB Open Source, based on Prometheus and Grafana. ScyllaDB Monitoring Stack 4.5.0 supports: Versions updates for Scylla Monitoring Stack 4.5.0 New Information in ScyllaDB Dashboards Overview Dashboard Change Instead of aggregating priority groups, each priority group will be shown as its own graph. Detailed Dashboard Change Added more information to the large cell, large rows, and large partition tables. Add a warning when a node is in a joining state for a long time, as other nodes stream data to the new node, based on the new token range ownership. As each node can hold tenths of terabytes, streaming might take many hours even a day. ScyllaDB actually makes sure streaming is under control, so it won’t add to the one line requests latency. It is a normal behavior that a node can take a few hours, sometimes even more than a day to join a cluster, but it’s something worth noting. So informative alerts were added to inform that a node is still in the joining state. The default marks are set to 5 hours and one day. Support taking Prometheus targets from directory #1995 Support Scylla Manager agent ports config parameter #1992 Support for Apple Silicon Valley Chips enhancement #1956 Remove reverse order read warnings #1498 --- ### Page: https://forum.scylladb.com/t/last-fortnight-in-scylladb-git-master-issue-201-2023-10-22/934 Title: Last fortnight in scylladb.git master (issue #201; 2023-10-22) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last two weeks. Commits in the 4e6fe34501…f181ac033a range are covered. There were 95 non-merge commits from 20 authors in that … Language: en Canonical URL: https://forum.scylladb.com/t/last-fortnight-in-scylladb-git-master-issue-201-2023-10-22/934 ## Headings Structure: H1: Last fortnight in scylladb.git master (issue #201; 2023-10-22) H3: Related topics ## Main Content: H1: Last fortnight in scylladb.git master (issue #201; 2023-10-22) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last two weeks. Commits in the 4e6fe34501…f181ac033a range are covered. There were 95 non-merge commits from 20 authors in that period. Some notable commits: Recently, we changed the schema version algorithm not to hash the entire schema as this causes slow performance with large numbers of tables. This has been reverted due to a regression. To generate efficient bloom filters, we estimate the number of partitions in the sstable we will produce. The estimation has been improved for data models where the partition keys dominate the on-disk size. The setup utility now works with disks that do not have UUIDs, such as those in some virtualized environments. Compaction will now avoid garbage-collecting tombstones that potentially delete data in commitlog. This prevents data resurrection in the event that a node crashes and replays commitlog. This is rare since generally commitlog data is relatively fresh and tombstones that delete such data would not be garbage collected for other reasons. Each node contains a local system.truncated table containing truncation records for other tables on the node. This table is now cleared after commitlog replay to avoid it being re-interpreted incorrectly in the case of two consecutive commitlog replays. The version string has been advanced to 5.5.0-dev to mark the start of the 5.4 stabilization cycle. Note the version string is not final and may change. Alternator, ScyllaDB’s implementation of the DynamoDB API, now responds more correctly to the DeleteTable API. The native nodetool implementation, which is slated to replace the Java-based nodetool, now supports the stop and compactionhistory commands. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/unable-to-search-same-string-in-two-columns-using-in-operator-in-scylladb/938 Title: Unable to Search Same String in Two Columns Using IN Operator in ScyllaDB - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello ScyllaDB Community, I am currently facing an issue while attempting to search the same string in two different columns using the IN operator in ScyllaDB. My aim is to achieve a query akin to column1 = 'c1' OR colu… Language: en Canonical URL: https://forum.scylladb.com/t/unable-to-search-same-string-in-two-columns-using-in-operator-in-scylladb/938 ## Headings Structure: H1: Unable to Search Same String in Two Columns Using IN Operator in ScyllaDB H3: Related topics ## Main Content: H1: Unable to Search Same String in Two Columns Using IN Operator in ScyllaDB H3: Related topics Hello ScyllaDB Community, I am currently facing an issue while attempting to search the same string in two different columns using the IN operator in ScyllaDB. My aim is to achieve a query akin to column1 = 'c1' OR column2 = 'c1' in other databases. However, as I’ve learned, the OR operator isn’t supported in ScyllaDB. I came across some discussions suggesting the use of the IN operator to circumvent this limitation. Following the guidance, I tried executing the following query: Unfortunately, I encountered an error stating: I have been trying to resolve this issue through various resources including the ScyllaDB-Users slack, GitHub, and other online platforms, but haven’t found a clear solution yet. The responses on the community forum have been quite slow and I am in urgent need of resolving this. I would immensely appreciate any guidance or alternative solutions to perform such a query in ScyllaDB. Has anyone faced a similar issue or have any insights on how to correctly use the IN operator for this scenario? Your help would expedite my project progress significantly. Thank you in advance! As you realized from the error, you can’t do multi-column relations without specifying clustering columns in their order. If we allowed that, it would be very expensive to scan through all non-key components. Doing a multi-column relation on a composite key, however, is another story, and can be circumvented by using two AND clauses. However, you are neither restricting your selection to a partition, nor a clustering column in your example, which makes me wonder what exactly you are trying to accomplish. If your example query is realistic and you are doing a full table scan, then why don’t you use the token() function instead? Even better, combine it with a query engine which will allow you to run your OR clause you want. Spark and Presto can easily accomplish this. If a full table scan is not what you want, then why don’t you simply issue two queries ? c1 || c2 can potentially retrieve: c1, c2 and (c1, c2). All you have to do is to ignore the latter if duplicates are a concern. In summary, you can’t accomplish what you want with an IN clause like this. i’m using secondary index it’s just simple i want rows which column1 value is s or column2 value is x. Unfortunately, this kind of query is not supported in ScyllaDB at all. You can issue two separate queries and piece together the result in the application. --- ### Page: https://forum.scylladb.com/t/plans-to-support-general-purpose-transactions-cep-15/939 Title: Plans to Support General Purpose Transactions? (CEP-15) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, with Cassandra 5.0 on the horizon, are there plans for ScyllaDB to also add support for ACID Transactions? https://cwiki.apache.org/confluence/display/CASSANDRA/CEP-15:+General+Purpose+Transactions Many Thanks. Kee… Language: en Canonical URL: https://forum.scylladb.com/t/plans-to-support-general-purpose-transactions-cep-15/939 ## Headings Structure: H1: Plans to Support General Purpose Transactions? (CEP-15) H3: Related topics ## Main Content: H1: Plans to Support General Purpose Transactions? (CEP-15) H3: Related topics Hi, with Cassandra 5.0 on the horizon, are there plans for ScyllaDB to also add support for ACID Transactions? https://cwiki.apache.org/confluence/display/CASSANDRA/CEP-15:+General+Purpose+Transactions Many Thanks. Keep up the amazing work. Stay safe! We are moving forward with stronger consistency of metadata (schema, topology). Next will be lightweight and then general transactions, but there is no time plan for these features yet. --- ### Page: https://forum.scylladb.com/t/last-2-weeks-in-scylla-cluster-tests-git-master-issue-28-2023-10-27/943 Title: Last 2 weeks in scylla-cluster-tests.git master (issue #28; 2023-10-27) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last 2 weeks. Commits in the 9d973065…63f27adc range are covered. There were 40 non-merge commits from 10 authors in… Language: en Canonical URL: https://forum.scylladb.com/t/last-2-weeks-in-scylla-cluster-tests-git-master-issue-28-2023-10-27/943 ## Headings Structure: H1: Last 2 weeks in scylla-cluster-tests.git master (issue #28; 2023-10-27) H3: Related topics ## Main Content: H1: Last 2 weeks in scylla-cluster-tests.git master (issue #28; 2023-10-27) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last 2 weeks. Commits in the 9d973065…63f27adc range are covered. There were 40 non-merge commits from 10 authors in that period. Some notable commits: Scylla log collected at the end of the tests with journalctl also has ms precision.. This log contains accurate time markers so are preferable for debugging. We improved data validation for nemeses which delete data, so we have more confidence in data integrity after running them. For resources/time savings, we don’t verify all the data, just specified by configuration sample. Validation is done during cluster health check - in between nemeses. We can also now configure the test scenario to be more intensive on tombstones, by using delete_rows nemesis flag. When using placement groups in performance tests, nodes added by nemesis are placed in the same placement group as other nodes. In the ‘v1.11.0*’ scylla-operator was added the possibility to use PodIPs for node-to-node and node-to-client connectivity. We added support of direct usage of the PodIP addresses for DB pods. We fixed and re-enabled the bare metal performance test. New SCT ability to configure disks by custom scripts to cover custom use cases. Also, with use of the new scylla_d_overrides_files option in test configuration we can override any file in /etc/scylla.d folder. Scripts and custom configurations are stored in an external private repository, which can be pulled in the jenkins pipeline with checkoutQaInternal groovy helper. Argus can show Gemini results similarly like shown in emails. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/unexpected-behavior-index-on-column-not-returning-results/944 Title: Unexpected Behavior: Index on Column Not Returning Results - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello everyone, I’ve recently encountered a peculiar issue and was hoping to get some insights or solutions from this community. Background: I have a database table with several columns. For optimization purposes, I ha… Language: en Canonical URL: https://forum.scylladb.com/t/unexpected-behavior-index-on-column-not-returning-results/944 ## Headings Structure: H1: Unexpected Behavior: Index on Column Not Returning Results H3: Related topics ## Main Content: H1: Unexpected Behavior: Index on Column Not Returning Results H3: Related topics I’ve recently encountered a peculiar issue and was hoping to get some insights or solutions from this community. Background: I have a database table with several columns. For optimization purposes, I have created separate secondary indexes on a few of these columns, and they work perfectly fine. However, I’ve run into a problem with the latest column I indexed. The Issue: After creating an index on a specific column (let’s call it ColumnX), when I perform a search based on that column, it returns no results. The strange part is, I am certain that the data I’m searching for exists in the table because when I search for the same data in a different field (say ColumnY), I can locate the desired entry. Additional Information: I’m baffled as to why this is happening, especially since I’ve indexed other columns in the same manner without any issues. Has anyone else encountered this problem or have any suggestions on how to fix it? Any help or insights would be greatly appreciated. Thank you in advance! This release is ancient, use a supported version >= 5.1 and retry after you run repair of the base table thanks but issued solve without doing anything i think because of adding data while creating index and i don’t know how to track index creation because i was able to see index thought that’s an issue. --- ### Page: https://forum.scylladb.com/t/need-help-with-transaction-modeling/946 Title: Need help with "transaction" modeling - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I apologize if I have the wrong category, this is my first topic Preface I have a relational sql database that works pretty well, but scaling it is a real pain So scylla came to my attention, in which I saw a potentia… Language: en Canonical URL: https://forum.scylladb.com/t/need-help-with-transaction-modeling/946 ## Headings Structure: H1: Need help with "transaction" modeling H3: Related topics ## Main Content: H1: Need help with "transaction" modeling H3: Related topics I apologize if I have the wrong category, this is my first topic Preface I have a relational sql database that works pretty well, but scaling it is a real pain So scylla came to my attention, in which I saw a potential salvation. To the point I studied the documentation of scylla, and figured out how to migrate everything but one item: In my relational database, I use transactions to allocate usage to data. To oversimplify, it looks like this: After reading probably all the articles on the subject of LWT and filtering in scylla, here’s the best I’ve come to: But the problem still remains. In order to “grab the rows” I need, I have to do a read first, because the secondary index doesn’t support update and delete operations, but even if it did, there’s no way to limit the rows that will be updated. In addition, the secondary index cannot be used with the counter type, which means there must be a read to increment, or decrement the value. The problem is that I don’t see any way to perform the read and subsequent update atomically. If don’t do it atomically, there’s no way to avoid data corruption. The only solution I see is to restrict simultaneous work with one partition at the service level, and work with the database without blocking. How can I solve my problem at the database level, and if there is a solution, how resource intensive is it? This is a very popular table in my database, and there will be a lot of similar operations. I would say this is the main type of load. Will scylla perform well in this type of operation? Thanks in advance for the answer Hi @cmandra, welcome to the ScyllaDB community! I don’t understand the data modeling you’ve came up with. Neither of the two indexes you created will allow you to retrieve without specifying id as in your SQL example. Therefore, using an index seems redundant as pretty much every partition within your proposed base table will always contain a single row. Thus, if you are always specifying id as your filtering clause (which is unlike the SQL transaction you demonstrated), you could well use ALLOW FILTERING instead. That said, your table also contains a column of type text (line), and this seems to be the main impediment for you to use a counter table, and not the actual Indexes you created. Lastly, we finally get to your main problem. At this point, I assume (given your previous feedback) that you had a chance to read and navigate through the code in Getting the Most out of Lightweight Transactions in ScyllaDB, where we demonstrate how a transfers table is used for a client to claim owning that transfer while using TTL. Happy to discuss further. Hi. Thank you very much for your reply! Yes, when I got to the tests I realized that I was wrong in some assumptions. Also the code of my table was incorrect, here is a more correct version: Sections will be partitioned by resource id, that’s about 1k-100k records, so ALLOW FILTERING doesn’t apply here. I read this article. Great implementation that looks like a solution to the problem, but my main problem is with the read by condition. I need the following condition to be satisfied: usages < 123 AND concurrent_usages < 10 But I don’t see a way to accomplish this with database tools without resorting to redundant reads. Here’s the next thing I came up with: Actually, this post doesn’t make any more sense initially, as I’ve encountered another problem, if I try to to apply the WHERE < clause to usages and concurrent_usages which are now part of the primary key I get an error: PRIMARY KEY column "concurrent_usages" cannot be restricted as preceding column "usages" is not restricted and there seems to be no way around this restriction. But, if this restriction could be bypassed, I would probably do something like this SELECT id FROM test.table_with_important_rows WHERE resource_id = 00000000-0000-0000-0000-0000-000000000000 AND usages < 123 AND concurrent_usages < 10 LIMIT 100 DELETE FROM test.table_with_important_rows WHERE resource_id = 00000000-0000-0000-0000-0000-000000000000 AND usages != -1 AND concurrent_usages != -1 AND id IN (...) IF usages < 123 AND concurrent_usages < 10 Oh… That won’t work either One or more errors occurred. (PRIMARY KEY column 'concurrent_usages' cannot have IF conditions) I think I’ve taken a wrong turn somewhere, please let me know if you see a solution to my problem. Thanks for your time! You can not have usages nor concurrent_usages part of your clustering key. This happens because key columns can not be updated. Therefore, your latter table isn’t going to work as you noticed. The first table you’ve shown seems to be closer than what you actually want. a PRIMARY KEY (resource_id, id) will result in all id’s being grouped inside a resource_id partition. In other words, every id will be a distinct row for a given resource_id. If I understood you correctly, this is what you want to represent with regards to your data organization. This brings us to the second part of your problem, related to your reads. As you have already realized, the base table doesn’t have a clustering key of neither usages nor concurrent_usages, and therefore doesn’t allow you to efficiently run WHERE usages < ? AND concurrent_usages < ? in an efficient way (given that your previous input was the need to walk down potentially through 100k records). You may overcome that by relying on a Materialized View. Like this: This will allow you to run queries like: SELECT * FROM x WHERE resource_id = ? AND usages < ? What about concurrent_usages? Well, you filter through it: SELECT * FROM usages_view WHERE resource_id = ? AND usages < ? AND concurrent_usages < ? ALLOW FILTERING The database will efficiently lookup your usages restriction and then, simply filter the matching concurrent_usages until it reaches your LIMIT. Bingo! And yes, you can have more than one view. In particular, if you have a need to restrict only by concurrent_usages, very likely another view would be more efficient rather than relying on the ALLOW FILTERING trick here. However, it is hard to say for certain what will work best and what not without actually testing with your data. After that is done, you go back to your base table and issue the mutation statements you need. Note: You have probably realized by now that ScyllaDB doesn’t do any sort of locking. That said, reads to your view may be inconsistent if concurrent clients modify the base table data. Ideally you want to ensure only a single client can update a given resource_id at a time, given this data modeling, and that way you can ensure you are conflict-free. If you would like to discuss more on potential other solutions and are willing to share additional details of your use case, feel free to DM me. Full disclosure, I am a ScyllaDB employee. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-202-2023-10-29/947 Title: Last week in scylladb.git master (issue #202; 2023-10-29) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the f181ac033a…227136ddf5 range are covered. There were 57 non-merge commits from 12 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-202-2023-10-29/947 ## Headings Structure: H1: Last week in scylladb.git master (issue #202; 2023-10-29) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #202; 2023-10-29) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the f181ac033a…227136ddf5 range are covered. There were 57 non-merge commits from 12 authors in that period. Some notable commits: Major compaction will now merge any sstables streamed in due to decommission or repair before starting compaction. This generates compacted sstables more in line with expectations. A bug in which a read retry after an allocation failure caused a crash was fixed. Tables which store data in object storage now default to UUID generation numbers. This allows dropping the extra UUID that we previously used just to avoid conflicts with integer generation numbers. Writing small sstable components to S3 is now more efficient. The automatic creation of internal keyspaces and tables has been streamlined, resulting in faster start-up. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/major-compaction-by-partition/952 Title: Major Compaction by Partition - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: We’re considering a major compaction in the next few days, due to massive partitions that we are going to shrink and then want to clean up soon afterwards. We believe they are causing performance issues on the cluster. W… Language: en Canonical URL: https://forum.scylladb.com/t/major-compaction-by-partition/952 ## Headings Structure: H1: Major Compaction by Partition H3: Related topics ## Main Content: H1: Major Compaction by Partition H3: Related topics We’re considering a major compaction in the next few days, due to massive partitions that we are going to shrink and then want to clean up soon afterwards. We believe they are causing performance issues on the cluster. When looking at the nodetool compact doc page, I see there is a --partition option. The help text isn’t very helpful here, but does that mean we can join the partition columns of the table to specify only that partition(s) to be compacted? Our primary key is similar to this: So, based on the doc page we could somehow combine foo, bar, and baz to compact a specific partition? We’re also confused as to how that even works at the SSTable level, since there should be multiple partitions per SSTable file. Hah! Good find. We actually removed the --partition option last week in docs: nodetool compact: remove unsupported partition option · scylladb/scylladb@70ba6b9 · GitHub I guess that answers your question FYI @Anna @Botond_Denes Yep, that does indeed answer all of my questions, thank you Felipe. --- ### Page: https://forum.scylladb.com/t/scylladb-in-action-out-now-in-meap/958 Title: ScyllaDB in Action out now in MEAP! - Announcements - ScyllaDB Community NoSQL Forum Meta Description: Hey there, Bo Ingram is writing a book on ScyllaDB! It’s available in the Manning Early Access Program (MEAP) with first 3 chapters. Check it out here: ScyllaDB in Action Cheers, Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-in-action-out-now-in-meap/958 ## Headings Structure: H1: ScyllaDB in Action out now in MEAP! H3: Related topics ## Main Content: H1: ScyllaDB in Action out now in MEAP! H3: Related topics Bo Ingram is writing a book on ScyllaDB! It’s available in the Manning Early Access Program (MEAP) with first 3 chapters. Check it out here: ScyllaDB in Action --- ### Page: https://forum.scylladb.com/t/how-to-make-monitoring-for-existing-grafana-prometheus-in-kubernetes/960 Title: How to make monitoring for existing Grafana & Prometheus in Kubernetes? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, I failed to find the info on how to make monitoring for existing Grafana & Prometheus in Kubernetes. Language: en Canonical URL: https://forum.scylladb.com/t/how-to-make-monitoring-for-existing-grafana-prometheus-in-kubernetes/960 ## Headings Structure: H1: How to make monitoring for existing Grafana & Prometheus in Kubernetes? H3: Related topics ## Main Content: H1: How to make monitoring for existing Grafana & Prometheus in Kubernetes? H3: Related topics Hi, I failed to find the info on how to make monitoring for existing Grafana & Prometheus in Kubernetes. with Scylla Operator we only support the managed monitoring in Monitoring | ScyllaDB Docs . I suppose you can copy the dashboards or make a similar setup based on that but that’s beyond an OSS level of support. --- ### Page: https://forum.scylladb.com/t/entrypoint-url-for-kubernetes-helm/961 Title: Entrypoint URL for Kubernetes Helm? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: What is the entry point for the ScyllaDB to use in the applications on Kubernetes Helm? Language: en Canonical URL: https://forum.scylladb.com/t/entrypoint-url-for-kubernetes-helm/961 ## Headings Structure: H1: Entrypoint URL for Kubernetes Helm? H3: Related topics ## Main Content: H1: Entrypoint URL for Kubernetes Helm? H3: Related topics What is the entry point for the ScyllaDB to use in the applications on Kubernetes Helm? Did you check the Accessing the database doc? There should be a service’s name following the -client convention. Depending where your clients sit, you may also be interested in Exposing ScyllaCluster | ScyllaDB Docs (note this is for v1.11, currently in release candidate) --- ### Page: https://forum.scylladb.com/t/release-scylladb-cloud-1-november-2023/962 Title: [RELEASE] ScyllaDB Cloud - 1 November 2023 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: We are happy to announce a new update to the ScyllaDB Cloud service. New features and enhancements: Cluster Resize enhancements: The Cluster Resize feature was enhanced by allowing you to select multiple DCs at once… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-cloud-1-november-2023/962 ## Headings Structure: H1: [RELEASE] ScyllaDB Cloud - 1 November 2023 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Cloud - 1 November 2023 H3: Related topics We are happy to announce a new update to the ScyllaDB Cloud service. New features and enhancements: --- ### Page: https://forum.scylladb.com/t/scylla-oss-5-2-9-fails-to-start-with-dynatrace-agent-installed/968 Title: Scylla OSS 5.2.9 fails to start with Dynatrace agent installed - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi ScyllaDB team, For instance monitoring we use Dynatrace Oneagent software. This software needs to start before ScyllaDB starts. So we install OneAgent, starte OneAgent service then ScyllaDB. But Scylla OSS 5.2.9 fai… Language: en Canonical URL: https://forum.scylladb.com/t/scylla-oss-5-2-9-fails-to-start-with-dynatrace-agent-installed/968 ## Headings Structure: H1: Scylla OSS 5.2.9 fails to start with Dynatrace agent installed H3: Related topics ## Main Content: H1: Scylla OSS 5.2.9 fails to start with Dynatrace agent installed H3: Related topics For instance monitoring we use Dynatrace Oneagent software. This software needs to start before ScyllaDB starts. So we install OneAgent, starte OneAgent service then ScyllaDB. But Scylla OSS 5.2.9 fails to start and dumps below log. Can someone help resolving this error please. systemd-coredump: Process 29201 (scylla) of user 991 dumped core.#012#012Stack trace of thread 29201:#012#0 0x00007f4baf32f2a0 n/a (liboneagentproc.so)#012#1 0x0000000004fb1cda _Unwind_RaiseException (scylla)#012#2 0x00007f4bada4674b __cxa_throw (libstdc++.so.6)#012#3 0x00007f4bada3a772 __cxa_guard_acquire.cold (libstdc++.so.6)#012#4 0x0000000004fb1c67 dl_iterate_phdr (scylla)#012#5 0x00007f4baf340a66 n/a (liboneagentproc.so)#012#6 0x00007f4baf346eaa n/a (liboneagentproc.so)#012#7 0x00007f4baf347130 n/a (liboneagentproc.so)#012#8 0x00007f4baf3437d4 dlsym (liboneagentproc.so)#012#9 0x0000000004fb1c7c dl_iterate_phdr (scylla)#012#10 0x00007f4baf340a66 n/a (liboneagentproc.so)#012#11 0x00007f4baf346eaa n/a (liboneagentproc.so)#012#12 0x00007f4baf347130 n/a (liboneagentproc.so)#012#13 0x00007f4baf3426cc fopen (liboneagentproc.so)#012#14 0x00007f4baeb96603 _gnutls_fips_mode_enabled (libgnutls.so.30)#012#15 0x00007f4baeb8630a _gnutls_rnd_preinit (libgnutls.so.30)#012#16 0x00007f4baeb7615d _gnutls_global_init (libgnutls.so.30)#012#17 0x00007f4baeb44fa3 lib_init (libgnutls.so.30)#012#18 0x00007f4baf548cde call_init (ld.so)#012#19 0x00007f4baf548dcc _dl_init (ld.so)#012#20 0x00007f4baf55f8e0 n/a (ld.so) We have seen problems with Dynatrace agent running alongside Scylla before. As it hooks to processes, it can cause not only Scylla to stall but in worse scenarios it may even cause the process to get hang. The general recommendation is simply not to run 3rd party processes along with Scylla. But if you really really must it, then contact your Dynatrace support team and ask for instructions on how to tell the agent not to hook to the Scylla process and also for the proper way to pin the agent to use specific CPUs, that Scylla is/will not actively (be) using. Newest version of Dynatrace OneAgent was fixed and now coexist with Scylla > 5.2. ● oneagent.service - Dynatrace OneAgent Loaded: loaded (/etc/systemd/system/oneagent.service; enabled; vendor preset: enabled) Active: active (running) since Fri 2024-06-21 10:28:03 CEST; 4min 49s ago Process: 1065 ExecStart=/opt/dynatrace/oneagent/agent/initscripts/oneagent start (code=exited, status=0/SUCCESS) Main PID: 1178 (oneagentwatchdo) Tasks: 50 (limit: 38413) Memory: 244.8M CGroup: /system.slice/oneagent.service ├─1178 /opt/dynatrace/oneagent/agent/lib64/oneagentwatchdog -bg -config=/opt/dynatrace/oneagent/agent/conf/watchdog.conf ├─1187 oneagentos -Dcom.compuware.apm.WatchDogTimeout=900 -watchdog.restart_file_location=/var/lib/dynatrace/oneagent/a> ├─1203 oneagentplugin -Dcom.compuware.apm.WatchDogTimeout=900 -Dcom.compuware.apm.WatchDogPort=50001 ├─1205 oneagentloganalytics -Dcom.compuware.apm.WatchDogTimeout=900 -Dcom.compuware.apm.WatchDogPort=50002 ├─1207 oneagentnetwork -Dcom.compuware.apm.WatchDogTimeout=900 -Dcom.compuware.apm.WatchDogPort=50003 ├─1273 /opt/dynatrace/oneagent/agent/lib64/oneagenteventstracer --logdir /var/log/dynatrace/oneagent/os --cputimeoffset> └─1958 /opt/dynatrace/oneagent/agent/lib64/oneagentebpfdiscovery --log-dir /var/log/dynatrace/oneagent/os/ --log-no-std> Jun 21 10:28:03 systemd[1]: Starting Dynatrace OneAgent… Jun 21 10:28:03 oneagent[1182]: 10:28:03 Dynatrace OneAgent service started. Jun 21 10:28:03 systemd[1]: Started Dynatrace OneAgent. ● scylla-server.service - Scylla Server Loaded: loaded (/lib/systemd/system/scylla-server.service; enabled; vendor preset: enabled) Drop-In: /etc/systemd/system/scylla-server.service.d └─capabilities.conf, dependencies.conf, sysconfdir.conf Active: active (running) since Fri 2024-06-21 10:28:14 CEST; 4min 46s ago Process: 861 ExecStartPre=/opt/scylladb/scripts/scylla_prepare (code=exited, status=0/SUCCESS) Main PID: 1069 (scylla) Status: “serving” Tasks: 8 (limit: 38413) Memory: 29.2G CGroup: /scylla.slice/scylla-server.slice/scylla-server.service └─1069 /usr/bin/scylla --log-to-syslog 1 --log-to-stdout 0 --default-log-level info --network-stack posix --io-properti> Jun 21 10:28:39 scylla[1069]: [shard 0:main] init - Scylla version 5.4.7-0.20240602.98139a8716ab initia Thanks for the update. But we have upgraded to 6.0.0 and we don’t see this issue anymore --- ### Page: https://forum.scylladb.com/t/scylladb-cloud-serverless-cluster-creation-fails-with-terraform/972 Title: ScyllaDB Cloud Serverless - Cluster creation fails with Terraform - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: The cluster creation fails with the following error - scylladbcloud_serverless_cluster.example: Creating... ╷ │ Error: error creating serverless cluster: Error "040702": General error creating the cluster (http status 2… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-cloud-serverless-cluster-creation-fails-with-terraform/972 ## Headings Structure: H1: ScyllaDB Cloud Serverless - Cluster creation fails with Terraform H3: Related topics ## Main Content: H1: ScyllaDB Cloud Serverless - Cluster creation fails with Terraform H3: Related topics The cluster creation fails with the following error - Hi there @outlander , Sorry for the late response. Is it possible that your serverless free trial is over or you don’t have credits? Can you manually create a new serverless cluster if you go to ScyllaDB Cloud ? --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-203-2023-11-05/973 Title: Last week in scylladb.git master (issue #203; 2023-11-05) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 227136ddf5…6cc5bcae80 range are covered. There were 127 non-merge commits from 22 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-203-2023-11-05/973 ## Headings Structure: H1: Last week in scylladb.git master (issue #203; 2023-11-05) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #203; 2023-11-05) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 227136ddf5…6cc5bcae80 range are covered. There were 127 non-merge commits from 22 authors in that period. Some notable commits: To fix an edge case, when a node joins the raft group, it now first waits for the first node to enter the NORMAL state.1 An ABBA-style deadlock was fixed for this scenario: if an Alternator table used strong write isolation and had a global secondary index, then index updates could wait on the base table update, and vice versa. This also happens with materialized views on a table updated using lightweight transactions (LWT). The embedded benchmarking tool, scylla perf-simple-query, now supports tablets. The test suite now verifies that nodes that shut down have not crashed while doing so. The native nodetool implementation now supports additional commands: cleanup, clearsnapshots, and listsnapshots. We still use the java-based nodetool by default. The option to control how replication strategies are allowed or denies has been changed. The container image now avoids installing “suggested” packages, to reduce its size. When using raft-based topology management, ScyllaDB will now roll back the last operation (e.g. bootstrap or decommission) automatically. When using raft topology mode, we will now automatically retry raft transactions on failure, so that nodes can be bootstrapped in parallel. Note that streaming is still serialized. Streaming now checks more closely if a table was dropped while it was streamed. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/new-scylladb-sample-app-video-streaming/975 Title: New ScyllaDB sample app: video streaming - Announcements - ScyllaDB Community NoSQL Forum Meta Description: We’ve recently published a new ScyllaDB sample app for video streaming. The app uses NextJS (with Blitz.jS) and Material-UI on the frontend with ScyllaDB Cloud serving requests on the backend. Have a look at the repo, a… Language: en Canonical URL: https://forum.scylladb.com/t/new-scylladb-sample-app-video-streaming/975 ## Headings Structure: H1: New ScyllaDB sample app: video streaming H3: Related topics ## Main Content: H1: New ScyllaDB sample app: video streaming H3: Related topics We’ve recently published a new ScyllaDB sample app for video streaming. The app uses NextJS (with Blitz.jS) and Material-UI on the frontend with ScyllaDB Cloud serving requests on the backend. Have a look at the repo, and see how ScyllaDB can be used to build a low-latency video streaming application with modern web tools like NextJS. Video streaming database schema used in the project: Hello, I need a similar scheme, but in the user history there will be 1000 records per day, that is, 350,000 records per year. I saw that in cassandra it is undesirable to have more than 100,000 entries per partition. Can I leave the same scheme in scylladb or do I need to add a “bucket” column, say, by month? As long as all the nodes have roughly the same amount of data you should be good to go. I’d suggest using some kind of ID as the partitioning key and another ID and/or the timestamp column as the clustering key depending on your queries. I’m happy to give feedback if you provide the schema you’re planning to use. CREATE TABLE tests ( id timeuuid, user_id uuid, status text, created_at timestamp, some_info…, PRIMARY KEY (id) ); CREATE TABLE user_tests ( user_id uuid, test_id timeuuid, status text, created_at timestamp, some_info…, PRIMARY KEY (user_id, created_at) ) WITH CLUSTERING ORDER BY(created_at DESC); Q1: select * from user_tests where user_id = ? limit 30; Q2: select * from user_tests where user_id = ? and created_at < ? and created_at > ? limit 30; And user_tests potentially may be 1000 per day, and need to store it for years. Yes that schema with those queries should work really well! We published a blog post detailing the data modeling process for the video streaming app: How to Build a Low-Latency Video Streaming App with ScyllaDB & NextJS - ScyllaDB --- ### Page: https://forum.scylladb.com/t/issues-in-scylla-installation/977 Title: Issues in Scylla Installation - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I tried to install the Scylla DB on my ubuntu 23.04. I followed the official documentation to install it but facing issues in the last command. The issue is given below: The following packages have unmet dependencies: s… Language: en Canonical URL: https://forum.scylladb.com/t/issues-in-scylla-installation/977 ## Headings Structure: H1: Issues in Scylla Installation H3: Related topics ## Main Content: H1: Issues in Scylla Installation H3: Related topics I tried to install the Scylla DB on my ubuntu 23.04. I followed the official documentation to install it but facing issues in the last command. The issue is given below: The following packages have unmet dependencies: scylla-tools-core : Depends: python (>= 2.7) but it is not installable or python2 but it is not installable E: Unable to correct problems, you have held broken packages. When I searched on the Internet, it shows that Scylla DB will run on Python3. but I am getting error in this and Even not able to install python 2 on the latest ubuntu. Can someone help to resolve the issue? TIA!! Which version are you trying to install? We have moved to python3 in a recent release. I tried to install ScyllaDB 5.2 version (from here: https://www.scylladb.com/download/#open-source). scylla-tools-core package was asking for previous python version. Will scylla-tools-core package work on python3 in recent release? Note: Using this way only, I was able to use scylla on Ubuntu 22.04 because Python2 package was there unlike Ubuntu 23.04 Looks like Python 3 support is only coming in the next release (5.4). In the meanwhile please use an OS which still has Python 2 available. See our OS compatibility page for guidance. --- ### Page: https://forum.scylladb.com/t/what-could-be-the-best-strategy-to-horizontally-upscale-the-scylla-cluster-in-the-anticipation-of-high-load-due-to-an-event/979 Title: What could be the best strategy to horizontally upscale the scylla cluster in the anticipation of high load due to an event - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: what could be the best strategy to horizontally upscale the scylla cluster in the anticipation of high load due to an event Language: en Canonical URL: https://forum.scylladb.com/t/what-could-be-the-best-strategy-to-horizontally-upscale-the-scylla-cluster-in-the-anticipation-of-high-load-due-to-an-event/979 ## Headings Structure: H1: What could be the best strategy to horizontally upscale the scylla cluster in the anticipation of high load due to an event H3: Related topics ## Main Content: H1: What could be the best strategy to horizontally upscale the scylla cluster in the anticipation of high load due to an event H3: Related topics what could be the best strategy to horizontally upscale the scylla cluster in the anticipation of high load due to an event Depends on your topology, but you will typically want to scale the cluster in increments according to your RF. The usual practice is to have RF=3 where each replica is located under a different availability zone. Thus, you would scale add one node to each AZ to sustain the load you are expecting. Unless you are planning to scale up your cluster, it also makes sense to remain under the same instance as your existing ones. And of course, remember to check out the docs on Adding a New Node Into an Existing ScyllaDB Cluster (Out Scale) | ScyllaDB Docs --- ### Page: https://forum.scylladb.com/t/realese-scylla-5-4-rc1-part-1/980 Title: [REALESE] Scylla 5.4 RC1 - part 1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce ScyllaDB Open Source 5.4 RC1, the first Release Candidate for the ScyllaDB Open Source 5.4 minor release. ScyllaDB 5.4 introduces full support for Repair Base Node Operations (RB… Language: en Canonical URL: https://forum.scylladb.com/t/realese-scylla-5-4-rc1-part-1/980 ## Headings Structure: H1: [REALESE] Scylla 5.4 RC1 - part 1 H2: Repair Based Node Operations (RBNO) H2: Node Level Metrics H2: Guardrails H2: Security H2: Deprecated and removed features H2: UDF / UDA - Preview H2: New nodetool implementation - Experimental H2: Object Storage Support - Experimental H3: Related topics ## Main Content: H1: [REALESE] Scylla 5.4 RC1 - part 1 H2: Repair Based Node Operations (RBNO) H2: Node Level Metrics H2: Guardrails H2: Security H2: Deprecated and removed features H2: UDF / UDA - Preview H2: New nodetool implementation - Experimental H2: Object Storage Support - Experimental H3: Related topics The ScyllaDB team is pleased to announce ScyllaDB Open Source 5.4 RC1, the first Release Candidate for the ScyllaDB Open Source 5.4 minor release. ScyllaDB 5.4 introduces full support for Repair Base Node Operations (RBNO), experimental consistent topology update, experimental S3 backend and many more improvements and bug fixes. Consistent schema management using Raft is now the default not only for new clusters, but also for clusters upgrading from an older version User Defined Function(UDF) feature is promoted to Preview. Find the ScyllaDB Open Source 5.4 repository for your Linux distribution here. ScyllaDB 5.4 RC1 Docker is also available. We encourage you to run ScyllaDB 5.4 release candidates on your test environments; this will help ensure that an upgrade to ScyllaDB 5.4 General Availability will proceed smoothly with your workload. Use the release candidate with caution; RC1 is not production-ready yet. You can help stabilize ScyllaDB Open Source 5.4 by reporting bugs here. Only the last two minor releases of the ScyllaDB Open Source project are supported. Once ScyllaDB Open Source 5.4 is officially released, only ScyllaDB Open Source 5.4 and ScyllaDB 5.2 will be supported, and ScyllaDB 5.1 will be retired. Note that we skipped 5.3 release. Repair Based Node Operations were introduced as an experimental feature in ScyllaDB Open Source 4.0. They use repair to stream data for node-operations like replace, bootstrap and others. In 5.4, RBNO is enabled by default for all operations: remove node,rebuild,bootstrap, and decommission. Replace node operation was already enabled by default. For more about Repair Based Node Operations see this Scylla Summit 2022 session by Asias He Most ScyllaDB metrics are per-shard, per-node, but not for a specific table. We now export some per-table metrics. These are exported once per node, not per shard, to reduce the number of metrics. #2198 Guardrails is a framework to protect ScyllaDB users and admins from common mistakes and pitfalls. In this release ScyllaDB includes a new guardrail on the replication factor. It is now possible to specify the minimum replication factor for new keyspaces via a new configuration item #8891. This matches the same functionality in Apache Cassandra #CASSANDRA-14557 The new RF guardrails include the following configuration: More Guardrails are expected in upcoming releases. See more on certificate-authentication docs. are initialized. See two new config parameters auth_superuser_name, auth_superuser_salted_password below. Wasm-based User Defined Functions (UDFs) and User Defined Aggregates (UDAs) were introduced as experimental in ScyllaDB 5.1, and we are now promoting them to Preview. The CQL syntax is compatible with Apache Cassandra. Examples: CREATE FUNCTION sample ( arg int ) …; CREATE FUNCTION sample ( arg text ) …; See UDF/UDA documentation for more information. Preview features aren’t production-ready, but are made available on a “preview” basis so that customers can get early access and provide feedback. Unlike Experimental features, we are committed to the backward compatibility of a preview feature API. We encourage users to use the preview feature on staging and testing clusters while we continue to extend the tests till they can graduate to production-ready. Below are improvements and bug fixes in UDF/UDA. The scylla executable can now act as nodetool by executing “scylla nodetool ”. With time the new scylla nodetool will replace the legacy Java nodetool completely. The new implementation is fully backward compatible with existing docs and legacy nodetool. The following commands are currently implemented: For the latest status of nodetool replacement progress see #15588 You can store an entire ScyllaDB Keyspace to Amazon S3, or a compatible object store. Enable the feature by: Enabling keyspace-storage-options experimental --experimental-features=keyspace-storage-options Setup the S3 endpoint parameter and credentials Create a keyspace with S3 storage option: CREATE KEYSPACE with STORAGE = { 'type': 'S3', 'endpoint': '$endpoint_name', 'bucket': '$bucket' } Note: at this phase, obj-storage is not ready for production, and should be used for testing only. Note2: there is no ALTER support for the STORAGE parameter. --- ### Page: https://forum.scylladb.com/t/realese-scylla-5-4-rc1-part-2/981 Title: [REALESE] Scylla 5.4 RC1 - part 2 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: More Improvements CQL API CQL table columns that have the list data type aren’t allowed to contain NULLs, but in certain situations list values in CQL literals or bind variables are allowed to contain NULLs (for exampl… Language: en Canonical URL: https://forum.scylladb.com/t/realese-scylla-5-4-rc1-part-2/981 ## Headings Structure: H1: [REALESE] Scylla 5.4 RC1 - part 2 H2: More Improvements H3: CQL API H3: Amazon DynamoDB Compatible API (Alternator) H3: Strongly Consistent Schema Management with Raft H3: Strongly Consistent Topology Updates with Raft - Experimental H3: Related topics ## Main Content: H1: [REALESE] Scylla 5.4 RC1 - part 2 H2: More Improvements H3: CQL API H3: Amazon DynamoDB Compatible API (Alternator) H3: Strongly Consistent Schema Management with Raft H3: Strongly Consistent Topology Updates with Raft - Experimental H3: Related topics CQL table columns that have the list data type aren’t allowed to contain NULLs, but in certain situations list values in CQL literals or bind variables are allowed to contain NULLs (for example, in LWT IF conditions that use the IN operator). The type system was relaxed to accept NULLs where this is allowed. Previously, these cases were handled by hard-to-maintain workarounds. The CQL USING TTL clause allows one to specify an INSERT or UPDATE’s time-to-live property, after which the cells are automatically deleted. TTL 0 was misinterpreted as the default TTL (which happens to be unlimited, usually) rather than an explicitly unlimited TTL. This is now fixed. #6447 The C-style cast syntax ((type) expression) can now be applied to bind variables ((type) ? or (type) :var) to explicitly specify the type of bind variables Example: blob_column = (blob)(int)12323 Error messages for incorrect usage of the CQL TOKEN() function have been improved. #13468 The check for altering permissions of functions in the system keyspace has been tightened. Error messages involving the CQL token function have been improved. Error messages involving CQL expressions will not be printed in a more user-friendly way. Previously they contained some debug information. Change Data Capture (CDC) exports updates to the database as a table containing changes. One option is to capture not only the change, but also the state of the row before it was changed. In some cases, in a lightweight transaction (LWT) change, the preimage could return the state of the row after the change instead of before the change. This is now fixed. #12098 The NetworkTopologyStrategy replication strategy will now reject an empty value for the replication factor. #13986 Materialized views require the “IS NOT NULL” qualifier on primary key elements, but also accept (and ignore) the qualifier on regular columns. The qualifier is now rejected when applied to regular columns. A configuration variable allows you to warn about the rejected clause, emit an error and fail the request, or ignore it. #10365 The count(column) function is supposed to only count cells where the column is not NULL. A regression caused count(column) to behave like count(*) for collection, tuple, and user-defined column types. This is now fixed. #14198. When performing the last-write-wins rule comparison, if the timestamp of the two versions being compared was equal, ScyllaDB first compared the cell value and then the expiration time (TTL). This is compatible with earlier versions of Cassandra. However, this could cause a NULL value to appear if the cell was overwritten with the same timestamp but a different TTL. The algorithm was changed to compare the cell value last, and check all the other metadata first, resulting in fewer surprising results. It is also compatible with current Cassandra versions. #14182 A GROUP BY query ought to return one row per group, except when all rows of a group are filtered out. However, ScyllaDB returned a row even for fully-filtered groups. This is now fixed, and ScyllaDB will not emit rows for filtered groups. #12477 In older versions of ScyllaDB, different clauses of CQL statements were processed using different code bases. ScyllaDB is gradually moving towards a single code base for processing expressions. It is now the SELECT clause’s turn, moving us closer to the goal of a unified expression syntax. As this is an internal refactoring, there are no user visible changes, apart from some names of fields in SELECT JSON statements changing (specifically, if those fields are function evaluations). A recent regression when using GROUP BY together with the ttl() and writetime() pseduo-functions was fixed. #14715 There is a new SELECT MUTATION_FRAGMENTS statement that allows seeing where the data that composes a selection comes from. Normally, cache, sstable, and memtable data are merged before output, but with this variant one can see the original source of the data. This is intended for forensics and is not a stable API. #11130 The CQL grammar incorrectly accepted nonsensical empty limit clauses such as SELECT * FROM tab LIMIT;. The errors were discovered later in processing, but with unhelpful error messages. They are now rejected. #14705. The CQL grammar incorrectly accepted nonsensical INSERT JSON statements such as INSERT INTO tab JSON;, causing a crash. This is now fixed. #14709 A mistake in function type inference, which could lead the CQL statements to claim there is ambiguity when in fact there is none, was fixed. The format of the timestamp data type is now compatible with Cassandra. #14518 In CQL, a few functions for dealing with counter types were added. #14501 A SELECT statement that has the DISTINCT keyword and also GROUP BY on clustering keys is now rejected. DISTINCT implies only selecting the partition key and static rows, so grouping on the clustering keys is nonsensical. #12479 When ALTERing a table, the compaction strategy options are now validated. #2336 A bug in the fromJson() CQL function when operating on NULL operands has been fixed #7912 The DESCRIBE statement now includes user defined types and functions #14170 The column names for SELECT CAST(b AS int) and similar expressions have been adjusted to match Cassandra. #14508 In some cases where a bind variable was used both for the partition key and to match a non-key column, ScyllaDB would not generate correct partition key routing for the driver. This is now fixed. #15374 A map value, when parsed from its JSON representation, did not parse the key correctly. This is now fixed. #7949 SSTable compression can be configured with a chunk size, with larger chunks trading less efficient I/O and higher latency for higher compression ratios. The chunk size is now capped at 128 kB, to avoid running out of memory. #9933 Alternator is ScyllaDB’s implementation of the DynamoDB API. Strongly Consistent Schema Management with Raft became the default for new clusters in ScyllaDB 5.2. In this release it is enabled by default when upgrading existing clusters. If you do not want to enable Raft, you should explicitly disable it in scylla.yaml of each node before the upgrade. #13980 Below are additional related fixes and updates: This release includes an experimental Strongly Consistent Topology Updates. To enable it, use the new consistent-topology-changes flag. Below are additional related fixes and updates: --- ### Page: https://forum.scylladb.com/t/realese-scylla-5-4-rc1-part-3/982 Title: [REALESE] Scylla 5.4 RC1 - part 3 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: Performance and stability The row cache will now purge expired tombstones before populating the cache, removing the performance impact of scanning tombstones. Note that non-expired tombstones are still loaded. ScyllaDB … Language: en Canonical URL: https://forum.scylladb.com/t/realese-scylla-5-4-rc1-part-3/982 ## Headings Structure: H1: [REALESE] Scylla 5.4 RC1 - part 3 H3: Performance and stability H3: Operations H3: Deployment and install H3: Tools H3: Configuration Updates H3: Admin REST API H3: Build H3: Monitoring, tracing and logging H3: Related topics ## Main Content: H1: [REALESE] Scylla 5.4 RC1 - part 3 H3: Performance and stability H3: Operations H3: Deployment and install H3: Tools H3: Configuration Updates H3: Admin REST API H3: Build H3: Monitoring, tracing and logging H3: Related topics The CQL shell, cqlsh, has been separated into its own repository. As part of that change, cqlsh is now compatible with Python 3. CQLSh is now available as a Docker image, and in PiPy, allowing you to easily use it when you do not need the entire ScyllaDB server, for example with Scylla Cloud. The port option in SSTableLoader was fixed. The cassandra-stress benchmarking tool’s -log hdrfile=… option now works with Java 11 Scylla process --list-tools option now correctly lists all tools invocable via the scylla binary. The JMX support application, used to support nodetool, now runs under Java 11. The scylla sstable tool now has more ways to obtain the schema. #10126. See Scylla SSTable docs for more info. A bug in the nodetool command to disable auto compaction has been fixed #13553 The nodetool checkAndRepairCdcStreams is used to align CDC streams with the cluster topology. It now works when topology is under Raft control. The nodetool refresh command gained the –primary-replica-only option. The sylla sstable tool now supports the scrub operation, enabling offline (and off-node) scrubbing of sstables. #14203 The cassandra-stress tool now supports the Java driver’s rack-aware policy. This can reduce cloud inter availability zone networking costs, with the downside of less even load balancing if care isn’t taken to balance the application. The setup utility supported an --online-discard switch to enable/disable online discard, but it did not actually work. This is now fixed. #14963 The nodetool stop RESHAPE command is supposed to stop the reshape operation, but in fact only aborted running reshape compactions, which were promptly restarted. It now aborts the entire operation as expected. #15058 The scylla.yaml configuration items are now documented in the documentation website. New and updated configuration options: It is now possible to disable configuration changes via the system.config virtual table using a configuration parameter. Use this option to prevent runtime configuration changes via CQL.#14355 task_ttl_in_seconds - Task Manager option: time for which information about finished tasks stays in memory. RF Guardrail config values (see above) Stream_plan_ranges_percentage is renamed to stream_plan_ranges_fraction Cache_index_pages is no enabled by default, with an index_cache_fraction value of 0.2 Index_cache_fraction is the maximum fraction of cache memory permitted for use by index cache. Clamped to the [0.0; 1.0] range. Must be small enough to not deprive the row cache of memory, but should be big enough to fit a large fraction of the index. The default value 0.2 means that at least 80% of cache memory is reserved for the row cache, while at most 20% is usable by the index cache. x_log2_compaction_groups option to controls static number of compaction groups per table per shard - is removed Live_updatable_config_params_changeable_via_cql - If set to true, configuration parameters defined with LiveUpdate option can be updated in runtime with CQL (more above) Enable_node_aggregated_table_metrics - Enable aggregated per node, per keyspace and per table metrics reporting, applicable if enable_keyspace_column_family_metrics is false. Default True. Enable_compacting_data_for_streaming_and_repair - Enable the compacting reader, which compacts the data for streaming and repair (load and stream included) before sending it to, or synchronizing it with peers. Can reduce the amount of data to be processed by removing dead data, but adds CPU overhead. Default: True. Table_digest_insensitive_to_expiry - When enabled, per-table schema digest calculation ignores empty partitions. Default: True. Schema_commitlog_segment_size_in_mb - ScyllaDB uses a separate commitlog, called the schema commitlog, for schema changes and topology operations in order to reduce the latency of these operations. The segmented size of the schema commitlog has been raised from 32MB to 128MB in order to avoid problems with large numbers of tables, as the entire schema must fit in a single segment. Stream_plan_ranges_percentage - Specify the percentage of ranges to stream in a single stream plan. Value is between 0 and 1. Default 0.1 #14191 alternator_describe_endpoints - Overrides the behavior of Alternator’s DescribeEndpoints operation. An empty value (the default) means DescribeEndpoints will return the same endpoint used in the request. The string ‘disabled’ disables the DescribeEndpoints operation. Any other string is the fixed value that will be returned by DescribeEndpoints operations. This was require to bypass AWS SDK issue When DynamoDB DescribeEndpoints is used, wrong scheme may be tacked on the result · Issue #2554 · aws/aws-sdk-cpp · GitHub Table_digest_insensitive_to_expiry - When enabled, per-table schema digest calculation ignores empty partitions. Default: True. Auth_certificate_role_queries - Regular expression used by CertificateAuthenticator to extract role name from an accepted transport authentication certificate subject info. See more in the Security section. Auth_superuser_name - Initial authentication super username. Ignored if authentication tables already contain a super user. Auth_superuser_salted_password - Initial authentication super user salted password. Create using mkpassword or similar. The hashing algorithm used must be available on the node host. Ignored if authentication tables already contain a super user password. strict_is_not_null_in_views - In materialized views, restrictions are allowed only on the view’s primary key columns. In old versions Scylla mistakenly allowed IS NOT NULL restrictions on columns which were not part of the view’s primary key. These invalid restrictions were ignored. This option controls the behavior when someone tries to create a view with such invalid IS NOT NULL restrictions. Can be true, false, or warn. Default: True. object_storage_config_file - part of the new experimental object store feature (above). Optionally, read object-storage endpoints config from file. “tablets” - new experimental flag. relabel_config_file - optionally, read relabel config from file. Schema_commitlog_directory - The directory where the schema commit log is stored. This is a special commitlog instance used for schema and system tables. For optimal write performance, it is recommended the commit log be on a separate disk partition (ideally, a separate physical device) from the data file directories. Nodeops_watchdog_timeout_seconds - Time in seconds after which node operations abort when not hearing from the coordinator. Default 120s. Nodeops_heartbeat_interval_seconds - Period of heartbeat ticks in node operations. Default 10. Query timeouts in configuration (e.g. read_request_timeout_in_ms) can now be hot-reloaded using SIGHUP. #12232 ScyllaDB has an error injection facility, used by QA to test error paths. It can now be enabled via configuration. Use with caution! The experimental flag used to enable consistent topology changes has been renamed from “raft” to "consistent-topology-changes. #14145 The schema commitlog size was accidentally set to 10TB, it’s now set to a reasonable size. The --max-io-requests init option, which has been obsolete for quite some time, was removed. It’s now possible to disable and enable tombstone compaction on a per-node basis using a REST API endpoint. This is useful if the user knows that all DELETEs were performed with CL=ALL and so there is no risk of data resurrection. The REST API that accepts sstable generation numbers now uses a string value, in preparation for using UUID generations. The type of the “generation” field of “sstable” in the return value of RESTful API entry point at “/storage_service/sstable_info” is changed from “long” to “string”. The API for performing sstable cleanup, and use by nodetool cleanup, will now wait for staging sstables to be cleaned up too. The hints synchronization point API allows an external user to wait for hints to replay. Misuse of the API cookie could lead to unbounded memory usage; the cookie is now protected with a checksum. #9405 The --experimental flag was removed. It was replaced some time ago with --experimental-features., which provides fine-grained control about which experimental features are enabled. There is a new REST API call to recalculate schema digests. It can be useful to heal some schema disagreement problems. #15380 Metrics updates below: --- ### Page: https://forum.scylladb.com/t/release-scylladb-cloud-8-november-2023/983 Title: [RELEASE] ScyllaDB Cloud - 8 November 2023 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: We are happy to announce a new update to the ScyllaDB Cloud service. New features and enhancements: We’ve added support for me-central1 (Doha) and me-central2 (Dammam) GCP regions. We’ve added support for a GCP devel… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-cloud-8-november-2023/983 ## Headings Structure: H1: [RELEASE] ScyllaDB Cloud - 8 November 2023 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Cloud - 8 November 2023 H3: Related topics We are happy to announce a new update to the ScyllaDB Cloud service. New features and enhancements: --- ### Page: https://forum.scylladb.com/t/how-to-disable-experimental-features-on-already-running-cluster/986 Title: How to disable experimental features on already running cluster - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I have a 5.2.9 cluster running. Currently in the scylla config I have the following experimental features enabled: experimental_features: - alternator-ttl - alternator-streams How can I disable these features on my… Language: en Canonical URL: https://forum.scylladb.com/t/how-to-disable-experimental-features-on-already-running-cluster/986 ## Headings Structure: H1: How to disable experimental features on already running cluster H3: Related topics ## Main Content: H1: How to disable experimental features on already running cluster H3: Related topics I have a 5.2.9 cluster running. Currently in the scylla config I have the following experimental features enabled: How can I disable these features on my cluster? I have tried removing them from scylla.yml and restart the service but the node will not join the cluster anymore with following error: (Feature 'ALTERNATOR_STREAMS' was previously enabled in the cluster but its support is disabled by this node. Set the corresponding configuration option to enable the support for the feature.) Enabling Alternator Streams means that the feature flag bit is set in gossip, so you can’t disable it any longer. But this shouldn’t be a problem, as you simply can avoid using it by not having any Alternator tables with Streams enabled. Just keep the ball rolling, or re-create the cluster with the feature disabled. Unrelated, yet worth to mention anyway: Alternator TTL is production ready on 5.2, so you no longer have to enable it manually on scylla.yaml --- ### Page: https://forum.scylladb.com/t/how-to-avoid-rust-drivers-prepared-values-limit/988 Title: How to avoid Rust driver's prepared values limit - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: The Rust driver’s session.execute seems to have a prepared values limited of 16. How does one use prepared statements to insert/update tables with numerous columns? For reference please see the update function below whic… Language: en Canonical URL: https://forum.scylladb.com/t/how-to-avoid-rust-drivers-prepared-values-limit/988 ## Headings Structure: H1: How to avoid Rust driver's prepared values limit H3: Related topics ## Main Content: H1: How to avoid Rust driver's prepared values limit H3: Related topics The Rust driver’s session.execute seems to have a prepared values limited of 16. How does one use prepared statements to insert/update tables with numerous columns? For reference please see the update function below which includes 16 values. If I add one additional column, error[E0277] is surfaced. Session::execute accepts anything that implements ValueList (warning: there is ongoing refactor, so soon the trait will be a bit different - but it doesn’t really matter in this case). Apart from tuples (up to length 16), there are implementations for slice and Vec, so you could use an array and pass it as slice - the problem is that arrays, unlike tuples, need to have all elements of the same type. To solve it, use CqlValue enum. Alternative solution would be to use our ValueList derive macro to implement ValueList trait for your struct - then you could just pass your struct to execute. That’s exactly the information this newby needed. With your instruction, I found the relevant information and examples in the driver documentation. Query values | ScyllaDB Docs Many thanks! --- ### Page: https://forum.scylladb.com/t/release-scylla-operator-1-11-0/989 Title: [RELEASE] Scylla Operator 1.11.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of Scylla Operator 1.11.0. Scylla Operator is an open-source project that helps ScyllaDB Open Source and ScyllaDB Enterprise users run ScyllaDB on Kubernetes. The Scy… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-operator-1-11-0/989 ## Headings Structure: H1: [RELEASE] Scylla Operator 1.11.0 H1: Notable changes H1: Supported versions H1: Upgrade instructions H2: Getting started with Scylla Operator H2: Related Links H3: Related topics ## Main Content: H1: [RELEASE] Scylla Operator 1.11.0 H1: Notable changes H1: Supported versions H1: Upgrade instructions H2: Getting started with Scylla Operator H2: Related Links H3: Related topics The ScyllaDB team is pleased to announce the release of Scylla Operator 1.11.0. Scylla Operator is an open-source project that helps ScyllaDB Open Source and ScyllaDB Enterprise users run ScyllaDB on Kubernetes. The Scylla Operator manages ScyllaDB clusters deployed to Kubernetes and automates tasks related to operating a ScyllaDB cluster, like installation, vertical and horizontal scaling, as well as rolling upgrades. Scylla Operator 1.11.0 improves stability and brings a few features. As with all of our releases, all API changes are backward compatible. For more changes and details check out the GitHub release notes. Upgrading from v1.10.x with kubectl apply doesn’t require any extra action, just take the manifest from v1.11.0 tag and substitute the released image. Using helm requires a mandatory manual step for every release because helm can’t handle CRDs updates. For details, see our upgrade documentation. We’ll welcome your feedback! Feel free to open an issue or reach out on the #scylla-operator channel in Scylla User Slack. We noticed an issue with upgrading from Operator v1.10 to v1.11.0, please wait with the upgrade until we release 1.11.1. ScyllaCluster’s created before the upgrade may get stuck on maintenance operations like rolling restarts, upgrades, scaling etc. ScyllaCluster’s created after the upgrade are not affected. 1.11.1 has been released, issue mentioned above is fixed there: --- ### Page: https://forum.scylladb.com/t/scylla-5-2-load-and-stream/991 Title: Scylla 5.2 Load and Stream - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello, I am trying to understand this feature. Nodetool refresh | ScyllaDB Docs What I don’t understand is if you are going to from 6 nodes(RF=3) to 4 nodes(RF=2), do you need to need to load data from all 6 nodes even… Language: en Canonical URL: https://forum.scylladb.com/t/scylla-5-2-load-and-stream/991 ## Headings Structure: H1: Scylla 5.2 Load and Stream H3: Related topics ## Main Content: H1: Scylla 5.2 Load and Stream H3: Related topics I am trying to understand this feature. Nodetool refresh | ScyllaDB Docs What I don’t understand is if you are going to from 6 nodes(RF=3) to 4 nodes(RF=2), do you need to need to load data from all 6 nodes even if the replication factor in 6 node cluster is 3? If we do need to load data from all 6 nodes into 4 node cluster, are there any risk of running out of space in the new cluster? Excellent question. I addressed Load and Stream specifics in https://www.scylladb.com/2023/09/18/5-more-intriguing-scylladb-capabilities-you-might-have-overlooked/ , so you may also want to check on that. if you are going to from 6 nodes(RF=3) to 4 nodes(RF=2), do you need to need to load data from all 6 nodes even if the replication factor in 6 node cluster is 3? Copying all 6 nodes indeed seem an overkill. But the real answer is that it depends. Are you dual-writing to both clusters? Do you expect all data present in the source cluster to match its target? Also, Do you use NetworkTopologyStrategy and spread the data to 3 AZs? If yes, then you can start dual-writing, run a repair job and once that repair job finishes snapshot your data from a single AZ and copy it over. Both cluster should be in sync afterwards. If we do need to load data from all 6 nodes into 4 node cluster, are there any risk of running out of space in the new cluster? All SSTable data is going to get streamed to its replicas, so you may want to let compaction pick up as you go through each Load and Stream step. @felipemendes good input. Should we add it to the docs? Do I need to disable compaction during load and stream? And then enable after each load and stream is completed. Not really. You may want to disable tombstones from getting compacted though, in case their gc_grace_seconds happen to expire. You can do so by setting tombstone_gc to repair. See Preventing Data Resurrection with Repair Based Tombstone Garbage Collection - ScyllaDB @tzach , that’s definitely a good idea. Ping me if anything --- ### Page: https://forum.scylladb.com/t/how-to-automatically-create-data-directories/994 Title: How to automatically create data directories? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: We use scylla operator to manage our scylla in Kubernetes. Every time we start the new node in the cluster we get the error during the execution of scylla_io_setup: /var/lib/scylla/commitlog was not found. Please check … Language: en Canonical URL: https://forum.scylladb.com/t/how-to-automatically-create-data-directories/994 ## Headings Structure: H1: How to automatically create data directories? H3: Related topics ## Main Content: H1: How to automatically create data directories? H3: Related topics We use scylla operator to manage our scylla in Kubernetes. Every time we start the new node in the cluster we get the error during the execution of scylla_io_setup: So we have to manually create this directories and rerun the script. Is it possible to automate it and create data directories automatically? Or may be there any other ways to do it? ScyllaDB is responsible for initializing working directory. What storage provisioner do you use? Please attach must-gather dump. I mean we run scylla on AWS EKS with scylla-operator. Every time we create a new node - scylla container starts and tries to execute scylla_io_setup before actual scylladb, which will create data directories. And of course it fails as no data directories are created. How can we tell scylla-operator to skip scylla-io-setup stage at all? --- ### Page: https://forum.scylladb.com/t/last-2-weeks-in-scylla-cluster-tests-git-master-issue-29-2023-11-10/995 Title: Last 2 weeks in scylla-cluster-tests.git master (issue #29; 2023-11-10) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 784d4ed2…8409417e range are covered. There were 30 non-merge commits from 8 authors in tha… Language: en Canonical URL: https://forum.scylladb.com/t/last-2-weeks-in-scylla-cluster-tests-git-master-issue-29-2023-11-10/995 ## Headings Structure: H1: Last 2 weeks in scylla-cluster-tests.git master (issue #29; 2023-11-10) H3: Related topics ## Main Content: H1: Last 2 weeks in scylla-cluster-tests.git master (issue #29; 2023-11-10) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 784d4ed2…8409417e range are covered. There were 30 non-merge commits from 8 authors in that period. Some notable commits: The ‘v1.11.0*’ scylla-operator allows configuring manual multiDC with externalSeeds config option. New quasi-multidc test was added by making a DB cluster out of the 2 separate ‘ScyllaCluster’ K8S resources created by the same Scylla-operator. Many different improvements to SCT framework was implemented to support multidc K8s clusters e.g add support of the multiple kubeconfig files (be able to talk to multiple K8S clusters), add ‘k8s_cluster’ attribute to all pod python classes or add region_name (k8s clusters) to loggers, so we can distinguish between different K8S clusters in the logs for easier debugging. Because the GCE network is now closed for the public by default, we added automatic ssh tunnels creation command to hydra to ease connection to db nodes for debugging purposes. with this change we can now do the following commands, e.g.: Also there’s a new gce-allow-public command which will open a GCE instance to the public communication, similar command as attach-test-sg for AWS See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-204-2023-11-12/996 Title: Last week in scylladb.git master (issue #204; 2023-11-12) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 6cc5bcae80…8d618bbfc6 range are covered. There were 95 non-merge commits from 17 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-204-2023-11-12/996 ## Headings Structure: H1: Last week in scylladb.git master (issue #204; 2023-11-12) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #204; 2023-11-12) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 6cc5bcae80…8d618bbfc6 range are covered. There were 95 non-merge commits from 17 authors in that period. Some notable commits: An sstable can be in different states: normal, staging (building views/indexes), quarantined. This is normally indicated by the directory the sstable is in, but for S3 backed sstables, it is now indicated by an entry in an internal table. S3 backed tables no longer create local directories. With raft topology coordination, a topology coordinator might lose the raft leadership. Errors due to this loss of leadership are now better handled. Alternator, ScyllaDB’s implementation of the DynamoDB API, now supports the ReturnValuesOnConditionCheckFailure feature. This makes handling contention more efficient. The native nodetool command now supports more commands: snapshot, drain, flush, disableautocompaction, enableautocompaction. It is not yet enabled by default. ScyllaDB uses two separate memory reservation systems for memtables: user, used for user writes, and system, used for ScyllaDB’s own writes. This ensures that a heavy write workload does not impact ScyllaDB’s internal housekeeping. The raft table was moved to the system memory reservation. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/load-and-stream-issue-with-memory-allocation/999 Title: Load and Stream issue with Memory Allocation - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hey guys, Running into issues while trying to load one of the tables. I’ve retried it multiple times by dropping the table and still getting same error. Not sure what is causing the issue. Other table works fine but thi… Language: en Canonical URL: https://forum.scylladb.com/t/load-and-stream-issue-with-memory-allocation/999 ## Headings Structure: H1: Load and Stream issue with Memory Allocation H3: Related topics ## Main Content: H1: Load and Stream issue with Memory Allocation H3: Related topics Running into issues while trying to load one of the tables. I’ve retried it multiple times by dropping the table and still getting same error. Not sure what is causing the issue. Other table works fine but this one table is giving us an issue. This is the error I’m getting Interesting. 5.2.9 does include a fix for a similar problem sstableloader/nodetool refresh: bad_alloc (seastar - Failed to allocate 536870912 bytes) · Issue #13491 · scylladb/scylladb · GitHub . Can you open a GitHub issue instead? And then please provide full logs. --- ### Page: https://forum.scylladb.com/t/network-requirement-about-cross-dc-link-for-scylladb-multi-dc-cluster/1000 Title: Network requirement about cross-dc link for scylladb multi-dc cluster? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi ScyllaDB team & everyone, We have scylladb multi-dc clusters and our network provider SLA of dropping packets in cross-dc link is <= 0.1% per month. This means in the worst case, we could have dropping packets rate c… Language: en Canonical URL: https://forum.scylladb.com/t/network-requirement-about-cross-dc-link-for-scylladb-multi-dc-cluster/1000 ## Headings Structure: H1: Network requirement about cross-dc link for scylladb multi-dc cluster? H3: Related topics ## Main Content: H1: Network requirement about cross-dc link for scylladb multi-dc cluster? H3: Related topics Hi ScyllaDB team & everyone, We have scylladb multi-dc clusters and our network provider SLA of dropping packets in cross-dc link is <= 0.1% per month. This means in the worst case, we could have dropping packets rate close to 3% one day in a month. Then if our cluster (at least RF=3 in each DC, CL=LOCAL_QUORUM) have 500000 writes/s, we doubt this worst case scenario could be a problem in scylladb, isn’t it? And would like to know if scylladb team have any basics or advanced cross-dc link network requirements for like dropping packets rate … etc? We are asking this question is because we face the similiar problem recently, dropping packets rate in cross-dc link reach close to 0.2% for an hour, then we saw many hints are written in the both DCs, and scylladb cluster spent almost a day to re-play all these hints generated in that hour. And currently our cluster only have 5000 writes/s, so we are worry about when write queris increase in the cluster, this problem will become more seriously. There is setting --max-hint-window-in-ms in scylladb, but not sure if this could help in this case. We also think about triggering repair job in whole cluster after incident, but since repair will produce more packets between DCs and repair multi-dc cluster need much more time, so in this repairing period we could face this dropping packets rate hike again and cause even more hints to be written (we faced this problem recently also). So what’s the recommendation by ScyllaDB team in this situation? Given your situation, it seems like it is doable to assume that not only you are connected through an unreliable network link, but also to a slow one, given that hints took long to get replayed, and repair can likewise take long. Ideally, hints replay shouldn’t take more than some minutes, so it is definitely strange it took almost a day. Occasional network failures aren’t a problem per se, given that you can - and SHOULD - always repair. You have the flexibility to break down a repair task per keyspace/table/replicas/token-range/etc, so even if repair fails (and notice ScyllaDB manager has retry mechanisms), you should still run it to completion in a regular basis. So what’s the recommendation by ScyllaDB team in this situation? Ideally, run your database behind a network you can trust and is fast enough for your needs, and handle occasional hiccups which will happen anyway. That said, there are other things you can do: Thanks for your response. The cross-dc link we are using which have SLA of average dropping packets rate in a month should be <= 0.1% but there is no guarantee in short period of time which mean it could higher than 0.1% like in my previous expression about worst case scenario as long as average dropping packets rate in a month still below 0.1%. And cross-dc bandwidth is not a problem currently, since our multi-dc scylladb cluster use 1Gbps at maximum in cross-dc link and available bandwidth in the link is at least 10Gbps. Because our data centers are in the east & west of US, so the latency is about 100ms in the link. Here are the hints graph when issues happened: As you can see there is large hints written in about a hour and then scylladb re-play (Hints sent) it in much slower pace. We alreay check other graphs and make sure there is no cql writes error within each DC, and cluster load is relatively low (<=25%) in the period, so the hints written here are most because of cross-dc sync. Originally, we think because this is hints for cross-dc (have larger latency) and scylladb re-play it in the background at lower priority so that’s why it re-play in such longer period, is this understanding not correct? Ideally, hints replay shouldn’t take more than some minutes, so it is definitely strange it took almost a day. This is true in our experience if drop packets rate hike last only few seconds or minutes, but is it still true when cross-dc link drop packets rate hike last more than a hour in multi-dc environment? Ideally, run your database behind a network you can trust and is fast enough for your needs, and handle occasional hiccups which will happen anyway. Occasional hiccups is not the problem we are concern about but we would like to know if scylladb have more specific network requirements such dropping packet rate or something … etc when people want to run multi-dc scylladb cluster in write query rate like 500000 or 1000000 writes/s? Sorry, “unreliable” or “fast enough” is a little too vague to us. Originally, we think because this is hints for cross-dc (have larger latency) and scylladb re-play it in the background at lower priority so that’s why it re-play in such longer period, is this understanding not correct? Well, as you’ve explained before the local DC nodes aren’t even overloaded, so there is little reason to lessen the hints priority. It shouldn’t take a day. This is true in our experience if drop packets rate hike last only few seconds or minutes, but is it still true when cross-dc link drop packets rate hike last more than a hour in multi-dc environment? Hints replay to a remote DC will definitely take longer given the RTT latency. Yet, a day definitely seems exaggerated. How long does a repair typically take, how large are your tables? Did you check the latency of a cross-DC query (this may be worth a try as you should expect an average being your RTT time, plus some overhead, but nothing very far from it). You may also shutdown a remote node for a while and check whether the situation reproduces. we would like to know if scylladb have more specific network requirements such dropping packet rate or something … etc when people want to run multi-dc scylladb cluster in write query rate like 500000 or 1000000 writes/s? Sorry, “unreliable” or “fast enough” is a little too vague to us. No, we don’t, and we recommend a network of 10Gbps or more Our architecture accepts the fact that failures and partitions can occur, and the cluster should be able to recover under that situation. I’d still check on the network side of things, as you clearly stated repairs also take a significant amount of time. Should you still feel stuck, please follow-up with an issue, provide thorough details on your setup, ScyllaDB version, and Prometheus data covering the timeframe for the reported incident. How long does a repair typically take, how large are your tables? Whole cluster repair (total 7 nodes) need ~6.5 hours, and data size per replica is 1.4 TB. Did you check the latency of a cross-DC query (this may be worth a try as you should expect an average being your RTT time, plus some overhead, but nothing very far from it). We don’t do cross-DC query directly, but we do monitor data sync cross-dc which meet our expectation which is like you said “RTT time, plus some overhead”. --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-2-10/1002 Title: [RELEASE] ScyllaDB 5.2.10 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.2.10, a bugfix release of the ScyllaDB 5.2 stable branch. ScyllaDB Open Source 5.2.10, like all past and future 5.x.y releases, is backward compatible and supports rolli… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-2-10/1002 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.2.10 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.2.10 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.2.10, a bugfix release of the ScyllaDB 5.2 stable branch. ScyllaDB Open Source 5.2.10, like all past and future 5.x.y releases, is backward compatible and supports rolling upgrades. Users are encouraged to upgrade to 5.2.10. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/nesuss-143421-vulnerability/1004 Title: Nesuss 143421 Vulnerability - Database Community - ScyllaDB Community NoSQL Forum Meta Description: Hello Scylla Community, I have a finding on my Scylla servers which is actually related to Apache Cassandra. It seems the scanner is seeing some code from Apache Cassandra that Scylla uses. Per Tenable, the finding is v… Language: en Canonical URL: https://forum.scylladb.com/t/nesuss-143421-vulnerability/1004 ## Headings Structure: H1: Nesuss 143421 Vulnerability H3: Related topics ## Main Content: H1: Nesuss 143421 Vulnerability H3: Related topics Hello Scylla Community, I have a finding on my Scylla servers which is actually related to Apache Cassandra. It seems the scanner is seeing some code from Apache Cassandra that Scylla uses. Per Tenable, the finding is valid as Scylla is using the same vulnerable code as Apache Cassandra. We are using the free version, so we cannot log a ticket for help. As of a few months ago, we were on the latest version of ScyllaDB. Can anyone confirm if they have this false finding as well and if so, what steps did you take to resolve? Does the newest version fix this, anyone know? In which Scylla version did you see this vulnerability ? The version we are running is 4.4.1. This release has been EOL for a long time now, we had many changes and updates since then, including vulnerability fixes. I recommend checking our latest release, which is 5.4.0 - https://www.scylladb.com/download/#open-source --- ### Page: https://forum.scylladb.com/t/data-modeling-question-save-space-using-lookup-tables/1005 Title: Data modeling question: save space using lookup tables? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I am in the process of finalizing my proof-of-concept of switching from PostgreSQL to ScyllaDB for my commercial project that I’ve been working on for the past 3 years. I have a working code for the data migration and m… Language: en Canonical URL: https://forum.scylladb.com/t/data-modeling-question-save-space-using-lookup-tables/1005 ## Headings Structure: H1: Data modeling question: save space using lookup tables? H3: Related topics ## Main Content: H1: Data modeling question: save space using lookup tables? H3: Related topics I am in the process of finalizing my proof-of-concept of switching from PostgreSQL to ScyllaDB for my commercial project that I’ve been working on for the past 3 years. I have a working code for the data migration and most of the main operations that the project currently does for the existing PostgreSQL database. I am a bit stuck on one dilemma. My database has been growing steadily by approximately 21-28 GB every day in the past year, which is roughly 50-70 million rows per day. Each row contains 12 numeric and timestamp fields, but also 4 text fields, each value varying between 1 and 6 UTF-8 characters long. Those text fields contain values from 2 different lists of strings, each of which contains fewer than 300 unique string values. 1 of those 4 fields is going to be part of the composite partition key (the other part of the partition key is a date) in my ScyllaDB table. In other words, a substantial percentage of my database contains repeated values in the NoSQL/denormalized schema. Obviously, in a relational database those fields would contain foreign keys to one of the two lookup tables. Although space is not an issue yet, but I anticipate it to become an issue in less than half a year. Would you recommend that I create and use (in the client side code) the lookup tables, or would I be better off letting ScyllaDB split the table between multiple servers based on partitions? I estimate that the cost of renting some extra servers (if the schema is not optimized for space savings) is roughly equivalent (or at least is comparable to) the cost of dealing with the increased code complexity if the schema is optimized for space. So, each row is something close to 0.3 - 0.5 Kb in size. How significant would database space savings be if I replace roughly 12 UTF8 characters with 4 short-integer foreign key values, but add a database overhead of 2 extra tables and all related compaction and other maintenance etc.? I do not have nearly enough experience with ScyllaDB to figure this out, and I would greatly appreciate a nudge in either of these 2 directions. Thank you very much in advance. Can you please present here the current schema plan? It’s hard to extract the tables from the text above. Hi, I replaced the real column names with more readable ones and sorted the list by type, which also added to readability: create table if not exists mykeyspace.mytablename ( key_1_string text, key_2_string text, id int, long1 bigint, long2 bigint, long3 bigint, long4 bigint, text1 text, text2 text, text3 text, text4 text, text5 text, text6 text, text7 text, text8 text, text9 text, double1 double, double2 double, double3 double, double4 double, int1 int, int2 int, int3 int, int4 int, primary key ((key_1_string, key_2_string), id) ) with comment = ‘One huge table with per-key1-per-key2 partitioning.’ and caching = {‘keys’: ‘ALL’, ‘rows_per_partition’: ‘ALL’} and compaction = {‘class’: ‘SizeTieredCompactionStrategy’} and compression = {‘sstable_compression’: ‘org.apache.cassandra.io.compress.LZ4Compressor’} and dclocal_read_repair_chance = 0 and speculative_retry = ‘99.0PERCENTILE’; In this modified schema, the values of key_2_string (which is part of the composite partitioning key) and text1, text2, and text3 look like ‘A’ or ‘BB’ or ‘ABC’ or ‘DDDD’ and sometimes 5 characters like ‘ABCDE’. All characters are Latin and case-insensitive (so I keep them all as capital letters). Those 4 values in each row can be replaced with a foreign key (with or without the real FK relationship, depending on whether I am using PostgreSQL or ScyllaDB), and those keys’ values will always fit in an int16 type. The remaining 6 TEXT columns always have 1-character values, so there is no need to optimize them. The questions I am asking for help with, are: How significant would database space savings be if I replace roughly 12 UTF8 characters with 4 short-integer foreign key values, but add a database overhead of 2 extra tables and all related compaction and other maintenance etc.? Is there a chance that a query that retrieves a whole partition (with paging, of course: some partitions have 2.5-3 million rows each) will speed up considerably because there will be less data (fewer bytes) to read and transfer? EDITED: I added a query performance-related question because currently my C+±based proof-of-concept code using ScyllaDB C++ driver working with a pretty decent local (VirtualBox) virtual machine (6 Gb DDR4 RAM, 4 CPU cores, the VM is dedicated to ScyllaDB and has nothing else except the OS) performs only about 5%-10% better than the same query running on PostgreSQL in a very “apples-to-apples” scenario. I am trying to optimize for speed first, and hopefully for storage (but that is not necessary, although desirable). Thank you! I think a lookup table makes sense if your data somehow follows a gaussian distribution, and the data is immutable. One aspect to keep in mind (and perhaps play with before trying any optimizations) is that data will get compressed only within a chunk. By default, chunk_length_in_kb is 4kB. See: Data Definition | ScyllaDB Docs On disk SSTables are compressed by block (to allow random reads). This defines the size (in KB) of the block. Bigger values may improve the compression rate, but increases the minimum size of data to be read from disk for a read. Therefore, the higher the chunk size, the better compression rates you may get - however reads will require more IOPS to retrieve a block. This may be fine (or not), specially if you need to frequently scan through several rows as in you mentioned in your question (2). https://www.scylladb.com/2017/08/01/compression-chunk-sizes-scylla/ also has a very nice table near the end, plus some previous benchmarks showing some results across different chunk sizes. Depending on your latency requirements, you may even play further with different compression algorithms. Notably here, the main tradeoff will be on CPU utilization. Compression in ScyllaDB, Part One - ScyllaDB discusses the basics of it, and part 2 at the bottom will discuss results and the impact of different compression algorithms using different chunk sizes. You may want to experiment with these to further find a good balance. One last thing that isn’t clear is whether you frequently update text1-text9 (or text1-text3, or text4-text9). If the answer is no, then you may also want to try an UDT? See If You Care About Performance Use UDT's As you may guess, considerably is hard to say. Though yes, it should enhance the query times. ScyllaDB C++ driver working with a pretty decent local (VirtualBox) virtual machine (6 Gb DDR4 RAM, 4 CPU cores, the VM is dedicated to ScyllaDB and has nothing else except the OS) performs only about 5%-10% better than the same query running on PostgreSQL in a very “apples-to-apples” scenario. Nitpick: There are indeed very little room for gains on commodity hardware. Reason it is relatively simple to maximize utilization with the aforementioned specs. You should expect a much higher % on our recommended instances. (~8G/vCPU and locally attached SSDs). So consider making it apple-to-apples in a more realistic scenario ;-)) I don’t even know where to begin, to thank you for your reply… Every paragraph… heck, every sentence is eye-opening and more than helpful. I have books on my shelf which have less of useful info than your reply. The most humbling thing for me is: your reply contains answers to questions that I did not even know I had to ask. Thanks to multiple pieces of advice that you gave me, it looks like I have my work cut out for me and just need to do several things now. The only thing I want to add is that my work on the local virtual machines includes tests only, and the production server is hardware+OS/dedicated with 64 Gb DDR4 RAM, and now I understand (thanks to your comment) why the speed improvement from PostgreSQL was less than I had expected. THANK YOU VERY MUCH!!! --- ### Page: https://forum.scylladb.com/t/how-in-the-rust-bindings-do-i-know-an-lwt-failed/1007 Title: How in the Rust bindings do I know an LWT failed? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Getting my feet wet with a small app with scylladb. I have a unique key I don’t want duplicated on insert so I have an LWT like below. I’m curious how I can know the LWT failed so I can return back a true/false so I can … Language: en Canonical URL: https://forum.scylladb.com/t/how-in-the-rust-bindings-do-i-know-an-lwt-failed/1007 ## Headings Structure: H1: How in the Rust bindings do I know an LWT failed? H3: Related topics ## Main Content: H1: How in the Rust bindings do I know an LWT failed? H3: Related topics Getting my feet wet with a small app with scylladb. I have a unique key I don’t want duplicated on insert so I have an LWT like below. I’m curious how I can know the LWT failed so I can return back a true/false so I can alert the client appropriately that the username is taken in case a race condition bypasses the basic client check. So the first row of any insert with ‘IF NOT EXISTS’ is always a boolean with the name ‘applied’. However, what makes this complicated is if the row DOES exist it will also return the existing row. This means the output format changes based on if it succeeds or not … The code above was the best way I could find to handle that. You must be playing with Cassandra, in Scylla the statement metadata is stable: How does Scylla LWT Differ from Apache Cassandra ? | ScyllaDB Docs --- ### Page: https://forum.scylladb.com/t/nodetool-gossipinfo/1010 Title: Nodetool gossipinfo - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, nodetool gossipinfo is returning some info about all tables as part of “X2”. What does the float value in X2 mean? cheers, Christian Language: en Canonical URL: https://forum.scylladb.com/t/nodetool-gossipinfo/1010 ## Headings Structure: H1: Nodetool gossipinfo H3: Related topics ## Main Content: H1: Nodetool gossipinfo H3: Related topics nodetool gossipinfo is returning some info about all tables as part of “X2”. What does the float value in X2 mean? It indicates the cache hit rate for the keyspace.table in question. We propagate this data via gossip for HWLB. Nonetheless, this is a good question and I agree X2 through X10 mean nothing. Will follow up with a doc issue. yes, its a bit hard to understand I assume that latency information is used for things like LatencyAwarePolicy, etc? I assume that latency information is used for things like LatencyAwarePolicy, etc? Nope, as it is exchanged through gossip (not to clients). The purpose of the cache hitrate is to enable Heat Weighted Load Balancing, so the coordinator node determines which replicas to query. I see, I thought this get also exposed to the clients. Thanks! --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-30-2023-11-18/1017 Title: Last week in scylla-cluster-tests.git master (issue #30; 2023-11-18) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 225293b9…80160f38 range are covered. During this period, we had 47 non-merge commits from … Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-30-2023-11-18/1017 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #30; 2023-11-18) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #30; 2023-11-18) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 225293b9…80160f38 range are covered. During this period, we had 47 non-merge commits from 10 authors. Here are some of the noteworthy commits: As we gather kallsyms, which often provide essential information for existing issues, we have developed a tool to decode kernel callstacks. Although it’s not yet ready to be integrated with an automatic decoder, the content is prepared for use as a standalone script for now. Due to InsufficientInstanceCapacity errors, certain tests will now be conducted in a different availability zone. For now, two tests have been relocated: large partition 200k pk and alternator short longevity. When adding new tests, please consider using an availability zone other than ‘a’ to minimize the likelihood of failure due to insufficient capacity. SCT now supports multiple K8S clusters. Scylla nodes can use dc_idx to map to the appropriate k8s_cluster . This has enabled multi K8S cluster creation in EKS. While there’s more work needed for full multi-dc longevity tests on K8s, we’re significantly closer to achieving this. Previously, installing Scylla from a repository would always select the latest version. With this change, you can select a specific version by adding the version after the colon in the scylla_repo URL. This will aid us in testing specific versions of Scylla in customer cases and reproducers. When adding new nodes, we can now specify the instance_type. This means that whoever is calling the .add_nodes() can define a different instance type than what is configured in the test YAML file. During node setup, if the package installation failed, it would cause the test to fail. We have now added retries for this process, enhancing stability. Furthermore, in most cases, package installation has been moved to the install_package node method to reduce code duplication and improve stability. The large collection test now verifies the appropriate log messages and system.large_cells content. Rolling upgrades have begun to use docker-based loaders, so we use a newer version of c-s and java-drivers. We have removed support for deprecated distros from the code (Ubuntu 16, Ubuntu 18, Debian 8, Debian 9). See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/unexpected-reshards-when-changing-root-volumes/1018 Title: Unexpected Reshards When Changing Root Volumes - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: We are working on a migration to Ubuntu from CentOS, and ran into an issue with nodes trying to reshard, because the CPU Set settings have changed since they were created. Background: We’ve run into an issue when tryin… Language: en Canonical URL: https://forum.scylladb.com/t/unexpected-reshards-when-changing-root-volumes/1018 ## Headings Structure: H1: Unexpected Reshards When Changing Root Volumes H3: Updating the Mode in perftune.yaml After a ScyllaDB Upgrade | ScyllaDB Docs H3: Related topics ## Main Content: H1: Unexpected Reshards When Changing Root Volumes H3: Updating the Mode in perftune.yaml After a ScyllaDB Upgrade | ScyllaDB Docs H3: Related topics We are working on a migration to Ubuntu from CentOS, and ran into an issue with nodes trying to reshard, because the CPU Set settings have changed since they were created. We’ve run into an issue when trying to upgrade nodes from CentOS to Ubuntu. The gist of what we are doing is using AWS EC2’s “Root Volume Replacement” to swap the CentOS root volume out for a new Ubuntu one, and handle any configuration before restarting the node. This has worked in all of our test clusters and all but 1 of our production clusters. That production cluster is using i3en.6xlarge instance types, with only 1 other cluster using i3en.xlarge (the rest are on i3*). The original AMI for most of the nodes on this cluster are an old 4.4.8 Scylla AMI, that has just be upgraded over the years. It was the last CentOS AMI available, so we were stuck there for a while. 2 nodes on the cluster have been rebuilt recently, and are running on a custom CentOS AMI that is based on 5.1+ Scylla AMIs. We have 1 node that has had the root volume replaced, using the above method. The node that had the root volume replacement is undergoing a reshard of all of the data. This was unexpected, because no other cluster has done this. Digging into things, we found that the offending node went from 24 cores to 22 cores in Prometheus metrics. CPU-2 seems to match other clusters, so the concern was more with why this cluster was using 24 cores to begin with. The following configurations were found on all of the 4.4.8 AMI based nodes: /etc/scylla.d/cpuset.conf /etc/scylla.d/perftune.yaml While the following was found on the newer nodes (2 new plus the root volume replaced node): /etc/scylla.d/cpuset.conf /etc/scylla.d/perftune.yaml I’ve dug around and can’t find a why, but I assume that something changed with how the perftune is run on the i3en instances between 4.4 and 5.1. My question boils down to, what is the recommended course of action here? This cluster is 39 nodes, with 400TB+ data on it, so resharding most of the nodes will probably take weeks, with a node being down the whole time. Is there a clean way we can keep the old CPU set for the old nodes for now? Will that have any negative impact? Our main concern is getting off of CentOS, so if we can do that quickly and worry about the reshards over a long time period that would be preferable, assuming that doesn’t cause any major issues. Another interesting side effect of this change was that the new node was under some super heavy load complaining about schema updates. We saw messages like the following to every other node in the cluster, and on every node in the cluster trying to hit the changed node: I assumed this was related to a schema mismatch, just without a failure. Running nodetool describecluster showed the same UUID version across the cluster, with no changes. The node was still working, but had high latency and overall performing way worse than the other nodes. We decided to go with a rolling restart, as I know that can fix schema mismatches, and that resolved the issue. This is mostly tangential to the main problem here, I am more just posting it since it is an outcome that I didn’t see explicitly stated anywhere. It seems that the change in CPU count and/or following reshard messed with the schema consolidation. As in the upgrade manual for 5.1, point to: ScyllaDB is an Apache Cassandra-compatible NoSQL data store that can handle 1 million transactions per second on a single server. You can save the previous cpu set file, and it should still work This change was done to reserve some of the cpus for network interface IRQ, and using same CPUs might affect performance of scylla Also the recommend way of upgrading is one release at a time, jumping between 4.4 → 5.1 might introduced compatibility issues, and isn’t tested procedure. Any change to number of CPU used means resharding, I don’t think there’s a way around it. Thanks for the documentation, that is helpful. Can you clarify some things though? As using different modes across one cluster is not recommended What’s the downside of this? Is this just a performance hit? As we are currently running with 3 nodes in sq_split (assuming that is the default) and the rest in mq mode. Also, the steps in the case of a changed cpuset (which I expect for our nodes) doesn’t do anything with the perftune.yaml backup, will Scylla recreate that with the restart? Just making sure I know all the ins and outs of this method, as this is only impacting our production cluster and we don’t have a good way to test this process in another cluster. To clarify on the version change, the nodes were upgraded one release at a time, it was the underlying AMI that jumped from 4.4 to 5.2. So, Scylla node setup and perftune were not rerun until we were recreating the root volumes of the nodes for this migration. Having a mixed setup, in which every node uses different cpuset is not just about performance But it’s a less common setup, that is less tested, there were a few issues with that mixed setup In previous releases. If there an actual reason for doing such a mix, one could, I wouldn’t recommend doing so just cause of some upgrading issue. As for inplace upgrade of the AMI root disk, it’s also something that isn’t tested We recommend in place upgrade of scylla and the system, Or replacing node with fresh nodes with newer AMI Ok thanks for the info Israel. I was wondering for the urgency on the mixed setup, and how long we would be safe to keep it mixed. It sounds like it’s just scary in the past and is just an unknown area for configuration. We will do the resharding, but we have apps that read from the cluster that are very read latency sensitive. They will fall behind if latency gets worse than 3 or 4 ms. Having a node down does that to us easily, so our plan is to do a few reshards each weekend and get towards the desired state over time. As for the upgrade process, the other options just seem worse in our opinion. In place OS changing sounds like torture, and replacing each node with a new one and decommissioning the old ones means shuffling the 100s of TB of data around the network and incurring massive costs. We understand it is untested with Scylla though. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-205-2023-11-19/1019 Title: Last week in scylladb.git master (issue #205; 2023-11-19) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 8d618bbfc6…eb674128ca range are covered. There were 77 non-merge commits from 17 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-205-2023-11-19/1019 ## Headings Structure: H1: Last week in scylladb.git master (issue #205; 2023-11-19) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #205; 2023-11-19) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 8d618bbfc6…eb674128ca range are covered. There were 77 non-merge commits from 17 authors in that period. Some notable commits: Internode remote procedure call metrics are now reported per (datacenter, verb) combination. This makes the latency metrics more meaningful. The start stop native transport API (used by nodetool enablebinary) mistakenly launched the listener in the streaming scheduling group, causing subsequent queries to run with reduced priority compared to maintenance operations such as repair and bootstrap. The listener is now correctly launched in the statement scheduling group. An edge case where a joining node is rejected from the cluster, but the rejection message times out, has been fixed. We now reject adding fields of type duration to a user-defined type ((UDT), if that UDT is used as a clustering key component. When writing an sstable, we estimate how many partitions it will have in order to size the bloom filter correctly. A few bugs were corrected for this estimation with the Time Window Compaction Strategy. The S3 object storage driver now tags partial object uploads, so that a lifecycle policy can delete them if they are abandoned due to a crash. A number of rare bugs in the row cache were fixed. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/does-scylla-support-scaling-multiple-nodes-at-the-same-time/1024 Title: Does scylla support scaling multiple nodes at the same time? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: The test found that when expanding two nodes at the same time into the cluster, one node will report an error. Other bootstrapping/leaving/moving nodes detected, cannot bootstrap while consistent_rangemovement is true (c… Language: en Canonical URL: https://forum.scylladb.com/t/does-scylla-support-scaling-multiple-nodes-at-the-same-time/1024 ## Headings Structure: H1: Does scylla support scaling multiple nodes at the same time? H3: Related topics ## Main Content: H1: Does scylla support scaling multiple nodes at the same time? H3: Related topics The test found that when expanding two nodes at the same time into the cluster, one node will report an error. Other bootstrapping/leaving/moving nodes detected, cannot bootstrap while consistent_rangemovement is true (check_for_endpoint_collision). What will happen if we change consistent_rangemovement to false? Support two nodes bootstrap at the same time, what will be the bad effects? Is this a production-ready solution? What will happen if we change consistent_rangemovement to false? Support two nodes bootstrap at the same time, what will be the bad effects? Is this a production-ready solution? You shouldn’t turn off this option, as it means you will disable consistent range movements (or enable inconsistent range movements ;-)) Consider you have nodes N1, N2, N3. You bootstrap N4 & N5, consider N4 started first and you started N5 before N4 finished. N4 will stream data from N1~N3 for the ranges it now owns. N5 bootstraps, places its vNodes into the ring similarly as N4 did. It will stream from N1~N4, but there may be ranges N4 hasn’t streamed yet. Worse, what if N4 fails? The general solution for these are Raft and Tablets. You may read ScyllaDB’s Path to Strong Consistency: A New Milestone - ScyllaDB - which includes a nice Q&A in the end which addresses your question. And for tablets: Why ScyllaDB is Moving to a New Replication Algorithm: Tablets - ScyllaDB Also related: https://github.com/apache/cassandra/blob/cassandra-3.9/doc/source/operating/topo_changes.rst#range-streaming --- ### Page: https://forum.scylladb.com/t/release-scylla-manager-3-2-4/1025 Title: [RELEASE] Scylla Manager 3.2.4 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces the release of ScyllaDB Manager 3.2.4 production-ready ScyllaDB Manager patch release of the stable ScyllaDB Manager 3.2 branch. ScyllaDB Manager is a centralized cluster administration and r… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-manager-3-2-4/1025 ## Headings Structure: H1: [RELEASE] Scylla Manager 3.2.4 H3: Backup H3: Repair H3: Other H3: Packaging H3: Related topics ## Main Content: H1: [RELEASE] Scylla Manager 3.2.4 H3: Backup H3: Repair H3: Other H3: Packaging H3: Related topics The ScyllaDB team announces the release of ScyllaDB Manager 3.2.4 production-ready ScyllaDB Manager patch release of the stable ScyllaDB Manager 3.2 branch. ScyllaDB Manager is a centralized cluster administration and recurrent tasks automation tool. This release fixes issues in Manager backup, in particular for Alternator, and repair. In addition to linux/amd64 images, we’ve begun publishing linux/arm64 docker images for scylla-manager and scylla-manager-agent (#3278). For a comprehensive list of commits included in the 3.2.4 release, please refer to this link. ScyllaDB Enterprise customers are encouraged to upgrade to ScyllaDB Manager 3.2.4 in coordination with the ScyllaDB support team. The new release includes upgrades of both ScyllaDB Manager Server and Agent. Scylla Manager 3.2.4 support the following Scylla Enterprise releases: And the following Open Source release (limited to 5 nodes see license): You can install and run Scylla Manager on Kubernetes using Scylla Operator. More here. --- ### Page: https://forum.scylladb.com/t/scylladb-entering-into-a-crashloop/1028 Title: Scylladb entering into a crashloop - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I tried to setup scylladb in a docker container using compose using this yml for docker compose version: '3' services: scylladb: image: scylladb/scylla ports: - "9042:9042" volumes: … Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-entering-into-a-crashloop/1028 ## Headings Structure: H1: Scylladb entering into a crashloop H3: Related topics ## Main Content: H1: Scylladb entering into a crashloop H3: Related topics I tried to setup scylladb in a docker container using compose using this yml for docker compose it keeps spitting errors out like it’s no tommorow this is running docker compose on windows with docker running on a virtual machine using WSL Is ./dev/scylla_data within WSL’s 9P filesystem protocol (ie: the /mnt/c mapping?), or is it within the virtual SDD file (default ext4.vhdx)? I don’t have a Windows machine available atm, but I would start by trying to start the container without the volume. It seems to be the cause of your problem. If 9P is what you are using, then it is simple: It isn’t POSIX compliant. If it isn’t, then it is likely something else I couldn’t figure just from these snippets. Fair enough, I tried on my work machine, and it seems to be okay with that volume on linux, grr why must windows be like this?? --- ### Page: https://forum.scylladb.com/t/release-scylladb-cloud-22-november-2023/1029 Title: [RELEASE] ScyllaDB Cloud - 22 November 2023 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: We are happy to announce a new update to the ScyllaDB Cloud service. New features and enhancements: Added support for Scylla Manager 3.2.4. We’ve enhanced the Choose your Cluster section inside the New Cluster wizard… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-cloud-22-november-2023/1029 ## Headings Structure: H1: [RELEASE] ScyllaDB Cloud - 22 November 2023 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Cloud - 22 November 2023 H3: Related topics We are happy to announce a new update to the ScyllaDB Cloud service. New features and enhancements: --- ### Page: https://forum.scylladb.com/t/counter-updates-timeouts-bad-alloc/1035 Title: Counter updates timeouts & bad_alloc - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, is there any known issue with counters and memory allocation in scylla? I just have seen a node causing lots of counter update errors: com.datastax.driver.core.exceptions.OperationTimedOutException: [/10.0.0.21:90… Language: en Canonical URL: https://forum.scylladb.com/t/counter-updates-timeouts-bad-alloc/1035 ## Headings Structure: H1: Counter updates timeouts & bad_alloc H3: Related topics ## Main Content: H1: Counter updates timeouts & bad_alloc H3: Related topics is there any known issue with counters and memory allocation in scylla? I just have seen a node causing lots of counter update errors: And some bad alloc errors: After a restart of that node everything was fine again. Is there perhaps some memory issue with counters in scylla? It looks like a bug. Please open an issue with all relevant info --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-206-2023-11-26/1038 Title: Last week in scylladb.git master (issue #206; 2023-11-26) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the eb674128ca…a472700309 range are covered. There were 94 non-merge commits from 18 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-206-2023-11-26/1038 ## Headings Structure: H1: Last week in scylladb.git master (issue #206; 2023-11-26) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #206; 2023-11-26) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the eb674128ca…a472700309 range are covered. There were 94 non-merge commits from 18 authors in that period. Some notable commits: During commitlog replay, we skip over corrupted sections. However if the corrupted section also has corrupt size, it can lead to a crash. This is now fixed. The bundled Prometheus node_exporter, used to expose operating system metrics, was updated to version 1.7.0 to address a security vulnerability. In some cases a materialized view table’s schema can be constructed when the base table’s schema is not yet known. We now avoid this illegal state. Previously, we added support for performing reconciliation (read repair) on partitions that are composed of a long sequence of tombstones without timeouts. This feature is now enabled only when the entire cluster is upgraded, to avoid warnings in the log. For S3 backed keyspaces, the endpoint configuration is now checked when the keyspace is created, to avoid surprises later. A bug in reconciliation (read repair) in conjunction with reverse queries and range tombstones, that could cause incorrect data to be returned from queries, has been fixed. The sstable validation tools, scylla sstable validate-checksum and scylla sstable validate, now returns output in json format. Node startup will now recalculate the schema digest fewer times during restart, resulting in faster startups. Tablets are a new, experimental way of distributing data across nodes and shards. There is now support for updating drivers about tablet topology, which can change quite frequently. When consistent cluster topology is enabled, we now reject a replace node operation if the node being replaced is not dead. Repair has gained a new mode for small tables that makes repair significantly faster. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/add-node-even-when-seed-is-unreachable/1040 Title: Add node even when seed is unreachable - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, Is there any way to add a new node to the cluster even when one of it’s configured seeds isn’t reachable but the rest are? e.g. typo… Language: en Canonical URL: https://forum.scylladb.com/t/add-node-even-when-seed-is-unreachable/1040 ## Headings Structure: H1: Add node even when seed is unreachable H3: Related topics ## Main Content: H1: Add node even when seed is unreachable H3: Related topics Hi, Is there any way to add a new node to the cluster even when one of it’s configured seeds isn’t reachable but the rest are? e.g. typo… In the above configuration, scylla-node3 has 4 invalid seeds: 3 from different subnets and unreachable, one from the local network (which is the GW), and node2 (the only valid entry). Note we haven’t specified an invalid FQDN here, as that would fail the DNS lookup. Thats great, when I originally tried it I failed on DNS lookup which wasn’t obvious from the logs. Thanks! I will quickly note for anyone visiting this topic in the future, that this is still a problem in the context of the first node of the cluster. see this issue. --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-2-11/1044 Title: [RELEASE] ScyllaDB 5.2.11 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.2.11, a bugfix release of the ScyllaDB 5.2 stable branch. ScyllaDB Open Source 5.2.11, like all past and future 5.x.y releases, is backward compatible and supports rolli… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-2-11/1044 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.2.11 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.2.11 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.2.11, a bugfix release of the ScyllaDB 5.2 stable branch. ScyllaDB Open Source 5.2.11, like all past and future 5.x.y releases, is backward compatible and supports rolling upgrades. Users are encouraged to upgrade to 5.2.11. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/release-scylladb-cloud-28-november-2023/1045 Title: [RELEASE] ScyllaDB Cloud - 28 November 2023 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: We are happy to announce a new update to the ScyllaDB Cloud service. New features and enhancements: Added HTTPS support for Scylla Alternator based clusters. The Certificate Authority public key to verify the server … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-cloud-28-november-2023/1045 ## Headings Structure: H1: [RELEASE] ScyllaDB Cloud - 28 November 2023 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Cloud - 28 November 2023 H3: Related topics We are happy to announce a new update to the ScyllaDB Cloud service. New features and enhancements: --- ### Page: https://forum.scylladb.com/t/automatic-down-nodes-removal/1046 Title: Automatic down nodes removal - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Does Scylla include an option to automatically remove inaccessible nodes after some time? Otherwise do you think periodically checking the cluster status and executing nodetool removenode a bad idea? I don’t want unhand… Language: en Canonical URL: https://forum.scylladb.com/t/automatic-down-nodes-removal/1046 ## Headings Structure: H1: Automatic down nodes removal H3: Related topics ## Main Content: H1: Automatic down nodes removal H3: Related topics Does Scylla include an option to automatically remove inaccessible nodes after some time? Otherwise do you think periodically checking the cluster status and executing nodetool removenode a bad idea? I don’t want unhandled issues that couldn’t be mitigated automatically and require manual intervention on some nodes to interfere with new nodes joining the cluster… It is recommended to call removenode But after 72h of a node not available it should be deleted from the gossip information on its own. It might interfere if you gonna try and reuse that ip address I would be careful doing automation based on the gossipinfo, to remove nodes I would also cross check with the cloud provider, to make sure you aren’t removing a node that was just out of contact for a few minutes. I see… Is there any way to slightly shorten the expire time? After briefly looking at the gossiper implementation, it seems like a hard coded value; But I would love to be proven wrong :') I’m not sure why it’s not configurable You can suggest it in an issue or even in a Pull request Ron Harel reacted to your message: --- ### Page: https://forum.scylladb.com/t/local-secondary-index-filtering/1050 Title: Local Secondary Index filtering - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I am learning Scylla, and trying out secondary indexes. I have found that local indexes are not working as they could. What I do: CREATE TABLE ks.t (pk int, ck int, v1 int, v2 int, PRIMARY KEY (pk, ck)); CREATE INDEX … Language: en Canonical URL: https://forum.scylladb.com/t/local-secondary-index-filtering/1050 ## Headings Structure: H1: Local Secondary Index filtering H3: Related topics ## Main Content: H1: Local Secondary Index filtering H3: Related topics I am learning Scylla, and trying out secondary indexes. I have found that local indexes are not working as they could. What I do: If I use ALLOW FILTERING, it does not use index. But it should. But if I query the index table (MATERIALIZED VIEW) then it works: As far as I understand, it is not available for global indexes, because then you would have a lot of requests to other nodes. But why does it not work for local indexes? Would it be also slow? Should I create a full MATERIALIZED VIEW for such a case? Would be very much usable with LIMIT. That’s an interesting question. In fact, we even discussed it under Local indexes could allow range queries on indexed elements · Issue #5547 · scylladb/scylladb · GitHub As you can see, it is not yet implemented. However, even though it is possible to implement such a feature, notice that the discussed problem statement still remains: Since the index does not hold contents for all columns, a range scan will require looking up the base table. Depending on the scan (such as large partitions involving several clustering keys, as mentioned in the linked issue), the results may not be what you’d expect. For example, in your provided output, notice how the index lookup with an equality clause returned all rows from the base table, whereas your view lookup misses the contents for column v2. You may still create a Materialized View including all (or only the relevant) columns, and then scan from it if that’s what you need. Or, depending on your query patterns and data distribution, maybe ALLOW FILTERING may turn out to be even better! As you are learning, follows a great talk to guide you down the path: https://www.youtube.com/watch?v=Ds9jbTeW0ks Thanks for your answer. But I am still concerned with ALLOW FILTERING not using an index. My guess is that it would work faster that way, because it’s already sorted. Currently it’s a scan of all the rows for a selected part of a PRIMARY KEY. I have checked that using a big table (1M rows), and having one indexed column filled only for 5 rows. And selecting by that column ... AND v > 0 ALLOW FILTERING in cqlsh, gives me a lot of empty responses: TRACING ON; gives a better understanding. Same quantity as if I would not filter by that extra column. Which means it scans all rows (a thousand rows for a selected part of a PRIMARY KEY). I did write it also on Github you have mentioned. But I am still concerned with ALLOW FILTERING not using an index. No one said you shouldn’t be, nor that ALLOW FILTERING is a magic solution to all problems. The linked resource provided earlier should provide you a guidance when ALLOW FILTERING makes sense, and when it doesn’t. I have checked that using a big table (1M rows), and having one indexed column filled only for 5 rows. I think you meant a partition with 1M rows with only few columns matching your restrictions. That means you’ve hit exactly one of the topics (and anti-patterns) covered in the reference material. In summary: Currently, if your data distribution involves partitions with many rows AND your restrictions turns to waste most of the scanned results, then you may still create a view to accomplish what you are after. And more: You don’t need to incur the full overhead of doubling your storage utilization, provided your view include only the columns you need. The only “drawback” is that such scans will need to go through the view, rather than the base table. Table with 1M rows , 1k per partition. You don’t need to incur the full overhead of doubling your storage utilization, provided your view include only the columns you need. But what if I need all of the columns… Quick example: Table news with fields id, title, published_at, views_count, text. And views_count is always increasing (not unique). How do I select 10 most viewed news? I would create an index, and select ordered with limit 10, as in Mysql, but Scylla does not allow to do it efficiently (without extra query and without disk space overhead (if I create a full materialized view)). Maybe I am missing something? (maybe not the best example for usage of Scylla, but anyway) In short: the problem is that we have an index table ks.t_v1_idx_index which gives me what I need, and I have to receive needed ids, and request the main table ks.t. But Scylla could do it efficiently without me sending it back (to the same node). You would say: “buy an extra storage and use MVs, it will be faster for Selects”. But selects can be cached, and also imagine that I can have a lot of updates to this table, which will be very slow if I would have a lot of different MVs containing the same fields. That’s why I would prefer a small MV with keys only (which is secondary index). I don’t think you are correctly stating the problem at hand. If you have 1K rows per partition, then the rest is irrelevant as you should restrict your partition when doing ALLOW FILTERING. 1K rows is just fine (that’s the default page size in most drivers, so you shouldn’t receive many empty pages as you stated earlier. cqlsh instead has a default page size of 100 - for obvious reasons) - although you don’t want to find yourself in a position where you hit few rows out of the grand total very frequently. Your provided example doesn’t state the problem either. It actually gives me more concerns out of the main subject: You are right that allowing an inequality clause on top of a LSI is doable, but I think you are missing the fact it is not a silver bullet solution, and it also carries its own trade-offs, and can fire back just as an ALLOW FILTERING clause would. A few: Regardless, I think that by now you should understand and have plenty of resources to further understand the implications of each choice at hand. Feature requests should be directed to GitHub, and we provided reference material for you to assess, understand how each solution at hand impacts performance, as well as advice on some of the options you may use. How you increase views_count? Are you doing a read before write? Why not use a counter instead? For example once an hour. Counters require separate table. But I anyway need to read it before write. Not a simple increment. And I have only one writer, so it must be fine. What makes you think you would be able to derive a top-k resultset out of this? Spoiler: ORDER BY on a secondary index is not supported. Is it not already sorted? I’m confused… I thought that’s the point of an index. But index table (MV) is sorted, and allows for X > 0, which means it’s not a hashmap. You mean that, the result will not be ordered as it’s in that MV. Hm… that’s interesting. Is this example anywhere close to what you actually want to achieve? Not actually. But it has that ordering column I want to use. Actually I want multiple ordering columns which are not unique and can be updated. If I have an MV with the same partition key as the main table, will it be stored on the same node? And also, can I have an MV to store only top 10 rows per partition? Like by ck in a descending order. And also tricky thing would be to combine multiple tables into one using something like MV. Is something like that possible? Because I am not sure if I must store a big set of data in a map field, keys of which I also want to use as index, so I would prefer a separate table for that, and have it combined with the main table for one of my queries. And also I would expect the secondary index to work with SELECT * FROM ks.t WHERE pk = ? AND v1 IN(?). For example v1 IN('', 'something') (where empty string, or ‘something’). --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-4-rc2/1052 Title: [RELEASE] ScyllaDB 5.4 RC2 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The Scylla team is pleased to announce ScyllaDB Open Source 5.4 RC2, the second Release Candidate for the Scylla Open Source 5.4 minor release. We encourage you to run ScyllaDB 5.4 release candidates on your test enviro… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-4-rc2/1052 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.4 RC2 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.4 RC2 H3: Related topics The Scylla team is pleased to announce ScyllaDB Open Source 5.4 RC2, the second Release Candidate for the Scylla Open Source 5.4 minor release. We encourage you to run ScyllaDB 5.4 release candidates on your test environments; this will help ensure that an upgrade to ScyllaDB 5.4 General Availability will proceed smoothly with your workload. Use the release candidate with caution; RC2 is not production-ready yet. You can help stabilize Scylla Open Source 5.4 by reporting bugs here. Only the last two minor releases of the ScyllaDB Open Source project are supported. Once Scylla Open Source 5.4 is officially released, ScyllaDB Open Source 5.4 and 5.2 will be supported, and ScyllaDB 5.1 will be retired. For a complete description of ScyllaDB 5.4 see ScyllaDB 5.4 RC1. Get ScyllaDB Open Source 5.4 (under “More Versions” for each distro) Updates and bug fixes since 5.4 RC1 (not including tests and docs updates) Alternator: Some sstables with large sizes left after TTL expiration, gc-grace-period and major compaction (tombstones are not deleted) #1191 Stability: failure detector apis need to call gossiper on shard 0 #15816 Stability: migration_manager: schema version correctness depends on order of feature enabling #16004 Stability: nodetool enablebinary starts the CQL server in the streaming group, instead of statement group #15485 Stability: Overloading scylla with materialized view writes can lead to deadlock #15844 Stability: raft topology: don’t register topology-on-raft RPCs in non-topology-on-raft mode. After this change, topology on raft RPCs are registered only if the experimental topology on raft mode is enabled #15862 Stability: raft: large delays between io_fiber iterations in schema change test. #15622 ScyllaDB uses two separate memory reservation systems for memtables: user, used for user writes, and system, used for ScyllaDB’s own writes. The root cause was Raft did not use the system reservation. Stability: read load failing after one node upgrade [bad_enum_set_mask (Bit mask contains invalid enumeration indices.)] #15795 Performance: Repairing a cluster after a restore causes severe reactor stalls throughout the cluster (due to expensive logging within do_repair_ranges() without yield) #14330 Install: scylla_post_install.sh: “[ $RHEL ]” does not work for RHEL, it only detects CentOS #16040 Stability: test_interrupt_build_process dtest failed with schema_registry - Tried to build a global schema for view ks.t_by_v2 with an uninitialized base info #14011 Stability: tests.topology_experimental_raft.test_raft_cluster_features.debug test is flaky. The root cause was error handling in the Raft coordinator. #15747 #15728 --- ### Page: https://forum.scylladb.com/t/how-to-enable-audit-logs-for-scylladb-how-many-types-of-queries-it-supports-and-how-to-see-the-audit-logs-for-the-executed-query-also-how-its-audit-log-looks-like/1054 Title: How to Enable Audit logs for ScyllaDB, how many types of queries it supports and how to see the audit logs for the executed query. Also, how its audit log looks like - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I want to see the Audit Logs for any Executed query and also how I can see those audit logs. Language: en Canonical URL: https://forum.scylladb.com/t/how-to-enable-audit-logs-for-scylladb-how-many-types-of-queries-it-supports-and-how-to-see-the-audit-logs-for-the-executed-query-also-how-its-audit-log-looks-like/1054 ## Headings Structure: H1: How to Enable Audit logs for ScyllaDB, how many types of queries it supports and how to see the audit logs for the executed query. Also, how its audit log looks like H3: Related topics ## Main Content: H1: How to Enable Audit logs for ScyllaDB, how many types of queries it supports and how to see the audit logs for the executed query. Also, how its audit log looks like H3: Related topics I want to see the Audit Logs for any Executed query and also how I can see those audit logs. @subrato see here, for answers ScyllaDB Auditing Guide | ScyllaDB Docs Note that Audit is a Scylla Enterprise-only feature. I want to see the Audit Logs for any Executed query and also how I can see those audit logs. This will have a significant performance impact. Every read (query) will be followed by a write (to the audit log). what will be the parameters of audit logs and how to see the audit logs of executed query?? Here is an example. I use the following config in scylla.yaml for Scylla Enterprise: Created a keyspace mykeyspace and a table heartrate_v10 run a SELECT commands: Looking at the audit table I see: Scylla Enterprise, we need to purchase this for enable audit log or it comes with by default feature of enable audit log. Audit is part of Scylla Enterprise. You need to configure it per the keyspace/table/audit level you need. Thank You so much for your help I am able to understand. I have one last question, suppose I want to capture all audit logs DDL, DML, DDL etc. for any keyspace/table, because every time It is not possible to configure for which Keyspace or table I want to see audit logs. So I want to capture all audit logs does not matter which keyspace or table. Then What would be the configuration?? First, it’s impossible to enable for all use cases. You must list all relevant use cases. Second, audits for every operation, including QUERY, will have a huge performance impact. Can you explain why you need such a level of audit? There may be alternative solutions. Actually, our use case is that we want to enable auditing for all operation, so that we will get all the information about the executed operations perform by user on ScyllaDB and capture those audit log on our system (Windows or Linux). --- ### Page: https://forum.scylladb.com/t/running-3-node-scylladb-in-docker/1057 Title: Running 3 node scylladb in docker - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hey, I’m having hard time starting 3 nodes setup in docker on my local env. 3rd node is reporting FATAL: Exception during startup, aborting: std::runtime_error (Could not setup Async I/O: Resource temporarily unavailab… Language: en Canonical URL: https://forum.scylladb.com/t/running-3-node-scylladb-in-docker/1057 ## Headings Structure: H1: Running 3 node scylladb in docker H3: Related topics ## Main Content: H1: Running 3 node scylladb in docker H3: Related topics Hey, I’m having hard time starting 3 nodes setup in docker on my local env. 3rd node is reporting The image is scylladb/scylla:5.2.0 and host machine is mac osx with intel processor. Any idea how to set this aio parameter on mac ? Try running sysctl -w kern.aiomax=1048576 ➜ ~ sudo sysctl -w kern.aiomax=1048576 Password: kern.aiomax: 90 sysctl: kern.aiomax=1048576: Invalid argument ➜ ~ sudo sysctl -w kern.aiomax="1048576" kern.aiomax: 90 sysctl: kern.aiomax=1048576: Invalid argument on macos it looks like -w is not valid I can set like this but the highest value it would let me set is 11000 (12000 and above said invalid) sudo sysctl kern.aiomax=11000 so stuck on the same thing. I might just try linux but would be nice if containers worked on macos I created a bug report about this problem: Docker on macOS fails with: "Could not setup Async I/O: Resource temporarily unavailable." · Issue #16806 · scylladb/scylladb · GitHub On a intel based chip with “2.8 GHz Quad-Core Intel Core i7” I’m able to start a 3 node cluster, each with --smp 1. $ sysctl -n hw.ncpu 8 sysctl -a | grep aio kern.aiomax: 90 kern.aioprocmax: 16 kern.aiothreads: 4 However, on a apple chip based (Apple M3 Pro), 2 nodes starts ok with --smp 1, but the 3 node has the issue of max-aio-nr. sysctl -a | grep aio kern.aiomax: 90 kern.aioprocmax: 16 kern.aiothreads: 4 I’m happy to share that there is an easy workaround that works also for macOS: You can start scylla with the following flag: “–reactor-backend=epoll” --- ### Page: https://forum.scylladb.com/t/does-gsi-synchronous-write-in-scylla-version-5-2-support-tables-created-by-alternator-interface/1058 Title: Does GSI synchronous write in Scylla version 5.2 support tables created by Alternator interface? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi all! Here is the situation. We initially set up a 5-node Scylla cluster using Docker, then utilized the Alternator interface to create a table. We ensured that the base table and the GSI table had only one common repl… Language: en Canonical URL: https://forum.scylladb.com/t/does-gsi-synchronous-write-in-scylla-version-5-2-support-tables-created-by-alternator-interface/1058 ## Headings Structure: H1: Does GSI synchronous write in Scylla version 5.2 support tables created by Alternator interface? H3: Related topics ## Main Content: H1: Does GSI synchronous write in Scylla version 5.2 support tables created by Alternator interface? H3: Related topics Hi all! Here is the situation. We initially set up a 5-node Scylla cluster using Docker, then utilized the Alternator interface to create a table. We ensured that the base table and the GSI table had only one common replica node in the cluster (for example, the base table replicas are on nodes 1, 2, and 3, while GSI table replicas are on nodes 3, 4, and 5). After that, GSI synchronous write was enabled for this table. Subsequently, we stopped the nodes where the GSI table replicas were located (for instance, nodes 3, 4, and 5). At this point, we were still able to successfully write data to the table. However, based on the logic of synchronous GSI, writing data at this stage should theoretically have been unsuccessful. So I am confused that does GSI synchronous write in Scylla version 5.2 support tables created by Alternator interface? There is no DynamoDB API for making a GSI synchronous, but I assume you went and changed its synchronous_updates flag via CQL. That should be fine, and should work - the “synchronous view updates” feature is part of the Scylla backend implementation, and both CQL and Alternator (DynamoDB API) frontends can use it. The “synchronous” update option was added as a contrast to the default “asynchronous” view updates - view updates that are done in the background and the user doesn’t wait for them. With synchronous updates, we do wait for the update we send over the network, but as you noticed there is one peculiarity: If we already know that the view replica is down, we don’t wait for it to come up - instead, we write a “view hint” - a record on disk that tells us to later try writing this view update. But at this point, Scylla does nothing more, and the code thinks it did what it needed to do, and returns. You’re right that the downside to this is that it is possible that a base-write with “synchronous view updates” can be done, and afterwards a read of the view - and the new data will be missing on the view, because although the view write got written to disk (so it’s durable), it’s not yet available. You are right that it does make sense that in this case the base write should fail instead of succeeding. I’ll open an issue about this in our bug tracker. By the way, if all three view replicas for a row are really down, you can’t actually confirm that the synchronous write didn’t do its job - you can’t read from the view and not see the new data because the read will fail. However, it is indeed possible that the dead nodes come back to life, and before the hints are replayed (which takes time), a read of the view happens and doesn’t see the new data. I opened an issue in Scylla’s bug tracker: Synchronous view updates may not guarante read-your-own-writes if nodes are down · Issue #16229 · scylladb/scylladb · GitHub --- ### Page: https://forum.scylladb.com/t/i-want-to-know-what-will-be-the-pricing-of-scylladb-enterprise-on-premise-edition/1061 Title: I want to know what will be the pricing of ScyllaDb Enterprise on-premise edition - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I have a doubt that, does ScyllaDB Enterprise on-premises have any trial edition, if not then what is the pricing details for Enterprise. Can someone tell me? Language: en Canonical URL: https://forum.scylladb.com/t/i-want-to-know-what-will-be-the-pricing-of-scylladb-enterprise-on-premise-edition/1061 ## Headings Structure: H1: I want to know what will be the pricing of ScyllaDb Enterprise on-premise edition H3: Related topics ## Main Content: H1: I want to know what will be the pricing of ScyllaDb Enterprise on-premise edition H3: Related topics I have a doubt that, does ScyllaDB Enterprise on-premises have any trial edition, if not then what is the pricing details for Enterprise. Can someone tell me? It depends on things like sizing and your specific use case. I suggest you get in touch via this page so that our experts can help you. --- ### Page: https://forum.scylladb.com/t/in-the-given-audit-log-enable-configuration-if-i-want-to-capture-all-audit-for-different-keyspaces-and-table-can-we-remove-last-two-parameter-so-that-it-will-capture-audit-logs-for-all-keyspaces/1062 Title: In the given audit log enable configuration, If I want to capture all audit for different keyspaces and table, can we remove last two parameter so that it will capture audit logs for all keyspaces? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: # audit setting # by default, Scylla does not audit anything. # It is possible to enable auditing to the following places: # - audit.audit_log column family by setting the flag to "table" audit: "syslog" # # List of st… Language: en Canonical URL: https://forum.scylladb.com/t/in-the-given-audit-log-enable-configuration-if-i-want-to-capture-all-audit-for-different-keyspaces-and-table-can-we-remove-last-two-parameter-so-that-it-will-capture-audit-logs-for-all-keyspaces/1062 ## Headings Structure: H1: In the given audit log enable configuration, If I want to capture all audit for different keyspaces and table, can we remove last two parameter so that it will capture audit logs for all keyspaces? H3: Related topics ## Main Content: H1: In the given audit log enable configuration, If I want to capture all audit for different keyspaces and table, can we remove last two parameter so that it will capture audit logs for all keyspaces? H3: Related topics Unfortunately, it looks like you need to explicitly list all tables in there, if you want them audited. @tzach is there a specific reason we don’t have a catch-all option? No specific reason. The assumption was that Audit, as a performance-consuming feature, would be used for specific keyspace and tables. --- ### Page: https://forum.scylladb.com/t/scylladb-university-live-training-event-december-5-2023/1063 Title: ScyllaDB University LIVE Training Event - December 5, 2023 - Announcements - ScyllaDB Community NoSQL Forum Meta Description: Next week, we’ll host the ScyllaDB University LIVE training event. It’s a half day of free online training by our top experts and engineers. The sessions will be interactive and won’t be available on demand. More info… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-university-live-training-event-december-5-2023/1063 ## Headings Structure: H1: ScyllaDB University LIVE Training Event - December 5, 2023 H3: Related topics ## Main Content: H1: ScyllaDB University LIVE Training Event - December 5, 2023 H3: Related topics Next week, we’ll host the ScyllaDB University LIVE training event. It’s a half day of free online training by our top experts and engineers. The sessions will be interactive and won’t be available on demand. More info here, hope to see you there! --- ### Page: https://forum.scylladb.com/t/re-c-c-api-cannot-insert-a-date-value-into-a-table-or-select-a-date-value-from-a-table/1064 Title: RE: C/C++ API - cannot insert a date value into a table or select a date value from a table - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi I am using the C/C++ API on Windows from Datastax. I have created a table that has several columns including some columns that are defined as time, timestamp and date. The time and timestamp columns are fine for in… Language: en Canonical URL: https://forum.scylladb.com/t/re-c-c-api-cannot-insert-a-date-value-into-a-table-or-select-a-date-value-from-a-table/1064 ## Headings Structure: H1: RE: C/C++ API - cannot insert a date value into a table or select a date value from a table H3: Related topics ## Main Content: H1: RE: C/C++ API - cannot insert a date value into a table or select a date value from a table H3: Related topics Hi I am using the C/C++ API on Windows from Datastax. I have created a table that has several columns including some columns that are defined as time, timestamp and date. The time and timestamp columns are fine for inserting and selecting, however when I try to either insert or do a select on the date column I get a bind error of CASS_ERROR_LIB_INVALID_VALUE_TYPE. After performing a select I use this code to get the data:: value = cass_row_get_column_by_name(m_row, “tDate”); cass_int32_t dte = 0; CassError rc = cass_value_get_int32(value, &dte); rc is set to CASS_ERROR_LIB_INVALID_VALUE_TYPE Have you tried with cass_uint32_t and cass_value_get_uint32? According to our documentation ( Data Types | ScyllaDB Docs ): Values of the date type are encoded as 32-bit unsigned integers representing a number of days with “the epoch” at the center of the range (2^31). Hi Yep I was using int32 instead of uint32, changed my code and it is all working now. --- ### Page: https://forum.scylladb.com/t/after-successfully-install-scylladb-enterprise-on-ubuntu-we-are-trying-to-update-configuration-file-scylla-yaml-and-after-saving-the-below-configuaration-we-tried-to-restart-scylla-but-getting-error/1066 Title: After Successfully Install ScyllaDB Enterprise on UBUNTU, we are trying to update configuration file scylla.yaml and after Saving the below configuaration, we tried to restart Scylla but getting error - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: # The name of the cluster. This is mainly used to prevent machines in # one logical cluster from joining another. cluster_name: 'Test Cluster' # This defines the number of tokens randomly assigned to this node on the ri… Language: en Canonical URL: https://forum.scylladb.com/t/after-successfully-install-scylladb-enterprise-on-ubuntu-we-are-trying-to-update-configuration-file-scylla-yaml-and-after-saving-the-below-configuaration-we-tried-to-restart-scylla-but-getting-error/1066 ## Headings Structure: H1: After Successfully Install ScyllaDB Enterprise on UBUNTU, we are trying to update configuration file scylla.yaml and after Saving the below configuaration, we tried to restart Scylla but getting error H3: Related topics ## Main Content: H1: After Successfully Install ScyllaDB Enterprise on UBUNTU, we are trying to update configuration file scylla.yaml and after Saving the below configuaration, we tried to restart Scylla but getting error H3: Related topics What’s the error? Share as many details as possible. After uncomment “Cluster name” and restarting i am getting below error Startup failed: exceptions::configuration_exception (Saved cluster name != configured name Demo cluster1) In scylla.yaml file I uncommented the Cluster name field and getting the below error: Startup failed: exceptions::configuration_exception (Saved cluster name != configured name Demo cluster1) Could you please tell me why it not working ?? --- ### Page: https://forum.scylladb.com/t/i-want-to-know-configuration-of-scylla-yaml-file-for-scylladb-enterprise-30days-trial-edition-on-ubuntu-so-that-i-can-capture-cql-query-audit-logs/1067 Title: I Want to know configuration of scylla.yaml file for ScyllaDB Enterprise 30days trial edition on UBUNTU so that i can Capture CQL query audit logs - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I Want to know configuration of scylla.yaml for enabling audit logging and capture all executed CQL queries. Language: en Canonical URL: https://forum.scylladb.com/t/i-want-to-know-configuration-of-scylla-yaml-file-for-scylladb-enterprise-30days-trial-edition-on-ubuntu-so-that-i-can-capture-cql-query-audit-logs/1067 ## Headings Structure: H1: I Want to know configuration of scylla.yaml file for ScyllaDB Enterprise 30days trial edition on UBUNTU so that i can Capture CQL query audit logs H3: Related topics ## Main Content: H1: I Want to know configuration of scylla.yaml file for ScyllaDB Enterprise 30days trial edition on UBUNTU so that i can Capture CQL query audit logs H3: Related topics I Want to know configuration of scylla.yaml for enabling audit logging and capture all executed CQL queries. It appears that you are simply starting several threads for every question or problem you face, without actually providing any context over what you are trying to do, nor providing any feedback on what you tried, which resources have you looked after, nor providing any logs or information that could enable us to assist. For example, here you are asking about Audit, which you have also done in: You have also posted on about problems with starting ScyllaDB, whereas here you are asking basically the same thing: And you have also posted an inquiry about pricing: At this point, it is clear that simply opening a question to every single question you may have will not scale. First and foremost, if you are evaluating ScyllaDB Enterprise, then you should get in touch with us so we can address all your questions promptly and work along with you during your trial period. You can get in touch with us here (there’s a large “Chat Now” button at the top of the page). Second, if you are not interested in talking to someone, then by all means, please get started with our already available community learning resources, such as: We will be more than glad to assist you as needed, have a great day! --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-207-2023-12-03/1069 Title: Last week in scylladb.git master (issue #207; 2023-12-03) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the a472700309…01e54f5b12 range are covered. There were 75 non-merge commits from 20 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-207-2023-12-03/1069 ## Headings Structure: H1: Last week in scylladb.git master (issue #207; 2023-12-03) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #207; 2023-12-03) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the a472700309…01e54f5b12 range are covered. There were 75 non-merge commits from 20 authors in that period. Some notable commits: A regression in IPv6 address formatting, which caused nodetool problems, was fixed. SELECT JSON now emits valid json for a column of type time. Tablets is a new, experiment way to distribute data across the cluster. A bug caused data movement across the cluster to halt schema changes (DDL statements). This is now fixed. It is now possible to create materialized views for a keyspace with tablets enabled. Such views may or may not work correctly. A graphical utility for visualizing tablet migrations is now available. A bug in how schema tables (the tables that store the schema for the database) are manipulated has been fixed. The bug is currently benign and is fixed to enable the further changes. A major compaction may now flush all tables under certain conditions; this is to reduce the probability that a tombstone will not be garbage-collected because it conflicts with data in commitlog. The experimental consistent cluster topology management now handles more failure cases automatically. Commitlog will now avoid going over its configured disk space size. To do so, it will flush memtables earlier. There is now a mechanism to deprecate configuration options. There is a new CQL statement, LIST EFFECTIVE SERVICE LEVEL, to query what service level attributes apply to a role. A role’s service level attributes may be a mix of different SERVICE LEVEL configurations. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/downsides-of-skip-wait-for-gossip-to-settle-option/1070 Title: Downsides of skip-wait-for-gossip-to-settle option - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi all, Could someone explain what are the possible downsides of using --skip-wait-for-gossip-to-settle 0 on a cluster with multiple nodes, if it should even be done? I’m currently using it on a single node cluster (te… Language: en Canonical URL: https://forum.scylladb.com/t/downsides-of-skip-wait-for-gossip-to-settle-option/1070 ## Headings Structure: H1: Downsides of skip-wait-for-gossip-to-settle option H3: Related topics ## Main Content: H1: Downsides of skip-wait-for-gossip-to-settle option H3: Related topics Hi all, Could someone explain what are the possible downsides of using --skip-wait-for-gossip-to-settle 0 on a cluster with multiple nodes, if it should even be done? I’m currently using it on a single node cluster (testing purposes) as it takes forever to start, is it normal? could it be caused by something else? thanks Once a node starts, all other endpoints are down. If you skip gossip, >=LOCAL_QUORUM queries may fail if a coordinator query hits the RPC of that node. Another issue: Make gossip synchronization on bootstrap more robust · Issue #2866 · scylladb/scylladb · GitHub That said, you may skip it for testing, as indeed waiting ~30s at every startup can be annoying. Keep it as is on production tho. --- ### Page: https://forum.scylladb.com/t/i-configured-the-scylladb-enterprise-on-amazonec2-linux-but-audit-logs-are-not-generating-i-added-picture-and-yaml-file-below/1073 Title: I Configured the ScyllaDB Enterprise on AmazonEC2 Linux but Audit Logs are not generating. I added Picture and YAML file below - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Below is my Scylladb.yaml file which I configured for enabling the audit logs. # Scylla storage config YAML ####################################### # This file is split to two sections: # 1. Supported parameters # 2. U… Language: en Canonical URL: https://forum.scylladb.com/t/i-configured-the-scylladb-enterprise-on-amazonec2-linux-but-audit-logs-are-not-generating-i-added-picture-and-yaml-file-below/1073 ## Headings Structure: H1: I Configured the ScyllaDB Enterprise on AmazonEC2 Linux but Audit Logs are not generating. I added Picture and YAML file below H3: Related topics ## Main Content: H1: I Configured the ScyllaDB Enterprise on AmazonEC2 Linux but Audit Logs are not generating. I added Picture and YAML file below H3: Related topics Below is my Scylladb.yaml file which I configured for enabling the audit logs. I also attached the picture where I am not able generate audit logs in Audit Table. In your example you’ve run a DDL statement. Check the documentation. This is, as you may guess, redundant and thus wrong. If you simply want to log all events in a keyspace, then just use audit_keyspaces. Hi, I am successfully Able to generate audit logs, but Error logs are not present here. --- ### Page: https://forum.scylladb.com/t/is-scylladb-a-one-shoe-fit-all/1075 Title: Is ScyllaDB a one shoe fit all? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: ScyllaDB supports OLTP use cases (so no need for rdbms) It supports document storage (so no need for mongodb) Columnar database (no need for cassandra). I might be missing something here, but it seems that ScyllaDB is … Language: en Canonical URL: https://forum.scylladb.com/t/is-scylladb-a-one-shoe-fit-all/1075 ## Headings Structure: H1: Is ScyllaDB a one shoe fit all? H3: Related topics ## Main Content: H1: Is ScyllaDB a one shoe fit all? H3: Related topics I might be missing something here, but it seems that ScyllaDB is covering all the use cases for RDBMS, MongoDB and Cassandra and is faster then the NoSQL counterparts. I still might use RDBMS for some use simple cases. But is ScyllaDB really a one shoe fit all DB? While I think there is no “one shoe fits all” solution when it comes to databases in general (because each use case is slightly different and the needs are different) we still see lots of ScyllaDB users who replace multiple tools with just ScyllaDB. In this sense, ScyllaDB CAN be a one-shoe-fits-all solution if your goals are aligned with what ScyllaDB can offer. ScyllaDB supports OLTP use cases (so no need for rdbms) Yes it does and would be a great choice for any use cases where you require low latency, high availability, and large scale It supports document storage (so no need for MongoDB) You can store JSONs in ScyllaDB and you can query it. Your data model won’t be as flexible but ScyllaDB will provide better and consistent performance. We have a great user story on when to use ScyllaDB vs MongoDB As far as comparing ScyllaDB to Cassandra, the interface is the same (CQL) but ScyllaDB is implemented with a different language (C++) and several benchmarks prove that ScyllaDB is more performant/less expensive than Cassandra. I still might use RDBMS for some use simple cases Yes I think that would make sense for a simple use case, and you can rely on ScyllaDB where performance and low latency matter. Note that ScyllaDB is not a Columnar DB, nor is Apache Cassandra. Both are considered “wide column store” (confusing term on its own). ScyllaDB is not the best DB for a pure analytic DB (OLAP), no real-time, use case. It is a good solution for mixing OLTP and OLAP, particularly with Workload Prioritization. ScyllaDB is not a Columnar DB What is it then? Can you explain the difference? I’m a bit confused, because here it says: A Wide Column Store, also known as a column store, column family store, columnar data store, or column store database A Wide Column Store, also known as a column store, column family store, columnar data store, or column store database IMHO, this is misleading, but since Wikipedia defines Wide-column_store as a type of Column-oriented_DBMS maybe its only me. If you can describe your use case, we can advise if it fits ScyllaDB well. --- ### Page: https://forum.scylladb.com/t/what-is-this-time-format-in-event-time-field-of-scylladb-audit-logs-please-somebody-tell-me-below-is-mentioned-audit-log/1078 Title: What is this time format in "event time" field of ScyllaDb audit logs. Please somebody tell me. Below is mentioned audit log - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: date | node | event_time | category | consistency | error | keyspace_name | operation … Language: en Canonical URL: https://forum.scylladb.com/t/what-is-this-time-format-in-event-time-field-of-scylladb-audit-logs-please-somebody-tell-me-below-is-mentioned-audit-log/1078 ## Headings Structure: H1: What is this time format in "event time" field of ScyllaDb audit logs. Please somebody tell me. Below is mentioned audit log H3: Related topics ## Main Content: H1: What is this time format in "event time" field of ScyllaDb audit logs. Please somebody tell me. Below is mentioned audit log H3: Related topics event_time is a timeuuid. Quoting from the cql types document: Version 1 UUID, generally used as a “conflict-free” timestamp. See Working with UUIDs for details --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-1-12/1079 Title: [RELEASE] ScyllaDB Enterprise 2022.1.12 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2022.1.12, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.1. Note that there is a newer Enterprise LTS release: ScyllaDB Enterpr… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-1-12/1079 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2022.1.12 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2022.1.12 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2022.1.12, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.1. Note that there is a newer Enterprise LTS release: ScyllaDB Enterprise 2023.1. While we will continue to support 2022.1 LTS, you can get additional features with 2023.1. Below is a list of performance and stability improvements and bug fixes, each with an open-source reference if available: --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-1-19/1080 Title: [RELEASE] ScyllaDB 5.1.19 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.1.19, a bugfix release of the ScyllaDB 5.1 stable branch. ScyllaDB Open Source 5.1.19, like all past and future 5.x.y releases are backward compatible and support rollin… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-1-19/1080 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.1.19 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.1.19 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.1.19, a bugfix release of the ScyllaDB 5.1 stable branch. ScyllaDB Open Source 5.1.19, like all past and future 5.x.y releases are backward compatible and support rolling upgrades. Note that the latest stable ScyllaDB Open Source release is 5.2, and you are encouraged to upgrade to it. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-4-0-part-1/1084 Title: [RELEASE] ScyllaDB 5.4.0 - Part 1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of ScyllaDB Open Source 5.4.0, a production-ready release of our open-source NoSQL database. ScyllaDB 5.4 introduces Repair Base Node Operations (RBNO) for all operat… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-4-0-part-1/1084 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.4.0 - Part 1 H2: Repair Based Node Operations (RBNO) H2: Node Level Metrics H2: Guardrails H2: Security H2: Deprecated and removed features H2: Deployment and install H2: UDF / UDA - Preview H2: New nodetool implementation - Experimental H2: Object Storage Support - Experimental H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.4.0 - Part 1 H2: Repair Based Node Operations (RBNO) H2: Node Level Metrics H2: Guardrails H2: Security H2: Deprecated and removed features H2: Deployment and install H2: UDF / UDA - Preview H2: New nodetool implementation - Experimental H2: Object Storage Support - Experimental H3: Related topics The ScyllaDB team is pleased to announce the release of ScyllaDB Open Source 5.4.0, a production-ready release of our open-source NoSQL database. ScyllaDB 5.4 introduces Repair Base Node Operations (RBNO) for all operations, experimental consistent topology update, experimental S3 backend and many more improvements and bug fixes. Consistent schema management using Raft will be enabled automatically on upgrade (more below). ScyllaDB is now supported on RHEL/Rocky 9, while RHEL/CentOS 7 support is deprecated. RHEL/Rocky 8, Ubuntu 20.04, 22.04, and Debian 10,11 support continue. Find the ScyllaDB Open Source 5.4 repository for your Linux distribution here. ScyllaDB 5.4 Docker is also available. Only the last two minor releases of the ScyllaDB Open Source project are supported. From now on, only ScyllaDB Open Source 5.4 and ScyllaDB 5.2 will be supported, and ScyllaDB 5.1 will be retired. (note there is no 5.3 release) Get ScyllaDB Open Source 5.4 as binary packages (RPM/DEB), AWS AMI, GCP Image and Docker image Upgrade from ScyllaDB 5.2 to ScyllaDB 5.4 Repair Based Node Operations were introduced as an experimental feature in ScyllaDB Open Source 4.0. They use repair to stream data for node-operations like replace, bootstrap and others. In 5.4, RBNO is enabled by default for all operations: remove node,rebuild,bootstrap, and decommission. Replace node operation was already enabled by default. See Repair Base Node Operations (RBNO) docs and the “Faster, Safer Node Operations with Repair vs Streaming” blog by Asias He Most ScyllaDB metrics are per-shard, per-node, but not for a specific table. We now export some per-table metrics. These are exported once per node, not per shard, to reduce the number of metrics. #2198 Guardrails is a framework to protect ScyllaDB users and admins from common mistakes and pitfalls. In this release ScyllaDB includes a new guardrail on the replication factor. It is now possible to specify the minimum replication factor for new keyspaces via a new configuration item #8891. This matches the same functionality in Apache Cassandra #CASSANDRA-14557 The new RF guardrails include the following configuration: More Guardrails are expected in upcoming releases. See more on certificate-authentication docs. Wasm-based User Defined Functions (UDFs) and User Defined Aggregates (UDAs) were introduced as experimental in ScyllaDB 5.1, and we are now promoting them to Preview. The CQL syntax is compatible with Apache Cassandra. Examples: CREATE FUNCTION sample ( arg int ) …; CREATE FUNCTION sample ( arg text ) …; See UDF/UDA documentation for more information. Preview features aren’t production-ready, but are made available on a “preview” basis so that users can get early access and provide feedback. Unlike Experimental features, we are committed to the backward compatibility of a preview feature API. Below are improvements and bug fixes in UDF/UDA. The scylla executable can now act as nodetool by executing “scylla nodetool ”. With time the new scylla nodetool will replace the legacy Java nodetool completely. The new implementation is fully backward compatible with the legacy nodetool. The following commands are currently implemented: For the latest status of nodetool replacement progress see #15588 You can store an entire ScyllaDB Keyspace to Amazon S3, or a compatible object store. Enable the feature by: –experimental-features=keyspace-storage-options CREATE KEYSPACE with STORAGE = { ‘type’: ‘S3’, ‘endpoint’: ‘$endpoint_name’, ‘bucket’: ‘$bucket’ } See Keyspace Storage Options docs. Note: at this phase, obj-storage is not ready for production, and should be used for testing only. Note2: there is no ALTER support for the STORAGE parameter. --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-4-0-part-2-cql-api-strongly-consistent-schema-management-alternator-and-more/1085 Title: [RELEASE] ScyllaDB 5.4.0 - Part 2: CQL API, Strongly Consistent Schema Management , Alternator and more - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: For 5.4 Part 1 More Improvements CQL API CQL table columns that have the list data type aren’t allowed to contain NULLs, but in certain situations list values in CQL literals or bind variables are allowed to contain NU… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-4-0-part-2-cql-api-strongly-consistent-schema-management-alternator-and-more/1085 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.4.0 - Part 2: CQL API, Strongly Consistent Schema Management , Alternator and more H2: More Improvements H3: CQL API H3: Amazon DynamoDB Compatible API (Alternator) H3: Strongly Consistent Schema Management with Raft H3: Strongly Consistent Topology Updates with Raft - Experimental H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.4.0 - Part 2: CQL API, Strongly Consistent Schema Management , Alternator and more H2: More Improvements H3: CQL API H3: Amazon DynamoDB Compatible API (Alternator) H3: Strongly Consistent Schema Management with Raft H3: Strongly Consistent Topology Updates with Raft - Experimental H3: Related topics Examples: blob_column = (blob)(int)12323 Alternator is ScyllaDB’s implementation of the DynamoDB API. Strongly Consistent Schema Management with Raft became the default for new clusters in ScyllaDB 5.2. In this release it is enabled by default when upgrading existing clusters. If you do not want to enable Raft, you should explicitly disable it in scylla.yaml of each node before the upgrade: consistent_cluster_management: false Source: Upgrade from 5.2 to 5.4 doc Below are additional related fixes and updates: This release includes an experimental Strongly Consistent Topology Updates. To enable it, use the new consistent-topology-changes flag. To enable, update the following in scylla.yaml Below are additional related fixes and updates: --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-4-0-part-3-monitoring-tools-performance-stability-and-more/1086 Title: [RELEASE] ScyllaDB 5.4.0 - Part 3 - Monitoring, Tools, Performance, Stability and more - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: For 5.4 Part 1 For 5.4 Part 2 Monitoring, tracing and logging Scylla Monitoring Stack released 4.5 and later will support ScyllaDB 5.4. metrics related updates below: There is a new metric for prepared statement cac… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-4-0-part-3-monitoring-tools-performance-stability-and-more/1086 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.4.0 - Part 3 - Monitoring, Tools, Performance, Stability and more H3: Monitoring, tracing and logging H3: Operations H3: Tools H3: Configuration Updates H3: Admin REST API H3: Performance and stability H3: Build H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.4.0 - Part 3 - Monitoring, Tools, Performance, Stability and more H3: Monitoring, tracing and logging H3: Operations H3: Tools H3: Configuration Updates H3: Admin REST API H3: Performance and stability H3: Build H3: Related topics For 5.4 Part 1 For 5.4 Part 2 Scylla Monitoring Stack released 4.5 and later will support ScyllaDB 5.4. metrics related updates below: As part of that change, cqlsh is now compatible with Python 3. CQLSh is now available as a Docker image, and in PiPy, allowing you to easily use it when you do not need the entire ScyllaDB server, for example with Scylla Cloud. See Scylla SSTable docs for more info. The scylla.yaml configuration items are now documented in the documentation website. New and updated configuration options: Index_cache_fraction is the maximum fraction of cache memory permitted for use by index cache. Clamped to the [0.0; 1.0] range. Must be small enough to not deprive the row cache of memory, but should be big enough to fit a large fraction of the index. The default value 0.2 means that at least 80% of cache memory is reserved for the row cache, while at most 20% is usable by the index cache. ScyllaDB uses two separate memory reservation systems for memtables: user, used for user writes, and system, used for ScyllaDB’s own writes. The root cause was Raft did not use the system reservation. Stability: read load failing after one node upgrade [bad_enum_set_mask (Bit mask contains invalid enumeration indices.)] #15795 Performance: Repairing a cluster after a restore causes severe reactor stalls throughout the cluster (due to expensive logging within do_repair_ranges() without yield) #14330 Install: scylla_post_install.sh: “[ $RHEL ]” does not work for RHEL, it only detects CentOS #16040 Stability: test_interrupt_build_process dtest failed with schema_registry - Tried to build a global schema for view ks.t_by_v2 with an uninitialized base info #14011 Stability: tests.topology_experimental_raft.test_raft_cluster_features.debug test is flaky. The root cause was error handling in the Raft coordinator. #15747 #15728 Stability: The mutation compactor now validates its input stream rather than the output stream. [IPv6 configuration] A node is stuck with “?U” status and Host ID is “null”, unclear reason #16039 Stability: assigning position_in_partition is not exception safe, can lead to incorrect data during memory stress #15822 Stability: Scylla cluster nodes utilizes 100% of CPU even with no load #12774, #13377, #7753 --- ### Page: https://forum.scylladb.com/t/select-with-static-columns-efficiently/1088 Title: Select * with static columns efficiently - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: For example I have the following table: (pk, ck1, ck2, s1 STATIC, s2 STATIC, v1, v2) I have multiple rows with the same pk, and if I request select * where pk = 1, it will return copies of static columns for each row. … Language: en Canonical URL: https://forum.scylladb.com/t/select-with-static-columns-efficiently/1088 ## Headings Structure: H1: Select * with static columns efficiently H3: Related topics ## Main Content: H1: Select * with static columns efficiently H3: Related topics For example I have the following table: (pk, ck1, ck2, s1 STATIC, s2 STATIC, v1, v2) I have multiple rows with the same pk, and if I request select * where pk = 1, it will return copies of static columns for each row. My question is: Is there any way to save wire? Maybe it’s just my perfectionism I see following options: send multiple select queries simultaneously (select all static, and select all non-static columns) Will It work slower? Or (I thought if it’s a columnar storage) it will be the same. Or it still will be slower because of pk hashing? it will not copy static columns in the response, you are wrong I think that for most CQL types, just using SELECT * is the best option. If your static columns are text or blob and you have really large values and a lot of small clustering row, then two separate queries (one for the static columns only, another for the rest), might make sense. I don’t know what is the size of values above which this makes sense. If I had to guess, I would say, in the MB range. Note that ScyllaDB internally uses a transfer format which stores static columns just once per partition. It is only from the coordinator to the application which uses the tabular format, which includes static columns on each row. Thanks. I wish it would not use the tabular format, and send the data in a more efficient way. --- ### Page: https://forum.scylladb.com/t/i-installed-scylladb-enterprise-30days-trail-on-amazon-ec2-instance-and-now-i-want-to-connect-dbvisualizer-to-our-this-scylladb-to-execute-query-i-install-datastax-cassandra-jar-in-dbvisulaizer-but-not-able-to-connect/1090 Title: I installed ScyllaDB Enterprise 30days Trail on Amazon EC2 instance and now I want to Connect dbvisualizer to our this ScyllaDb to execute query. I install DataStax cassandra jar in Dbvisulaizer but not able to connect - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I change rpc_address and local_address to Publicip of EC2 instance but after restarting getting below error. Startup failed: std::system_error (error system:99, posix_listen failed for address 35.173.42.14:9180: Cannot … Language: en Canonical URL: https://forum.scylladb.com/t/i-installed-scylladb-enterprise-30days-trail-on-amazon-ec2-instance-and-now-i-want-to-connect-dbvisualizer-to-our-this-scylladb-to-execute-query-i-install-datastax-cassandra-jar-in-dbvisulaizer-but-not-able-to-connect/1090 ## Headings Structure: H1: I installed ScyllaDB Enterprise 30days Trail on Amazon EC2 instance and now I want to Connect dbvisualizer to our this ScyllaDb to execute query. I install DataStax cassandra jar in Dbvisulaizer but not able to connect H3: Related topics ## Main Content: H1: I installed ScyllaDB Enterprise 30days Trail on Amazon EC2 instance and now I want to Connect dbvisualizer to our this ScyllaDb to execute query. I install DataStax cassandra jar in Dbvisulaizer but not able to connect H3: Related topics I change rpc_address and local_address to Publicip of EC2 instance but after restarting getting below error. --- ### Page: https://forum.scylladb.com/t/release-scylla-manager-3-2-5/1091 Title: [RELEASE] Scylla Manager 3.2.5 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces the release of ScyllaDB Manager 3.2.5 production-ready ScyllaDB Manager patch release of the stable ScyllaDB Manager 3.2 branch. ScyllaDB Manager is a centralized cluster administration and r… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-manager-3-2-5/1091 ## Headings Structure: H1: [RELEASE] Scylla Manager 3.2.5 H3: Related topics ## Main Content: H1: [RELEASE] Scylla Manager 3.2.5 H3: Related topics The ScyllaDB team announces the release of ScyllaDB Manager 3.2.5 production-ready ScyllaDB Manager patch release of the stable ScyllaDB Manager 3.2 branch. ScyllaDB Manager is a centralized cluster administration and recurrent tasks automation tool. This release fixes issues in Manager backup, in particular for Alternator, and repair. In this version, the team improved the agent’s memory utilization (#3298) during the backup task. The issue with memory utilization was noticeable in clusters utilizing the LCS compaction strategy, which generate many small SSTables files ScyllaDB Enterprise customers are encouraged to upgrade to ScyllaDB Manager 3.2.5 in coordination with the ScyllaDB support team. The new release includes upgrades of both ScyllaDB Manager Server and Agent. Scylla Manager 3.2.5 support the following Scylla Enterprise releases: And the following Open Source release (limited to 5 nodes see license): You can install and run Scylla Manager on Kubernetes using Scylla Operator. More here. --- ### Page: https://forum.scylladb.com/t/what-is-considered-a-large-cluster/1098 Title: What is considered a large cluster? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, out of curiosity: What is considered to be a large ScyllaDB cluster? (in terms of number of nodes) Is anyone using RF > 3 to avoid problems with availability? regards, Christian Language: en Canonical URL: https://forum.scylladb.com/t/what-is-considered-a-large-cluster/1098 ## Headings Structure: H1: What is considered a large cluster? H3: Related topics ## Main Content: H1: What is considered a large cluster? H3: Related topics Hi, out of curiosity: What is considered to be a large ScyllaDB cluster? (in terms of number of nodes) Is anyone using RF > 3 to avoid problems with availability? We know of a cluster of a user of ours, that has 180 nodes. That is definitely considered large and in some ways it is pushing ScyllaDB, just by the sheer size of that cluster. That said, I don’t know at what number of nodes a cluster is considered large. I would say somewhere in the low-mid double digit number of nodes (30-50). I’ve also seen a cluster with many data centers, geographically dispersed across multiple continents. Number of nodes is one measure of large, number of DC’s is another. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-31-2023-12-09/1099 Title: Last week in scylla-cluster-tests.git master (issue #31; 2023-12-09) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last 3 weeks. Commits in the 91aec87b…7c7279b7 range are covered. There were 76 non-merge commits from 13 authors in… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-31-2023-12-09/1099 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #31; 2023-12-09) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #31; 2023-12-09) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last 3 weeks. Commits in the 91aec87b…7c7279b7 range are covered. There were 76 non-merge commits from 13 authors in that period. Some notable commits: Exposed shard-aware ports in AWS security group fixing issue with running tests from local machine with backend on AWS. We added 2 CI multidc jobs for the EKS backend - 3 and 12 hours long. New util for decoding kallsyms It can be called from command line as follow: Scylla-Manager now works in MultiDC K8S tests. SCT supports MultiDC on GKE backend and 2 AWS-like tests were added. In scylla-operator-v1.11.0 was added the possibility to gather K8S logs by using it’s binary. It is required for support requests and bug reports. ‘must-gather’ operator command is now used in SCT to gather these logs to a separate archive. SCTConfiguration class now can detect the scylla version early with get_version_based_on_conf method. We can use this information for test configuration before db nodes are up. Argus now collects and presents SCT version based on git repository details: url, branch, sha. latte is a rust base stress tool that works with scylla rust driver, it claimed to be faster and more efficient than cassandra-stress NoSQLBench, and tlp-stress. We introduced support for it and can stress Scylla using this tool. From time to time we upgrade Node exporter, which is in charge of OS metrics from Scylla nodes. Now we check it’s availability in artifact tests. SCT enables user encryption by default when the test is configured with KMS - which is enabled by default if enterprise and AWS. “EndOfQuotaNemesis” was created.. This Nemesis is configuring XFS quota for scylla user on a Scylla node and verifies proper handling of this error. When testing on the Docker backend we can now set specific docker network, this enables testing with components that are started with docker-compose (e.g. Kafka). See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-208-2023-12-10/1100 Title: Last week in scylladb.git master (issue #208; 2023-12-10) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 01e54f5b12…d62a5fc60b range are covered. There were 146 non-merge commits from 18 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-208-2023-12-10/1100 ## Headings Structure: H1: Last week in scylladb.git master (issue #208; 2023-12-10) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #208; 2023-12-10) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 01e54f5b12…d62a5fc60b range are covered. There were 146 non-merge commits from 18 authors in that period. Some notable commits: Commitlog format has changed to individually checksum each disk sector. This allows discrimination between three types of sectors: those containing valid data, those corrupted, and those that contain data from a previous use of the commitlog segment file. This should reduce false-positive warnings about commitlog corruptions during replay. Note the format is not compatible with previous versions, so upgrades should flush memtables first. Tablets are a new, experimental way of distributing data across the cluster. A table that uses tablets and has a materialized view will now correctly and consistently pair base-table tablets with corresponding view-table tablets, reducing chances for losing consistency compared to vnode-based tables. When tablets are migrated to a different node, they will not prevent other tablets from initiating migration, increasing concurrency. It’s now possible to use cmake to build ScyllaDB from the configure.py entry point. The scylladb-kernel-conf package tunes the Linux kernel scheduler via sysfs to improve latency. These tunings were lost in Linux 5.13+ due to kernel changes. They are now restored. ScyllaDB can store changes to its own metadata using a separate commitlog, to prevent it waiting for user data to commit. This separate commitlog is now mandatory. There is now a script to assist in decoding UUID-based sstable generation numbers. ScyllaDB detects internal stalls using a stall detector. It can do so using a timer, or using a more hardware performance counter, which can also track kernel stalls. It now sets up permissions for itself to use the performance counter. The native nodetool implementation, scylla nodetool, now supports more commands: decommission, rebuild. removenode, logging level commands, move, and refresh. It is not yet used by default. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/is-there-is-any-way-to-connect-third-party-tool-db-visualizer-to-scylla-db-on-aws-ec2-instance/1103 Title: Is there is any way to connect Third party tool DB Visualizer to Scylla dB on AWS Ec2 instance - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I am Trying to Connect Visualizer with Scylldb which I installed in Amazon Linux UBUNTU, and I want to connect to my database through DB Visualizer and want to execute CQL queries. But I Don’t know how and which driver … Language: en Canonical URL: https://forum.scylladb.com/t/is-there-is-any-way-to-connect-third-party-tool-db-visualizer-to-scylla-db-on-aws-ec2-instance/1103 ## Headings Structure: H1: Is there is any way to connect Third party tool DB Visualizer to Scylla dB on AWS Ec2 instance H3: Related topics ## Main Content: H1: Is there is any way to connect Third party tool DB Visualizer to Scylla dB on AWS Ec2 instance H3: Related topics I am Trying to Connect Visualizer with Scylldb which I installed in Amazon Linux UBUNTU, and I want to connect to my database through DB Visualizer and want to execute CQL queries. But I Don’t know how and which driver we can use ? You can use the “Cassandra DataStax” driver in DB Visualizer. ScyllaDB is Cassandra-compatible so that should work correctly. --- ### Page: https://forum.scylladb.com/t/scylladb-column-storage-on-disk/1106 Title: ScyllaDB column storage on disk - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I’m new to ScyllaDB and trying to understand how data is stored at the storage level. I’m exploring the use of ScyllaDB for OLAP workloads, where analytical tasks could benefit from a columnar storage structure, storing … Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-column-storage-on-disk/1106 ## Headings Structure: H1: ScyllaDB column storage on disk H3: Related topics ## Main Content: H1: ScyllaDB column storage on disk H3: Related topics I’m new to ScyllaDB and trying to understand how data is stored at the storage level. I’m exploring the use of ScyllaDB for OLAP workloads, where analytical tasks could benefit from a columnar storage structure, storing entire columns contiguously in memory. As ScyllaDB is a wide column store, I’m wondering if it stores column family data together on disk. Can someone clarify and guide me to relevant documentation? Any help is much appreciated. Thank you ScyllaDB does not have a columnar data format. Data is storage is partition and row oriented. See sstables-directory-structure.md for more information about our on-disk format. --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-2-15/1107 Title: [RELEASE] ScyllaDB Enterprise 2022.2.15 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2022.2.15, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.2 Feature Release.. Note the latest ScyllaDB Enterprise release is 202… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-2-15/1107 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2022.2.15 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2022.2.15 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2022.2.15, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.2 Feature Release.. Note the latest ScyllaDB Enterprise release is 2023.1 LTS, and you are encouraged to upgrade to it in coordination with the ScyllaDB Support team. The following issues are fixed in this release (with an open-source reference, if available): --- ### Page: https://forum.scylladb.com/t/does-scylladb-enterprise-support-auditing-of-error-audit-logs-or-not-i-was-trying-to-execute-some-error-cql-queries-but-unable-to-see-the-error-logs-in-audit-audit-table/1108 Title: Does ScyllaDB Enterprise support auditing of error audit logs or not? I was trying to execute some Error CQL queries but unable to see the Error logs in audit.audit_table - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: As you can see only success audit logs are here, but Error audit logs are not appearing. date | node | event_time | category | consistency | error | keyspace_nam… Language: en Canonical URL: https://forum.scylladb.com/t/does-scylladb-enterprise-support-auditing-of-error-audit-logs-or-not-i-was-trying-to-execute-some-error-cql-queries-but-unable-to-see-the-error-logs-in-audit-audit-table/1108 ## Headings Structure: H1: Does ScyllaDB Enterprise support auditing of error audit logs or not? I was trying to execute some Error CQL queries but unable to see the Error logs in audit.audit_table H3: Related topics ## Main Content: H1: Does ScyllaDB Enterprise support auditing of error audit logs or not? I was trying to execute some Error CQL queries but unable to see the Error logs in audit.audit_table H3: Related topics As you can see only success audit logs are here, but Error audit logs are not appearing. What do you mean by “Error audit logs”? Can you please give examples? Suppose I execute some error query like… Wanna create table which already exist so it should generate the audit logs with “ERROR” column “true” and also syntax error audit logs are also not generating. Does that means it did not support audit logs for error queries.?? I am getting only error audit logs for Authorization error type when an unauthorised user try to execute some query then it will through an error and it is capturing this error audit log only. No, it’s not supported for CQL queries. You can report an issue if you think it’s a valid use case you’re interested in such functionality and someone will examine this possibility. Yeah it is a major problem. It is a valid use case. I am using this as a security purpose in my project by sending this audit logs to system so that I can able to see what is going on. Suppose someone execute something which might affect our database but as I said it did not audit those error logs then it could be a problem, how can we know that what are the activity which was performed. How can I raise this issue so that officials could take a look into it? Failed requests are not executed in the database and therefore do not affect the database (other than the added load of processing them). How can I raise this issue so that officials could take a look into it? Normally, feature requests for enterprise-only features go through the ScyllaDB support channels, established during the sales process. I do not know how this works for a enterprise trial subscription, there is no established process for this. Normally, pre-deployment issues like this are worked out during the POC. You can contact sales to get a POC going. Ok thanks for this info. --- ### Page: https://forum.scylladb.com/t/efficient-use-of-cache-with-different-numbers-of-columns-on-select-queries/1109 Title: Efficient use of cache with different numbers of columns on Select queries - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I have some table: CREATE TABLE tests ( a timeuuid, b double, c text, d timestamp, other columns... PRIMARY KEY (a) ); and I need different columns for different cases. For example: Q1: select a, b from tests… Language: en Canonical URL: https://forum.scylladb.com/t/efficient-use-of-cache-with-different-numbers-of-columns-on-select-queries/1109 ## Headings Structure: H1: Efficient use of cache with different numbers of columns on Select queries H3: Related topics ## Main Content: H1: Efficient use of cache with different numbers of columns on Select queries H3: Related topics and I need different columns for different cases. For example: Q1: select a, b from tests where a = ? limit 1; Q2: select a, c from tests where a = ? limit 1; Q3: select c, b from tests where a = ? limit 1; Q4: select b from tests where a = ? limit 1; Im not using BYPASS CACHE. Different cases have different throughput. What will be more effective: make different statements for specific required fields or make one general statement, for example: select a, b, c from tests where a = ? limit1; Will it be more efficient to take data from the cache in one request? I want to understand whether different requests will be stored in separate caches or will they be grouped? Need I worry about it ?) There is a single unified cache per table, on each replica. Whether you have 1 prepared statements or 10 prepared statements reading the same partition, it makes no difference from the cache POV. BTW, you don’t need the limit 1, your partition will have exactly one row anyway, since you have no clustering columns. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-32-2023-12-16/1111 Title: Last week in scylla-cluster-tests.git master (issue #32; 2023-12-16) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This brief update highlights some intriguing commits to the scylla-cluster-tests.git master from the past week. The range of commits covered is 3d323707…d3c93ecd. During this period, there were 18 non-merge commits made… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-32-2023-12-16/1111 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #32; 2023-12-16) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #32; 2023-12-16) H3: Related topics This brief update highlights some intriguing commits to the scylla-cluster-tests.git master from the past week. The range of commits covered is 3d323707…d3c93ecd. During this period, there were 18 non-merge commits made by 5 authors. Here are some of the notable changes: A new ‘hydra’ version was released, featuring updated k8s components. In the OSS->Enterprise upgrade path, we now recover system tables from a snapshot while performing node rollback. This change was made to test the upgrade path and avoid the malformed_sstable_exception error. We also updated the Scylla-manager version in K8s tests to 3.2.5. The latest version now supports ARM instance types. Work has been outlined for restructuring the Jenkins jobs, which will allow us to have greater control and the ability to extend it when necessary. As COMPACT STORAGE is being deprecated, we have removed its setting in the code. However, we will continue to test it in functional tests. Lastly, we are gearing up for testing Scylla Kafka connectors on a large scale and have prepared a design for testing it. Stay tuned for the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-209-2023-12-17/1112 Title: Last week in scylladb.git master (issue #209; 2023-12-17) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the d62a5fc60b…10a11c2886 range are covered. There were 106 non-merge commits from 18 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-209-2023-12-17/1112 ## Headings Structure: H1: Last week in scylladb.git master (issue #209; 2023-12-17) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #209; 2023-12-17) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the d62a5fc60b…10a11c2886 range are covered. There were 106 non-merge commits from 18 authors in that period. Some notable commits: The auto-snapshot feature automatically creates a snapshot on TRUNCATE or DROP TABLE. It is now ignored on object storage tables, since they don’t implement snapshots yet. The gossip-based schema propagation code uses a hash of the schema metadata to check if nodes have the same schema, since changes to the schema can happen in different nodes independently. This schema calculation can be time consuming in a cluster with thousands of tables. In consistent schema mode, therefore, we no longer compute a hash of the entire schema but instead generate a schema version using a timeuuid. This speeds up operations on clusters with many tables. There is now an API for starting a compaction job without waiting for it. The native nodetool command (invoked using scylla nodetool) now implements the scrub command. The token_metadata class is an internal data structure containing state about nodes (principally tokens). It is now indexed by host IDs rather than node IP addresses, since Raft keeps track of IDs. This is mostly transparent to users but has implications on upgrades. The cqlsh configuration file is now in its correct place for the container image. The scylla sstable command now has a shard-of subcommand, to map an sstable to a shard. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/i-want-to-read-audit-log-through-syslog-of-scylla-db-enterprise-on-ubuntu-22-04-i-want-to-send-to-external-server-how-could-i-achieve-that-any-changes-is-needed-in-configuration-or-not/1113 Title: I want to read audit log through SYSLOG of Scylla DB enterprise on UBUNTU 22.04, I want to send to external server how could I achieve that, any changes is needed in Configuration or not? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Only this Particular link I found for syslog nothing else. ScyllaDB Auditing Guide | ScyllaDB Docs Language: en Canonical URL: https://forum.scylladb.com/t/i-want-to-read-audit-log-through-syslog-of-scylla-db-enterprise-on-ubuntu-22-04-i-want-to-send-to-external-server-how-could-i-achieve-that-any-changes-is-needed-in-configuration-or-not/1113 ## Headings Structure: H1: I want to read audit log through SYSLOG of Scylla DB enterprise on UBUNTU 22.04, I want to send to external server how could I achieve that, any changes is needed in Configuration or not? H3: Related topics ## Main Content: H1: I want to read audit log through SYSLOG of Scylla DB enterprise on UBUNTU 22.04, I want to send to external server how could I achieve that, any changes is needed in Configuration or not? H3: Related topics Only this Particular link I found for syslog nothing else. ScyllaDB Auditing Guide | ScyllaDB Docs This is the only needed configuration on Scylla side. You can define a third party tool like syslog-ng to stream these logs somewhere else. --- ### Page: https://forum.scylladb.com/t/smallest-sensible-memory-footprint-for-small-database/1116 Title: Smallest sensible memory footprint for small database? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I have a small system with a small number of users that is currently on Cassandra, I use it for the safety of the replication (2 instances on two cheap VPS’s). There is no more than 300Mb of data in the entire thing with… Language: en Canonical URL: https://forum.scylladb.com/t/smallest-sensible-memory-footprint-for-small-database/1116 ## Headings Structure: H1: Smallest sensible memory footprint for small database? H3: Related topics ## Main Content: H1: Smallest sensible memory footprint for small database? H3: Related topics I have a small system with a small number of users that is currently on Cassandra, I use it for the safety of the replication (2 instances on two cheap VPS’s). There is no more than 300Mb of data in the entire thing with maybe a max of 1 million records and maybe an average of 60 queries per minute only in peak times, and it’s not really going to grow much over time (I regularly clean out unnecessary data). Would it be ok, or inadvisable, to try and switch to Scylla DB with a 512Mb memory limit (on a 1Gb vps). I’ve been using Cassandra for a while now, and its in need of a second upgrade, which I never enjoy, so its the time to think about Scylla DB again. I see no reason, why your workload wouldn’t work with ScyllaDB. There is a good chance that all or most of your data would fit into cache and thus be served quickly. Question: why would you want to limit memory to 512MB on a 1GB VPS? Why not let ScylaDB use all the memory? Thanks, thats great! The number was just an estimate. For now I just wanted to check if this basic hypothetical setup was not crazy before I started investigating and testing our setup with Scylla DB in more detail. Perhaps we can allocate more. (We have a simple rust web app on each server, smtp, and imap as well. I just guesstimated what might work.) Ah I see. Note that ScyllaDB does not like neighbors much, by default it acts like it owns the machine it runs on. There are options to tune this. For starters, I recommend partitioning the CPU cores between ScyllaDB and the other applications, making sure ScyllaDB does not have to share CPU cores with other long running server apps. You can use the --cpuset command line option for ScyllaDB to restrict it to certain CPU cores. You can check the documentation of the other services you wish to run on how to achieve the same, or you can use systemd (I’m almost sure it has some options for this) or taskset to achieve this if they don’t have native support for this. Thankyou, thats helpful, I’ll make sure to look into those settings. Another note. As far as I know, the minimum amount of memory we test ScyllaDB with, is 256MB/CPU core. Below this amount of memory, you might hit unexpected problems. A more comfortable amount is 512MB/CPU core. Just something to keep in mind when planning resource distribution. Don’t forget to configure swap on you machines, otherwise the kernel will kill ScyllaDB as the largest memory consumer, if other services accidentally consume all remaining memory. --- ### Page: https://forum.scylladb.com/t/memory-issue-for-scylla-5-2-4-0-20230623-cebbf6c5df2b/1119 Title: Memory issue for scylla-5.2.4-0.20230623.cebbf6c5df2b - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi Team, We are getting continuous errors for the memory for scylla 5.2.4. Please find the logs below Dec 19 14:05:35 ip-10-8-40-9 scylla: [shard 0] storage_proxy - Exception when communicating with 10.8.40.9, to read… Language: en Canonical URL: https://forum.scylladb.com/t/memory-issue-for-scylla-5-2-4-0-20230623-cebbf6c5df2b/1119 ## Headings Structure: H1: Memory issue for scylla-5.2.4-0.20230623.cebbf6c5df2b H3: Related topics ## Main Content: H1: Memory issue for scylla-5.2.4-0.20230623.cebbf6c5df2b H3: Related topics We are getting continuous errors for the memory for scylla 5.2.4. Please find the logs below Dec 19 14:05:35 ip-10-8-40-9 scylla: [shard 0] storage_proxy - Exception when communicating with 10.8.40.9, to read from system_distributed.service_levels: std::bad_alloc (std::bad_alloc) Dec 19 14:05:39 ip-10-8-40-9 scylla: [shard 0] storage_proxy - Exception when communicating with 10.8.40.9, to read from system_auth.roles: std::bad_alloc (std::bad_alloc) Dec 19 14:05:39 ip-10-8-40-9 scylla: [shard 0] storage_proxy - Exception when communicating with 10.8.40.129, to read from system_auth.roles: std::bad_alloc Dec 19 14:05:45 ip-10-8-40-9 scylla: [shard 0] storage_proxy - Exception when communicating with 10.8.40.9, to read from system_distributed.service_levels: std::bad_alloc (std::bad_alloc) Dec 19 14:05:51 ip-10-8-40-9 scylla: [shard 0] storage_proxy - Exception when communicating with 10.8.40.9, to read from system_auth.roles: std::bad_alloc (std::bad_alloc) Dec 19 14:05:51 ip-10-8-40-9 scylla: [shard 0] storage_proxy - Exception when communicating with 10.8.40.129, to read from system_auth.roles: std::bad_alloc Dec 19 14:05:52 ip-10-8-40-9 scylla: [shard 0] storage_proxy - Exception when communicating with 10.8.40.9, to read from system.batchlog: std::bad_alloc (std::bad_alloc) Dec 19 14:05:52 ip-10-8-40-9 scylla: [shard 0] batchlog_manager - Exception in batch replay: exceptions::read_failure_exception (Operation failed for system.batchlog - received 0 responses and 1 failures from 1 CL=ONE.) Dec 19 14:05:55 ip-10-8-40-9 scylla: [shard 0] storage_proxy - Exception when communicating with 10.8.40.9, to read from system_distributed.service_levels: std::bad_alloc (std::bad_alloc) Dec 19 14:06:00 ip-10-8-40-9 scylla: [shard 0] storage_proxy - Exception when communicating with 10.8.40.9, to read from system_auth.roles: std::bad_alloc (std::bad_alloc) Dec 19 14:06:00 ip-10-8-40-9 scylla: [shard 0] storage_proxy - Exception when communicating with 10.8.40.129, to read from system_auth.roles: std::bad_alloc Dec 19 14:06:00 ip-10-8-40-9 scylla: [shard 0] storage_proxy - Exception when communicating with 10.8.40.9, to read from system_auth.roles: std::bad_alloc (std::bad_alloc) Dec 19 14:06:00 ip-10-8-40-9 scylla: [shard 0] storage_proxy - Exception when communicating with 10.8.40.129, to read from system_auth.roles: std::bad_alloc Dec 19 14:06:05 ip-10-8-40-9 scylla: [shard 0] storage_proxy - Exception when communicating with 10.8.40.9, to read from system_distributed.service_levels: std::bad_alloc (std::bad_alloc) Dec 19 14:06:06 ip-10-8-40-9 scylla: [shard 0] storage_proxy - Exception when communicating with 10.8.40.9, to read from system_auth.roles: std::bad_alloc (std::bad_alloc) Dec 19 14:06:06 ip-10-8-40-9 scylla: [shard 0] storage_proxy - Exception when communicating with 10.8.40.129, to read from system_auth.roles: std::bad_alloc Dec 19 14:06:15 ip-10-8-40-9 scylla: [shard 0] storage_proxy - Exception when communicating with 10.8.40.9, to read from system_distributed.service_levels: std::bad_alloc (std::bad_alloc) Dec 19 14:06:21 ip-10-8-40-9 scylla: [shard 0] storage_proxy - Exception when communicating with 10.8.40.9, to read from system_auth.roles: std::bad_alloc (std::bad_alloc) Dec 19 14:06:21 ip-10-8-40-9 scylla: [shard 0] storage_proxy - Exception when communicating with 10.8.40.129, to read from system_auth.roles: std::bad_alloc Dec 19 14:06:24 ip-10-8-40-9 scylla: [shard 0] storage_proxy - Exception when communicating with 10.8.40.9, to read from system_auth.roles: std::bad_alloc (std::bad_alloc) Dec 19 14:06:24 ip-10-8-40-9 scylla: [shard 0] storage_proxy - Exception when communicating with 10.8.40.129, to read from system_auth.roles: std::bad_alloc Dec 19 14:06:25 ip-10-8-40-9 scylla: [shard 0] storage_proxy - Exception when communicating with 10.8.40.9, to read from system_distributed.service_levels: std::bad_alloc (std::bad_alloc) Please help us in rectifying the problem. Regards, Soyal Badkur We would require more details. A backtrace would be helpful. Does the system crash? If so please send the coredump. Are you using large collections? Please see the documentation for reporting a problem. You can open a support ticket if you are a customer and a GitHub issue otherwise. --- ### Page: https://forum.scylladb.com/t/release-scylladb-rust-driver-0-11-0/1121 Title: [RELEASE] ScyllaDB Rust Driver 0.11.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce ScyllaDB Rust Driver 0.11.0, an asynchronous CQL driver for Rust, optimized for Scylla, but also compatible with Apache Cassandra! Some interesting statistics: over 1,018k dow… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-rust-driver-0-11-0/1121 ## Headings Structure: H1: [RELEASE] ScyllaDB Rust Driver 0.11.0 H2: Notable changes H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Rust Driver 0.11.0 H2: Notable changes H3: Related topics The ScyllaDB team is pleased to announce ScyllaDB Rust Driver 0.11.0, an asynchronous CQL driver for Rust, optimized for Scylla, but also compatible with Apache Cassandra! Some interesting statistics: The main feature of this PR is refactor of serialization API. Old traits and structs (Value, ValueList, BatchValues - renamed to LegacyBatchValues, SerializedValues - renamed to LegacySerializedValues) are replaced by new ones (SerializeCql, SerializeRow, new BatchValues, new SerializedValues). There are wrappers and helper implementations provided, designed to aid in gradually migrating to new API - see the migration guide in the book for more information. Old traits and structs will be removed in one of future versions. New serialization API has a benefit of type safety - now, if you use wrong type for a bind marker in a query, you will get an understandable error locally (without actually executing the query on the database) instead of cryptic deserialization error from Scylla, or worse - silent data corruption. API cleanups / breaking changes: New features / enhancements: CI / developer tool improvements: Congrats to all contributors and thanks everyone for using our driver! The source code of the driver can be found here: The official crates.io registry entry is here: Thank you for your attention, please do not hesitate to contact us if you have any questions, issues, feature requests, or are simply interested in our driver! Contributors since the last release: @Lorak FYI, I moved the post of “Release Notes” Category --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-1-13/1123 Title: [RELEASE] ScyllaDB Enterprise 2022.1.13 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2022.1.13, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.1. Note that there is a newer Enterprise LTS release: ScyllaDB Enterpr… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-1-13/1123 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2022.1.13 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2022.1.13 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2022.1.13, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.1. Note that there is a newer Enterprise LTS release: ScyllaDB Enterprise 2023.1. While we will continue to support 2022.1 LTS, you can get additional features with 2023.1. Below is a list of performance and stability improvements and bug fixes, each with an open-source reference if available: --- ### Page: https://forum.scylladb.com/t/what-is-sql-server/1125 Title: What is SQL Server - Database Community - ScyllaDB Community NoSQL Forum Meta Description: SQL Server is a relational database management system (RDBMS) developed by Microsoft. It stores and retrieves data as requested by other software applications, supporting structured query language (SQL) for managing and … Language: en Canonical URL: https://forum.scylladb.com/t/what-is-sql-server/1125 ## Headings Structure: H1: What is SQL Server H3: Related topics ## Main Content: H1: What is SQL Server H3: Related topics SQL Server is a relational database management system (RDBMS) developed by Microsoft. It stores and retrieves data as requested by other software applications, supporting structured query language (SQL) for managing and manipulating data. It offers features for data storage, retrieval, and analysis, making it a powerful solution for managing databases. @Anurag_Sharma please do not spam the forum with non related topics. Keep it ScyllaDB related. How do I migrate data from SQL Server to ScyllaDB? You’ll have to migrate your data model, the data itself, and make changes to the application (the business function). There are a few ways to do this, offline migration (also called Cold Migration) or zero downtime migration (also called Hot Migration or online migration). You can learn more in the Migrating to ScyllaDB lesson on ScyllaDB University. If you have specific questions feel free to ask here. --- ### Page: https://forum.scylladb.com/t/license-of-scylladb-oss-installation/1128 Title: License of ScyllaDB OSS installation - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, Where can I find for ScyllaDB OSS license? I see this license at scylladb/LICENSE.AGPL at master · scylladb/scylladb · GitHub. Is this license just for Git repo? Or also a software license for installation and use i… Language: en Canonical URL: https://forum.scylladb.com/t/license-of-scylladb-oss-installation/1128 ## Headings Structure: H1: License of ScyllaDB OSS installation H3: Related topics ## Main Content: H1: License of ScyllaDB OSS installation H3: Related topics Where can I find for ScyllaDB OSS license? I see this license at scylladb/LICENSE.AGPL at master · scylladb/scylladb · GitHub. Is this license just for Git repo? Or also a software license for installation and use in a machine. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-210-2023-12-24/1131 Title: Last week in scylladb.git master (issue #210; 2023-12-24) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 10a11c2886…2590274f95 range are covered. There were 86 non-merge commits from 16 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-210-2023-12-24/1131 ## Headings Structure: H1: Last week in scylladb.git master (issue #210; 2023-12-24) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #210; 2023-12-24) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 10a11c2886…2590274f95 range are covered. There were 86 non-merge commits from 16 authors in that period. Some notable commits: Tablets is a new, experimental method of managing replication and distribution. With tablets, each token range has its own sstables and memtables. There is now support for splitting such ranges when a table grows. This support is not yet wired into the tablet load balancer. Alternator, ScyllaDB’s implementation of the DynamoDB API, now supports tablets. Secondary indexes now support tables with tablets. Change data capture (CDC) is now rejected on tables using tablets, as CDC was not yet integrated with tablets. The ALTER KEYSPACE statement now rejects changing a keyspace from vnodes to tablets or back. Using tablets must be decided during keyspace creation. Schema management using Raft is now mandatory. Clusters will now switch to Raft on upgrade. Until now, only new clusters were created with Raft while upgrades had to opt in. The nodetool command will now invoke the native nodetool implementation for most commands, falling back to the Java-based nodetool and scylla-jmx for a few unimplemented commands. A crash on a table drop that is concurrent with streaming has been fixed. ScyllaDB will now listen on a maintenance socket in addition to the normal CQL port. The maintenance socket is active even when the node is not yet fully joined, allowing troubleshooting access. The maintenance socket uses unix-domain sockets, not TCP. Support for cqlsh is not yet enabled. A Seastar update improves performance in debug mode. A few Rust dependencies were updated to address minor security vulnerabilities. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2023-1-3/1132 Title: [RELEASE] ScyllaDB Enterprise 2023.1.3 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2023.1.3, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2023.1 LTS Release. 2023.1.3 patch release includes multiple minor bug fixes… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2023-1-3/1132 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2023.1.3 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2023.1.3 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2023.1.3, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2023.1 LTS Release. 2023.1.3 patch release includes multiple minor bug fixes. You are encouraged to upgrade to it in coordination with the ScyllaDB Support team. The following issues are fixed in this release (with an open-source reference, if available): --- ### Page: https://forum.scylladb.com/t/scylla-manager-restore-happens-very-slowly/1135 Title: Scylla Manager -- Restore happens very slowly - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello, I was testing a restore of my production data backup stored on a s3 bucket where data size is 170GB with RF=3 onto a separate cluster. So, net data would be 510GB. To my surprise, the restore is taking really a l… Language: en Canonical URL: https://forum.scylladb.com/t/scylla-manager-restore-happens-very-slowly/1135 ## Headings Structure: H1: Scylla Manager -- Restore happens very slowly H3: Related topics ## Main Content: H1: Scylla Manager -- Restore happens very slowly H3: Related topics Hello, I was testing a restore of my production data backup stored on a s3 bucket where data size is 170GB with RF=3 onto a separate cluster. So, net data would be 510GB. To my surprise, the restore is taking really a lot of time, around 9-10 hrs. S3 VPC endpoint is enabled and data download speeds are very quick when done manually. Not sure if this is the expected time to restore, please let me know if I might be doing something wrong here or is it an expected time period ? If yes, how can we maintain a good RTO ? Hi, I suggest opening a new issue for the scylla-manager, including the manager version you’re using and the logs. It’s indeed sound too long for the restore. Worth mention the size of original cluster and the new cluster you’re trying to restore to. --- ### Page: https://forum.scylladb.com/t/error-starting-up-scylla-in-github-actions/1136 Title: Error starting up Scylla in Github Actions - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I get an error while starting Scylla with github actions, it worked for months and suddenly not anymore, any ideas? Scylla 0.11.1 · Jasperav/Catalytic@e9a6f39 · GitHub Language: en Canonical URL: https://forum.scylladb.com/t/error-starting-up-scylla-in-github-actions/1136 ## Headings Structure: H1: Error starting up Scylla in Github Actions H3: Related topics ## Main Content: H1: Error starting up Scylla in Github Actions H3: Related topics I get an error while starting Scylla with github actions, it worked for months and suddenly not anymore, any ideas? Scylla 0.11.1 · Jasperav/Catalytic@e9a6f39 · GitHub In scylla rust driver CI we use slightly different healt check: --health-cmd "cqlsh --debug scylladb" --health-interval 5s --health-retries 10. Notice that we explicitly pass scylladb to cqlsh. I see that it was changed in this commit: CI: use container hostname in healthcheck · scylladb/scylla-rust-driver@3be1149 · GitHub in response to scylla issue: Docker: can not connect to Scylla 5.4 with CQLSh · Issue #16329 · scylladb/scylladb · GitHub From the issue it seems that this should be fixed in 5.4.1, so try upgrading to this if you don’t want to change healthcheck command. --- ### Page: https://forum.scylladb.com/t/can-the-scylla-operator-be-installed-namespace-scoped/1137 Title: Can the scylla operator be installed namespace scoped? - Kubernetes Operator - ScyllaDB Community NoSQL Forum Meta Description: Hi, I have a requirement where in a single kubernetes cluster multiple teams can have separate scylla operator deployments and their own management. I am trying to make the operator namespace scoped but not finding any s… Language: en Canonical URL: https://forum.scylladb.com/t/can-the-scylla-operator-be-installed-namespace-scoped/1137 ## Headings Structure: H1: Can the scylla operator be installed namespace scoped? H3: Related topics ## Main Content: H1: Can the scylla operator be installed namespace scoped? H3: Related topics Hi, I have a requirement where in a single kubernetes cluster multiple teams can have separate scylla operator deployments and their own management. I am trying to make the operator namespace scoped but not finding any such namespaces to watch switch in the crd. Any inputs related to the same would be helpful. it is not possible to install the operator as a regular user. CRDs are not namespaced but cluster scoped and managed by cluster administrators. Hence the operator (or any operator based on CRDs) is not a regular app but more of a cluster extension. Plus the CRD needs to be in sync with the operator version (±1) so you can’t really handle the CRD and the operator deployment independently. That said the operator is multitenant, so your teams can create ScyllaClusters independently, just as they do say for Deployments or StatefulSets. The k8s cluster is managed by our own so we have the ability to deploy CRDs. But we still want the operator scope to be limited within namespaces of each team. I agree that a single operator can manage all the scylla clusters independently but the requirement is to let the teams deploy their own operator + scylla cluster in their own namespace with one time admin role for the CRD installations for each team. CRDs have the option to be namespace scoped according to CRD Scope | Operator SDK. This “scope” field is also present in all the scylla CRDs: NodeConfig, ScyllaCluster, ScyllaOperatorConfig, ScyllaDBMonitoring with value as Cluster. Can’t this be changed to Namespaced? CRDs have the option to be namespace scoped according CustomResourceDefinitions are always cluster scoped. This “scope” field is also present in all the scylla CRDs: NodeConfig, ScyllaCluster, ScyllaOperatorConfig, ScyllaDBMonitoring with value as Cluster. Custom Resources can be both, but that’s not relevant to my point about CRDs. But again suppose we have the rights to deploy the CRDs with admin permission since that’s a one time activity. Then we want the scylla CRs to be namespace scoped and only require roles instead of clusterroles. Is that possible? since that’s a one time activity not at all, the CRD has to be kept in sync with the operator binary/version and updated. It is also why the idea of namespaced scoped operator with cluster scoped CustomResourceDefinition lacks ground. Then we want the scylla CRs to be namespace scoped the ones related to the operator are rightfully cluster scoped as there can be only one operator. the ones that users use, like ScyllaCluster, are namespace scoped. the CRD has to be kept in sync with the operator binary/version and updated That’s only if and when the team wants to update the operator, and if so that will be very less often and can be handled with the help of the cluster admins. So again the question remains what stops from giving an option to deploy the operator as namespace/list of namespaces scoped and not only one in the whole cluster? --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-2-16/1138 Title: [RELEASE] ScyllaDB Enterprise 2022.2.16 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2022.2.16, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.2 Feature Release. Note the latest ScyllaDB Enterprise release is 2023… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-2-16/1138 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2022.2.16 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2022.2.16 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2022.2.16, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.2 Feature Release. Note the latest ScyllaDB Enterprise release is 2023.1 LTS, and you are encouraged to upgrade to it in coordination with the ScyllaDB Support team. The following issues are fixed in this release (with an open-source reference, if available): --- ### Page: https://forum.scylladb.com/t/release-scylla-operator-1-11-1/1140 Title: [RELEASE] Scylla Operator 1.11.1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of Scylla Operator 1.11.1 :rocket: Scylla Operator 1.11.1 brings a couple of bug fixes. As with all of our releases, all API changes are backward compatible. Notabl… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-operator-1-11-1/1140 ## Headings Structure: H1: [RELEASE] Scylla Operator 1.11.1 H3: Related topics ## Main Content: H1: [RELEASE] Scylla Operator 1.11.1 H3: Related topics The ScyllaDB team is pleased to announce the release of Scylla Operator 1.11.1 Scylla Operator 1.11.1 brings a couple of bug fixes. As with all of our releases, all API changes are backward compatible. Upgrade instructions Upgrading from v1.10.0 or v1.11.0 with kubectl apply doesn’t require any extra action, just take the manifest from v1.11.1 tag and substitute the released image. Using helm requires a mandatory manual step for every release because helm can’t handle CRDs updates. For details, see our upgrade documentation. Regards, Scylla Operator Team --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-211-2023-12-31/1142 Title: Last week in scylladb.git master (issue #211; 2023-12-31) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 2590274f95…f1dea4bc8a range are covered. There were 46 non-merge commits from 10 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-211-2023-12-31/1142 ## Headings Structure: H1: Last week in scylladb.git master (issue #211; 2023-12-31) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #211; 2023-12-31) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 2590274f95…f1dea4bc8a range are covered. There were 46 non-merge commits from 10 authors in that period. Some notable commits: The initial_tablets keyspace option, used with the experimental tablets distribution algorithm, is now stored in system_schema.scylla_keyspaces to improve flexibility and compatility. Support for RHEL 7 and its variants (e.g. CentOS 7) has been removed. RHEL 7 will reach end-of-life in June 2024. A regression in SELECT * GROUP BY has been fixed. The regression involved interaction between wildcard selection and GROUP BY. We now allow EXECUTE permissions on native functions. A previous tightening of permission enforcement caused a regression when executing native functions. On Ubuntu, the installer now handles conflicts between a system process updating apt metadata and the installer itself. The fencing mechanism prevents a read or write from accessing an outdated replica. The mechanism is now disabled on local tables, as these can never be out of date. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/unit-pricing-breakdown/1144 Title: Unit Pricing Breakdown - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi — Has anyone computed unit economics per data stored (e.g., per GB)? Does Scylla become more cheap per unit cost with the more data stored (e.g., $2/GB at 100GB becomes $1/GB at 200GB)? And/or, is there a breakdown … Language: en Canonical URL: https://forum.scylladb.com/t/unit-pricing-breakdown/1144 ## Headings Structure: H1: Unit Pricing Breakdown H3: Related topics ## Main Content: H1: Unit Pricing Breakdown H3: Related topics Hi — Has anyone computed unit economics per data stored (e.g., per GB)? Does Scylla become more cheap per unit cost with the more data stored (e.g., $2/GB at 100GB becomes $1/GB at 200GB)? And/or, is there a breakdown of Tiered pricing, if one exists? Hi @kai See Pricing information for ScyllaDB for picing data and pricing questions. --- ### Page: https://forum.scylladb.com/t/unable-to-connect-to-scylladb-5-4-in-self-managed-aws-instance/1147 Title: Unable to connect to ScyllaDB 5.4 in self-managed AWS instance - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I am trying to setup ScyllaDB 5.4 on my own AWS account but when I try to connect through SSH, nothing happens. It timeouts after awhile. I followed the guide here: Launch ScyllaDB on AWS | ScyllaDB Docs. These are my e… Language: en Canonical URL: https://forum.scylladb.com/t/unable-to-connect-to-scylladb-5-4-in-self-managed-aws-instance/1147 ## Headings Structure: H1: Unable to connect to ScyllaDB 5.4 in self-managed AWS instance H3: Related topics ## Main Content: H1: Unable to connect to ScyllaDB 5.4 in self-managed AWS instance H3: Related topics I am trying to setup ScyllaDB 5.4 on my own AWS account but when I try to connect through SSH, nothing happens. It timeouts after awhile. I followed the guide here: Launch ScyllaDB on AWS | ScyllaDB Docs. These are my exact steps, note that I can connect with the same security groups and subnets when I install a plain AWS AMI instance with a ec2-user as username, so that is configured well: { “scylla_yaml”: { “cluster_name”: “test-cluster”, “seed_provider”: [{“class_name”: “org.apache.cassandra.locator.SimpleSeedProvider”}], }, “post_configuration_script”: “#! /bin/bash\nyum install cloud-init-cfn”, “start_scylla_on_first_boot”: true } Now I wait until I see 2/2 checks are completed in AWS EC2 instance overview. After that, I tried connecting to both the public IP4 address and Public IPv4 DNS with username scyllaadm, root and ec2-user. They all freeze. I try connecting like this: ssh -i KPNew.pem scyllaadm@34.243.209.220. Am I missing something? This is the system log: gist:c3c31fdbb32f83c1127070a360ac65c9 · GitHub Can you provide the region/ami id you are trying to use? Did you have the same connection issue with and without your user-data? Hi Jasper, Please provide output of “ssh -vvv -i …” I am facing the same issue. Similar logs as attached in the original post. Furthermore, I can’t connect with EC2 serial console either, stuck at a Login: prompt where I don’t know the username and password despite trying scyllaadm and many common passwords AMI: ami-0d313f4fc54cf77a2 Region: us-west-2 @bneigher one more question, what is the type of the key that was used ? RSA ? can you try the same with ed25519 key ? hi @Jasper_Visser can you help me to connect scylladb using aws ec2 instance with 3 node --- ### Page: https://forum.scylladb.com/t/according-to-scylladb-document-for-syslog-it-mentioned-that-we-can-generate-logs-in-centos-but-i-am-able-to-generate-it-in-ubuntu-also-may-be-document-need-to-be-updated/1148 Title: According to ScyllaDb document, For Syslog it mentioned that we can generate logs in CentOS, but I am able to generate it in UBUNTU also. May be Document need to be updated - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Syslog reference Link: ScyllaDB Auditing Guide | ScyllaDB Docs Syslog Configuration I did in rsyslog.conf file in SCYLLADB Please verify also: # /etc/rsyslog.conf configuration file for rsyslog # # For more informatio… Language: en Canonical URL: https://forum.scylladb.com/t/according-to-scylladb-document-for-syslog-it-mentioned-that-we-can-generate-logs-in-centos-but-i-am-able-to-generate-it-in-ubuntu-also-may-be-document-need-to-be-updated/1148 ## Headings Structure: H1: According to ScyllaDb document, For Syslog it mentioned that we can generate logs in CentOS, but I am able to generate it in UBUNTU also. May be Document need to be updated H3: How to setup Rsyslog server on Ubuntu 22.04 H3: Related topics ## Main Content: H1: According to ScyllaDb document, For Syslog it mentioned that we can generate logs in CentOS, but I am able to generate it in UBUNTU also. May be Document need to be updated H3: How to setup Rsyslog server on Ubuntu 22.04 H3: Related topics Syslog reference Link: ScyllaDB Auditing Guide | ScyllaDB Docs Syslog Configuration I did in rsyslog.conf file in SCYLLADB Please verify also: I followed this below link for configuring SYSLOG. How to configure rsyslog server on ubuntu 20. Steps to follow. 1- check the availability of rsyslog server. Step 2- open the conf file. Est. reading time: 4 minutes please if anyone have official link for syslog configuration in SCYLLDB then please mention here also. Below is the image of the syslog which I able to generate in UBUNTU. But In documentation it was mentioned we can only do it in CentOS. --- ### Page: https://forum.scylladb.com/t/does-scylladb-support-api-or-not-and-can-we-capture-audit-logs-for-rest-api-command-also-or-not-please-give-some-reference-link-regarding-this/1149 Title: Does SCYLLADB Support API or not and can we capture audit logs for REST API command also or not? Please give Some reference link regarding this - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: It is possible to Generate audit logs for REST API in SCYLLADB? Please mention reference link for The REST API and how many type of REST API it will support with Example how we can execute REST API in SCYLLADB? Language: en Canonical URL: https://forum.scylladb.com/t/does-scylladb-support-api-or-not-and-can-we-capture-audit-logs-for-rest-api-command-also-or-not-please-give-some-reference-link-regarding-this/1149 ## Headings Structure: H1: Does SCYLLADB Support API or not and can we capture audit logs for REST API command also or not? Please give Some reference link regarding this H3: Related topics ## Main Content: H1: Does SCYLLADB Support API or not and can we capture audit logs for REST API command also or not? Please give Some reference link regarding this H3: Related topics It is possible to Generate audit logs for REST API in SCYLLADB? Please mention reference link for The REST API and how many type of REST API it will support with Example how we can execute REST API in SCYLLADB? Hello @subrato, ScyllaDB does have a REST API interface, for admin/support access. By default it is served on port 10000 on each node and typically it is not exposed publicly, so that one needs to first SSH into the node to access the API, so sshd audit logs can be considered instead for security purposes. As for the REST API calls themselves, many of them print a log message to the system log, but not all of them. Actually I am using SCYLLADB Enterprise trail In Amazon EC2 UBUNTU 22.04 and I want to know is it possible to Run REST API and generate audit logs in SCYLLA-AUDIT file where other DDL,DML,DCL,AUTH,ADMIN query audit log stored. Could you please prove some reference url for this that’s would be helpful. It would be really helpful for me if you can provide these details. --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-2-12/1152 Title: [RELEASE] ScyllaDB 5.2.12 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.2.12, a bugfix release of the ScyllaDB 5.2 stable branch. ScyllaDB Open Source 5.2.12, like all past and future 5.x.y releases, is backward compatible, and supports roll… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-2-12/1152 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.2.12 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.2.12 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.2.12, a bugfix release of the ScyllaDB 5.2 stable branch. ScyllaDB Open Source 5.2.12, like all past and future 5.x.y releases, is backward compatible, and supports rolling upgrades. Note that the latest ScyllaDB Open Source stable release is 5.4 and you are encouraged to upgrade to it. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/doest-restapi-and-cli-audit-logs-support-for-scylladb-enterprise-trial/1154 Title: Doest RESTAPI and CLI audit logs support for SCYLLADB Enterprise Trial - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Please give me reference link. Language: en Canonical URL: https://forum.scylladb.com/t/doest-restapi-and-cli-audit-logs-support-for-scylladb-enterprise-trial/1154 ## Headings Structure: H1: Doest RESTAPI and CLI audit logs support for SCYLLADB Enterprise Trial H3: Related topics ## Main Content: H1: Doest RESTAPI and CLI audit logs support for SCYLLADB Enterprise Trial H3: Related topics Please give me reference link. Subrato, you already asked this on a different topic; please don’t start multiple threads about the same question, as it’s cluttering the forum. --- ### Page: https://forum.scylladb.com/t/reuse-pagingstate-in-another-execution-in-scylla-with-java/1156 Title: Reuse PagingState in another execution in Scylla with Java - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Is there possible to store and transform into a String, a base64 or anything that I can return in JSON, to populate the object in another execution. I found this example, but I’m not getting how to convert and decovert … Language: en Canonical URL: https://forum.scylladb.com/t/reuse-pagingstate-in-another-execution-in-scylla-with-java/1156 ## Headings Structure: H1: Reuse PagingState in another execution in Scylla with Java H3: Related topics ## Main Content: H1: Reuse PagingState in another execution in Scylla with Java H3: Related topics Is there possible to store and transform into a String, a base64 or anything that I can return in JSON, to populate the object in another execution. I found this example, but I’m not getting how to convert and decovert the ByteBuffer. I have clients with milions of registers, which I need to paginate my queries, what I need is return in json an indicator that must have more registries, and the next request, the client send me this indicator, if there anyway to use number page, for example the client inform me the limit of registries in 1000 and want the 5º page I think it will be the best for the user experience. If there is another way to paginate the ScyllaDB Query in different executions, I appreciate if somebody can share. Unfortunately, I don’t think this will work the way you want it to. The paging state in general is an opaque blob from the driver’s perspective. The information it contains is not meant to be interpreted by drivers and clients. It also cannot be used to, for example, fetch the 5th page in an array of pages, each with 100 rows. To implement paging the way you describe, you can use LIMIT N in combination of discarding elements you don’t need from the result-set. --- ### Page: https://forum.scylladb.com/t/seeing-too-many-open-files-error-in-4-6/1157 Title: Seeing "too many open files" error in 4.6 - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: hello, I think I might be seeing this bug in my test bed where I am using 4.6.5 release and the open file descriptors just keeps on increasing for scylla read/write. I am using STCS and I see a lot of small sstables cre… Language: en Canonical URL: https://forum.scylladb.com/t/seeing-too-many-open-files-error-in-4-6/1157 ## Headings Structure: H1: Seeing "too many open files" error in 4.6 H3: Related topics ## Main Content: H1: Seeing "too many open files" error in 4.6 H3: Related topics hello, I think I might be seeing this bug in my test bed where I am using 4.6.5 release and the open file descriptors just keeps on increasing for scylla read/write. I am using STCS and I see a lot of small sstables created. Over the period of 1 week, the open file descriptors just keeps on increasing exponentially. So, much so that I get “Too many open files” error. And eventuall scylla just stops taking any client request. Is anyone encountering this 4.6 release? or do you recommend to upgrade to 5 release? I don’t see a point upgrading since this bug attached is still open. Any suggestions? Hello Ken, re #8170, how is resharding involved in your case? Did you configure STCS with non-default values? What is min_threshold set to? Also, how many files are we talking about? It might be possible that you just need to raise the limit to accommodate for scylla’s needs. The reason is that for performance-related reasons, scylla keeps two open files per SSTable for their data and index components. Hello @bhalevy I am using default values with STCS. Didn’t change min_threshold value or set inside scylla.yaml. I increased ulimit to 1.6 million but after few days I see scylla open file reaches that limit. I’m on 4.6.11. I don’t understand this behavior For the last couple of days, I have stopped running repairs and the situation seems to be under control. I don’t see scylla’s growing number of open file descriptors. So, looks like that running nodetool repair if run manually causes the situation to get worst. Do we need to run nodetool repair? Or does the system take care of keeping the data consistent across replicas? --- ### Page: https://forum.scylladb.com/t/how-to-quickly-release-the-capacity-of-scylladb/1158 Title: How to quickly release the capacity of scylladb? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I created a table through the Alterator interface and filled it with a large amount of data. Now I need to clean up the database space, so I deleted this table through AWS. aws dynamodb delete-table --table-name ycsb T… Language: en Canonical URL: https://forum.scylladb.com/t/how-to-quickly-release-the-capacity-of-scylladb/1158 ## Headings Structure: H1: How to quickly release the capacity of scylladb? H3: Related topics ## Main Content: H1: How to quickly release the capacity of scylladb? H3: Related topics I created a table through the Alterator interface and filled it with a large amount of data. Now I need to clean up the database space, so I deleted this table through AWS. Through nodetool status, it can be seen that the cluster’s capacity has decreased, but the actual disk usage space has not decreased. I found that its data directory is still there. I am also unable to perform compact on this table because it no longer exists. So how to quickly release cluster space? Do we have to wait for the tombstone time to expire and automatically clean it up? This has been too long. Thanks! By default, scylla takes a snapshot of a table right before it’s deleted, in case the data needs to be recovered. To reclaim the storage space, the snapshot needs to be removed. --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-4-1/1160 Title: [RELEASE] ScyllaDB 5.4.1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.4.1, a bugfix release of the ScyllaDB 5.4 stable branch. ScyllaDB Open Source 5.4.1, like all past and future 5.x.y releases, is backward compatible and supports rolling… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-4-1/1160 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.4.1 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.4.1 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.4.1, a bugfix release of the ScyllaDB 5.4 stable branch. ScyllaDB Open Source 5.4.1, like all past and future 5.x.y releases, is backward compatible and supports rolling upgrades. Users are encouraged to upgrade to 5.4.1. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/transactional-outbox-or-alternative-pattern-with-scylladb/1162 Title: Transactional outbox or alternative pattern with scylladb - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: How to implement a transactional outbox pattern or different pattern with scylladb to write the event to scylladb first and then process it in micro batches to mitigate the dual write problem and garuntee event publishin… Language: en Canonical URL: https://forum.scylladb.com/t/transactional-outbox-or-alternative-pattern-with-scylladb/1162 ## Headings Structure: H1: Transactional outbox or alternative pattern with scylladb H3: Related topics ## Main Content: H1: Transactional outbox or alternative pattern with scylladb H3: Related topics How to implement a transactional outbox pattern or different pattern with scylladb to write the event to scylladb first and then process it in micro batches to mitigate the dual write problem and garuntee event publishing AFAIU the Transactional Outbox pattern in microservices, one can use ScyllaDB CDC to track updates to a table (e.g., service 2 in Transactional outbox). Note that ScyllaDB does not support multi-table transactions, so there is no way to insert data to two tables (Order table and Outbox table) in one transaction. --- ### Page: https://forum.scylladb.com/t/last-3-weeks-in-scylla-cluster-tests-git-master-issue-33-2024-01-05/1164 Title: Last 3 weeks in scylla-cluster-tests.git master (issue #33; 2024-01-05) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This brief report highlights some notable commits to scylla-cluster-tests.git master from the past three weeks, specifically within the 9016c20b…f5bc3186 range. During this period, we had 89 non-merge commits from 13 au… Language: en Canonical URL: https://forum.scylladb.com/t/last-3-weeks-in-scylla-cluster-tests-git-master-issue-33-2024-01-05/1164 ## Headings Structure: H1: Last 3 weeks in scylla-cluster-tests.git master (issue #33; 2024-01-05) H3: Related topics ## Main Content: H1: Last 3 weeks in scylla-cluster-tests.git master (issue #33; 2024-01-05) H3: Related topics This brief report highlights some notable commits to scylla-cluster-tests.git master from the past three weeks, specifically within the 9016c20b…f5bc3186 range. During this period, we had 89 non-merge commits from 13 authors. Here are some key updates: The Operator’s NodeConfig CRD is now reused for static volume provisioner, enabling the configuration of SSD disks into RAID arrays. This change allows us to simplify the SCT code and eliminate redundant pods in EKS and GKE. We have started testing Scylla ARM docker images with functional tests and one longevity in EKS environment. The EKS module can now detect ARM instance types and select the appropriate image for VM. Developers can add a config file that enables tablets in Jenkins jobs, making it easier to enable this feature in any SCT test/longevity. The Scylla-operator upgrade test has been enhanced with new checks, including verifying new ‘UID’ values, ‘status.conditions’, and scylla-operator images. The monitoring branch was updated to 4.6 by default. We can now run the must-gather binary on any possible arch on K8S, allowing us to collect logs from ARM instances. The disrupt_memory_stress was disabled to avoid uncontrolled situations with this nemesis. Please take note of the new issues and pull-requests templates when working with GitHub, and adhere to the new guidelines. In some cases, such as twcs, we have a table with many tombstones and our method for getting a list of ks/cf with data could be inaccurate. The cfstats utility is now used, which may be slower but is more accurate. The list of nemesis used in longevities is now cycled, so it never ends and fills the entire test run. This eliminates the need to fine-tune cases based on their running length. We’ve been defaulting to syslog-ng for some time, and rsyslog was filling the disk even when not needed, complicating node config. We have removed rsyslog from SCT. Improvements were made to NodeBootstrapAbortManager , correcting timeouts for operations to prevent races between threads and ensure all operations finish as expected. We’ve introduced a new script that allows us to exclude incorrect performance results from ES history by modifying the result’s job-name. The scale-cluster test has been stabilized by reducing the load, as the previous load was too heavy for the disk to manage. The audit test cases now default to syslog. This change was made because the table option was causing issues and could lead to query failures when auditing with CL=ONE. To streamline our workflow, any created or reopened SCT issues are now automatically added to the QA board, as we primarily work with boards. In certain scenarios, we may need to adjust timeouts relative to the number of nodes. To facilitate this, we’ve started to collect the number of database nodes when measuring operation times. Stay tuned for the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-monitoring-stack-4-6-0/1166 Title: [RELEASE] ScyllaDB Monitoring Stack 4.6.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of ScyllaDB Monitoring Stack 4.6.0 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB Enterprise and ScyllaDB Open Source, based on Prometheus … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-monitoring-stack-4-6-0/1166 ## Headings Structure: H1: [RELEASE] ScyllaDB Monitoring Stack 4.6.0 H1: Deprecation warning H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Monitoring Stack 4.6.0 H1: Deprecation warning H3: Related topics The ScyllaDB team is pleased to announce the release of ScyllaDB Monitoring Stack 4.6.0 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB Enterprise and ScyllaDB Open Source, based on Prometheus and Grafana. ScyllaDB Monitoring Stack 4.6.0 supports: Versions updates for ScyllaDB Monitoring Stack 4.6.0 New Information in ScyllaDB Dashboards Overview Dashboard Change Detailed Dashboard Change Per-Table Dashboard Change There was a long inconsistency when adding ports to the Prometheus target files, which is now resolved. Most users use the default port numbers for: ScyllaDB, node_exporter, and ScyllaDB manager agents and only use one target file. If that is your case, you do not need to change anything. Scylla Monitoring will take the port numbers from the files if the files are set explicitly and contain ports. For example, if you are specifying explicitly the node_exporter file (the -n flag to start-all.sh), you should ensure there are only IPs in the file with no ports or that you are using the correct node_exporter port number (by default, 9100). So, if you just copy the scylla_servers file and use that with the ScyllaDB ports, it would no longer work. $ cat node_exporter_servers.yml Most likely, you can use a single target file for scylla_server and do not need to explicitly add a node_exporter or ScyllaDB manager agent file. IPv6 users no longer need to add port numbers to the target files. It is recommended to use a single file (do not explicitly set the node_exporter file) without ports. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-212-2024-01-07/1167 Title: Last week in scylladb.git master (issue #212; 2024-01-07) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the f1dea4bc8a…7e84e03f52 range are covered. There were 60 non-merge commits from 11 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-212-2024-01-07/1167 ## Headings Structure: H1: Last week in scylladb.git master (issue #212; 2024-01-07) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #212; 2024-01-07) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the f1dea4bc8a…7e84e03f52 range are covered. There were 60 non-merge commits from 11 authors in that period. Some notable commits: Internal updates to system.peers ensure the host_id column is filled in correctly; this is required for Raft to operate correctly. Note all released versions already fill in this column. Communication of tablet routing information to the drivers has been streamlined. Communication of tablet routing information to the drivers is now negotiated; older drivers will not receive tablet routing information they will not use. An edge case in converting JSON numbers close to the maximum integer has been fixed. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/upcoming-scylladb-summit-2024/1169 Title: Upcoming ScyllaDB Summit 2024 - Announcements - ScyllaDB Community NoSQL Forum Meta Description: Next month, we will host our annual ScyllaDB Summit 2024. The event is online (and free) and includes two days of keynotes, technical talks, hands-on labs, and community building. It’s a great place to share your experi… Language: en Canonical URL: https://forum.scylladb.com/t/upcoming-scylladb-summit-2024/1169 ## Headings Structure: H1: Upcoming ScyllaDB Summit 2024 H3: Related topics ## Main Content: H1: Upcoming ScyllaDB Summit 2024 H3: Related topics Next month, we will host our annual ScyllaDB Summit 2024. The event is online (and free) and includes two days of keynotes, technical talks, hands-on labs, and community building. It’s a great place to share your experience, learn from others, connect with your peers, and discuss everything high-performance database-related. You can read more about it here. Hope to see you there! ScyllaDB Summit is starting tomorrow! You can still save your spot here. At the event’s start, @brother21913 and I will host two live training labs where you can get hands-on practice and improve your NoSQL skills. See you tomorrow! --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-2-13/1177 Title: [RELEASE] ScyllaDB 5.2.13 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.2.13, a bugfix release of the ScyllaDB 5.2 stable branch. ScyllaDB Open Source 5.2.12, like all past and future 5.x.y releases, is backward compatible, and supports roll… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-2-13/1177 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.2.13 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.2.13 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.2.13, a bugfix release of the ScyllaDB 5.2 stable branch. ScyllaDB Open Source 5.2.12, like all past and future 5.x.y releases, is backward compatible, and supports rolling upgrades. Note that the latest ScyllaDB Open Source stable release is 5.4 and you are encouraged to upgrade to it. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/what-is-the-difference-between-nodetool-remove-node-and-decommission/1180 Title: What is the difference between nodetool remove node and decommission? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: ScyllaDB have two similar operations, nodetool decommission and nodetool removenode when should each one be used? Language: en Canonical URL: https://forum.scylladb.com/t/what-is-the-difference-between-nodetool-remove-node-and-decommission/1180 ## Headings Structure: H1: What is the difference between nodetool remove node and decommission? H3: Related topics ## Main Content: H1: What is the difference between nodetool remove node and decommission? H3: Related topics ScyllaDB have two similar operations, nodetool decommission and nodetool removenode when should each one be used? nodetool decommission removes a live node from the cluster, while nodetool removenode removes a dead node. Both operations are used to remove a node from the cluster; both operations include data streaming. Decommission is preferred since it uses the removed node for streaming; use it when possible. --- ### Page: https://forum.scylladb.com/t/getting-failed-mounting-raid-volume-is-my-data-gone/1181 Title: Getting 'Failed mounting RAID volume!' - Is my data gone? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I am running a single self-managed Scylla instance in AWS. After running 1 week without any problems, I got this message from AWS: EC2 has detected degradation of the underlying hardware hosting your Amazon EC2 instanc… Language: en Canonical URL: https://forum.scylladb.com/t/getting-failed-mounting-raid-volume-is-my-data-gone/1181 ## Headings Structure: H1: Getting 'Failed mounting RAID volume!' - Is my data gone? H3: SSD instance store volumes - Amazon Elastic Compute Cloud H3: Related topics ## Main Content: H1: Getting 'Failed mounting RAID volume!' - Is my data gone? H3: SSD instance store volumes - Amazon Elastic Compute Cloud H3: Related topics I am running a single self-managed Scylla instance in AWS. After running 1 week without any problems, I got this message from AWS: EC2 has detected degradation of the underlying hardware hosting your Amazon EC2 instance (instance-ID: i-XX). I needed to stop-and-start my instance. After spinning the instance up again and sshed to the instance, I see it fails to boot because of this error: Failed mounting RAID volume! Scylla has aborted startup because of a missing RAID volume. Is there anything I can do to restart scylla with the same dataset as before the restart? As we spoke in Users Slack, no – your data is gone, particularly if you are using locally attached SSDs: The data on an SSD instance volume persists only for the life of its associated instance. SSD instance store volumes - Amazon Elastic Compute Cloud Thanks @felipemendes, here is the discussion from the Slack channel: Felipe: Nope, AWS allocated you a new set of local SSDs. You can recreate the array and then start the replace node procedure. Jasper: You can recreate the array and then start the replace node procedure. what do you mean by that? So the data is lost? Felipe: in that particular node, yes. You do have other replicas, no? Jasper: No, no other replica’s atm, it’s just for development though Some Amazon EC2 instances support solid state drives (SSD) to deliver high random I/O performance. Jasper: Ok and do you have any reference to You can recreate the array and then start the replace node procedure.? Felipe: https://opensource.docs.scylladb.com/stable/operating-scylla/procedures/cluster-management/rebuild-node.htmlhttps://opensource.docs.scylladb.com/stable/kb/raid-device.htmlDepending on your ScyllaDB version you may need to manually delete the systemd .mount resources in your /etc and follow with a systemctl daemon-reload afterwards. --- ### Page: https://forum.scylladb.com/t/using-bare-metal-servers-for-scylladb/1183 Title: Using bare-metal servers for ScyllaDB - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello! Happy New Year 2024 and Merry Christmas! I have a question: please tell me which server models and vendors are recommended for use for deploying ScyllaDB on bare-metal? I have read the system requirements, but per… Language: en Canonical URL: https://forum.scylladb.com/t/using-bare-metal-servers-for-scylladb/1183 ## Headings Structure: H1: Using bare-metal servers for ScyllaDB H3: Related topics ## Main Content: H1: Using bare-metal servers for ScyllaDB H3: Related topics Hello! Happy New Year 2024 and Merry Christmas! I have a question: please tell me which server models and vendors are recommended for use for deploying ScyllaDB on bare-metal? I have read the system requirements, but perhaps there is a list of reliable models? Thank you very much! Hi, We don’t have any recommended models or vendors. If you mostly interested in performance, you should care about number of cores, cores speed, enough RAM and fast NVMe drives. --- ### Page: https://forum.scylladb.com/t/not-able-to-create-kafka-source-connector/1188 Title: Not able to create Kafka source connector - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I am trying to create a kafka source connector with scylladb-cdc-source -connector. I have enable CDC on the table I want link this connector to, but I getting following error when I try to create a connector. Connector… Language: en Canonical URL: https://forum.scylladb.com/t/not-able-to-create-kafka-source-connector/1188 ## Headings Structure: H1: Not able to create Kafka source connector H3: Related topics ## Main Content: H1: Not able to create Kafka source connector H3: Related topics I am trying to create a kafka source connector with scylladb-cdc-source -connector. I have enable CDC on the table I want link this connector to, but I getting following error when I try to create a connector. Any help is appreciated. The correct name of username configuration option is scylla.user, not scylla.username. The connector should start correctly after correcting: --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-34-2024-01-13/1190 Title: Last week in scylla-cluster-tests.git master (issue #34; 2024-01-13) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This brief update highlights some noteworthy commits to the scylla-cluster-tests.git master from the past week, specifically those within the 9c5b107f…990f3ed8 range. During this period, there were 17 non-merge commits … Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-34-2024-01-13/1190 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #34; 2024-01-13) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #34; 2024-01-13) H3: Related topics This brief update highlights some noteworthy commits to the scylla-cluster-tests.git master from the past week, specifically those within the 9c5b107f…990f3ed8 range. During this period, there were 17 non-merge commits contributed by 5 authors. Here are some of the significant changes: The many-clients test, which is based on i3.metal machines (72 cores), has been improved by incorporating an authenticator into the c-s commands. The connectionsPerHost c-s option was also used to simulate 100k connections per node. In the event that scylla-bench encounters a failure during data validation, SCT will now generate a CRITICAL event and terminate the test. When commitlog-use-hard-size-limit is activated in Scylla, SCT initiates a thread that continuously checks the “free segments” and “commitlog directory” metrics throughout the test. This ensures that the commitlog does not exceed the set limit and that the free segments do not reduce to zero. Several recent fixes have been implemented to ensure the hydra utility is compatible with podman. This week, the restoration of the monitor stack was successfully resolved. Stay tuned for more updates in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/how-to-solve-row-count-in-a-table-time-out/1194 Title: How to solve row count in a table time out - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: When I perform a row count on my Table I get a timeout error: select count (*) from <Tablename>; Is giving an error OperationTimedOut: errors={‘x.x.x.x’: ‘Client request timeout. See Session.execute_async’}, last_host=… Language: en Canonical URL: https://forum.scylladb.com/t/how-to-solve-row-count-in-a-table-time-out/1194 ## Headings Structure: H1: How to solve row count in a table time out H3: Related topics ## Main Content: H1: How to solve row count in a table time out H3: Related topics When I perform a row count on my Table I get a timeout error: select count (*) from ; Is giving an error OperationTimedOut: errors={‘x.x.x.x’: ‘Client request timeout. See Session.execute_async’}, last_host=x.x.x.x How can I solve this? If you’re using cqlsh, try increasing the “request timeout” when starting cqlsh. I think it should work by providing it in the cqlsh command: cqlsh --request-timeout=6000 --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-213-2024-01-15/1195 Title: Last week in scylladb.git master (issue #213; 2024-01-15) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 7e84e03f52…423234841e range are covered. There were 97 non-merge commits from 14 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-213-2024-01-15/1195 ## Headings Structure: H1: Last week in scylladb.git master (issue #213; 2024-01-15) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #213; 2024-01-15) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 7e84e03f52…423234841e range are covered. There were 97 non-merge commits from 14 authors in that period. Some notable commits: Consistent topology Raft operations will now be aborted during shutdown, preventing it from hanging. Consistent topology will now reject nodetool removenode if the node is still alive. In consistent topology mode, the system will now track nodes that require the cleanup operation and issue it automatically if needed during decommission. The operator may still issue nodetool cleanup earlier to prevent a later decommission from taking too long. Host ID to host IP translation is now done just-in-time on each query rather than when updating topology, to prevent problems with removed nodes not having IP addresses. Queries to local table will now no longer be automatically parallelized to avoid shutdown problems. Local tables typically have little data and don’t benefit from parallelization. Tasks are used to track compaction work for the REST API. Internally-triggered compaction tasks are now removed immediately after completion so as not to consume memory. The me sstable format is now mandatory. The system can still read older formats. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/ubuntu-linux-error-starting-3-node-scylladb-on-docker/1197 Title: Ubuntu Linux - Error starting 3-node ScyllaDB on Docker - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I am trying to create a 3-node ScyllaDB using Docker running on Ubuntu Linux 22.04. I am executing docker run commands as a non-root user. This is from the S101 Scylla DB Essentials Lab Overview and Setup instructions… Language: en Canonical URL: https://forum.scylladb.com/t/ubuntu-linux-error-starting-3-node-scylladb-on-docker/1197 ## Headings Structure: H1: Ubuntu Linux - Error starting 3-node ScyllaDB on Docker H3: Related topics ## Main Content: H1: Ubuntu Linux - Error starting 3-node ScyllaDB on Docker H3: Related topics I am trying to create a 3-node ScyllaDB using Docker running on Ubuntu Linux 22.04. I am executing docker run commands as a non-root user. This is from the S101 Scylla DB Essentials Lab Overview and Setup instructions. I have set fs.aio-max-nr to 1048576 in /etc/sysctl.conf and have run sysctl -p /etc/sysctl.conf. root@M93p:/etc# sysctl -p /etc/sysctl.conf fs.aio-max-nr = 1048576 The lab provides docker run syntax to create three nodes (Node_X, Node_Y and Node_Z). I am able to create Node_X and Node_Y, but when creating Node_Z, I get the following error: FATAL: Exception during startup, aborting: std::runtime_error (Could not setup Async I/O: Resource temporarily unavailable. The most common cause is not enough request capacity in /proc/sys/fs/aio-max-nr. Try increasing that number or reducing the amount of logical CPUs available for your application) lscpu on my workstation shows: Architecture: x86_64 CPU op-mode(s): 32-bit, 64-bit Address sizes: 39 bits physical, 48 bits virtual Byte Order: Little Endian CPU(s): 4 On-line CPU(s) list: 0-3 Vendor ID: GenuineIntel Model name: Intel(R) Core™ i5-4570 CPU @ 3.20GHz CPU family: 6 Model: 60 Thread(s) per core: 1 Core(s) per socket: 4 The workstation has 24GB of memory and I am taking the defaults on Docker for the three containers. Any suggestions on a work-around for this issue? Thanks for reporting, please send the logs. I am executing a docker run to create three containers for ScyllaDB Nodes: Node_X, Node_Y and Node_Z: I have set fs.aio-max-nr to double the recommended size – this is running on a Intel 4-core workstation with 24GB memory. gjjohnson@M93p:~$ sudo sysctl -w fs.aio-max-nr=2560000 fs.aio-max-nr = 2560000 I have set this kernel parameter in /etc/sysctl.conf and have run a sysctl -p sudo sysctl -p /etc/sysctl.conf Node_X and Node_Y startup successfully with the following message: init - Scylla version 5.2.0-0.20230427.429b696bbc1b initialization completed. I wait for an initialization complete message in Node_X before running Node_Y and then wait for Node_Y to complete initialization before running Node_Z. When I attempt to start Node_Z it fails with the following message: FATAL: Exception during startup, aborting: std::runtime_error (Could not setup Async I/O: Resource temporarily unavailable. The most common cause is not enough request capacity in /proc/sys/fs/aio-max-nr. Try increasing that number or reducing the amount of logical CPUs available for your application) Since I get the same message after increasing fs.aio-max-nr, I decided to try to limit each container to one CPU. Again, Node_X and Node_Y start successfully and Node_Z fails with the same error. Let me know if you also need the logs from Node_X and Node_Y. Recently it was reported the same also for macOS here - Docker on macOS fails with: "Could not setup Async I/O: Resource temporarily unavailable." · Issue #16806 · scylladb/scylladb · GitHub. However, if it happens also in other platforms it may get higher priority. Can you please add your input in that issue as well? I read the MacOS issue and it is the same error, but the suggested Kernel parameters mentioned for the MacOS are different than the Ubuntu Kernel parameters. Ok – I used cloud.scylladb.com to run the exercises in S101: ScyllaDB Essentials. I’m now working on S110: The Mutant Monitoring System and the lab setup is using docker-compose to create a local cluster. I’ve removed all of my previous containers, run git clone to download all of the required configuration files and code samples, and have re-started Docker desktop and have run docker-compose up -d. I am still encountering the same error starting one of the three nodes in the cluster. My kernel parameters are: gjjohnson@M93p:~/scylla-code-samples/mms$ sudo sysctl -p /etc/sysctl.conf [sudo] password for gjjohnson: fs.aio-max-nr = 2097152 fs.inotify.max_user_watches = 524288 I tried fs.aio-max-nr with the suggested value and was not able to start all three nodes, so I doubled the value to 2,097,152. There was no documentation on setting fs.inotify.max_user_watches, but in the setup document for for S110: The Mutant Monitoring System, the output from sysctl -p /etc/sysctl.conf shows: ubuntu $ sysctl -p /etc/sysctl.conf fs.inotify.max_user_watches = 524288 fs.aio-max-nr = 1048576 So, I set both parameters on my system. I have 24GB of memory on this workstation and have set the Docker Desktop memory limit to 15GB. On the node that fails to start, scylla-node3, the error message is: 024-01-24 00:25:27 FATAL: Exception during startup, aborting: std::runtime_error (Could not setup Async I/O: Resource temporarily unavailable. The most common cause is not enough request capacity in /proc/sys/fs/aio-max-nr. Try increasing that number or reducing the amount of logical CPUs available for your application) 2024-01-24 00:25:28 Traceback (most recent call last): 2024-01-24 00:25:28 File “/opt/scylladb/scripts/libexec/scylla-housekeeping”, line 196, in 2024-01-24 00:25:28 args.func(args) 2024-01-24 00:25:28 File “/opt/scylladb/scripts/libexec/scylla-housekeeping”, line 122, in check_version 2024-01-24 00:25:28 current_version = sanitize_version(get_api(‘/storage_service/scylla_release_version’)) 2024-01-24 00:25:28 ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ 2024-01-24 00:25:28 File “/opt/scylladb/scripts/libexec/scylla-housekeeping”, line 80, in get_api 2024-01-24 00:25:28 return get_json_from_url(“http://” + api_address + path) 2024-01-24 00:25:28 ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ 2024-01-24 00:25:28 File “/opt/scylladb/scripts/libexec/scylla-housekeeping”, line 75, in get_json_from_url 2024-01-24 00:25:28 raise RuntimeError(f’Failed to get “{path}” due to the following error: {retval}') 2024-01-24 00:25:28 RuntimeError: Failed to get “http://localhost:10000/storage_service/scylla_release_version” due to the following error: 2024-01-24 00:25:31 FATAL: Exception during startup, aborting: std::runtime_error (Could not setup Async I/O: Resource temporarily unavailable. The most common cause is not enough request capacity in /proc/sys/fs/aio-max-nr. Try increasing that number or reducing the amount of logical CPUs available for your application) I would really like to resolve this issue so I can proceed with the rest of this class. I’m trying to get a workaround so at least developer-mode won’t face this issue. I’m happy to share that there is an easy workaround that works also on macOS: You can start scylla with the following flag: “–reactor-backend=epoll” I was able to successfully start a 3-node cluster with your suggested change. So, for S110: The Mutant Monitoring System lab, the 3-node cluster is built using Docker Compose. Running the updated commands outside of Docker Compose will start the nodes successfully, but they will not be correctly configured for the labs. I edited my local copy of ~/scylla-code-samples/mms/docker-compose.yml and after deleting the nodes on my cluster, I was able to rebuild them correctly using docker-compose. The updated docker-compose.yml file is: scylla-node1: container_name: scylla-node1 image: scylladb/scylla:5.2.0 restart: always command: --seeds=scylla-node1,scylla-node2 --memory 750M --api-address 0.0.0.0 --reactor-backend=epoll volumes: scylla-node2: container_name: scylla-node2 image: scylladb/scylla:5.2.0 restart: always command: --seeds=scylla-node1,scylla-node2 --memory 750M --api-address 0.0.0.0 --reactor-backend=epoll volumes: scylla-node3: container_name: scylla-node3 image: scylladb/scylla:5.2.0 restart: always command: --seeds=scylla-node1,scylla-node2 --memory 750M --api-address 0.0.0.0 --reactor-backend=epoll volumes: networks: web: driver: bridge I don’t guarantee this is the best way to set up docker-compose.yml since I am not a Docker expert, but this does create a working 3-node cluster on my workstation. Thanks to everyone at Scylla for your help with this work-around! GJ Johnson --- ### Page: https://forum.scylladb.com/t/i-have-saw-multiple-integration-in-scylladb-enterprise-but-nowhere-is-mentioned-that-is-it-possible-to-fetch-audit-logs-for-those-integration/1199 Title: I have saw multiple integration in ScyllaDB Enterprise. But nowhere is mentioned that is it possible to fetch audit logs for those integration - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Does Scylla-audit also store audit logs from other integration also. I did not find any reference regarding that. Language: en Canonical URL: https://forum.scylladb.com/t/i-have-saw-multiple-integration-in-scylladb-enterprise-but-nowhere-is-mentioned-that-is-it-possible-to-fetch-audit-logs-for-those-integration/1199 ## Headings Structure: H1: I have saw multiple integration in ScyllaDB Enterprise. But nowhere is mentioned that is it possible to fetch audit logs for those integration H3: Scylla Integrations and Connectors | ScyllaDB Docs H3: Related topics ## Main Content: H1: I have saw multiple integration in ScyllaDB Enterprise. But nowhere is mentioned that is it possible to fetch audit logs for those integration H3: Scylla Integrations and Connectors | ScyllaDB Docs H3: Related topics Does Scylla-audit also store audit logs from other integration also. I did not find any reference regarding that. What do you mean by integration? Integration with what? ScyllaDB is an Apache Cassandra-compatible NoSQL data store that can handle 1 million transactions per second on a single server. Does ScyllaDB also store audit logs from this integration also. I did not find any reference regarding that. not find any reference regarding that. No, there aren’t such built-in integrations at the moment. --- ### Page: https://forum.scylladb.com/t/does-experimental-feature-consistent-topology-changes-support-online-modification/1200 Title: Does experimental feature `consistent-topology-changes` support online modification? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Now I have a up and normal 5.4.0 cluster. After adding configuration consistent-topology-changes to them, only the seed node is functioning properly. The other two nodes are not displayed in the cluster, but their lo… Language: en Canonical URL: https://forum.scylladb.com/t/does-experimental-feature-consistent-topology-changes-support-online-modification/1200 ## Headings Structure: H1: Does experimental feature `consistent-topology-changes` support online modification? H3: Related topics ## Main Content: H1: Does experimental feature `consistent-topology-changes` support online modification? H3: Related topics Now I have a up and normal 5.4.0 cluster. The consistent-topology-changes experimental feature does not yet support upgrades. I.e. if you start with a cluster without the feature, and then try to enable it, the cluster will end up in some undefined state. To get a functioning cluster with consistent-topology-changes, you need to bootstrap it from scratch with this option. But remember that this is an experimental feature, it’s not production ready. It’s under heavy development with backward incompatible changes happening every week. It’s fine to play with it with clusters that you’re going to discard. --- ### Page: https://forum.scylladb.com/t/will-compaction-clean-up-all-existing-tombstones-on-the-node/1202 Title: Will `Compaction` clean up all existing tombstones on the node? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I have a question. Compaction will clean the tombstone. When we proactively execute Compaction, will it clean up all the tombstones that have already been generated? Or is it cleaning up some of the tombstones that have … Language: en Canonical URL: https://forum.scylladb.com/t/will-compaction-clean-up-all-existing-tombstones-on-the-node/1202 ## Headings Structure: H1: Will `Compaction` clean up all existing tombstones on the node? H3: Related topics ## Main Content: H1: Will `Compaction` clean up all existing tombstones on the node? H3: Related topics I have a question. Compaction will clean the tombstone. When we proactively execute Compaction, will it clean up all the tombstones that have already been generated? Or is it cleaning up some of the tombstones that have already been produced? When compaction occurs, the data will be expunged completely and the corresponding disk space recovered. I see this sentence here. Tombstones typically have a data it shadows. For example, consider: INSERT INTO ks.t (key, val) VALUES (0, 0); At a later time you then issue: DELETE FROM ks.t WHERE key=0; In that case, your data may live on a given SSTable, whereas the tombstone will live in another. When compaction compacts the SSTable containing the data along with the SSTable containing the tombstone, your data will be evicted. If the tombstone has already expired (ie: Past gc_grace_seconds), then the tombstone will be evicted as well. Otherwise, the tombstone will be kept, to ensure you have a time to repair in case that tombstone didn’t reach other nodes as part of your original write. If I execute nodetool compact on node A, will all generated tombstones on that node be cleared? Does Command nodetool compact only clean up expired tombstones?That is to say, tombstones can only be cleared after they expire. If the value of gc_grace_seconds is 10 days, then the tombstone will only be cleared after 10 days. Whether or not I have executed nodetool compact within these 10 days? Thanks! When you run a major (nodetool compact), all expired tombstones will be evicted as they will be compacted along with the data they shadow. Non-expired tombstones will still be kept, but this shouldn’t be a problem as tombstones are really small, and since you already performed a major, then it means that data it shadows got evicted. As a result, as soon as the tombstone expires after your major, compaction will eventually evict it. --- ### Page: https://forum.scylladb.com/t/node-failed-to-join-the-cluster-when-i-use-the-experimental-feature-consistent-topology-changes/1203 Title: Node failed to join the cluster when I use the experimental feature "consistent-topology-changes" - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Now I am starting a cluster with experimental feature “consistent-topology-changes” . For the first eleven nodes , I enabled the experimental feature successfully . However , when it comes to the twelve node , it failed … Language: en Canonical URL: https://forum.scylladb.com/t/node-failed-to-join-the-cluster-when-i-use-the-experimental-feature-consistent-topology-changes/1203 ## Headings Structure: H1: Node failed to join the cluster when I use the experimental feature "consistent-topology-changes" H3: Related topics ## Main Content: H1: Node failed to join the cluster when I use the experimental feature "consistent-topology-changes" H3: Related topics Now I am starting a cluster with experimental feature “consistent-topology-changes” . For the first eleven nodes , I enabled the experimental feature successfully . However , when it comes to the twelve node , it failed to join the cluster . I am confused why does this situation happen ? Here are error logs of seed node and failed-to-join node: Jan 17 16:24:32 HOSTNAME-485 scylla[158302]: [shard 0:stat] schema_tables - Schema version changed to 187398f8-6575-3b4d-b98b-5602749c6ec3 Jan 17 16:24:33 HOSTNAME-485 scylla[158302]: [shard 0:stat] schema_tables - Schema version changed to 187398f8-6575-3b4d-b98b-5602749c6ec3 Jan 17 16:24:33 HOSTNAME-485 scylla[158302]: [shard 0:stat] schema_tables - Schema version changed to 187398f8-6575-3b4d-b98b-5602749c6ec3 Jan 17 16:24:33 HOSTNAME-485 scylla[158302]: [shard 0:stat] schema_tables - Schema version changed to 187398f8-6575-3b4d-b98b-5602749c6ec3 Jan 17 16:24:33 HOSTNAME-485 scylla[158302]: [shard 0:stat] schema_tables - Schema version changed to 187398f8-6575-3b4d-b98b-5602749c6ec3 Jan 17 16:24:33 HOSTNAME-485 scylla[158302]: [shard 0:stat] schema_tables - Schema version changed to 187398f8-6575-3b4d-b98b-5602749c6ec3 Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:stat] schema_tables - Schema version changed to 187398f8-6575-3b4d-b98b-5602749c6ec3 Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:stat] schema_tables - Schema version changed to 187398f8-6575-3b4d-b98b-5602749c6ec3 Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down group 0 service Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down group 0 service was successful Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down storage service notifications Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down storage service notifications was successful Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down system distributed keyspace Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down system distributed keyspace was successful Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down migration manager notifications Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down migration manager notifications was successful Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down storage service notifications Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down storage service notifications was successful Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down sstables loader API Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down sstables loader API was successful Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down sstables loader Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down sstables loader was successful Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down repair API Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down repair API was successful Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down messaging service API Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down messaging service API was successful Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down storage service messaging Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down storage service messaging was successful Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down cdc log service Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down cdc log service was successful Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down CDC Generation Management service Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down CDC Generation Management service was successful Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down repair service Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:main] task_manager - Stopping module repair Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:main] task_manager - Unregistered module repair Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 2:main] task_manager - Stopping module repair Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 6:main] task_manager - Stopping module repair Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 7:main] task_manager - Stopping module repair Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 13:main] task_manager - Stopping module repair Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 17:main] task_manager - Stopping module repair ……………… ……………… Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down storage_service was successful Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down tablet allocator Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down tablet allocator was successful Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down direct_failure_detector Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down direct_failure_detector was successful Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down fd_pinger Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down fd_pinger was successful Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down raft_address_map Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down raft_address_map was successful Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down gossiper Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:main] gossip - My status = UNKNOWN Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:main] gossip - No local state or state is in silent shutdown, not announcing shutdown Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:main] gossip - Disable and wait for gossip loop started Jan 17 16:24:34 HOSTNAME-485 scylla[158302]: [shard 0:goss] rpc - client 10.249.138.159:50784 msg_id 180: exception “gate closed” in no_wait handler ignored Jan 17 16:24:35 HOSTNAME-485 scylla[158302]: [shard 0:goss] rpc - client 10.249.138.157:51552 msg_id 181: exception “gate closed” in no_wait handler ignored Jan 17 16:24:35 HOSTNAME-485 scylla[158302]: [shard 0:goss] rpc - client 10.249.138.154:63312 msg_id 183: exception “gate closed” in no_wait handler ignored Jan 17 16:24:36 HOSTNAME-485 scylla[158302]: [shard 0:goss] rpc - client 10.249.138.158:50496 msg_id 183: exception “gate closed” in no_wait handler ignored Jan 17 16:24:36 HOSTNAME-485 scylla[158302]: [shard 0:stre] gossip - failure_detector_loop: Finished main loop Jan 17 16:24:36 HOSTNAME-485 scylla[158302]: [shard 0:main] gossip - Gossip is now stopped Jan 17 16:24:36 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down gossiper was successful Jan 17 16:24:36 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down messaging service Jan 17 16:24:36 HOSTNAME-485 scylla[158302]: [shard 0:main] messaging_service - Shutting down nontls server Jan 17 16:24:36 HOSTNAME-485 scylla[158302]: [shard 0:main] messaging_service - Shutting down tls server Jan 17 16:24:36 HOSTNAME-485 scylla[158302]: [shard 0:main] messaging_service - Shutting down tls server - Done Jan 17 16:24:36 HOSTNAME-485 scylla[158302]: [shard 0:main] messaging_service - Stopping client for address: 10.249.138.156:0 Jan 17 16:24:36 HOSTNAME-485 scylla[158302]: [shard 0:main] messaging_service - Stopping client for address: 10.249.138.158:0 Jan 17 16:24:36 HOSTNAME-485 scylla[158302]: [shard 0:main] messaging_service - Stopping client for address: 10.249.138.155:0 Jan 17 16:24:36 HOSTNAME-485 scylla[158302]: [shard 0:main] messaging_service - Stopping client for address: 10.249.138.162:0 Jan 17 16:24:36 HOSTNAME-485 scylla[158302]: [shard 0:main] messaging_service - Stopping client for address: 10.249.138.160:0 Jan 17 16:24:36 HOSTNAME-485 scylla[158302]: [shard 0:main] messaging_service - Stopping client for address: 10.249.138.157:0 Jan 17 16:24:36 HOSTNAME-485 scylla[158302]: [shard 0:main] messaging_service - Stopping client for address: 10.249.138.154:0 Jan 17 16:24:36 HOSTNAME-485 scylla[158302]: [shard 0:main] messaging_service - Stopping client for address: 10.249.138.161:0 Jan 17 16:24:36 HOSTNAME-485 scylla[158302]: [shard 0:main] messaging_service - Stopping client for address: 10.249.138.159:0 Jan 17 16:24:36 HOSTNAME-485 scylla[158302]: [shard 0:main] messaging_service - Stopping client for address: 10.249.138.163:0 Jan 17 16:24:36 HOSTNAME-485 scylla[158302]: [shard 0:main] messaging_service - Stopping client for address: 10.249.138.153:0 Jan 17 16:24:36 HOSTNAME-485 scylla[158302]: [shard 0:main] messaging_service - Stopped clients Jan 17 16:24:36 HOSTNAME-485 scylla[158302]: [shard 0:main] messaging_service - Stopping client for address: 10.249.138.153:0 ………… ………… Jan 17 16:24:36 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down compaction_manager was successful Jan 17 16:24:36 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down task_manager Jan 17 16:24:36 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down task_manager was successful Jan 17 16:24:36 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down service_memory_limiter Jan 17 16:24:36 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down service_memory_limiter was successful Jan 17 16:24:36 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down sst_dir_semaphore Jan 17 16:24:36 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down sst_dir_semaphore was successful Jan 17 16:24:36 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down migration manager notifier Jan 17 16:24:36 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down migration manager notifier was successful Jan 17 16:24:36 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down prometheus API server Jan 17 16:24:36 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down prometheus API server was successful Jan 17 16:24:36 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down API server Jan 17 16:24:36 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down API server was successful Jan 17 16:24:36 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down sighup Jan 17 16:24:36 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down sighup was successful Jan 17 16:24:36 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down configurables Jan 17 16:24:36 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Shutting down configurables was successful Jan 17 16:24:36 HOSTNAME-485 scylla[158302]: [shard 0:main] init - Startup failed: seastar::timed_out_error (timedout) What scylla version are you using? I suggest opening a new issue for it in Issues · scylladb/scylladb · GitHub I am using scylla 5.4.0 . I’ve opened a new issue in GitHub : Node failed to join the cluster when I use the experimental feature “consistent-topology-changes” · Issue #16869 · scylladb/scylladb · GitHub . --- ### Page: https://forum.scylladb.com/t/kafka-how-to-extract-nested-field-in-json-structure/1204 Title: Kafka: How to extract nested field in JSON structure? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello, I am using scylladb-cdc-source-connector to publish cdc messages to my kafka topic. I am noticing that in JSON messages some keys have standard values whereas some has nested object which has “value” key inside co… Language: en Canonical URL: https://forum.scylladb.com/t/kafka-how-to-extract-nested-field-in-json-structure/1204 ## Headings Structure: H1: Kafka: How to extract nested field in JSON structure? H3: Related topics ## Main Content: H1: Kafka: How to extract nested field in JSON structure? H3: Related topics Hello, I am using scylladb-cdc-source-connector to publish cdc messages to my kafka topic. I am noticing that in JSON messages some keys have standard values whereas some has nested object which has “value” key inside containing actual value. I think it has to do with how the connector is processing primary key and non-primary key columns (non-primary key columns having nested values). How can I access nested fields of JSON record to put “city” field as a key of the record? Suppose my Kafka message data is : KEY: {“id”: 4, “name”: “John”} VALUE: { “before”: null, “after”: { “id”: 4, “name”: “John”, “city”: { “value”: “New York City” } }, “source”: { …some source cofig }, “op”: “c”, “ts_ms”: 1623834752982, “transaction”: null } Value: { “before”: null, “after”: { “id”: 4, “name”: “John”, “city”: “New York City” // this part is changed }, “source”: { …some source cofig }, “op”: “c”, “ts_ms”: 1623834752982, “transaction”: null } As per my research I came to understand that there is no transform operation/ SMT in kafka that deal with nested values (correct me if I am wrong). Is there anything I can do with scylladb structure or kafka message to get desired results? The reason I need this in this particular format is we are currently switching from MySql to Scylladb and we already have sink connector and ES setup to consume message with this particular format. So we are trying to have minimal changes. Any help is appreciated. Thank You! @piotr do you know how to do it? This is a quite complicated issue. I am not sure if this is possible without “ruining” the rest of the record or writing your own custom SMT, but there may be a workaround. Bad news is that as you said, most of the SMTs currently do not support accessing nested fields. Good news is this seems to be something that is being worked on. You can read all the details in a neatly organized proposal here: https://cwiki.apache.org/confluence/display/KAFKA/KIP-821%253A+Connect+Transforms+support+for+nested+structures There is a list of affected and non-affected SMTs inside. As you can see ValueToKey is affected. What you can do now is try to apply a series of transformations that will bring the field you want to the top-level. There are two transformations that may be of use here: Flatten and ExtractField. You can chain the uses of ExtractField to reach the nested field you need, but that will unfortunately discard everything else. Flatten also may be problematic, because a lot of fields will change. For reference here’s a “before and after” of a single flatten transformation: Fields that were additionally enclosed in “value” structure will have “value” added to their names. Keep in mind that if it suits your use case there is a ScyllaExtractNewState transformer that can help you with discarding the “value” part. If you decide to try out the chain of SMTs route, keep in mind that CDC connector produces also heartbeat messages to a different topic. For example if you decide to use ExtractField SMT, the usual messages will have an “after” field, but heartbeat messages will not and connector will ultimately crash. To make it work you need to filter out the correct messages, for example by specifying a predicate for your transformations that matches to a concrete topic name. Here’s an example how it would look like in a .properties configuration file: So to sum it up, you can chain: ScyllaExtractNewState, ExtractField_1, …, ExtractField_n, ValueToKey to achieve desired result. Alternatively you can try Flatten, some kind of Rename SMT for your field and then ValueToKey I’ve run a test for your use case @Aditya_Mathur and I think this may be sufficient: After applying ScyllaExtractNewRecordState the messages for my table already look like this: After applying ValueToKey, the key looks like this: So it turns out we don’t even need to chain ExtractFields. I was expecting the message structure to be a little different before that. You will need to adjust topic prefix in the predicate field to match your topic and field in ValueToKey to match city field instead of v1 --- ### Page: https://forum.scylladb.com/t/segmentation-fault-on-shard-x/1206 Title: Segmentation fault on shard x - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello there! We have a problem with our scylla cluster. Could you help me understand what exactly happens? Logs at the end Some context: For some time (maybe even from the very beginning but it didn’t happen so often)… Language: en Canonical URL: https://forum.scylladb.com/t/segmentation-fault-on-shard-x/1206 ## Headings Structure: H1: Segmentation fault on shard x H3: Related topics ## Main Content: H1: Segmentation fault on shard x H3: Related topics Hello there! We have a problem with our scylla cluster. Could you help me understand what exactly happens? Logs at the end Some context: For some time (maybe even from the very beginning but it didn’t happen so often), we experience random restarts of our scylla pods. Usually there is some warning about oversized allocation but it’s “normal” as it happens all the time. We don’t have any way to see what exact query caused segmentation faults, but even if we found out, it still shouldn’t happen right? Is it known problem? Can you deduct what went wrong based on logs? Setup: 3-node scylla cluster in Azure Kubernetes Service (3x Standard_D8as_v5 with 8 vCPUs and 32GiB RAM) scylla server version: 5.1.18 AKS Node image: AKSUbuntu-2204gen2containerd-202312.06.0 ################# /etc/scylla.d/io_properties.yaml ################# /etc/scylla/scylla.yaml ################# ulimit -a logs: Scylla build_id: 724b167bb6412a9242e392d39188abf36a8c259e Opened github issue - Segmentation fault on shard x · Issue #16841 · scylladb/scylladb · GitHub I guess this post can be taken down if needed --- ### Page: https://forum.scylladb.com/t/scylladb-summit-free-virtual-agenda-announced/1208 Title: ScyllaDB Summit (Free + Virtual) Agenda Announced - Announcements - ScyllaDB Community NoSQL Forum Meta Description: Join us at ScyllaDB Summit, a free 2-day community event that’s intentionally virtual, highly interactive, and purely technical. You will hear about your peers’ experiences and discover new ways to alleviate your own la… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-summit-free-virtual-agenda-announced/1208 ## Headings Structure: H1: ScyllaDB Summit (Free + Virtual) Agenda Announced H3: Database Insights from Disney, Discord, Expedia & More at ScyllaDB Summit H3: Related topics ## Main Content: H1: ScyllaDB Summit (Free + Virtual) Agenda Announced H3: Database Insights from Disney, Discord, Expedia & More at ScyllaDB Summit H3: Related topics Join us at ScyllaDB Summit, a free 2-day community event that’s intentionally virtual, highly interactive, and purely technical. You will hear about your peers’ experiences and discover new ways to alleviate your own latency, throughput, and cost pains. Learn how your peers at Disney, Discord, Expedia, Supercell, Digital Turbine, ShareChat, Zee, Paramount, and more are tackling their toughest database challenges. --- ### Page: https://forum.scylladb.com/t/scylladb-timeout-error-comes-while-writing-reading-from-to-table/1210 Title: ScyllaDB timeout error comes while writing/reading from/to table - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: We are using ScyallDB version 3.0.8 in a production environment with 3 nodes cluster. Each node has 32-core CPU 128GB RAM. ScyllaDB has 1 DB with an hourwise partition table. Each hour table max contains 30M rows, with … Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-timeout-error-comes-while-writing-reading-from-to-table/1210 ## Headings Structure: H1: ScyllaDB timeout error comes while writing/reading from/to table H3: Related topics ## Main Content: H1: ScyllaDB timeout error comes while writing/reading from/to table H3: Related topics We are using ScyallDB version 3.0.8 in a production environment with 3 nodes cluster. Each node has 32-core CPU 128GB RAM. ScyllaDB has 1 DB with an hourwise partition table. Each hour table max contains 30M rows, with each row having 4 columns containing email information ( TxId, from_address,to_address, maile_content). Each Row has max 50KB size. We have been facing a timed-out error in some time while reading/writing to the disk. I have configured the timed out error. in scyall.yml is 10 seconds. Below is my nodetool proxyhistograms output Percentile Read Latency Write Latency Range Latency CAS Read Latency CAS Write Latency View Write Latency (micros) (micros) (micros) (micros) (micros) (micros) 50% 378.00 542.50 2546.00 0.00 0.00 0.00 75% 496.00 169026.50 3107.00 0.00 0.00 0.00 95% 943.75 464572.00 4862.75 0.00 0.00 0.00 98% 1112.50 549202.00 10827.00 0.00 0.00 0.00 99% 2569.75 1820080.00 31794.50 0.00 0.00 0.00 Min 64.00 66.00 1604.00 0.00 0.00 0.00 Max 12925.00 10000085.00 2013221.00 0.00 0.00 0.00 why I am getting timedout error, I am looking for read and write latency <=1 milli seconds. How can I achiev this? Why are you using such an old ScyllaDB version? Please try with 5.4 (or the latest) First, as Guy wrote above, this is a very very old scylladb version. Latest OSS release is 5.4.1. I suggest you to use that (though, your upgrade path isn’t going to be simple). (Maybe a backup and restore would be easier, but still not sure clean it would be). Second, how does your reads look like? Third, IIUC, you have partitions with 30M rows, with 50K each row? Does it mean you have 1.43 TB size of a partition?? In that case, your data model is not designed well and your target latency impossible. You must redesigning your data model (based on how would want to read the data). Forth, how big is user dataset? how much data do you save? I guess you’re not using TWCS, I’m not even sure it was introduced in this release. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-35-2024-01-19/1211 Title: Last week in scylla-cluster-tests.git master (issue #35; 2024-01-19) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This brief report highlights some intriguing commits to the scylla-cluster-tests.git master over the past week, specifically within the 086c12d2…6a910e04 range. During this period, we had 27 non-merge commits from 5 aut… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-35-2024-01-19/1211 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #35; 2024-01-19) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #35; 2024-01-19) H3: Related topics This brief report highlights some intriguing commits to the scylla-cluster-tests.git master over the past week, specifically within the 086c12d2…6a910e04 range. During this period, we had 27 non-merge commits from 5 authors. Here are some of the notable changes: For those interested in the nemesis list that SisyphusMonkey is about to create, you can obtain it by running the following command: hydra nemesis-list --backend aws --config test-cases/PR-provision-test-docker.yaml This command will output the list of nemesis based on the configuration files used. Check out the changes here. Tests running on K8S-EKS backend have transitioned to using the real S3 service for Scylla-manager backups, moving away from Minio. This change enables us to start running ‘mgmt restore’ nemesis, and a job for this has been added. You can see these updates here and here. The EndOfQuota Nemesis has been updated to read all mount options and search for the expected value. This nemesis can now be run on all backends. Previously, only EventsFilter had the option to set extra time before eviction. Now, DbEventsFilter has been granted the same capability. See the change here. Lastly, we’ve begun to collect runner metrics using node_exporter and have integrated it into the monitoring stack. This will assist us in understanding the load on the sct-runner, identifying issues with the test code, and determining if we need to choose larger instances/disks for certain tests. A new image version for sct-runner (1.7) was released with this change. Looking forward to updating you on the next batch of changes in the scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/help-with-scylla-monitoring-and-helm-deployment/1214 Title: Help with scylla-monitoring and Helm deployment - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hey everyone, I have a question. So I have a deployment that includes ScyllaDB, that I do using helmfile (this is the helmfile). I already have Prometheus and Grafana in the deployment, because I’m using the kube-promet… Language: en Canonical URL: https://forum.scylladb.com/t/help-with-scylla-monitoring-and-helm-deployment/1214 ## Headings Structure: H1: Help with scylla-monitoring and Helm deployment H3: Related topics ## Main Content: H1: Help with scylla-monitoring and Helm deployment H3: Related topics Hey everyone, I have a question. So I have a deployment that includes ScyllaDB, that I do using helmfile (this is the helmfile). I already have Prometheus and Grafana in the deployment, because I’m using the kube-prometheus-stack helm chart. I was trying to add monitoring to ScyllaDB (from here), because IIUC ScyllaDB already logs some Prometheus metrics and if I enable the monitoring I can have access to the metrics and dashboards. Is that correct? Based on that assumption, I added this: To my scylla.values.yaml, hoping that that would do it, but when I try deploying locally with kind I get this error: Which makes me think that I need to create the ServiceMonitor, or does that just mean that it’s failing to create the ServiceMonitor for some reason? TL;DR my question is, given the current state of my deployment, how do I make my existing Prometheus deployment scrape the ScyllaDB metrics? And also, how can I get access to the dashboards that scylla-monitoring provides based on said metrics? I also tried posting on the ScyllaDB slack but got no answer. Thank you so much for any help! Hello, have you tried to check: Monitoring | ScyllaDB Docs. I am not sure if reusing the same monitoring stack for scylla and some other project is easy to achieve. Typically we deploy it as a standalone thing, optinally with metric export (to another prometheus). --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-214-2024-01-21/1216 Title: Last week in scylladb.git master (issue #214; 2024-01-21) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 423234841e…b1ba904c49 range are covered. There were 99 non-merge commits from 17 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-214-2024-01-21/1216 ## Headings Structure: H1: Last week in scylladb.git master (issue #214; 2024-01-21) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #214; 2024-01-21) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 423234841e…b1ba904c49 range are covered. There were 99 non-merge commits from 17 authors in that period. Some notable commits: If the experimental tablets feature is enabled, newly-created keyspaces will use tablets. This facilitates testing. One can opt-out by using the TABLETS = property map. nodetool repair can now repair keyspaces that use tablets, with some limitations. The materialized view building code now works with tablets. The nodetool refresh command now automatically uses load-and-stream for tables that use tablets. Since tablets are more dynamic, it’s not possible for the operator to place sstables in nodes that own them, so using load-and-stream is safer. The DESCRIBE CLUSTER statement now ignores keyspaces that use tablets, since each table within them has its own token range ownership mapping. The bundled Python driver was updated to a version that supports tablets. Many tests now run with tablets enabled. Cleanup is used to remove data from nodes after it was migrated to other nodes. Cleanup now ignores keyspaces that never migrate data away. When recovering after a crash, we skip commitlog replay for data that we know was captured in sstables. However, we must ignore sstables that were not created on this node, as the commitlog positions they refer to are invalid on this node. There is now a new table that tracks topology changes and their completion status. The table is active when experimental consistent topology is enabled. When using consistent topology, a failed attempt to join a node will now shut it down rather than hanging. A bug in the rebuild command when experimental consistent topology is enabled was fixed. The service level mechanism works by polling the internal tables used to represent it. The polling interval can now be configured. This is useful to speed up tests. The bundled cqlsh command now works with unix-domain sockets as well as TCP. This allows it to use the new administrative CQL interface. There is now new option for testing code coverage. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/auditing-in-scylla-db-cloud/1219 Title: Auditing in Scylla Db Cloud - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi Team, May I know the process to enable auditing in Scylla Db Cloud? Language: en Canonical URL: https://forum.scylladb.com/t/auditing-in-scylla-db-cloud/1219 ## Headings Structure: H1: Auditing in Scylla Db Cloud H3: Related topics ## Main Content: H1: Auditing in Scylla Db Cloud H3: Related topics Hi Team, May I know the process to enable auditing in Scylla Db Cloud? Hi Soumalya, are you referring to enabling audit logs in the ScyllaDB instance? If so, please take a look at How to Enable Audit logs for ScyllaDB, how many types of queries it supports and how to see the audit logs for the executed query. Also, how its audit log looks like - ScyllaDB - ScyllaDB Community NoSQL Forum. I think he is trying to ask about scylladb cloud. Does it means Scylladb cloud is like an instance ?? ScyllaDB Cloud is a Scylla DB (Enterprise) managed by Scylla. Classification: Public We understand that it is kind of a managed service for Scylla DB(Enterprise) but we are unable to find any steps to enable auditing in Scylla DB Cloud(we know how to enable audit logging in Enterprise). Can you please provide some steps for the same? Hi Soumalya, thanks for the clarification. We do not currently support the ability to enable Audit logging from the ScyllaDB Cloud App. Can you explain a bit more about your specific use case? Hi Michael, we are not sure what Scylla Db cloud app is. We are planning to use Scylla Db cloud as a database where we communicate to this database by executing some queries. We are eager to know if these queries can generate some kind of audit logs in Scylla Db cloud. This is our use case. You are welcome to try out ScyllaDB Cloud by running our Free Trial. You can start directly from https://cloud.scylladb.com. You can also take a look at our Lab - Lab: Getting Started with ScyllaDB Cloud - ScyllaDB University Hi Michael, we have created a cluster but unable to see any settings to enable auditing as mentioned in this below link: ScyllaDB Cloud Security Concepts | ScyllaDB Docs Hi Soumalya, the mention of Auditing in this doc refers to internal auditing for Scylla personnel. We do not support displaying auditing information for customers today inside Scylla Cloud --- ### Page: https://forum.scylladb.com/t/scylladb-monitoring-on-datadog/1221 Title: ScyllaDB Monitoring on datadog - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Monitoring stack details on single instance: Scylla-Monitoring Stack version: 4.4.5 OS: Ubuntu 22.04.2 LTS Prometheus : v2.44.0 Grafana: 9.5.5 datadog agent : v7.50.3 Scylla node details: : OS: Ubuntu 20.04.6 LTS n… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-monitoring-on-datadog/1221 ## Headings Structure: H1: ScyllaDB Monitoring on datadog H3: Related topics ## Main Content: H1: ScyllaDB Monitoring on datadog H3: Related topics Monitoring stack details on single instance: Scylla-Monitoring Stack version: 4.4.5 Scylla node details: : I have allowed required port on security group. Can someone suggest here. The Scylla Monitoring stack holds far too many metrics for the Datadog to handle. Therefore, we are marking a small subset of metrics with a label (which used to be level and now is dd) and using that for the Datadog integration. For the node_exported metrics, we are using the Prometheus relabel config to add the label. I suggest you update to the latest 4.6. There are a few changes we made. If there are more metrics you need, you can edit prometheus/prometheus.yml.template Look for the node_exporter section’s relabel config, it adds a label, you can add more metrics over there. Okay , thanks for information. --- ### Page: https://forum.scylladb.com/t/how-do-we-pass-a-number-to-the-mintimeuuid-cql-function/1223 Title: How do we pass a number to the mintimeuuid() cql function - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I am testing out migrating some code from go to rust. It’s been mostly productive and painless, I haven’t become stuck at any point except for this one. I was wondering if someone could point me in the right direction. W… Language: en Canonical URL: https://forum.scylladb.com/t/how-do-we-pass-a-number-to-the-mintimeuuid-cql-function/1223 ## Headings Structure: H1: How do we pass a number to the mintimeuuid() cql function H3: Related topics ## Main Content: H1: How do we pass a number to the mintimeuuid() cql function H3: Related topics I am testing out migrating some code from go to rust. It’s been mostly productive and painless, I haven’t become stuck at any point except for this one. I was wondering if someone could point me in the right direction. We are selecting for a subset of records in certain timeblock partitions using the mintimeuuid, like this: However, when we try to pass an i64 into mintimeuuid() in the same way we can in clash (and in gocql), we get an error, and I can’t work out the appropriate solution, how do we do this with the ScyllaDB driver? failed: BadQuery(SerializationError(SerializationError(BuiltinSerializationError { rust_name: “(i64, i64, i64)”, kind: ColumnSerializationFailed { name: “arg0(system.mintimeuuid)”, err: SerializationError(BuiltinTypeCheckError { rust_name: “i64”, got: Timestamp, kind: MismatchedType { expected: [BigInt] } }) } }))) I am testing out different ways to build a query that might hint that an i64 parameter is expected. Interestingly, this query (below) seems to make my docker instance restart, which seems odd. Although I might be on a slightly old version (Scylla version 5.4.0~dev-0.20230801.37b548f46365) After updating to the latest 5.4.1, this cal query doesn’t crash ScyllaDB, but it still fails: failed: BadQuery(SerializationError(SerializationError(BuiltinSerializationError { rust_name: "(i64, i64, i64)", kind: ColumnSerializationFailed { name: "arg0(system.mintimeuuid)", err: SerializationError(BuiltinTypeCheckError { rust_name: "i64", got: Timestamp, kind: MismatchedType { expected: [BigInt] } }) } }))) Wait, nope, I can still crash the ScyllaDB:lastest (refreshed a few days ago) by simply typing this query inside the docker terminal: And for completeness, it’s still possible to do this query inside the docker terminal but not inside the rust driver. The rust driver seems not able to pass a bigint through to the scylladb instance: For now the only workaround I can find is to pass the i32/int values inside the query string rather than as a query parameter: @Jay please open an issue in the ScyllaDB bug tracker, with the reproducer for the crash. We will want to look into this and fix it. Regarding Rust driver aspect: The error would be more readable if you used normal print, not debug print. The error says that you passed an i64 Rust values to a place that serializes to Timestamp CQL type. This is not possible, you must pass one of supported types: You can find this information in documentation: --- ### Page: https://forum.scylladb.com/t/user-defined-queries-are-not-working-in-scylladb-enterprise-cql-shell-and-through-an-error-i-mentioned-in-description/1224 Title: User defined Queries are not working in ScyllaDb Enterprise CQL shell and through an error I mentioned in description - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: CREATE OR REPLACE FUNCTION div(dividend double, divisor double) RETURNS NULL ON NULL INPUT RETURNS double LANGUAGE LUA AS 'return dividend/divisor;'; InvalidRequest: Error from server: code=2200 [Invalid query] … Language: en Canonical URL: https://forum.scylladb.com/t/user-defined-queries-are-not-working-in-scylladb-enterprise-cql-shell-and-through-an-error-i-mentioned-in-description/1224 ## Headings Structure: H1: User defined Queries are not working in ScyllaDb Enterprise CQL shell and through an error I mentioned in description H3: Related topics ## Main Content: H1: User defined Queries are not working in ScyllaDb Enterprise CQL shell and through an error I mentioned in description H3: Related topics I tried to change the Scylla.yaml file but enable_user_defined_functions this property was not there . So I go to official Documentation : Functions | ScyllaDB Docs What scylla version are you using? I think that recently this feature moved from “experimental” to “preview” mode (new mode). Not sure though if it should be enabled by default or not. could be an issue. Latest scylla enterprise 2023.1 version I am using in UBUNTU 22.04 In this release it’s still experimental. You should be able to use it in experimental flags. have you tried it? No I am not able to enable this, I had tried in scylla.yaml file by uncomment “experimental_feature” field and set “udf” to this as you can see in above screenshot, But still not working. Could you please help me how can I achieve this? Did you spell udf correctly in the config file? Yeah I just use the commented part as you can see in screenshot . I used it like this in config file: experimental_features: udf How to mention in config file can you please tell me if I am wrong? There are two separate options to enable UDF, and both have to be enabled for UDF to work, see below: Yeah, I was able to execute this query, but audit logs are not populating. Looks like the CREATE OR REPLACE FUNCTION statement does not trigger an audit event. Only statement to which we explicitly added audit events will log a corresponding audit event. The statement that are supported w.r.t. audit is driven by customer demand for these (given that this is an enterprise-only feature). Does that means audit logs for Function will not be generated?? Functions are executed as part of other statements, so we will generate audit logs depending on whether said statement generates audit logs or not. I did not get your point. Suppose I run this below command, but it did not generate audit-log. Yes, I don’t think we added audit support for this statement, so it will not generate audit log. But these should be also audited it could be an issue, Only USER DEFINED QUERIES are not generating audit logs. File an issue in scylla repository and someone will look into fixing it if it gets prioritzed. Where can I do this in GitHub?? Github Issue was created - https://github.com/scylladb/scylla-enterprise/issues/3863 I am not able to see anything in the mentioned GitHub issue link which you have provided. It is showing page not found. If you wish to track progress of this issue, please contact our support for enterprise clients. --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-2-14/1225 Title: [RELEASE] ScyllaDB 5.2.14 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.2.14, a bugfix release of the ScyllaDB 5.2 stable branch. ScyllaDB Open Source 5.2.12, like all past and future 5.x.y releases, is backward compatible, and supports roll… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-2-14/1225 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.2.14 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.2.14 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.2.14, a bugfix release of the ScyllaDB 5.2 stable branch. ScyllaDB Open Source 5.2.12, like all past and future 5.x.y releases, is backward compatible, and supports rolling upgrades. Note that the latest ScyllaDB Open Source stable release is 5.4 and you are encouraged to upgrade to it. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/does-scylla-implement-the-sstablerepairedset-tool/1226 Title: Does Scylla implement the `sstablerepairedset` tool? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I saw a tool called sstablerepairedset in cassandra, which can mark different SSTable files with a mark indicating whether they have been repaired or not. But I haven’t seen any related implementation in Scylla, except… Language: en Canonical URL: https://forum.scylladb.com/t/does-scylla-implement-the-sstablerepairedset-tool/1226 ## Headings Structure: H1: Does Scylla implement the `sstablerepairedset` tool? H3: Related topics ## Main Content: H1: Does Scylla implement the `sstablerepairedset` tool? H3: Related topics I saw a tool called sstablerepairedset in cassandra, which can mark different SSTable files with a mark indicating whether they have been repaired or not. But I haven’t seen any related implementation in Scylla, except in Java submadel. I tried running it, but it reported the following error: Can scylla also repair only SSTables marked as unrepaired like cassandra. Thanks! No, we don’t support this tool or this use case. What are you trying to achieve with that? sstablerepairdset is used in conjunction with Cassandra’s incremental repair functionality, which we did not implement in ScyllaDB (nor do we have any plans to implement it). The flag this tool sets is ignored by ScyllaDB and therefore using the tool itself makes no sense. --- ### Page: https://forum.scylladb.com/t/scylladb-datadog-integration-issue/1229 Title: ScyllaDB Datadog integration Issue - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: We are getting the issue post integrating the scylladb on datadog. Not getting the metric value from node exporter only the names. What could be the issue? Current monitoring stack version: 4.4.5 , do we required monit… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-datadog-integration-issue/1229 ## Headings Structure: H1: ScyllaDB Datadog integration Issue H3: Related topics ## Main Content: H1: ScyllaDB Datadog integration Issue H3: Related topics We are getting the issue post integrating the scylladb on datadog. Not getting the metric value from node exporter only the names. What could be the issue? Current monitoring stack version: 4.4.5 , do we required monitoring stack 5.x for this integration. The latest monitoring version is 4.6.1; though it’s unnecessary, upgrading to the latest version is a good idea. Datadog is limited in the number of metrics it can handle compared to the amount collected by Prometheus. For that reason, we mark some of the metrics (both from ScyllaDB and node_exporter) with a specific tag. You can see that configuration under: prometheus/prometheus.yml.template in the job_name: node_exporter Maybe there are metrics you need that aren’t there. You can add them manually, open an issue for Scylla monitoring, and I’ll add them in the next release. Is next release done ? Have these things added in that release. Or still it is in planning. --- ### Page: https://forum.scylladb.com/t/what-is-the-impact-of-starting-the-data-integrity-check-of-the-sstable-file-on-the-cluster/1231 Title: What is the impact of starting the data integrity check of the SStable file on the cluster? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I see that parameter enable_sstable_data_integrity_check can control whether to enable integrity checks on SSTables. Its default value is false, and turning it on will have an impact on performance. I saw its implementa… Language: en Canonical URL: https://forum.scylladb.com/t/what-is-the-impact-of-starting-the-data-integrity-check-of-the-sstable-file-on-the-cluster/1231 ## Headings Structure: H1: What is the impact of starting the data integrity check of the SStable file on the cluster? H3: Related topics ## Main Content: H1: What is the impact of starting the data integrity check of the SStable file on the cluster? H3: Related topics I see that parameter enable_sstable_data_integrity_check can control whether to enable integrity checks on SSTables. Its default value is false, and turning it on will have an impact on performance. I saw its implementation in the commit of sstables: introduce file interposer for integrity check. Its check occurs before writing the SSTable file, after writing the SSTable file, and after reading the SSTable file. It should be inferred from this that it has had an impact on write performance. May I ask if it has any other impacts? Thanks! It has a decent impact on performance. It’s not recommended to enable it by default for a production cluster / node. In case of a severe issue, a ScyllaDB developer may suggest to enable temporarily this in order to do some validations to the sstable. If you want to have sstables integrity check once in a while, you can use “nodetool scrub” command with its options. The best way to get integrity checking for your data on disk is to enable Sstable compression, when creating your tables (can be enabled later via alter table as well). Sstable compression stores checksums next to the data in the Sstable data file, and these checksums are checked every time the data is read. Tables have compression on by default, so unless you disabled compression for your tables when creating them, you already have integrity checking. If you want to have sstables integrity check once in a while, you can use “nodetool scrub” command with its options. Currently, this only checks checksums on compressed sstables We have plans to change this in the near future, such that nodetool scrub --mode=VALIDATE can be used to force a checksum check on all Sstables, compressed or not. Hi! I saw the addition of function scylla sstable scrub in 5.4, what is the difference between it and function nodetool scrub? Currently, this only checks checksums on compressed sstables And what is the meaning of compressed sstables? Thanks! I saw the addition of function scylla sstable scrub in 5.4, what is the difference between it and function nodetool scrub? scylla-sstable scrub is just an off-line version of nodetool scrub, meaning that you don’t need a running ScyllaDB process to do the scrub. The use-case this was developed for is when sstables in a backup are found to be corrupt and the backup cannot be restored because ScyllaDB refuses the sstables. In this case, the sstables can be fixed, before loading them to ScyllaDB. And what is the meaning of compressed sstables? Simply, sstables for tables, for which sstable-compression is enabled. You can tell whether an sstable is compressed or not, by checking whether the CompressionInfo.db component file exists or not. If you want to have sstables integrity check once in a while, you can use “nodetool scrub” command with its options. Hi, I would like to know what impact my long-term continuous scheduling of nodetool scrub will have on the cluster? Does it have the same impact on the cluster as compaction? Because I see in the code that it is a task scheduled by CompactionManager. Yes, scrub is just a type of compaction and it will have the same effect as an ongoing compaction. How do you use Scrub in production to ensure data security? If I schedule repairs during the cycle, is it still necessary to schedule Scrub? @bo_li please start a new discussion for follow up questions. That would make it easier for other community members to find. Currently we only use scrub we we identify bad sstables, that scrub is able to fix. We also use nodetool scrub --mode=VALIDATE, to find bad sstables. Currently the scrub is quite limited in what it can fix and identify. But we have plans to develop scrub to be more capable, in detecting bad sstables and as a general tool to make sure bad sstables are identified and qurantined. ok, I have started another topic here. --- ### Page: https://forum.scylladb.com/t/is-there-no-way-to-close-raft-after-opening-it/1232 Title: Is there no way to close Raft after opening it? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Once enabled, Raft cannot be disabled on your cluster. The cluster nodes will fail to restart if you remove the Raft feature. Why can’t it close? If I really want to close Raft now, can I do so by deleting some table i… Language: en Canonical URL: https://forum.scylladb.com/t/is-there-no-way-to-close-raft-after-opening-it/1232 ## Headings Structure: H1: Is there no way to close Raft after opening it? H3: Related topics ## Main Content: H1: Is there no way to close Raft after opening it? H3: Related topics Once enabled, Raft cannot be disabled on your cluster. The cluster nodes will fail to restart if you remove the Raft feature. If I really want to close Raft now, can I do so by deleting some table information from the system’s record Raft? Thanks! I don’t think so and it was never tested. Why would you want to turn it off? We are in the process of using Raft for more and more things. Starting with ScyllaDB 5.2, Raft is mandatory to use, especially if you want to benefit from the improvements we developed on top of it. So I don’t recommend trying to turn it off, instead please report any problems you experience related to it. --- ### Page: https://forum.scylladb.com/t/which-snitch-or-replication-strategy-should-i-use/1233 Title: Which snitch or replication strategy should I use? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Can I use SimpleSnitch and SimpleStrategy if I have a single datacenter? Language: en Canonical URL: https://forum.scylladb.com/t/which-snitch-or-replication-strategy-should-i-use/1233 ## Headings Structure: H1: Which snitch or replication strategy should I use? H3: Related topics ## Main Content: H1: Which snitch or replication strategy should I use? H3: Related topics Can I use SimpleSnitch and SimpleStrategy if I have a single datacenter? One should always use “NetworkTopologyStrategy” as it saves you from troubles later. I think we even wants/wanted to deprecate SimpleStrategy. Regarding Snitch, it depends where you run your cluster. e.g. AWS, GCP and Azure have special snitches. --- ### Page: https://forum.scylladb.com/t/updated-workload-prioritization-lesson-on-scylladb-university/1234 Title: Updated Workload Prioritization Lesson on ScyllaDB University - University and Training - ScyllaDB Community NoSQL Forum Meta Description: We recently updated the Workload Prioritization and Attributes lesson on ScyllaDB University. Workload prioritization is a feature that allows running Online Transaction Processing (OLTP) and Online Analytical Processi… Language: en Canonical URL: https://forum.scylladb.com/t/updated-workload-prioritization-lesson-on-scylladb-university/1234 ## Headings Structure: H1: Updated Workload Prioritization Lesson on ScyllaDB University H3: Related topics ## Main Content: H1: Updated Workload Prioritization Lesson on ScyllaDB University H3: Related topics We recently updated the Workload Prioritization and Attributes lesson on ScyllaDB University. Workload prioritization is a feature that allows running Online Transaction Processing (OLTP) and Online Analytical Processing (OLAP) workloads on the same cluster. The traditional approach involves segregating these workloads. In the lesson, you’ll learn about ScyllaDB’s solution to efficiently run both on the same cluster. The Workload Prioritization mechanism allows users to define and prioritize different workloads based on configured shares. This ensures fair distribution of system resources, preventing a single job from monopolizing resources and affecting other tasks. The lesson includes examples of configuring Workload Prioritization, demonstrating its impact on latency and throughput in various scenarios. Some of these features are relevant for the Enterprise version only. The change reflects recent updates to the feature and elaborates on this topic. Do you have any feedback or thoughts? Something else you’d like to see? This would be a great place to discuss the lesson and Workload Prioritization in general. --- ### Page: https://forum.scylladb.com/t/how-do-you-use-query-with-more-than-15-parameters-of-mixed-type/1235 Title: How do you use query() with more than 15 parameters of mixed type? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Unless I misunderstand something, it appears to me that once you go over 15 parameters for a query, we hit a SerializeRow trait problem: | |_____^ the trait `SerializeRow` is not implemented for `(String, &String, &Str… Language: en Canonical URL: https://forum.scylladb.com/t/how-do-you-use-query-with-more-than-15-parameters-of-mixed-type/1235 ## Headings Structure: H1: How do you use query() with more than 15 parameters of mixed type? H3: Related topics ## Main Content: H1: How do you use query() with more than 15 parameters of mixed type? H3: Related topics Unless I misunderstand something, it appears to me that once you go over 15 parameters for a query, we hit a SerializeRow trait problem: If I remove a parameter, any parameter, then the query works. Is my understanding correct? This is implemented using a trait that has a maximum number of options? (Full example code below). How do we solve this? I have removed as many parameters as I can to try and get this to work, but this is the minimum set I need to create a person record. I guess I could break this query up into two separate database calls, but two database round trips seems less than ideal? Check out Query values | ScyllaDB Docs. You can use any type that implements SerializeRow as values for a query. You can use the #[derive(SerializeRow)] macro to make a struct serializable, if the field names match the table structure. I swear I looked in the documentation and didn’t notice that. I don’t know how I missed it. Thanks! (Although I did spend a lot more time clicking through the example code in the repo than I did reading the html pages.) I’d like to add one note to this thread: if you want to see documentation for SerializeRow macro it is temporarily only available in scylla-cql crate docs, not in scylla crate docs. This is because of our mistake and will be fixed in next release (so in 1-2 days probably). Here’s the current link: SerializeRow in scylla_cql::macros - Rust --- ### Page: https://forum.scylladb.com/t/how-to-ensure-a-batch-query-reaches-the-correct-partition/1237 Title: How to ensure a Batch query reaches the correct partition? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi all. Completely new to Cassanda/Scylla type databases. Here’s my question: Let’s say I’m trying to model a simple event-sourcing database using ScyllaDB: create table test_keyspace.event_streams ( stream_id int,… Language: en Canonical URL: https://forum.scylladb.com/t/how-to-ensure-a-batch-query-reaches-the-correct-partition/1237 ## Headings Structure: H1: How to ensure a Batch query reaches the correct partition? H3: Related topics ## Main Content: H1: How to ensure a Batch query reaches the correct partition? H3: Related topics Hi all. Completely new to Cassanda/Scylla type databases. Here’s my question: Let’s say I’m trying to model a simple event-sourcing database using ScyllaDB: So partitioning is based on stream_id. I want to perform batch event inserts, where all events have the same stream_id, and I want this query to be sent to the right nodes (the ones matching the stream_id partition), how can I ensure that? Can I ensure the query above reaches only the nodes that are part of the partition of stream_id = 42? Also, can anyone help me understand LOGGED vs UNLOGGED better? ScyllaDB takes care of this automatically, you don’t have to do anything for a query to be only sent to the nodes it should be sent to., Note that your driver will send the query to a single ScyllaDB node (the coordinator – this is chosen according to the session’s load-balancing policy), the coordinator will take care of only contacting replicas which are affected by the query. BTW why do you want to use a BATCH? In general, individual queries are better. Logged batches are written into a system table, before they are executed. Upon failure, they will be retried until they succeed. To add to this from Rust driver perspective (I see that the question has a rust tag): Rust driver will base it’s choice of coordinator node for the batch on the first statement of it. This is a heuristic to get some form of shard-awareness for batches. This heuristic is perfect for your scenario @bayov , assuming you only use prepared statements in batches - which you definitely should! Unprepared statements containing values in batches have a massive performance penalty in Rust driver. Hi! thanks for the response! I’m trying to model event sourcing using ScyllaDB. I want persisted events to represent immutable facts in my system. Sometimes, a single action in the system will generate more than one event (probably always less than a dozen though). I want to atomically commit them to the DB to ensure that either all of these events become immutable facts, or none of them do. My desired transaction boundary is a single stream ID, so I won’t try to achieve atomic writes to multiple different stream IDs, only to one. And a stream ID should always exist in a single partition, so it should be possible to guarantee this transaction boundary. Does that make sense? What is the use-case for UNLOGGED mode? Is it simply to save network bandwidth and perform multiple queries together, but without any transactional-guarantees? Yes, BATCH inserts to a single partition are indeed applied as a single mutation. A single mutation is applied atomically, it either succeed or fails, there is no in-between. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-36-2024-01-27/1238 Title: Last week in scylla-cluster-tests.git master (issue #36; 2024-01-27) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This brief report highlights some significant commits to the scylla-cluster-tests.git master from the past week, covering commits in the d311c6c1…cf79ff75 range. During this period, 24 non-merge commits were made by 5 s… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-36-2024-01-27/1238 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #36; 2024-01-27) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #36; 2024-01-27) H3: Related topics This brief report highlights some significant commits to the scylla-cluster-tests.git master from the past week, covering commits in the d311c6c1…cf79ff75 range. During this period, 24 non-merge commits were made by 5 software engineers in test and 1 software engineer, with some notable changes: We have resolved an issue regarding the restoration of monitoring when ports were occupied. Now, in such cases, a random port is chosen. The upgrade of scylla-bench with gocql support for tablets was reverted due to a scylla-bench issue. Support for RackAwareRoundRobinPolicy, introduced in scylladb/java-driver#227, was added to the cassandra-stress thread. This is automatically used when we set up multiple racks, allowing the loader to spread across multiple AZs. A new test was added to fully utilize this feature. The network interface configuration has been refined to better test Scylla networking with multiple NIC/IP combinations. Now, it’s easier to define which Scylla endpoint is mapped to which interface and IP. The subject of the cloud resources usage email has been reverted back to static. Please ensure you correctly assign the “important” label when your name is listed, so you don’t overlook any lingering resources. Metrics from the sct-runner revealed an issue with high CPU usage that could lead to a ‘Bad packet length’ problem. To combat this, we updated the sct-runner instance and began to trim db log lines longer than 5K to prevent searches and regexes from taking too much time. Finally, the new tablets_initial_scale_factor scylla.yaml option is now supported and can be used in the append_scylla_yaml test configuration. Stay tuned for the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/feedback-and-ideas-about-the-scylladb-community/1239 Title: Feedback and Ideas about the ScyllaDB Community - Database Community - ScyllaDB Community NoSQL Forum Meta Description: I recently wrote a blog post about the ScyllaDB Community. I see the community’s involvement in shaping the future of ScyllaDB as highly important. This Forum and the Slack Channel are key communication channels for te… Language: en Canonical URL: https://forum.scylladb.com/t/feedback-and-ideas-about-the-scylladb-community/1239 ## Headings Structure: H1: Feedback and Ideas about the ScyllaDB Community H3: Related topics ## Main Content: H1: Feedback and Ideas about the ScyllaDB Community H3: Related topics I recently wrote a blog post about the ScyllaDB Community. I see the community’s involvement in shaping the future of ScyllaDB as highly important. This Forum and the Slack Channel are key communication channels for technical discussions, showcasing use cases, and real-time interactions. This would be a great place to discuss the blog post and how we can help the community grow. Any thoughts? --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-215-2024-01-28/1240 Title: Last week in scylladb.git master (issue #215; 2024-01-28) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the b1ba904c49…fe3bc00045 range are covered. There were 149 non-merge commits from 23 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-215-2024-01-28/1240 ## Headings Structure: H1: Last week in scylladb.git master (issue #215; 2024-01-28) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #215; 2024-01-28) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the b1ba904c49…fe3bc00045 range are covered. There were 149 non-merge commits from 23 authors in that period. Some notable commits: There is a new operating mode, maintenance mode. In this mode the node does not communicate with clients or other nodes, and only listens on the maintenance socket and the REST API. It can be used to fix damaged nodes. Tablets now support the removenode and replacenode operations. When replaying commitlog mutations after a crash, we now ignore mutations to tablets that were migrated away from this node, preventing data resurrection. Writing sstables no longer blocks migrations of unrelated tablets. The schema of the system.tablets table no longer includes the keyspace name in the partition key, as it is redundant. When using consistent topology, we now drain requests from the decommissioning node, to prevent it continuing to use an outdated topology. In experimental consistent topology, we no longer remove an endpoint if it is being replaced by another node with the same IP address. CQL tracing now records two additional fields: the user who initiated the query, and the reader concurrency semaphore that is in charge of admitting the query on the replica. We now invalidate materialized view prepared statements when the schema of the base table changes. This prevents a stale prepared statement from returning incorrect results. CQL multi-column comparisons now have better NULL checks. The DROP TYPE IF EXISTS statement now works if the type’s keyspace doesn’t exist. Metrics now support sending Prometheus histograms as native histograms, with a sufficiently new version of Prometheus. The native histograms consume less bandwidth than the text-based histograms. The scylla sstable tool now supports loading the schema of materialized views and indexes. The hint manager is now started later in the boot process, until we have better information about other nodes in the cluster. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/can-we-make-an-enum-serializable-so-it-can-go-to-and-from-a-cqlvalue-in-rust/1242 Title: Can we make an enum serializable so it can go to and from a CqlValue (in rust)? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I am exploring how to use FromRow and SerializeCql which seems to support all of the basic types. I would like to support using enums (for example see the PersonStatus below. Is it possible to coerce an enum to work wi… Language: en Canonical URL: https://forum.scylladb.com/t/can-we-make-an-enum-serializable-so-it-can-go-to-and-from-a-cqlvalue-in-rust/1242 ## Headings Structure: H1: Can we make an enum serializable so it can go to and from a CqlValue (in rust)? H3: value.rs - source H3: Related topics ## Main Content: H1: Can we make an enum serializable so it can go to and from a CqlValue (in rust)? H3: value.rs - source H3: Related topics I am exploring how to use FromRow and SerializeCql which seems to support all of the basic types. I would like to support using enums (for example see the PersonStatus below. Is it possible to coerce an enum to work within FromRow and SerializeCql, are there any examples anywhere? What is the alternative if using enum is not possible? Would we just implement a full “deserialize struct” and “serialize struct” trait to translate and spit out the values for the scylla driver? Looking at the driver code for how how serializing works gives me some hint that perhaps it might be possible to implement a serialization trait for the rust driver, but I’m not sure. Could an enum in the end users code go ahead and implement something like: Source of the Rust file `src/types/serialize/value.rs`. You won’t be able to accomplish this using derive macros and I don’t see how they could be extended to work on arbitrary enums. You’ll need to do this manually. I assume your enum won’t store any data in any variants. First you need to decide what is the type you want to use in database for it (e.g. Int, corresponding to i32 Rust type). Then you’ll need to write a SerializeCql and FromCqlVal implementations for your enum. Those implementations can be simple and just delegate to i32 implementation. Here’s the full example from your code. I removed ContactType field from Contact because you didn’t provide definition for it. I used num-traits and num-derive crates to convert from i32 to PersonStatus. I also derived Clone and Copy for PersonStatus to be able to cast it to i32. I opened an issue in Rust Driver to extend derive macros for your use case (enums without data): Serialization / Deserialization derive macros should work for enums without data · Issue #923 · scylladb/scylla-rust-driver · GitHub --- ### Page: https://forum.scylladb.com/t/getting-rclone-issues-while-running-scylla-manager-commands/1245 Title: Getting rclone issues while running scylla-manager commands - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello, We are getting rclone exceptions while running scylla manager commands in GCP VM’s. Exceptions : root@scyl-deng-xp-ab1v1-prd-ase1:/home/chaitanyatondlekar# scylla-manager-agent download-files -L gcs:gcs-deng-ab… Language: en Canonical URL: https://forum.scylladb.com/t/getting-rclone-issues-while-running-scylla-manager-commands/1245 ## Headings Structure: H1: Getting rclone issues while running scylla-manager commands H3: Related topics ## Main Content: H1: Getting rclone issues while running scylla-manager commands H3: Related topics Hello, We are getting rclone exceptions while running scylla manager commands in GCP VM’s. Exceptions : @Michal_Leszczynski can you please assist? Hello, looks like some config mismatch, see rclone.awsRegionFromMetadataAPI part. Perhaps rcone thinks it’s AWS cluster not GCP? this is a know Scylla Manager issue (see `scylla-manager-agent check-location` command fails to handle region · Issue #3691 · scylladb/scylla-manager · GitHub) that will be fixed in next releases. When there is no region specified in agent’s AWS config, Scylla Manager tries to fetch this information and fails (because of running on GCP). This indeed prints an error log, but it shouldn’t have an effect on the actual procedure. So does SM work correctly despite this error in logs? that will be fixed in next releases. Exact which scylla manager release it will be fixed. ? We have tried with 3.0 and 3.1 both and both had same exceptions. This is impacting actual procedure while doing restoration through ansible. I am following this : Restore | ScyllaDB Docs I hope it will be a part of 3.2.6 but I don’t know if it will be backported to older 3.0, 3.1. Generally speaking, the advised way of restoring data (present since SM 3.1) is via SM restore task (Restore | ScyllaDB Docs). --- ### Page: https://forum.scylladb.com/t/how-to-start-scyalladb-in-docker-by-enabling-passwordauthenticator/1248 Title: How to start scyalladb in docker by enabling PasswordAuthenticator - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello, I am new to scyalladb and I dont see documentation on starting scyalladb on docker with PasswordAuthenticator with username and password as variables. Is there a way to do that ? Language: en Canonical URL: https://forum.scylladb.com/t/how-to-start-scyalladb-in-docker-by-enabling-passwordauthenticator/1248 ## Headings Structure: H1: How to start scyalladb in docker by enabling PasswordAuthenticator H3: Related topics ## Main Content: H1: How to start scyalladb in docker by enabling PasswordAuthenticator H3: Related topics I am new to scyalladb and I dont see documentation on starting scyalladb on docker with PasswordAuthenticator with username and password as variables. Is there a way to do that ? Appreciate any response on this issue ? You can follow the steps here: Enable Authentication | ScyllaDB Docs Note that there is even separate instructions for a Docker deployment. Thank you @Botond_Denes I have seen that. The solution for docker is to use volume where you can provide the config.yaml to overwrite. I was more interested in environment variables, all I need is to enable authentication and I dont need to provide the whole set of config. --- ### Page: https://forum.scylladb.com/t/new-migration-lesson-on-scylladb-university/1250 Title: New Migration Lesson on ScyllaDB University - University and Training - ScyllaDB Community NoSQL Forum Meta Description: The Migrating to ScyllaDB lesson on ScyllaDB University has been updated. Many users ask about migration and what’s the best way to do it. The answer depends on your use case, existing database, and specific requirement… Language: en Canonical URL: https://forum.scylladb.com/t/new-migration-lesson-on-scylladb-university/1250 ## Headings Structure: H1: New Migration Lesson on ScyllaDB University H3: Related topics ## Main Content: H1: New Migration Lesson on ScyllaDB University H3: Related topics The Migrating to ScyllaDB lesson on ScyllaDB University has been updated. Many users ask about migration and what’s the best way to do it. The answer depends on your use case, existing database, and specific requirements. There are different migration methods. The lesson covers common migration scenarios including migrating from DynamoDB to ScyllaDB and from Cassandra to ScyllaDB. It includes typical use cases, best practices, practical tips, and a hands-on example. Any questions or input? This would be a great place to discuss the lesson and migration in general. --- ### Page: https://forum.scylladb.com/t/i-found-a-bug-in-scylladb-enterprising-auditing-where-for-create-table-drop-table-these-types-of-queries-sometimes-generating-duplicate-audit-logs/1251 Title: I found a Bug in Scylladb enterprising auditing, where for CREATE TABLE, DROP TABLE, these types of queries sometimes generating duplicate audit logs - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Executed Queries: CREATE TABLE loads(ID int PRIMARY KEY); DROP TABLE loads; Audit Logs: 2024-01-30T11:28:57.577649+00:00 ip-172-31-63-35.ec2.internal scylla-audit: "172.31.63.35:0", "DDL", "ONE", "loads", "mykeyspace"… Language: en Canonical URL: https://forum.scylladb.com/t/i-found-a-bug-in-scylladb-enterprising-auditing-where-for-create-table-drop-table-these-types-of-queries-sometimes-generating-duplicate-audit-logs/1251 ## Headings Structure: H1: I found a Bug in Scylladb enterprising auditing, where for CREATE TABLE, DROP TABLE, these types of queries sometimes generating duplicate audit logs H3: Related topics ## Main Content: H1: I found a Bug in Scylladb enterprising auditing, where for CREATE TABLE, DROP TABLE, these types of queries sometimes generating duplicate audit logs H3: Related topics As You can see in above audit logs only difference is timestamp milliseconds, if they are duplicating then why changing timestamps. Hello, plase make sure you are looking at the correct timestamp, it should be event_time column in the audit table. What you posted doesn’t look like a direct SELECT result. How does your SELECT query look like? I did not mention SELECT query here. Only CREATE and DROP TABLE. Actually, I am generating audit log in syslog not in audit table and in syslog format is like this. Below is audit setting: And I change the time format template in rsyslog.conf file for better precise timestamp. Ok, I see. Are you sure your are not calling CREATE and DROP in a loop? Unfortunatelly this time is not taken directly from query but generated at the time audit log is written. This can cause order to be slightly different, so in theory it can mean that there was a following sequence: CREATE TABLE loads … DROP TABLE loads… CREATE TABLE loads … DROP TABLE loads… I am executing these queries inside CQL shell one by one and then inside scylla-audit.log file duplicate logs are populated like this as I mentioned. I only execute these queries once. I cannot reproduce it locally by working with the latest enterprise release - which release do you use? Also, can you please send us your rsyslog.conf file? I was using the default one. Also, where do you have this scylla-audit.log and is this the file you are reading? I was just looking at the logs with “journalctl” I suspect there may be some configuration issue, not necessarily on scylla side. I am using latest Scylladb Enterprise version in AWS EC2 Linux UBUNTU 22.04. This below template is used for precise time in audit logs. These belwo three-line template will generate Scylla-audit.log file in var/log/hostname directory. I am facing this issue even also for Successful Login using cqlsh u -cassandra p -cassandra. In this case also duplicate audit-logs are generated on successful login in CQLSHELL. Thanks. I can confirm that the issue reproduces on our end (though a bit differently). I’ll open a github issue to have it fixed and will link it here. Please could you also look into this issue also [User defined Queries are not working in ScyllaDb Enterprise CQL shell and through an error I mentioned in description] Here’s the Github Issue for the bug discussed in this conversation - https://github.com/scylladb/scylla-enterprise/issues/3861 For the other bug I filed a separate bug report and linked it in the cited forum thread. I am not able to see anything in the mentioned GitHub issue link which you have provided. It is showing that page not found If you wish to track progress of this issue, please contact our enterprise client’s support. --- ### Page: https://forum.scylladb.com/t/release-scylladb-cloud-30-january-2024/1252 Title: [RELEASE] ScyllaDB Cloud - 30 January 2024 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: We are happy to announce a new update to the ScyllaDB Cloud service. AWS Transit Gateway Integration Now Available: We are excited to announce the availability of our new AWS Transit Gateway feature, offering you more … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-cloud-30-january-2024/1252 ## Headings Structure: H1: [RELEASE] ScyllaDB Cloud - 30 January 2024 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Cloud - 30 January 2024 H3: Related topics We are happy to announce a new update to the ScyllaDB Cloud service. AWS Transit Gateway Integration Now Available: We are excited to announce the availability of our new AWS Transit Gateway feature, offering you more flexibility in your network connectivity choices as part of our Premium tier offering. Our new AWS Transit Gateway integration streamlines network connections, allowing for the easy linking of multiple networks via a single transit gateway. With AWS Transit Gateway, any network connected to the transit gateway can communicate seamlessly with others also linked to it, simplifying network management and connectivity. This premium feature enhances the capabilities available to our users, ensuring robust and efficient network handling. However, if your needs are better met by direct, individual links, VPC Peering is still available and fully supported. This means you have the choice to add either Transit Gateway or VPC Peering connections, depending on your specific requirements. For Bring Your Own Account (BYOA) customers looking to leverage the Transit Gateway feature, it’s important to contact Support for necessary updates to your IAM policy. This ensures that the integration functions smoothly within your existing infrastructure. Updated UI for Cluster Connections: In our latest UI update, after creating a cluster, you will now find a “Connections” tab on the Cluster Details page. This replaces the previously labeled “VPC Peering” tab. The new “Connections” tab is your central place to view all connections to your cluster. From here, you can easily add a new connection and select the type of connection you need, whether it’s VPC Peering or Transit Gateway. This streamlined approach makes managing your network connections more intuitive and efficient. To add a new Transit Gateway connection, follow our guide here. To migrate an existing VPC Peering connection to a Transit Gateway connection, follow our migration guide here. For more information, see our Cluster Connections docs. --- ### Page: https://forum.scylladb.com/t/how-do-i-use-query-safely/1254 Title: How do I use `query()` safely? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: In a Rust codebase I work with, we use scylladb::Session::query() to retrieve some fixed-size results. However, sometimes we get the error: Tombstones processed by unpaged query exceeds limit of 10000 (configured via q… Language: en Canonical URL: https://forum.scylladb.com/t/how-do-i-use-query-safely/1254 ## Headings Structure: H1: How do I use `query()` safely? H3: Related topics ## Main Content: H1: How do I use `query()` safely? H3: Related topics In a Rust codebase I work with, we use scylladb::Session::query() to retrieve some fixed-size results. However, sometimes we get the error: I’ve investigated this error and it seems that: As a caller of the API, I can’t know statically how many tombstones my query will traverse. Does this mean that the only way to avoid triggering this error is to always use paged queries? If that’s true, I don’t see what the use case for unpaged queries is — doesn’t this imply that the user must always use paged queries, regardless of the size of their results? Hello, you can check this article on how to reduce number of tombstones: How to flush old tombstones from a table | ScyllaDB Docs. You can statistically assess how many tombstones there are based on your write load and compaction rate. --- ### Page: https://forum.scylladb.com/t/managing-ip-addresses-of-kubernetes-worker-nodes/1256 Title: Managing IP addresses of Kubernetes worker nodes - Kubernetes Operator - ScyllaDB Community NoSQL Forum Meta Description: Could you please explain how to manage changes in the IP addresses of Kubernetes worker nodes? The monitoring stack, connectors, and other clients rely on these IPs, but they need to be updated every time the node pool u… Language: en Canonical URL: https://forum.scylladb.com/t/managing-ip-addresses-of-kubernetes-worker-nodes/1256 ## Headings Structure: H1: Managing IP addresses of Kubernetes worker nodes H3: Related topics ## Main Content: H1: Managing IP addresses of Kubernetes worker nodes H3: Related topics Could you please explain how to manage changes in the IP addresses of Kubernetes worker nodes? The monitoring stack, connectors, and other clients rely on these IPs, but they need to be updated every time the node pool undergoes maintenance. With hostNetworking and ephemeral IP addresses, you need a script when your client starts that scrapes those from the kubernetes API. I’d recommend not using hostNetworking mode though. With Service IPs, or Pod IPs this becomes more managable to scrape, but even better is to use the IP for service -client that proxies to all nodes, and the client will discover the other nodes on its own. Thanks. How is this service configured and can it be exposed as load balancer? I found the documentation about exposing ScyllaCluster, but the client service is not mentioned. --- ### Page: https://forum.scylladb.com/t/does-scylla-reorder-columns-in-newly-created-tables/1257 Title: Does Scylla reorder columns in newly created tables? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I am creating tracking.tracking_data as: CREATE TABLE tracking_data ( first_name text, last_name text, timestamp timestamp, location varchar, speed double, heat double, telepathy_powers int, primary key((first_n… Language: en Canonical URL: https://forum.scylladb.com/t/does-scylla-reorder-columns-in-newly-created-tables/1257 ## Headings Structure: H1: Does Scylla reorder columns in newly created tables? H3: Related topics ## Main Content: H1: Does Scylla reorder columns in newly created tables? H3: Related topics I am creating tracking.tracking_data as: CREATE TABLE tracking_data ( first_name text, last_name text, timestamp timestamp, location varchar, speed double, heat double, telepathy_powers int, primary key((first_name, last_name), timestamp)) WITH CLUSTERING ORDER BY (timestamp DESC) AND COMPACTION = {‘class’: ‘TimeWindowCompactionStrategy’, ‘base_time_seconds’: 3600, ‘max_sstable_age_days’: 1}; But when I describe the table, the columns are in a different sequence: cqlsh:tracking> describe keyspace tracking; CREATE KEYSPACE tracking WITH replication = {‘class’: ‘NetworkTopologyStrategy’, ‘DC1’: ‘3’} AND durable_writes = true; CREATE TABLE tracking.tracking_data ( first_name text, last_name text, timestamp timestamp, heat double, location text, speed double, telepathy_powers int, PRIMARY KEY ((first_name, last_name), timestamp) ) WITH CLUSTERING ORDER BY (timestamp DESC) AND bloom_filter_fp_chance = 0.01 AND caching = {‘keys’: ‘ALL’, ‘rows_per_partition’: ‘ALL’} AND comment = ‘’ AND compaction = {‘base_time_seconds’: ‘3600’, ‘class’: ‘TimeWindowCompactionStrategy’, ‘max_sstable_age_days’: ‘1’} AND compression = {‘sstable_compression’: ‘org.apache.cassandra.io.compress.LZ4Compressor’} AND crc_check_chance = 1.0 AND dclocal_read_repair_chance = 0.0 AND default_time_to_live = 0 AND gc_grace_seconds = 864000 AND max_index_interval = 2048 AND memtable_flush_period_in_ms = 0 AND min_index_interval = 128 AND read_repair_chance = 0.0 AND speculative_retry = ‘99.0PERCENTILE’; Good question. AFAIK in ScyllaDB (and Apache Cassandra), there is no guarantee regarding the order of the columns unless they are part of the Primary Key. The logic behind this has to do with the way that data is represented under the hood. This is (at some level) explained in the Basic Data Modeling lesson in Scylla University (mostly in the Importance of Primary Key Selection lesson and the Importance of the Clustering Key lesson). That means you should not rely on a particular order for the storage. Instead, you can define how the data will be presented at the application level. Internally, columns are ordered alphabetically, except primary key columns, those are ordered as they were specified in the initial table definition. Ok, thanks for the clarification. --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2023-1-4/1258 Title: [RELEASE] ScyllaDB Enterprise 2023.1.4 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2023.1.4, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2023.1 LTS Release. 2023.1.4 patch release includes multiple minor bug fixes… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2023-1-4/1258 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2023.1.4 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2023.1.4 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2023.1.4, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2023.1 LTS Release. 2023.1.4 patch release includes multiple minor bug fixes. You are encouraged to upgrade to it in coordination with the ScyllaDB Support team. The following issues are fixed in this release (with an open-source reference, if available): --- ### Page: https://forum.scylladb.com/t/what-is-the-meaning-of-the-content-in-sstable/1264 Title: What is the meaning of the content in SSTable? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I tested writing and deleting a record normally and found that its SSTable content is as follows. write: [ { "partition" : { "key" : [ "user6284781860667377211" ], "position" : 0 }, "rows" : [… Language: en Canonical URL: https://forum.scylladb.com/t/what-is-the-meaning-of-the-content-in-sstable/1264 ## Headings Structure: H1: What is the meaning of the content in SSTable? H3: Related topics ## Main Content: H1: What is the meaning of the content in SSTable? H3: Related topics I tested writing and deleting a record normally and found that its SSTable content is as follows. write: When I do a lot of writing and deleting, I find that some properties of the records do not exist when reading data. I have found the SSTable content information for this record. There are three pieces of information in total. So, what I want to ask is, should a normal write record include both liveness_info and deletion_info, and cells is not empty. Is there a problem with the first and second types of records? Seeing time happening simultaneously. Is the second type caused by data loss? Thanks! Before 2024-01-31, what will happen if you read this row ? I guess there are two possibilities, one is reading 1 (containing cells data), and the other is reading 2 (not containing cells data), because the tstamp of liveness_info of 1 and 2 is exactly the same(“2024-01-28T08:01:22.490042Z”). Looks like this is a BUG? This is C*'s sstabledump and I’m not that familiar with the output-format of this tool, but I think that liveness_info is the row_marker. The row_marker is inserted only by INSERT statements, but not UPDATE statements. A row not having this is completely fine, this is not a data loss. For a more clear output, you can try scylla-sstable, which is our ScyllaDB-native reimplementation of sstabledump. This tool is available in 5.1 already, but best to use the latest 5.2 or 5.4 for the best experience. Hi! I want to know what the second situation represents? liveness_info time is later than deletion_info time, but cells are empty. Does it mean that this record is alive but the attribute is empty? But the first type of record and the second type of record are the same record, which is a bit contradictory. Because the cells of the first type of record are not empty. In (1) you have a row marker and the row has cells. In (2) the row has a row tombstone which covers the cells so they are removed, but it doesn’t cover the row-marker. In (3) there is a new row tombstone which now also covers the row-marker (and so the row marker is also removed). The row marker (liveness_info) does not apply to the row’s content, only the row itself. A row with a live row marker, but no cells, will appear in CQL as an empty row, with just the clustering and static columns having any value. --- ### Page: https://forum.scylladb.com/t/release-scylla-operator-1-11-2/1265 Title: [RELEASE] Scylla Operator 1.11.2 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of Scylla Operator 1.11.2 :rocket: Scylla Operator 1.11.2 brings a bug fix. As with all of our releases, all API changes are backward compatible. Notable changes I… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-operator-1-11-2/1265 ## Headings Structure: H1: [RELEASE] Scylla Operator 1.11.2 H3: Related topics ## Main Content: H1: [RELEASE] Scylla Operator 1.11.2 H3: Related topics The ScyllaDB team is pleased to announce the release of Scylla Operator 1.11.2 Scylla Operator 1.11.2 brings a bug fix. As with all of our releases, all API changes are backward compatible. For more changes and details check out the GitHub release notes. Upgrade instructions Upgrading from v1.10.0 or v1.11.1 with kubectl apply doesn’t require any extra action, just take the manifest from v1.11.2 tag and substitute the released image. Using helm requires a mandatory manual step for every release because helm can’t handle CRDs updates. For details, see our upgrade documentation . Regards, Scylla Operator Team --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-37-2024-02-02/1266 Title: Last week in scylla-cluster-tests.git master (issue #37; 2024-02-02) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This brief report highlights some fascinating commits to the scylla-cluster-tests.git master from the past week. The commits covered range from 818882db to 6833e66b. During this period, 5 Software Developers in Test and… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-37-2024-02-02/1266 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #37; 2024-02-02) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #37; 2024-02-02) H3: Related topics This brief report highlights some fascinating commits to the scylla-cluster-tests.git master from the past week. The commits covered range from 818882db to 6833e66b. During this period, 5 Software Developers in Test and 1 Software Developer made 22 non-merge commits. Here are some noteworthy changes: We have started collecting and displaying syslog-ng stats in the monitor, enabling us to track logging rates and drops. We’ve switched to ‘gp3’ EBS disks from ‘gp2’ for all K8S nodes we provision on EKS backend. These disks are both faster and more cost-effective. We’re in the process of moving cloud cleanup procedures to the SCT codebase. The first step was to periodically remove unused Azure instances and Resource Groups. All cleanup scripts use the keep tag, where alive prevents deletion and an integer value specifies the number of hours to retain the instance or Resource Group. Having recently started using ‘launch template’ resources in EKS, we now delete these too after test runs. We’ve extracted some logic for creating a stress command, allowing for more widespread use. We’ve added support for the cql-stress stress tool to SCT. This tool, similar to cassandra-stress but implemented in Rust, allows us to test rust-based drivers on a larger scale. Currently, we support read, write, counter_read, and counter_write commands. A new test based on this tool has been added. For more information, visit: GitHub - scylladb/cql-stress Stay tuned for the next update on last week’s changes in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-rust-driver-0-12-0/1267 Title: [RELEASE] ScyllaDB Rust Driver 0.12.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce ScyllaDB Rust Driver 0.12.0, an asynchronous CQL driver for Rust, optimized for Scylla, but also compatible with Apache Cassandra! Some interesting statistics: over 1,156k dow… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-rust-driver-0-12-0/1267 ## Headings Structure: H1: [RELEASE] ScyllaDB Rust Driver 0.12.0 H2: Changes H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Rust Driver 0.12.0 H2: Changes H3: Related topics The ScyllaDB team is pleased to announce ScyllaDB Rust Driver 0.12.0, an asynchronous CQL driver for Rust, optimized for Scylla, but also compatible with Apache Cassandra! Some interesting statistics: API cleanups / breaking changes: New features / enhancements: CI / developer tool improvements: Congrats to all contributors and thanks everyone for using our driver! The source code of the driver can be found here: The official crates.io registry entry is here: Thank you for your attention, please do not hesitate to contact us if you have any questions, issues, feature requests, or are simply interested in our driver! Contributors since the last release: --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-216-2024-02-04/1270 Title: Last week in scylladb.git master (issue #216; 2024-02-04) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the fe3bc00045…017a574b16 range are covered. There were 147 non-merge commits from 17 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-216-2024-02-04/1270 ## Headings Structure: H1: Last week in scylladb.git master (issue #216; 2024-02-04) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #216; 2024-02-04) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the fe3bc00045…017a574b16 range are covered. There were 147 non-merge commits from 17 authors in that period. Some notable commits: If a table that uses tablets detects that the average tablet size is larger than a configured amount, it will split all tablets for that table. This allows a growing table to use more shards, and to have reasonably-sized tablets for migration. Alternator, ScyllaDB’s implementation of the DynamoDB API, will now enable tablets if tablets themselves (an experimental feature) are enabled. When tablets are enabled, we will now refuse resharding (changing the number of shards in a node) as this is not implemented yet. There is now a metric for tablet count that can be used to observe the load balancer. In consistent topology (experimental), an ambiguity between nodes that were decommissioned and nodes that failed bootstrap was fixed. In consistent topology, a failed rebuild command will result in an error instead of an infinite retry loop. In consistent topology, streaming will now run under the streaming scheduling group rather than the gossip scheduling group. This ensures streaming has limited I/O and CPU consumption. The rewritesstables command will now execute in the streaming/maintenance group, reducing its impact on the rest of the system. There is now an API to trigger a Raft snapshot of cluster metadata. This is useful to work around a bug in the 5.2 Raft integration. A Raft snapshot will be triggered if we detect that we’re bootstrapping from an older ScyllaDB cluster. A crash in the mintimeuuid() function has been fixed. The system will now scan sstable files during startup in a way that prevents fragmentation of the kernel inode and dentry caches. This helps reduce memory pressure on systems with many sstables. A cleanup operation will now flush memtables, to prevent data that managed to stay in memtables from not being cleaned up. The scylla sstable tools that use Lua scripts now load the os and math libraries. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-4-2/1272 Title: [RELEASE] ScyllaDB 5.4.2 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.4.2, a bugfix release of the ScyllaDB 5.4 stable branch. ScyllaDB Open Source 5.4.2, like all past and future 5.x.y releases, is backward compatible and supports rolling… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-4-2/1272 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.4.2 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.4.2 H4: Inserted data only becomes available after restart H4: row_cache: update _prev_snapshot_pos even if apply_to_incomplete() is preempted H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.4.2, a bugfix release of the ScyllaDB 5.4 stable branch. ScyllaDB Open Source 5.4.2, like all past and future 5.x.y releases, is backward compatible and supports rolling upgrades. Users are encouraged to upgrade to 5.4.2. Issue fixed in this release: Does this include the fix for this issue: *Installation details* Scylla version: 5.4.1 Cluster size: 1 Node Platform: D…ocker on Kubernetes After Upgrading Scylla from 5.2.11 via 5.4.0 to 5.4.1, we started observing missing data in Scylla. Every once in a while, an `INSERT` is committed successfully but the inserted values are not visible until Scylla is restarted. As far as we can tell, version 5.2.11 was not affected while 5.4.1 is. 5.4.0 was not long enough in operation to make a reliable statement. The suspected bug occurs extremely rarely, making it hard for us to reproduce. Out of ~800M inserts per day, only a dozen are affected. In our running environment, I was able to perform the following steps: 1. Execute the normal `INSERT` workload using multiple concurrent clients. 2. Wait until a client notices that a `SELECT` on a previously `INSERT`ed key returns no columns 3. Shut down all clients to make sure there is no unnecessary load on our ScyllaDB 4. (optional) wait as long as you want for ScyllaDB to (process backlog/perform compactions/achieve consistency/whatever...) 5. Execute query: `SELECT * FROM tenant_6e10XXXX_XXXX_XXXX_XXXX_XXXXXXXXXXXX.edgestore WHERE key = `. The result is empty. 6. Restart ScyllaDB and wait until it is operational. 7. Execute the same query again and notice the `INSERT`ed data suddenly became available: |key|column1|value| |--|--|--| ||0xXX|0xXXXXXXXXXXXXXXXX| ||0xXX|0xXXXXXXXXXXXXXXXXXXXX| ||0xXXXX|0xXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX| ||0xXXXX|0xXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX| ||0xXXXX|0xXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX| ||0xXXXX|0xXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX| ||0xXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX|0xXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX| `X` obviously masks potentially private hexadecimal data which is not relevant to this report. What I already tried instead of restarting Scylla: - `nodetool refresh` - `nodetool flush` - `nodetool rebuild` and `nodetool repair` even though I'm aware both shouldn't make any difference on a single node cluster. It looks to me like Scylla has somehow recognized the transaction as completed even though it has not been persisted as expected. There seems to be some procedure which is triggered either by the shutdown or by the startup that picks this transaction up and replays it so the its modifications actually become available. **Appendix**: The Scylla [logs](https://github.com/scylladb/scylladb/files/13918121/logs-without-compactions.txt) captured during this operation. I removed all compaction logs to reduce the file to a reasonable size. No. the commit is backported to 5.4 and will be part of the next 5.4 patch release (5.4.3) Commit e81fc1f095f0d39f4bbc6503960450b7fda3ba1b accidentally broke the control f…low of row_cache::do_update(). Before that commit, the body of the loop was wrapped in a lambda. Thus, to break out of the loop, `return` was used. The bad commit removed the lambda, but didn't update the `return` accordingly. Thus, since the commit, the statement doesn't just break out of the loop as intended, but also skips the code after the loop, which updates `_prev_snapshot_pos` to reflect the work done by the loop. As a result, whenever `apply_to_incomplete()` (the `updater`) is preempted, `do_update()` fails to update `_prev_snapshot_pos`. It remains in a stale state, until `do_update()` runs again and either finishes or is preempted outside of `updater`. If we read a partition processed by `do_update()` but not covered by `_prev_snapshot_pos`, we will read stale data (from the previous snapshot), which will be remembered in the cache as the current data. This results in outdated data being returned by the replica. (And perhaps in something worse if range tombstones are involved. I didn't investigate this possibility in depth). Note: for queries with CL>1, occurences of this bug are likely to be hidden by reconciliation, because the reconciled query will only see stale data if the queried partition is affected by the bug on on *all* queried replicas at the time of the query. Fixes #16759 Closes scylladb/scylladb#17138 (cherry picked from commit ed98102c45a522393cd3bf478a0b6a712d192167) --- ### Page: https://forum.scylladb.com/t/can-we-use-tombstone-gc-mode-repair-for-gsi-tables/1273 Title: Can we use tombstone_gc = {'mode': 'repair'} for GSI tables? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I have two questions here about tombstone_gc = {‘mode’: ‘repair’} for GSI tables. I am not sure whether tombstone_gc = {‘mode’: ‘repair’} is supported for GSI tables or not. How can I ensure that the tombstone_gc mode … Language: en Canonical URL: https://forum.scylladb.com/t/can-we-use-tombstone-gc-mode-repair-for-gsi-tables/1273 ## Headings Structure: H1: Can we use tombstone_gc = {'mode': 'repair'} for GSI tables? H3: Related topics ## Main Content: H1: Can we use tombstone_gc = {'mode': 'repair'} for GSI tables? H3: Related topics I have two questions here about tombstone_gc = {‘mode’: ‘repair’} for GSI tables. AFAIK, tombstone_gc = {'mode': 'repair'} should work with any table, that can be repaired and the materialized view that is backing the GSI certainly can be repaired. @nyh are you aware of any potential gotchas around this? Yes, a materialized view is a normal table under the hood, but one cannot modify it by ALTER TABLE and you need to do ALTER MATERIALIZED VIEW. I believe, but did not test myself that you can modify this tombstone_gc with ALTER MATERIALIZED VIEW. I also believe (but didn’t test) that if you already have a base table with tombstone_gc and later create a view, this view will inherit this setting. But both beliefs should be tested (I’ll do this later). Beyond that, there is a separate question of whether users should repair view tables. My view (sorry for the pun) is that they should - they should repair base tables and then view tables. There are known rare inconsistencies (“ghost rows”) that may arise from repairing view tables, but I think the benefits of such repair - including tombstone GC - greatly outweigh the risks. But this is still an open discussion. CC @Konstantin_Osipov @Yaniv_Kaul We’ve recently added (via [Backport 5.2] schema: add scylla specific options to schema description by Jadw1 · Pull Request #16786 · scylladb/scylladb · GitHub ) a reasonable way to see tombstone_gc for tables. I wrote a test and confirmed (on latest ScyllaDB master) that: The tombstone_gc option can be set on a materialized view either when it’s first created with CREATE MATERIALIZED VIEW or when later with ALTER MATERIALIZED VIEW The current setting is correctly printed with the “DESC ” statement. This is the new server-side describe statement, that replaced the older statement that used to live in CQL (and didn’t support tombstone_gc or other non-standard options). The gc_mode set on a base table is not inherited by its views - you need to set it separately for each view (again, while creating the view or later). I don’t know if this is a bug or a feature, but it shouldn’t cause any problems once you know you need to set it per view separately. There are known rare inconsistencies (“ghost rows”) that may arise from repairing view tables Thanks very much for your reply. However, I am still confused about the rare inconsistencies (“ghost rows”) which may arise from repairing view tables. Can you give me some specific examples or issues about this? Can you give me some specific examples or issues about “ghost rows” which arise from repairing view tables? Thanks! @nyh When I enabled the gc_mode=repair function of the base table and view table, there was a phenomenon of missing certain attributes in the records of my view table. During this period, the cluster is scheduling repairs normally. Under what circumstances does this phenomenon occur and is it related to function gc_mode=repair ? Thanks! --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-1-14/1274 Title: [RELEASE] ScyllaDB Enterprise 2022.1.14 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2022.1.14, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.1. Note that ScyllaDB Enterprise 2022.1 and 2022.2 will support EOL in… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-1-14/1274 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2022.1.14 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2022.1.14 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2022.1.14, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.1. Note that ScyllaDB Enterprise 2022.1 and 2022.2 will support EOL in June 2024. You are encouraged to upgrade to ScyllaDB Enterprise LTS 2023.1 and soon-to-be-released Scylla Enterprise 2024.1 in coordination with the Scylla Support team. Below is a list of performance and stability improvements and bug fixes, each with an open-source reference if available: --- ### Page: https://forum.scylladb.com/t/multidc-cluster-observing-100-cpu-when-connecting-scylla-with-kafka-source-connector-to-read-cdc-logs/1276 Title: MultiDC cluster ,Observing 100% cpu when connecting scylla with kafka source connector to read cdc logs - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: We have our multi dc setup with 3 node in dc1 and 3 in dc2 . Although both the DCs are in the same region n same subnet, This is done as we require separate clusters for reading n writing data. Our setup creates a new t… Language: en Canonical URL: https://forum.scylladb.com/t/multidc-cluster-observing-100-cpu-when-connecting-scylla-with-kafka-source-connector-to-read-cdc-logs/1276 ## Headings Structure: H1: MultiDC cluster ,Observing 100% cpu when connecting scylla with kafka source connector to read cdc logs H3: Related topics ## Main Content: H1: MultiDC cluster ,Observing 100% cpu when connecting scylla with kafka source connector to read cdc logs H4: MultiDC cluster, observing 100% CPU when connecting kafka debezium source conector H3: Related topics We have our multi dc setup with 3 node in dc1 and 3 in dc2 . Although both the DCs are in the same region n same subnet, This is done as we require separate clusters for reading n writing data. Our setup creates a new table everyday at 12 midnight with cdc enabled in DC1 At the same time we also create a kafka source connector to consume cdc logs from DC2 everyday Issue: At around 12 When creating a new source con , we observe scylladb servers on dc2 consumes 100% cpu. We increased cpu from 16cores to 32 but still same behavior Once the kafka connector creates its topic and start reading data from cdc log ,scylladb cpu cools down The logs in syslog shows reader_concurrency_semaphores for that time period. Any expert thoughts is appreciated Thanks in advance The cpu utilization itself isn’t always an indication of a problem, however the reader_concurrency_semaphore may be a stronger indication. For how long do you have this high utlization and the reader_conccurency_semaphore errors? Also, how exactly are your CDC configured? Full logs would be more helpful, but sounds like a better place for it would be a Github issue. Thanks a lot for responding Already raised this on github Here is the link ,along with logs and other necessary details you need We have our multi dc setup with 3 node in dc1 and 3 in dc2 . Although both the D…Cs are in the same region n same subnet, This is done as we require separate clusters for reading n writing data. Our setup creates a new table everyday at 12 midnight with cdc enabled in DC1 At the same time we also create a kafka source connector to consume cdc logs from DC2 everyday Issue: At around 12 When creating a new source con , we observe scylladb servers on dc2 consumes 100% cpu. We increased cpu from 16cores to 32 but still same behavior Once the kafka connector creates its topic and start reading data from cdc log ,scylladb cpu cools down Logs: The logs in syslog shows reader_concurrency_semaphores for that time period. Any expert thoughts is appreciated Thanks in advance The reader_concurreny_semaphore log only comes for 15 mins when a new kafka debezium source connector spawns. Cpu also goes high during that time . Exact details are available in the link, please go through it and help in understandng the exact issue We manage to reduce 15 mins of High CPUby tweaking scylla.query.time.window.size in kafka source connector , but still no clue about 100% cpu utilisation Also it is observed , our 6 node dc2 cluster and 4 nodes kafka connector, there is a total of 40000 active connection on each scylladb node Any clue about this?? --- ### Page: https://forum.scylladb.com/t/cannot-remove-unreachable-dead-nodes-from-my-cluster/1278 Title: Cannot remove unreachable dead nodes from my cluster - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Two nodes in my ScyllaDB cluster (which has total 4 nodes) went down because of corrupted storage. One of them was the seed node. I did a rolling restart changing the seed node among the remaining two available nodes. B… Language: en Canonical URL: https://forum.scylladb.com/t/cannot-remove-unreachable-dead-nodes-from-my-cluster/1278 ## Headings Structure: H1: Cannot remove unreachable dead nodes from my cluster H3: Related topics ## Main Content: H1: Cannot remove unreachable dead nodes from my cluster H3: Related topics Two nodes in my ScyllaDB cluster (which has total 4 nodes) went down because of corrupted storage. One of them was the seed node. I did a rolling restart changing the seed node among the remaining two available nodes. But now I’m unable to remove the dead nodes from cluster. nodetool removenode command gets stuck indefnitely. This is my nodetool status: scylla logs are filled with these errors: I’m unable to add any new nodes to my cluster because gossip says it can’t add new nodes until the status of any node is UNKNOWN Is there a way I can forcefully remove the node from my cluster? I see this was discussed also in #10292, in comments 1, 2 and 3. I will copy the solution here, for search-ability: First, when a node is unreachable. It is much better to run the replace a dead procedure than running removenode. When more than one nodes are down, you can use ignore_dead_nodes_for_replace option to ignore the peer down node when running replace. With replace, you can add 2 more nodes back. I tried replace node procedure with the two config params you mentioned, but faced an error that looked like #13865 But debugging that issue nudged me towards the Raft manual recovery procedure guide, using which I was able to recover my cluster! (removenode worked when all UN nodes where in recovery mode) Thank you helping me out @asias! --- ### Page: https://forum.scylladb.com/t/recursive-functions/1281 Title: Recursive functions - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I want to model and query an n-ary tree. Is there a way to query recursively? Something like the PSQL: WITH RECURSIVE included_parts(sub_part, part) AS ( SELECT sub_part, part FROM parts WHERE part = 'our_product'… Language: en Canonical URL: https://forum.scylladb.com/t/recursive-functions/1281 ## Headings Structure: H1: Recursive functions H3: Related topics ## Main Content: H1: Recursive functions H3: Related topics I want to model and query an n-ary tree. Is there a way to query recursively? Something like the PSQL: No, you will have to iterate in the client and query each subsequent level yourself. --- ### Page: https://forum.scylladb.com/t/optimize-deletion-through-partitioning/1283 Title: Optimize Deletion Through Partitioning - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I have the following schema that I tried to tailor to optimize data deletion by leveraging partitioning on the expiry and bucket fields. I thought this approach was helpful when I wanted to bulk delete data based on the … Language: en Canonical URL: https://forum.scylladb.com/t/optimize-deletion-through-partitioning/1283 ## Headings Structure: H1: Optimize Deletion Through Partitioning H3: How Level Infinite Implemented CQRS and Event Sourcing on Top of Apache... H3: How Numberly Replaced Kafka with a Rust-Based ScyllaDB Shard-Aware Application H3: Related topics ## Main Content: H1: Optimize Deletion Through Partitioning H3: How Level Infinite Implemented CQRS and Event Sourcing on Top of Apache... H3: How Numberly Replaced Kafka with a Rust-Based ScyllaDB Shard-Aware Application H3: Related topics I have the following schema that I tried to tailor to optimize data deletion by leveraging partitioning on the expiry and bucket fields. I thought this approach was helpful when I wanted to bulk delete data based on the expiry date and possibly distribute the data across different buckets for balanced storage and efficiency, as everything in a specific partition will be marked as deleted(or not, at least I thought it would.) Queries are offloaded to materialized views designed for particular access patterns that use id and expiry compared to the date they were made. I am not sure whether this is recognized in the realm of distributed databases that the base table’s purpose is to optimize deletion and not directly used for retrieval of its data? I need to deal with time-based data need to be purged after a certain period, even if the data is never used through the materialized view. I appreciate any advice or suggestions. It sounds like TTL would be relevant for your use case. The Expiring Data with TTL (Time to Live) lesson on ScyllaDB University would be a good starting point. You might also want to look at Time-Window Compaction Strategy (TWCS). It sounds like TTL would be relevant for your use case. The Expiring Data with TTL (Time to Live) lesson on ScyllaDB University would be a good starting point. You might also want to look at Time-Window Compaction Strategy (TWCS) . Thank you for the suggestion to explore TTL and TWCS in ScyllaDB University. These features are indeed powerful for managing data with natural expiration patterns, and I appreciate the guidance towards these resources. However, my use case involves a dynamic expiry value that varies depending on the some_relevant_id , which introduces a level of complexity that might not align perfectly with the standard application of TWCS. While TWCS is optimized for scenarios with predictable, uniform expiration patterns, my dataset features varied TTLs, making it challenging to leverage TWCS for efficient compaction and storage management directly. My concern revolves around ensuring efficient data deletion that doesn’t affect read operations performed, but this design choice has led me to rely on materialized views to handle my query needs, particularly for accessing data by id, which is not the primary key in my base table schema. Given this context, I’m curious about the broader community’s experience and best practices regarding this approach. Specifically: Thank you for your guidance, and I look forward to any further discussion or advice that can help refine our strategy. Please correct me if I am wrong, here are my guesses. Configuring my records table to use the Time Window Compaction Strategy with a 1-day window size might offer indirect benefits to query performance. TWCS optimizes the compaction process within the base table by organizing data into time-based windows. This leads to more efficient storage and management of data, which indirectly benefits materialized views by ensuring they are updated and maintained with less overhead. Operations on the ScyllaDB, such as querying, inserting, updating, and deleting data, engage both partitioning logic and SSTable mechanics. For instance, an update operation involves finding the right partition and then creating a new SSTable to reflect the update. By minimizing write amplification in the base table through efficient compaction, TWCS may indirectly reduce the load on the database during updates. This results in faster and more efficient updates to materialized views. As for the TTL, is deliberately setting slightly different TTL values even for the records within the same partition will prevent all the data from expiring, or being marked for deletion, mitigate the potential performance issues associated with a large volume of data expiring simultaneously? @hanishi you should be using TWCS when your access patterns involve append-only TTL data with very few (or none at all) updates. You should NOT use it when manually deleting data. Also notice that the compaction strategy of your choice will primarily be used for your base table only. That is, by default, once you create a view your view table is going to default to STCS, though you can manually change it after your view gets created. To be quite honest, I don’t understand what exactly you are trying to accomplish. Your base table feeds an underlying view, but you are primarily querying the view? Is it correct to assume that you want to delete all records with a given expiry date? If yes, then why don’t you do it the other way around (the base table follows your view schema, and then you have a view which has your existing base table schema)? Another thing you can do, as @Guy mentioned is to rely on TTL and get rid of the expiry and bucket columns. You insert the data with a specific time-to-live period, if that period ever changes, you upsert the record with a new TTL. If you don’t do anything, then past X period of time your entries should get evicted. Let us know whether this helps in any way. On top of that, as you asked about delete specific patterns, maybe these talks may shed you some light: A use-case study of why Level Infinite, a division of Tencent, uses ScyllaDB as the state store of our Proxima Beta gaming platform’s service architecture. How we leverage time window compaction strategy (TWCS) to power a distributed queue-like event... How Numberly used Rust & ScyllaDB to replace Kafka, streamlining the way all its AdTech components send and track messages (whatever their form). @felipemendes Thank you for your insights, I really appreciate. The expiry date is critical for analytics for understanding or predicting the volume of data that will be deleted in the future I need to keep this field as it serves a different purpose than TTL, acting as a piece of actionable data rather than just a mechanism for data deletion. However, I can still make use of TTL to handle the automatic deletion of data. When inserting records, the TTL value based on the expiry date can be set to delete the rows automatically. I tried to tailor to optimize data deletion by leveraging partitioning on the expiry and bucket fields, because all rows are to be deleted eventually on that expiry date or later. Regarding the design of the records table we’re managing with ScyllaDB, it’s important to note that most of the records in the database are never accessed and only exist until their expiry date. This was the rationale behind my initial approach. However, as long as deletion by TTL remains efficient and does not negatively impact read operations on the base table, we could proceed with a simple base table structure without the need for any attached materialized views. The adoption of TWCS is anticipated to offer significant benefits, particularly in the context of our reconsidered base table schema outlined below: Here, the bucket_id could be derived from the id by applying a consistent hash function and modulo operation to limit the number of buckets. This way, although id itself remains unique, several id values share the same partition which will mitigate creation of vast amount of small partitions, each will only contain the data for that specific id. (Please correct me if I am wrong about this) I still need to filter the record by id when retrieving data I need though. Could you please take a moment to evaluate this rearranged records table? Your expertise and feedback would be immensely helpful in ensuring it’s optimized for our use case. Thank you in advance for your time and assistance. --- ### Page: https://forum.scylladb.com/t/error-creating-table-line-1-63-no-viable-alternative-at-input-text/1285 Title: Error creating table:line 1:63 no viable alternative at input 'text' - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Here is my code for the Create table query Session.Query("CREATE TABLE IF NOT EXISTS todo (id uuid PRIMARY KEY, username text UNIQUE, title text, description text, status text, created_at timestamp, updated_at timestamp… Language: en Canonical URL: https://forum.scylladb.com/t/error-creating-table-line-1-63-no-viable-alternative-at-input-text/1285 ## Headings Structure: H1: Error creating table:line 1:63 no viable alternative at input 'text' H3: Related topics ## Main Content: H1: Error creating table:line 1:63 no viable alternative at input 'text' H3: Related topics Here is my code for the Create table query And i am encountering the type mismatch error- “Connected to DB successfully 2024/02/07 11:52:34 Error creating table:line 1:63 no viable alternative at input ‘text’” the output states that my database connection is successful but the create table query has a type mismatch. Please help CQL does not have an UNIQUE column attribute. Remove this and you should be fine. If you want a column to be unique, make it part of the primary key, keys are unique by definition: It worked, thanks a lot! --- ### Page: https://forum.scylladb.com/t/do-i-ever-need-to-disable-the-scylladb-cache-to-use-less-memory/1287 Title: Do I ever need to disable the ScyllaDB cache to use less memory? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Is it possible to disable the cache? If so, how do I do it, and when? Language: en Canonical URL: https://forum.scylladb.com/t/do-i-ever-need-to-disable-the-scylladb-cache-to-use-less-memory/1287 ## Headings Structure: H1: Do I ever need to disable the ScyllaDB cache to use less memory? H3: ScyllaDB CQL Extensions | ScyllaDB Docs H3: Data Definition | ScyllaDB Docs H3: Related topics ## Main Content: H1: Do I ever need to disable the ScyllaDB cache to use less memory? H3: ScyllaDB CQL Extensions | ScyllaDB Docs H3: Data Definition | ScyllaDB Docs H3: Related topics Is it possible to disable the cache? If so, how do I do it, and when? Normally, one should not need to disable the cache. The cache uses memory unused for other purposes. When demand for memory increases, ScyllaDB frees memory by evicting items from the cache (in LRU order). If demand for memory is very high, the cache can be fully evicted. If one still wants to disable the cache, they can do it by adding enable_cache: false to scylla.yaml, or alternatively, adding the --enable-cache=0 command-line parameter to ScyllaDB. Its best to let ScyllaDB control how much memory is allocated internally to each function. Note that you can by-pass cache for a specific query ScyllaDB is an Apache Cassandra-compatible NoSQL data store that can handle 1 million transactions per second on a single server. And control the cache per table ScyllaDB is an Apache Cassandra-compatible NoSQL data store that can handle 1 million transactions per second on a single server. Between these two options, you have a granular control for using the cache. --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-2-17/1288 Title: [RELEASE] ScyllaDB Enterprise 2022.2.17 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2022.2.17, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.2 Feature Release. Note the latest ScyllaDB Enterprise release is 2023… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-2-17/1288 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2022.2.17 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2022.2.17 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2022.2.17, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.2 Feature Release. Note the latest ScyllaDB Enterprise release is 2023.1 LTS, and very soon 2024.1 LTS, and you are encouraged to upgrade to it in coordination with the ScyllaDB Support team. The following issues are fixed in this release (with an open-source reference, if available): --- ### Page: https://forum.scylladb.com/t/i-am-trying-to-execute-below-queries-but-getting-exceptions-unable-to-fix-these/1290 Title: I am Trying to Execute below queries but Getting Exceptions, unable to fix these - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: CREATE KEYSPACE mykeyspace4 WITH replication = {'class': 'NetworkTopologyStrategy', 'replication_factor' : 3, 'DC2': 2} Exception: `ConfigurationException: Unrecognized strategy option {DC1} passed to org.apache.cas… Language: en Canonical URL: https://forum.scylladb.com/t/i-am-trying-to-execute-below-queries-but-getting-exceptions-unable-to-fix-these/1290 ## Headings Structure: H1: I am Trying to Execute below queries but Getting Exceptions, unable to fix these H3: Related topics ## Main Content: H1: I am Trying to Execute below queries but Getting Exceptions, unable to fix these H3: Related topics (1) You are using a non-existent Datacenter in the keyspace definition, this is why the statement fails. The error message is confusing and I opened an issue for it. (2) tombstone_gc can only be enabled with tables that are part of a keyspace, which has at least replication_factor of 2 or more. If your keyspace has replication_factor of 1 (which is really not recommended), just set gc_grace_seconds = 0 in your schema options, to speed up tombstone purging. (3) Looks like you need to give ScyllaDB read/write rights to /etc/scylla/data_encryption_keys/. (4) The STORAGE option is a new and experimental feature, which is only available in the latest OSS releases. I suppose the version you are using either doesn’t have this feature, or the experimental feature is not enabled. How to Enable this (3) Looks like you need to give ScyllaDB read/write rights to /etc/scylla/data_encryption_keys/ . any idea? How to add STORAGE option in Experimental feature. I am using below mentions conf. in scylla.yaml: 4- Alter Keyspace with Storage not working: (2) You need to use chown or chmod to change the permissions on said directory such that the user under which ScyllaDB runs (this depends on the installation) has read/write access to it. (3) You need to add the keyspace-storage-options to the list of enabled experimental features. (4) We do not support changing the storage options on a keyspace 1- Even though i give write permission to directory using sudo chmod +w /etc/scylla/ but still: 2- Below queries showing errors: --- ### Page: https://forum.scylladb.com/t/bootstrap-repair-of-us-west-nodes-takes-time-in-multi-dc-cluster/1291 Title: Bootstrap repair of us-west nodes takes time in multi-dc cluster - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, We have a Multi-DC ScyllaDB cluster. With 10 nodes in AWS us-east-1 and 10 nodes in us-west-2. We use scylla-ansible-role to bring up new clusters. We have observed that bootstrap of nodes in usw2 takes lot of time … Language: en Canonical URL: https://forum.scylladb.com/t/bootstrap-repair-of-us-west-nodes-takes-time-in-multi-dc-cluster/1291 ## Headings Structure: H1: Bootstrap repair of us-west nodes takes time in multi-dc cluster H3: Related topics ## Main Content: H1: Bootstrap repair of us-west nodes takes time in multi-dc cluster H3: Related topics We have a Multi-DC ScyllaDB cluster. With 10 nodes in AWS us-east-1 and 10 nodes in us-west-2. We use scylla-ansible-role to bring up new clusters. We have observed that bootstrap of nodes in usw2 takes lot of time than in use1. Looks like table repair during bootstrap is taking long time. Any pointer on how to debug and fix this. Please provide the relevant logs that show what is taking time. Do the nodes in different DCs have differing shard count? Are you using RBNO based bootstrap? Here are few lines from logs. Looks like the repair during bootstrap takes a lot of time in us-west-2. In US-EAST-1 it took 13 sec where as in US-WEST-2 it took about 27 min. in US-EAST-1: Feb 22 13:06:07 x.y.z.ec2.internal scylla[2966]: [shard 0:stre] repair - bootstrap_with_repair: started with keyspaces={system_traces, system_distributed_everywhere, system_distributed, system_auth}, nr_ranges_total=9179 Feb 22 13:06:20 x.y.z.ec2.internal scylla[2966]: [shard 0:stre] repair - bootstrap_with_repair: finished with keyspaces={system_traces, system_distributed_everywhere, system_distributed, system_auth} in US-WEST-2: Feb 22 13:07:34 a.b.c.ec2.internal scylla[3055]: [shard 0:stre] repair - bootstrap_with_repair: started with keyspaces={system_traces, system_distributed_everywhere, system_distributed, system_auth}, nr_ranges_total=9421 Feb 22 13:34:28 a.b.c.ec2.internal scylla[3055]: [shard 0:stre] repair - bootstrap_with_repair: finished with keyspaces={system_traces, system_distributed_everywhere, system_distributed, system_auth} @avikivity , I have attached some log for reference in this thread. Can you provide some pointer please. The attached logs do not contain any information w.r.t. to what might be the cause of the slowness. I didn’t found any error in scylla-server logs. Let me know where else to check. Do the nodes in different DCs have differing shard count? Are you using RBNO based bootstrap? Can you please answer this? The answer might provide a lead. Nodes in different DC are of same EC2 types. These nodes got same number of shards. I think it is using RBNO based bootstrap. I am using scylla-ansible-roles git repo to create scylla cluster. In the repo scylla.yaml template is at scylla-ansible-roles/ansible-scylla-node/templates/scylla.yaml.j2 at master · scylladb/scylla-ansible-roles · GitHub I don’t see any config for RBNO in this template. So I think it is RBNO(default approach). Could this be because there is only 1 seed and this seed is in us-east-1? Seeds are only used when joining the cluster, they are not used afterwards. The fact that small tables take a lot of time to repair, much more than what would expect, is a known problem and we recently merged a pull request improving this. That said, I don’t know why those small system tables take so much more time to stream in one DC, compared with the other. How did you configure the replication of system_auth? This is a keyspace that the user is expected to adjust as the cluster is expanded? I hope this new PR will reduce some time consumption. After all nodes in cluster is up and running I run a script which Yes, it is repair that takes a long time for tiny tables. We have recently moved node-operations to use repair behind the scenes (hence the name RBNO) and now node operations are affected too. For reference, this is the PR: repair: Introduce small table optimization by asias · Pull Request #15974 · scylladb/scylladb · GitHub --- ### Page: https://forum.scylladb.com/t/scylladb-enterprise-release-2024-1-0/1292 Title: ScyllaDB Enterprise Release 2024.1.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: ScyllaDB Enterprise Release 2024.1.0 The ScyllaDB team is pleased to announce the release of ScyllaDB Enterprise 2024.1.0 LTS, a production-ready ScyllaDB Enterprise Long Term Support Major Release. With 2024.1 LTS out,… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-enterprise-release-2024-1-0/1292 ## Headings Structure: H1: ScyllaDB Enterprise Release 2024.1.0 H1: ScyllaDB Enterprise Release 2024.1.0 H2: Performance Improvements H3: ScyllaDB Enterprise 2024.1 vs ScyllaDB 2023.1 H3: Latency tests H3: ScyllaDB Enterprise 2024.1 vs ScyllaDB Open Source 5.4 H2: Encryption at Rest (EaR) Enhancements H3: Amazon KMS Integration for Encryption at Rest H3: Transparent Data Encryption H2: Repair Based Node Operations (RBNO) H2: Node Level Metrics H2: Guardrails H2: Security H3: Encryption at transit, TLS certificates: H3: Default superuser name and password H3: FIPS Tolerant H2: Deprecated and removed features H2: Strongly Consistent Schema Management with Raft H3: Related topics ## Main Content: H1: ScyllaDB Enterprise Release 2024.1.0 H1: ScyllaDB Enterprise Release 2024.1.0 H2: Performance Improvements H3: ScyllaDB Enterprise 2024.1 vs ScyllaDB 2023.1 H4: Throughput tests H3: Latency tests H3: ScyllaDB Enterprise 2024.1 vs ScyllaDB Open Source 5.4 H2: Encryption at Rest (EaR) Enhancements H3: Amazon KMS Integration for Encryption at Rest H3: Transparent Data Encryption H2: Repair Based Node Operations (RBNO) H2: Node Level Metrics H2: Guardrails H2: Security H3: Encryption at transit, TLS certificates: H3: Default superuser name and password H3: FIPS Tolerant H2: Deprecated and removed features H2: Strongly Consistent Schema Management with Raft H3: Related topics The ScyllaDB team is pleased to announce the release of ScyllaDB Enterprise 2024.1.0 LTS, a production-ready ScyllaDB Enterprise Long Term Support Major Release. With 2024.1 LTS out, ScyllaDB Enterprise 2022.1 and 2022.2 will support EOL in June 2024. More information on ScyllaDB Long Term Support (LTS) policy is available here. The ScyllaDB Enterprise 2024.1 release is based on ScyllaDB Open Source 5.4, and introduces significant performance improvements, additional Encryption At Rest (EaR) functionality, Repair Base Node Operations (RBNO) for all operations, and many more improvements and bug fixes. Consistent schema management using Raft, introduced in 2023.1, will be enabled automatically on upgrade (more below). In addition, 2024.1 includes enhancements to Encryption at Rest (EaR), including Amazon KMS integration, and default encryption at rest, first introduced in 2023.1.2. Together, these improvements allow you to easily use your own key for a cluster wide EaR. ScyllaDB Enterprise customers are encouraged to upgrade to ScyllaDB Enterprise 2023.1, and are welcome to contact our Support Team with questions. 2024.1 include many performance improvements which translate to: 2024.1 has more than x1.5 higher throughput compared to 2023.1! In some cases, this can translate to %35 reduction in the number of vCPU required to support a similar load, which means a similar reduction in vCPU cost. Latency tests are done in 50% of the max throughput tested. As demonstrated below, the latency (both mean and p99) is 33% lower, even with the higher throughput. instance_type_db: i3.2xlarge instance_type_loader: c4.2xlarge Through tests: cassandra-stress [mixed|read|write] no-warmup cl=QUORUM duration=50m -schema ‘replication(factor=3)’ -mode cql3 native -rate threads=100 -pop ‘dist=gauss(1..30000000,15000000,1500000)’ ScyllaDB Enterprise 2024.1 is based on ScyllaDB Open Source 5.4, but includes Enterprise only performance improvement features. As shown below, throughput gain is significant, while latency is lower. Test setups and parameters are equal to the above Enterprise tests. Scylla Enterprise has supported Encryption at Rest (EaR) for a long time. So far, one could store the keys for EaR locally, in an encrypted table, or in an external KMIP server. This release Release adds: Scylla Enterprise has supported Encryption at Rest (EaR) for a long time. So far, one could store the keys for EaR locally, in an encrypted table, or an external KMIP server. Release 2023.1.2 added the ability to use Amazon KMS keys. ScyllaDB can now use Customer Managed Key (CMK), stored in KMS, to create, encrypt, and decrypt Data Keys (DEK), which are then used to encrypt and decrypt the data in storage, such as SSTables, Commit logs, Batches, and hints logs. KMS creates DEK from CMK DEK (plain text version) is used to encrypt the data at rest. Diagrams from: AWS KMS keys - AWS Key Management Service Before using KMS, you need to set KMS as a key provider and validate that ScyllaDB nodes have permission to access and use the CMK you created in KMS. Once you do that, you can use the CMK in the CREATE and ALTER TABLE commands with KmsKeyProviderFactory, as follows CREATE TABLE myks.mytable (……) WITH scylla_encryption_options = { ‘cipher_algorithm’ : ‘AES/CBC/PKCS5Padding’, ‘secret_key_strength’ : 128, ‘key_provider’: ‘KmsKeyProviderFactory’, ‘kms_host’: ‘my_endpoint’ Where “my_key” point to a section in scylla.yaml aws_use_ec2_credentials: true aws_use_ec2_region: true master_key: alias/MyScyllaKey You can also use the KMS provider to encrypt System level data. See more examples and info here. Transparent Data Encryption (TDE), adds a way to define Encryption at Rest parameters per cluster, not only per table. This allows the system administrator to enforce encryption of all tables using the same master key, for example, from KMS, without specifying the encryption parameter per table. For example, with the following in scylla.yaml, all tables will be encrypted using encryption parameters of my-kms1 user_info_encryption: key_provider: KmsKeyProviderFactory, See more examples and info here. RBNO provides a more robust, reliable, and safer data streaming for node operations like node-replace and node-add/remove. In particular, a failed node operation can resume from the point it stopped – without sending data that has already been synced. In addition, with RBNO enabled, you don’t need to repair before or after node operations, such as replace or removenode. Repair Based Node Operations were introduced as an experimental feature in ScyllaDB Open Source 4.0. They use repair to stream data for node-operations like replace, bootstrap and others. In 2024.1, RBNO is enabled by default for all operations: remove node, rebuild, bootstrap and decommission. Replace node operation was already enabled by default in 2023.1. See Repair Base Node Operations (RBNO) docs and Scylla Summit 2022 session by Asias He Most ScyllaDB metrics are per-shard, per-node, but not for a specific table. We now export some per-table metrics. These are exported once per node, not per shard, to reduce the number of metrics. #2198 Guardrails is a framework to protect ScyllaDB users and admins from common mistakes and pitfalls. In this release ScyllaDB includes a new guardrail on the replication factor. It is now possible to specify the minimum replication factor for new keyspaces via a new configuration item #8891. This matches the same functionality in Apache Cassandra #CASSANDRA-14557 The new RF guardrails include the following configuration: More Guardrails are expected in upcoming releases. In addition to the EaR Enhancements above, the following security features introduced in 2024.1: It is now possible to use TLS certificates to authenticate and authorize a user to ScyllaDB. The system can be configured to derive the user role from the client certificate and derive the permissions the user has from that role. #10099 See more on certificate-authentication docs. It is now possible to specify the initial superuser name and password (salted) in scylla.yaml config or command line. Note that config values become redundant as soon as auth tables are initialized. See two new config parameters auth_superuser_name, auth_superuser_salted_password below. ScyllaDB Enterprise can now run on FIPS enabled Ubuntu, using libraries that were compiled with FIPS enabled, such as OpenSSL, GnuTLS, and more. Strongly Consistent Schema Management with Raft became the default for new clusters in ScyllaDB Enterprise 2023.1. In this release it is enabled by default when upgrading existing clusters. Note that when you have a two-DC cluster with the same number of nodes in each DC, the cluster will lose the quorum if one of the DCs is down. We recommend configuring three DCs per cluster to ensure that the cluster remains available and operational when one DC is down. See docs for more on Quorum Requirement. If you are not sure, contact the ScyllaDB Support team for advice. If you do not want to enable Raft, you should explicitly disable it in scylla.yaml of each node before the upgrade: consistent_cluster_management: false Source: Upgrade from 2023.1 to 2024.1 doc --- ### Page: https://forum.scylladb.com/t/scylladb-enterprise-release-2024-1-0-deployment-and-more-improvements/1293 Title: ScyllaDB Enterprise Release 2024.1.0 - Deployment and more improvements - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: See 2024.1 release notes Deployment and install ScyllaDB Enterprise 2024.1 is officially supported on Rocky / RHEL 9. RHEL / CentOS 7 support is deprecated and won’t be supported going forward. ScyllaDB installation wi… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-enterprise-release-2024-1-0-deployment-and-more-improvements/1293 ## Headings Structure: H1: ScyllaDB Enterprise Release 2024.1.0 - Deployment and more improvements H2: Deployment and install H2: More Improvements H3: Streaming: Add stream_plan_ranges_fraction H3: CQL API H3: Amazon DynamoDB Compatible API (Alternator) H3: Performance and stability H3: Operations H3: Related topics ## Main Content: H1: ScyllaDB Enterprise Release 2024.1.0 - Deployment and more improvements H2: Deployment and install H2: More Improvements H3: Streaming: Add stream_plan_ranges_fraction H3: CQL API H3: Amazon DynamoDB Compatible API (Alternator) H3: Performance and stability H3: Operations H3: Related topics See 2024.1 release notes This option allows user to change the number of ranges to stream in batch per stream plan. Currently, each stream plan streams 10% of the total ranges. The default value is the same as before: 10% of total ranges. #14191 Examples: blob_column = (blob)(int)12323 ScyllaDB uses two separate memory reservation systems for memtables: user, used for user writes, and system, used for ScyllaDB’s own writes. The root cause was Raft did not use the system reservation. Stability: read load failing after one node upgrade [bad_enum_set_mask (Bit mask contains invalid enumeration indices.)] #15795 Performance: Repairing a cluster after a restore causes severe reactor stalls throughout the cluster (due to expensive logging within do_repair_ranges() without yield) #14330 Install: scylla_post_install.sh: “[ $RHEL ]” does not work for RHEL, it only detects CentOS #16040 Stability: test_interrupt_build_process dtest failed with schema_registry - Tried to build a global schema for view ks.t_by_v2 with an uninitialized base info #14011 Stability: tests.topology_experimental_raft.test_raft_cluster_features.debug test is flaky. The root cause was error handling in the Raft coordinator. #15747 #15728 Stability: The mutation compactor now validates its input stream rather than the output stream. [IPv6 configuration] A node is stuck with “?U” status and Host ID is “null”, unclear reason #16039 Stability: assigning position_in_partition is not exception safe, can lead to incorrect data during memory stress #15822 Stability: Scylla cluster nodes utilizes 100% of CPU even with no load #12774, #13377, #7753 Performance: Major compaction will now merge any sstables streamed in due to decommission or repair before starting compaction. This generates compacted sstables more in line with expectations. #11915. Performance: To generate efficient bloom filters, we estimate the number of partitions in the sstable we will produce. The estimation has been improved for data models where the partition keys dominate the on-disk size. #15726. Stability: repair should handle abort_requested_exception mode gracefully #15710 Stability: row_cache::row_cache() isn’t exception-safe #15632 Stability: Recently, we changed the schema version algorithm not to hash the entire schema as this causes slow performance with large numbers of tables. This has been reverted due to a regression. #15530. Stability: Compaction will now avoid garbage-collecting tombstones that potentially delete data in commitlog. This prevents data resurrection in the event that a node crashes and replays commitlog. This is rare since generally commitlog data is relatively fresh and tombstones that delete such data would not be garbage collected for other reasons. #14870 Performance: Bloom filter efficiency can be reduced after node operation. When writing an sstable, ScyllaDB estimates how many partitions it will have in order to size the bloom filter correctly. In some cases, the estimation was suboptimal for TWCS. #15704 Stability: commitlog replay can cause abort due to over-extended skip. During commitlog replay, ScyllaDB skips over corrupted sections. However if the corrupted section also has corrupt size, it can lead to a crash. #15269 Stability: compaction_manager::perform_cleanup does not handle condition_variable_timed_out, than may cause nodetool cleanup to fail with exit status 2. #15669 Stability: tasks: dangling reference to task’s child pointer #16380 Stability: ICS is incorrectly calculating GC before on its own, without taking into account the GC mode and other factors. This might lead to ICS wrongly assuming data can be GCed, but then compaction down the road realizes the data cannot be GCed Stability ICS does not respect staleness condition for sstable runs possibly shadowing data in memtable Stability: ICS cross-tier tombstone compaction can be delayed indefinitely and doesn’t respect ‘tombstone_compaction_interval’ Stability: ICS is not honoring tombstone GC mode when scheduling compaction jobs Stability: Fix a race condition in LDAP setup Stability: A regression in IPv6 address formatting, which caused nodetool problems, like breaking when there is an Alternator GSI in the database #16153, or cause a node to be stuck with “?U” status and Host ID is "null #16039 Stability: a rare case in which after a node restart, during a short window of time, in which workload prioritizations are not yet propagated to the node, a node can assume wrong priorities, and crash with OOM. Stability: Adding a column to the base table should invalidate prepared statements for views #16392 Stability: nodes crashing during repair operations (due to no reader-closing on unexpected exception) #16606 Stability: tombstone might not be garbage-collected due to conflicts with data in commitlog. #15777 --- ### Page: https://forum.scylladb.com/t/scylladb-enterprise-release-2024-1-0-tools-configuration-admin-api-and-monitoring/1294 Title: ScyllaDB Enterprise Release 2024.1.0 - Tools, Configuration, Admin API and Monitoring - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: Scylla Enterprise 2024.1 Release Notes. Tools The CQL shell, cqlsh, has been separated into its own repository. As part of that change, cqlsh is now compatible with Python 3. CQLSh is now available as a Docker imag… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-enterprise-release-2024-1-0-tools-configuration-admin-api-and-monitoring/1294 ## Headings Structure: H1: ScyllaDB Enterprise Release 2024.1.0 - Tools, Configuration, Admin API and Monitoring H3: Tools H3: Monitoring, tracing and logging H3: Admin REST API H3: Configuration H2: Additional bug fixes H3: Related topics ## Main Content: H1: ScyllaDB Enterprise Release 2024.1.0 - Tools, Configuration, Admin API and Monitoring H3: Tools H3: Monitoring, tracing and logging H3: Admin REST API H3: Configuration H2: Additional bug fixes H3: Related topics Scylla Enterprise 2024.1 Release Notes. As part of that change, cqlsh is now compatible with Python 3. CQLSh is now available as a Docker image, and in PiPy, allowing you to easily use it when you do not need the entire ScyllaDB server, for example with Scylla Cloud. See Scylla SSTable docs for more info. Scylla Monitoring Stack released 4.6 and later supports ScyllaDB Enterprise 2024.1 metrics related updates below: The scylla.yaml configuration items are now documented in the documentation website. New and updated configuration options: Index_cache_fraction is the maximum fraction of cache memory permitted for use by index cache. Clamped to the [0.0; 1.0] range. Must be small enough to not deprive the row cache of memory, but should be big enough to fit a large fraction of the index. The default value 0.2 means that at least 80% of cache memory is reserved for the row cache, while at most 20% is usable by the index cache. The following issues have been fixed on top of what was fixed in Scylla Open Source 5.4.0, with open source reference if available. In addition, all relevant bug fixes from 2023.1.x are fixed in 2024.1.0 --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-38-2024-02-09/1295 Title: Last week in scylla-cluster-tests.git master (issue #38; 2024-02-09) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the d92d7cc7…fba16698 range are covered. There were 7 non-merge commits from 6 Software Engin… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-38-2024-02-09/1295 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #38; 2024-02-09) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #38; 2024-02-09) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the d92d7cc7…fba16698 range are covered. There were 7 non-merge commits from 6 Software Engineers in Test and 1 Software Engineer in that period. Some notable commits: We’ve enabled the use of Scylla Manager in each of the Scylla tenants, thanks to the fix of an S3 access issue. This enhancement increases testing coverage for Scylla Manager in a multi-tenant environment. After addressing the issue discovered in scylla-bench v0.1.19, we bumped its version to v0.1.20 utilizing the gocql driver that supports Scylla’s ‘tablets’ feature. Several Gemini tests will be running on bigger instances to mitigate issues caused by overloading the cluster. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-217-2024-02-11/1298 Title: Last week in scylladb.git master (issue #217; 2024-02-11) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 017a574b16…7a710425f0 range are covered. There were 115 non-merge commits from 19 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-217-2024-02-11/1298 ## Headings Structure: H1: Last week in scylladb.git master (issue #217; 2024-02-11) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #217; 2024-02-11) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 017a574b16…7a710425f0 range are covered. There were 115 non-merge commits from 19 authors in that period. Some notable commits: A bug in the row cache could cause a query to return incomplete data if, under moderate load, the query was preempted. This is now fixed. Note that QUORUM reads would usually see the missing data completed from the other replica, so the bug is mostly visible with consistency level ONE or similar. A bug was fixed where a deletion in a base table that affected a large number of materialized view rows to be updated could cause some updates to be missed. A possible crash in the REST API when stopping compaction on a keyspace was fixed. The CQL date type is now more resistant to overflow. An inadvertant protocol change that caused streaming in a mixed version cluster to fail has been rectified. The REST API for moving a tablet now validates that replication strategy constraints (for example, rack placement) are not violated. A keyspace that has tablets enabled can now only be created if the cluster has sufficient nodes to satisfy the replication factor constraint. The describe_ring API enumerates what fraction of the token space each node is responsible for. It now supports tablets, as each table has its own node to token range mapping. Streaming now has additional protection for the case where the streamed table is dropped during the process. Streaming is now more careful to avoid locking cluster metadata, allowing more tablet migrations to proceed in parallel. The native nodetool command now supports the describering command. The REST API for moving a tablet now allows concurrent tablet migration. An additional REST API that works on an arbitrary keyspace was determined to be unused and redundant, and therefore removed. There is now a procedure for upgrading a cluster from using gossip-managed to raft-managed, and for recovering a cluster that lost its raft quorum. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/unexpected-exception-when-pinging/1301 Title: Unexpected exception when pinging - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hey, Recently I’ve deployed many scylla nodes on AWS but I’m constantly running into WARN direct_failure_detector - unexpected exception when pinging a84dc681-ddee-4c0e-b140-4c686c1d7297: seastar::rpc::unknown_verb_erro… Language: en Canonical URL: https://forum.scylladb.com/t/unexpected-exception-when-pinging/1301 ## Headings Structure: H1: Unexpected exception when pinging H3: Related topics ## Main Content: H1: Unexpected exception when pinging H3: Related topics Hey, Recently I’ve deployed many scylla nodes on AWS but I’m constantly running into WARN direct_failure_detector - unexpected exception when pinging a84dc681-ddee-4c0e-b140-4c686c1d7297: seastar::rpc::unknown_verb_error (unknown verb) (for many different hosts) which on some nodes result in shutdown and on others just spams. Has anyone else encountered this? Thank you for reporting this, it is likely an issue we need to fix. That said, it might be better to move this discussion to the ScyllaDB issue tracker, can you please open a new issue and add the following details (in addition to your description above): What ScyllaDB version are you running? Are all nodes running the same version? Do you see any other errors in the logs? Thanks for the quick reply! All my nodes (~100) are running version 5.2.9, with lightweight resources: 2Gi of memory and 2 vCPUs. I know it might be weird but that’s to simulate a very specific workload… I also saw multiples of: I suspect it’s some kind of resource problem as I am trying to have them really slim, but didn’t saw something indicative of that. @kbr you might want to take a look at this, AFAIK you were chasing this issue with IP address translation. We also need to get to the bottom of this unknown verb error. Like I said above, it is best to open an issue in the ScyllaDB issue tracker because the forum is not the best venue for investigating problems. You can use this link to open an issue: https://github.com/scylladb/scylladb/issues/new --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-4-3/1302 Title: [RELEASE] ScyllaDB 5.4.3 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.4.3, a bugfix release of the ScyllaDB 5.4 stable branch. ScyllaDB Open Source 5.4.3, like all past and future 5.x.y releases, is backward compatible and supports rolling… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-4-3/1302 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.4.3 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.4.3 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.4.3, a bugfix release of the ScyllaDB 5.4 stable branch. ScyllaDB Open Source 5.4.3, like all past and future 5.x.y releases, is backward compatible and supports rolling upgrades. Users are encouraged to upgrade to 5.4.3. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-2-18/1303 Title: [RELEASE] ScyllaDB Enterprise 2022.2.18 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2022.2.18, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.2 Feature Release. Note the latest ScyllaDB Enterprise release is 2024… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-2-18/1303 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2022.2.18 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2022.2.18 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2022.2.18, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.2 Feature Release. Note the latest ScyllaDB Enterprise release is 2024.1 LTS, and you are encouraged to upgrade to it in coordination with the ScyllaDB Support team. Get ScyllaDB Enterprise 2022.2.18 (customers only, or 30-day evaluation) Upgrade from 2022.1.x to 2022.2.y Upgrade from ScyllaDB Enterprise 2022.2.x to 2024.1.y Upgrade from ScyllaDB Open Source 5.1 to ScyllaDB 2022.2 The following issues are fixed in this release (with an open-source reference, if available): --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-2-15/1306 Title: [RELEASE] ScyllaDB 5.2.15 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.2.15, a bugfix release of the ScyllaDB 5.2 stable branch. ScyllaDB Open Source 5.2.15, like all past and future 5.x.y releases, is backward compatible, and supports roll… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-2-15/1306 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.2.15 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.2.15 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.2.15, a bugfix release of the ScyllaDB 5.2 stable branch. ScyllaDB Open Source 5.2.15, like all past and future 5.x.y releases, is backward compatible, and supports rolling upgrades. Note that the latest ScyllaDB Open Source stable release is 5.4 and you are encouraged to upgrade to it. Issue fixed in this release: CQL: Server-side DESCRIBE should include scylladb-specific options like paxos_grace_seconds, and tombstone_gc #12389 Correctness: Inserted data only becomes available after restart. The root cause is a very rare bug in the cache, which is hidden (in most cases) by replication and reconciliation #16759. Thanks to user rngcntr for reporting and helping debug it. Stability: Adding a column to the base table should invalidate prepared statements for views #16392 Stability: Bootstrap fails due to inconsistent schema after group0 catch up when table was recreated in the past, in an older version of Scylla #16683 Stability: database_test may fail due to ignored exceptional future on table drop #14971 Stability: nodes crashing during repair operations (due to no reader-closing on unexpected exception) #16606 Stability: Raft system storage needs to save the snapshot descriptor and truncate the log atomically. #9603 The original fix proved to be incomplete. Stability: raft: large delays between io_fiber iterations in schema change test. #15622. ScyllaDB uses two separate memory reservation systems for memtables: user, used for user writes, and system, used for ScyllaDB’s own writes. The root cause was Raft did not use the system reservation. Stability: Task manager api responds with Internal Server Error #14914 Stability: tasks: dangling reference to task’s child pointer #16380 Stability: When using Raft for topology and schema changes, ScyllaDB does not force the schema and topology to be transferred to new nodes. #14066 Stability: mintimeuuid() call with large negative timestamps may crash ScyllaDB. #17035 Setup: scylla_raid_setup: faillback to other paths when UUID not available #13803 --- ### Page: https://forum.scylladb.com/t/reader-concurrency-semaphore-multiprocessing-timeout-cpu-overload-100/1309 Title: Reader_concurrency_semaphore: Multiprocessing, timeout, CPU overload 100% - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi all. Can you give me some advice? [cqlsh 5.0.1 | Cassandra 3.0.8 | CQL spec 3.3.1 | Native protocol v4] docker-compose exec scylla_node scylla --version 5.2.7-0.20230821.e0ebc95025d1 I’m using multiprocessing + asy… Language: en Canonical URL: https://forum.scylladb.com/t/reader-concurrency-semaphore-multiprocessing-timeout-cpu-overload-100/1309 ## Headings Structure: H1: Reader_concurrency_semaphore: Multiprocessing, timeout, CPU overload 100% H3: Related topics ## Main Content: H1: Reader_concurrency_semaphore: Multiprocessing, timeout, CPU overload 100% H3: Related topics Hi all. Can you give me some advice? [cqlsh 5.0.1 | Cassandra 3.0.8 | CQL spec 3.3.1 | Native protocol v4] I’m using multiprocessing + asyncio for parallelism and asynchronous requests without blocking. Database queries are executed cyclically. Each process has its own cluster + session. Everything works fine for a while, but then the SELECT queries fail. Scylla runs on a single node, the configuration is simple (--memory=8G, --smp=1). I looked at the load - there are enough resources on the server. The Scylla container is constantly running at 100+% CPU. If you run it in only one process, everything works without errors. CPU load 90+%. Do I understand correctly that the problem is the use of multiprocessing and incorrect configuration (do you need to allocate so much CPU so that it does not exceed 100% of the load on the database?)? Seems like you are overloading your node. You either need to reduce load (use single loader process), or give ScyllaDB more than 1 CPUs so it can spread the load. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-218-2024-02-18/1311 Title: Last week in scylladb.git master (issue #218; 2024-02-18) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 7a710425f0…e132ffdb60 range are covered. There were 92 non-merge commits from 20 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-218-2024-02-18/1311 ## Headings Structure: H1: Last week in scylladb.git master (issue #218; 2024-02-18) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #218; 2024-02-18) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 7a710425f0…e132ffdb60 range are covered. There were 92 non-merge commits from 20 authors in that period. Some notable commits: The in-memory representation of cells has been optimized; tables with very small cells should be a significant improvement in memtable and cache density. The native nodetool now supports the describecluster getendpoint, gossipinfo compactionstats, and viewbuildinfo commands. A bug in repair that could have caused problems with mixed-version clusters has been fixed. Unpaged queries will no longer fail when they reach the tombstone limit. The tombstone limit ends a page in an unpaged query, but unpaged queries are allowed to scan tombstones until they time out of reach the memory limit. It is strongly recommended to use paged queries. Alternator is ScyllaDB’s implementation of the DynamoDB API. A correctness problem when using a Alternator Global Secondary Index that has two key columns has been fixed. The topology code will now issue updates to topology tables concurrently, improving performance on large clusters. The topology code ignores nodes that are being removed. This has been changed to cooperate with tablet migration. During shutdown, we now wait for compactions to complete when shutting down system tables to avoid updates to system.compaction_history from racing with its shutdown. Alternator, ScyllaDB’s implementation of the DynamoDB API, will now reject attempts to enable TTL if running with tablets, as that is not implemented yet. The natural_endpoints API, used to find which nodes serve a particular key, now supports tablets. There are now more example scripts for use with the scylla sstables command. The topology code is now more careful to forget old node IP address after an address change. The nodetool upgradesstables command now correctly handles keyspaces that use tablets. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2023-1-5/1313 Title: [RELEASE] ScyllaDB Enterprise 2023.1.5 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2023.1.5, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2023.1 LTS Release. 2023.1.5 patch release includes multiple minor bug fixes… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2023-1-5/1313 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2023.1.5 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2023.1.5 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2023.1.5, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2023.1 LTS Release. 2023.1.5 patch release includes multiple minor bug fixes. ScyllaDB Enterprise’s latest Long Term Support (LTS) is 2024.1. While we continue to support 2023.1, you are encouraged to upgrade to it in coordination with the ScyllaDB Support team. The following issues are fixed in this release (with an open-source reference, if available): --- ### Page: https://forum.scylladb.com/t/how-to-use-feast-feature-store-with-scylladb-cloud/1315 Title: How to use Feast feature store with ScyllaDB Cloud? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Feast.dev is an open source feature store tool that integrates with Cassandra/ScyllaDB. Here’s an example configuration with Feast and ScyllaDB Cloud: project: scylla_feature_repo registry: data/registry.db provider: lo… Language: en Canonical URL: https://forum.scylladb.com/t/how-to-use-feast-feature-store-with-scylladb-cloud/1315 ## Headings Structure: H1: How to use Feast feature store with ScyllaDB Cloud? H3: Related topics ## Main Content: H1: How to use Feast feature store with ScyllaDB Cloud? H3: Related topics Feast.dev is an open source feature store tool that integrates with Cassandra/ScyllaDB. Here’s an example configuration with Feast and ScyllaDB Cloud: For more information about the integration and limitations, check out the Feast documentation. For an example feature store app with ScyllaDB, check out this sample app. --- ### Page: https://forum.scylladb.com/t/unavailableexception-x-required-but-only-0-alive/1318 Title: UnavailableException (x required but only 0 alive) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi to everyone, We are trying to migrate from cassandra to scylla 4.6 version. For scylla we have a cluster with 4 nodes. Memory : 100 GB per instance Disk : 1Tb Cpu pinning has been done. Compaction strategy for a… Language: en Canonical URL: https://forum.scylladb.com/t/unavailableexception-x-required-but-only-0-alive/1318 ## Headings Structure: H1: UnavailableException (x required but only 0 alive) H3: Related topics ## Main Content: H1: UnavailableException (x required but only 0 alive) H3: Related topics Hi to everyone, We are trying to migrate from cassandra to scylla 4.6 version. For scylla we have a cluster with 4 nodes. Memory : 100 GB per instance Disk : 1Tb Cpu pinning has been done. Compaction strategy for all tables as of now : Size Tiered. RF is 4. Since all nodes in same data centre we are using SimpleSnitch. com.datastax.driver.core.exceptions.UnavailableException: Not enough replicas available for query at consistency LOCAL_ONE (1 required but only 0 alive) Getting this error intermittently rarely but causes a lot of write errors in that time. Scylla yaml also attached. Grafana screenshots attached for that time with all node view. What could be the causes for this error, if all nodes are up and RF is also 4? Scylla 4.6 is quite an old release at this point. It is possible that the problem you are suffering from was fixed in more current ScyllaDB releases. Please upgrade to 5.2 (or even better, 5.4) and report back whether the problem persists. Thanks @Botond_Denes , will upgrade and replicate I see you’re using PasswordAuthenticator, please also make sure to alter RF accordingly, it may cause similar issues with default RF (see Enable Authentication | ScyllaDB Docs) --- ### Page: https://forum.scylladb.com/t/how-much-space-does-a-scylla-pod-would-require-to-deploy-scylla-db-in-k8s/1320 Title: How much space does a scylla pod would require to deploy scylla db in k8s? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I am trying to deploy scyllaDB as a k8s service. Currently, trying to create a cluster with 3 pods. As of now, I attached 15GB storage to each of the pod - total of 45 GB to scylla cluster. but I am still getting below … Language: en Canonical URL: https://forum.scylladb.com/t/how-much-space-does-a-scylla-pod-would-require-to-deploy-scylla-db-in-k8s/1320 ## Headings Structure: H1: How much space does a scylla pod would require to deploy scylla db in k8s? H3: Related topics ## Main Content: H1: How much space does a scylla pod would require to deploy scylla db in k8s? H3: Related topics I am trying to deploy scyllaDB as a k8s service. Currently, trying to create a cluster with 3 pods. As of now, I attached 15GB storage to each of the pod - total of 45 GB to scylla cluster. but I am still getting below error: Any idea about optimum and minimum volume to be mounted to each pod? This error is not about amount of storage required, but number of maximum aio operations, which is a kernel property. As message suggests, you should increase /proc/sys/fs/aio-max-nr on the k8s Node. Or use Scylla Operator where this value can be set as part of ScyllaCluster.spec.sysctls. Setting it to 30000000 should be enough. --- ### Page: https://forum.scylladb.com/t/release-scylladb-cloud-21-february-2024/1322 Title: [RELEASE] ScyllaDB Cloud - 21 February 2024 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: We are happy to announce a new update to the ScyllaDB Cloud service. New features and enhancements: We’ve added support for 6,000GB storage on n2-highmem-8 instances in all ScyllaDB Cloud supported regions. We’ve… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-cloud-21-february-2024/1322 ## Headings Structure: H1: [RELEASE] ScyllaDB Cloud - 21 February 2024 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Cloud - 21 February 2024 H3: Related topics We are happy to announce a new update to the ScyllaDB Cloud service. New features and enhancements: We’ve added support for 6,000GB storage on n2-highmem-8 instances in all ScyllaDB Cloud supported regions. We’ve added support for 3,000GB/6,000GB/9,000GB storage on n2-highmem-64 instances in all ScyllaDB Cloud supported regions. Added support for ScyllaDB Enterprise versions 2024.1.0, 2023.1.5, 2022.2.17, and 2022.2.18. Added support for Scylla Monitoring Stack 4.6.2. --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-1/1323 Title: [RELEASE] ScyllaDB Enterprise 2024.1.1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2024.1.1, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2024.1 LTS Release. 2024.1.1 patch release includes multiple minor bug fixes… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-1/1323 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2024.1.1 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2024.1.1 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2024.1.1, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2024.1 LTS Release. 2024.1.1 patch release includes multiple minor bug fixes. The following issues are fixed in this release (with an open-source reference, if available): --- ### Page: https://forum.scylladb.com/t/scylladb-university-live-march-19-2024/1324 Title: ScyllaDB University Live - March 19, 2024 - Announcements - ScyllaDB Community NoSQL Forum Meta Description: ScyllaDB University Live is taking place next month. It’s a half day of free online training with some of our top engineers and experts. We will cover topics like data modeling, getting started with ScyllaDB, building a… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-university-live-march-19-2024/1324 ## Headings Structure: H1: ScyllaDB University Live - March 19, 2024 H3: Related topics ## Main Content: H1: ScyllaDB University Live - March 19, 2024 H3: Related topics ScyllaDB University Live is taking place next month. It’s a half day of free online training with some of our top engineers and experts. We will cover topics like data modeling, getting started with ScyllaDB, building a real-world application, and more. The sessions will also include hands-on labs, and you’ll have a chance to learn by doing. Save your spot, hope to see you there! This is happening next week. You can still save your spot here. --- ### Page: https://forum.scylladb.com/t/seastar-httpd-example/1326 Title: Seastar httpd example - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I am building simple httpd application example in apps. I wonder if http source in seastar handles json body when sending post request and if not if there are some projects that do so : scylladb database, etc? Thanks Language: en Canonical URL: https://forum.scylladb.com/t/seastar-httpd-example/1326 ## Headings Structure: H1: Seastar httpd example H3: Related topics ## Main Content: H1: Seastar httpd example H3: Related topics I am building simple httpd application example in apps. I wonder if http source in seastar handles json body when sending post request and if not if there are some projects that do so : scylladb database, etc? Hello, don’t know about seastar example but for sure we handle json and post in alternator (dynamodb api), see scylladb/alternator/server.cc at 7cb1c10fed4874bdce752852fdc4a0ee1abfb8d2 · scylladb/scylladb · GitHub Thanks for information, I got inspiration from server part of scylladb and I wanted to test server performance using seastar library. The idea was to bench with the oio-sds (GitHub - open-io/oio-sds: High Performance Software-Defined Object Storage for Big Data and AI, that supports Amazon S3 and Openstack Swift) open source solution where a proxy server targets an sqlite3 database as backend. Here my idea was to build a proxy that writes into foundationdb database: GitHub - laitassou/proxy I am disappointed because . both solutions struggle at about 300 write operation per seconds: Any idea because I am far from httpd server performance here HTTPD benchmark · scylladb/seastar Wiki · GitHub I know I am not using any dpdk etc? but I expect much more. Here is my startup program: Any idea or advice ? Thanks --- ### Page: https://forum.scylladb.com/t/storage-with-i4i-instance-for-scylla/1327 Title: Storage with i4i instance for scylla - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi Community, I am trying to setup scylla 3 node cluster on production on i4i instance. I am a little apprehensive about using ephemeral storage as a primary disk. I have read @Felipe Cardeneti Mendes answer to this her… Language: en Canonical URL: https://forum.scylladb.com/t/storage-with-i4i-instance-for-scylla/1327 ## Headings Structure: H1: Storage with i4i instance for scylla H3: Related topics ## Main Content: H1: Storage with i4i instance for scylla H3: Related topics Hi Community, I am trying to setup scylla 3 node cluster on production on i4i instance. I am a little apprehensive about using ephemeral storage as a primary disk. I have read @Felipe Cardeneti Mendes answer to this here : Top Mistakes with ScyllaDB: Storage - ScyllaDB : Locally-attached disk mythbusting But I’m still not sure if a node goes down or is stopped and started, other than backup through replication what are my options for data recovery? Also I want to know would it be suggested to use an ebs disk attached with i4i for backup. If yes how will that work? Read and write through nvme ssd and backup on ebs or is it not possible. If attaching ebs with nvme disk wouldn’t be utilising i4i’s cabability, then what would be the next best bet for instance type for scylla with persistant storage? Hi, we generally don’t recommend setups other than local NVMs, from our experience it’s very unlikely that all instances would disappear at once. You can also use RF=5 or multi-dc. If you still want to experiment you can read how discord currently does it, but it’s a complex setup and may be more likely you mismanage it than the problem you want to protect from. See How Discord Stores Trillions of Messages some ideas. --- ### Page: https://forum.scylladb.com/t/creation-of-keyspaces-while-deploying-scylladb-on-k8s/1329 Title: Creation of Keyspaces while deploying ScyllaDB on k8s? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I was going through the steps of deploying scyllaDB on kubernetes and I successfully deployed it on k8s. Now, I want to automate the creation of keyspaces in database (if not possible, create some presumed keyspaces. Is… Language: en Canonical URL: https://forum.scylladb.com/t/creation-of-keyspaces-while-deploying-scylladb-on-k8s/1329 ## Headings Structure: H1: Creation of Keyspaces while deploying ScyllaDB on k8s? H3: Related topics ## Main Content: H1: Creation of Keyspaces while deploying ScyllaDB on k8s? H3: Related topics I was going through the steps of deploying scyllaDB on kubernetes and I successfully deployed it on k8s. Now, I want to automate the creation of keyspaces in database (if not possible, create some presumed keyspaces. Is there any way that It can be done while deploying database on k8s? At this point you should follow the same steps you’d use with a regular ScyllaDB deployment, we don’t automate keyspace creation any further. You’ll find kubectl exec helpful for running cqlsh on your ScyllaDB nodes. --- ### Page: https://forum.scylladb.com/t/which-is-the-right-choice-for-an-app-dating-scylladb-or-mongodb/1330 Title: Which is the right choice for an app dating ScyllaDB or MongoDB? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: MongoDB Advantages for Dating Apps: Flexible Schema: MongoDB’s ability to store semi-structured data without a fixed schema is particularly beneficial for dating apps, which often need to accommodate a wide variety of… Language: en Canonical URL: https://forum.scylladb.com/t/which-is-the-right-choice-for-an-app-dating-scylladb-or-mongodb/1330 ## Headings Structure: H1: Which is the right choice for an app dating ScyllaDB or MongoDB? H3: Related topics ## Main Content: H1: Which is the right choice for an app dating ScyllaDB or MongoDB? H3: Related topics MongoDB Advantages for Dating Apps: Flexible Schema: MongoDB’s ability to store semi-structured data without a fixed schema is particularly beneficial for dating apps, which often need to accommodate a wide variety of user profile information and adapt to new features. Real-time Updates: MongoDB’s support for real-time updates through change streams is crucial for dating apps, where users expect to see new matches, messages, and other updates instantly. Geospatial Queries: Many dating apps use location-based features, and MongoDB’s built-in support for geospatial data makes it easier to implement these features. JSON Support: MongoDB stores data in a format similar to JSON, which is widely used in web development. This makes it easier for developers to work with the data, as they can use the same data structures and formats on the server and client sides. Here are my questions: How does ScyllaDB compare to MongoDB in terms of handling semi-structured data and schema flexibility? Are there any third-party solutions or workarounds needed to achieve similar functionality? How does ScyllaDB’s data storage and retrieval process compare to MongoDB’s JSON-like format in terms of performance and ease of use for web development? Are there any specific considerations or third-party tools that need to be used to work with JSON data effectively in ScyllaDB? Does ScyllaDB offer built-in support for geospatial queries, and if not, what are the recommended approaches or third-party solutions for implementing location-based features in dating apps? I’m looking for insights from the community on how ScyllaDB can be optimized for dating apps, especially in areas where MongoDB might have an advantage. Any advice, best practices, or experiences shared would be greatly appreciated. Mongo = Document Scylla = Big column You can build the same things? yes Is it the same way to work and think? No Could you please do a favor for me and give me more information and some keywords or books to help me understand and implement it? I am learning back-end. I have 1.5 years with Elixir and Phoenix, but I started to learn with ScyllaDB University, and it’s clear even to me. I am a beginner. I think that Scylla Devhub and University contain all you need to know about Scylla and how to model and administrate data. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-219-2024-02-25/1333 Title: Last week in scylladb.git master (issue #219; 2024-02-25) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the e132ffdb60…57c408ab5d range are covered. There were 104 non-merge commits from 14 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-219-2024-02-25/1333 ## Headings Structure: H1: Last week in scylladb.git master (issue #219; 2024-02-25) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #219; 2024-02-25) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the e132ffdb60…57c408ab5d range are covered. There were 104 non-merge commits from 14 authors in that period. Some notable commits: Change Data Capture (CDC) has the concept of “generations” that describe changes to cluster topology. Each time topology changes, a new generation is creates with a new set of CDC streams for the application to listen to. We now allow writes to be written to the previous generation’s streams, to accomodate clients with unsynchronized clocks, or with lightweight transactions (LWT) which can have the same effect as unsynchronized clocks. This also works with consistent topology. The native nodetool command now supports the following subcommands: The range_to_endpoint_map API, that describes the ring topology for a keyspace, now supports tablets. Note it needs an additional table parameter for that. The repair history table records the last time token ranges were repaired on a particular node; this in turn helps tombstone garbage collection determine whether tombstones can be purged or not. The table is now updated in tablet mode as well. The repair history table now loads faster. A hang when a decommissioned node is restarted was fixed; the operation now correctly fails. The maintenance port is a unix-domain socket that can accept CQL statements even when the node is degraded. It now accepts connections from users in the server’s group as a simple authentication mechanism. We now recover from tablet migration failures. In general, failures are retried, unless the node is being removed. Materialized view memory accounting is now more accurate. This reduces the risk of running out of memory in materialized view intensive workloads. Bundled tools now use the Seastar epoll reactor backend rather than linux-aio; this reduces the risk of startup failures. The SELECT * FROM mutation_fragments() statement allows inspecting the underlying data storage in tables. It now supports tablets. Cluster topology can be locked to avoid queries and operations from having the topology change under their feet; there is now a detector to see if locks are held for too long. The removenode --force operation has been deprecated for gossip mode and disabled for consistent topology mode as unsafe. The cluster_status virtual table now works in consistent topology mode. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/cant-update-udt-using-gocqlx/1336 Title: Can't update UDT using gocqlx - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi so I have a usecase in which a user can create a secondary index on any table by calling an API. And I’m recording these indexes in a common table using map<text, frozen> where indexes is a User Defined Type: CREATE … Language: en Canonical URL: https://forum.scylladb.com/t/cant-update-udt-using-gocqlx/1336 ## Headings Structure: H1: Can't update UDT using gocqlx H3: Related topics ## Main Content: H1: Can't update UDT using gocqlx H3: Related topics Hi so I have a usecase in which a user can create a secondary index on any table by calling an API. And I’m recording these indexes in a common table using map where indexes is a User Defined Type: When the API is called the secondary index is created. This part is working fine. But updating the row in the common tables is not. All the data in the UDT is stored as null: Here’s my gocqlx code: I hope the code is self-explanatory enough but all I’m doing is creating a map of string->indexes(UDT) to put in the common table and updating the common table with it. Side note: why is the documentation on gocqlx so bad? Hi, this question was answered in this github issue: Can’t update UDT using gocqlx · Issue #264 · scylladb/gocqlx · GitHub. In short you have to add cql: struct tags to map specify the CQL field name to be mapped to a struct field. Information about it can be found in gocql documentation: gocql/doc.go at 34fdeebefcbf183ed7f916f931aa0586fdaa1b40 · gocql/gocql · GitHub. --- ### Page: https://forum.scylladb.com/t/i-am-not-able-to-execute-triggers-query-in-cql-shell/1339 Title: I am not able to Execute Triggers Query in CQL shell - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Query: CREATE TRIGGER myTrigger ON t USING 'org.apache.cassandra.triggers.InvertedIndex'; Exception: Does ScyllaDb support triggers or not?? Language: en Canonical URL: https://forum.scylladb.com/t/i-am-not-able-to-execute-triggers-query-in-cql-shell/1339 ## Headings Structure: H1: I am not able to Execute Triggers Query in CQL shell H3: Related topics ## Main Content: H1: I am not able to Execute Triggers Query in CQL shell H3: Related topics Does ScyllaDb support triggers or not?? No, triggers are not supported by ScyllaDB. Thanks For the information… --- ### Page: https://forum.scylladb.com/t/scylla-introduces-json-support-to-select-and-insert-statements-but-not-able-to-understand-how-to-execute-them/1341 Title: Scylla introduces JSON support to SELECT and INSERT statements. But not able to understand how to execute them? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi @denesb, I am not able to understand how to execute JSON queries. Any example would be helpful along with prerequisite queries. Language: en Canonical URL: https://forum.scylladb.com/t/scylla-introduces-json-support-to-select-and-insert-statements-but-not-able-to-understand-how-to-execute-them/1341 ## Headings Structure: H1: Scylla introduces JSON support to SELECT and INSERT statements. But not able to understand how to execute them? H3: Related topics ## Main Content: H1: Scylla introduces JSON support to SELECT and INSERT statements. But not able to understand how to execute them? H3: Related topics I am not able to understand how to execute JSON queries. Any example would be helpful along with prerequisite queries. Note, that object keys are always string in JSON, even if in this case, the map represented by the JSON object is map. --- ### Page: https://forum.scylladb.com/t/cannot-remove-dn-node-from-my-cluster/1342 Title: Cannot remove DN node from my cluster - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: When I use ‘nodetool removenode’ to remove a DN node from my cluster(with 143 nodes and scylla version 5.1.15), I encountered an error and the node failed to be removed finally. Here is the error message: [HOSTNAME-01 … Language: en Canonical URL: https://forum.scylladb.com/t/cannot-remove-dn-node-from-my-cluster/1342 ## Headings Structure: H1: Cannot remove DN node from my cluster H3: Related topics ## Main Content: H1: Cannot remove DN node from my cluster H3: Related topics When I use ‘nodetool removenode’ to remove a DN node from my cluster(with 143 nodes and scylla version 5.1.15), I encountered an error and the node failed to be removed finally. Here is the error message: I’m curious why does this happen and how can I solve this problem. I see it in the code because the obtained ops_uuids is empty, but the command is correct. Is this a bug? @Botond_Denes I don’t know what that means. @Asias_He? In any case, this may be a bug that was fixed in later releases, 5.1 is not supported anymore. --- ### Page: https://forum.scylladb.com/t/what-is-internode-compression-and-how-does-it-work/1343 Title: What is internode compression and how does it work? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: What is internode compression? How does it work? How is it configured? Language: en Canonical URL: https://forum.scylladb.com/t/what-is-internode-compression-and-how-does-it-work/1343 ## Headings Structure: H1: What is internode compression and how does it work? H3: Related topics ## Main Content: H1: What is internode compression and how does it work? H3: Related topics What is internode compression? How does it work? How is it configured? By default, ScyllaDB uses LZ4 compression for internode communication, which helps in reducing network bandwidth usage and improving overall performance. Setting internode_compression in ScyllaDB impacts how data is compressed when transferred between nodes in a cluster. Changing this setting can have several impacts, such as more CPU overhead and memory usage because of compression and decompression during the communication among nodes. There is no negative impact on the network side, AFAIK. It is an optional configuration to set. You can read more about it in the Documentation. To elaborate, compression algorithms vary in compression speed, decompression speed, and compression ratio. Generally, there is a tradeoff between the LZ4 compression algorithm (also used by Cassandra), which offers a relatively low (about 25%) reduction in size, and more efficient algorithms such as ZSTD, which offer better compression rates. The tradeoff is that using ZSTD requires more time to compress (and decompress) the transmitted data. --- ### Page: https://forum.scylladb.com/t/questions-about-scylla-java-driver-4-x-shardaware-feature-and-nat-issue/1346 Title: Questions about scylla java driver 4.x ShardAware feature and NAT issue - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, I want to make clear my understanding about scylla java driver 4.x. = Background = We have a service application running in k8s with auto scale deployment in AWS VPC1, which uses Cassandra java driver 3.x to acces… Language: en Canonical URL: https://forum.scylladb.com/t/questions-about-scylla-java-driver-4-x-shardaware-feature-and-nat-issue/1346 ## Headings Structure: H1: Questions about scylla java driver 4.x ShardAware feature and NAT issue H3: Related topics ## Main Content: H1: Questions about scylla java driver 4.x ShardAware feature and NAT issue H3: Related topics I want to make clear my understanding about scylla java driver 4.x. = Background = We have a service application running in k8s with auto scale deployment in AWS VPC1, which uses Cassandra java driver 3.x to access Scylla enterprise cluster which runs in AWS VPC2. So actually there is a k8s NAT reside between the two systems. Because Scylla java driver can provide Shard-Aware feature comparing with Cassandra driver, so we decide to migrate application to use Scylla java driver. Now we have chosen Scylla java driver 4.15.0.1, and tested the migrated application to connect Scylla enterprise cluster port 19042, all looks good. = question = We are not sure whether the Shard-Aware feature works or not in our case. From the driver’s public doc, I cannot find answers. And there are no useful posts on internet to clarify. Why we have such a question? 1, ScyllaDB can provide normal service even Shard-Aware feature not work. 2, Scylla java driver 4.x doc does not state it is shard aware. 3, No public doc clarifies the relationship between Shard-Aware feature and the different serving port (9042 vs 19042). We are confused if we should use 19042 or 9042. 4, There is a public post (Connect Faster to ScyllaDB with a Shard-Aware Port - ScyllaDB) saying that Shard-Aware feature does not work when the application is behind a NAT, but from Scylla java driver 4.x code, the logic looks like the Shard-Aware feature can work even there is a NAT. So we are totally confused right now. In other words, we want to know which statements are true: 1, Scylla java driver 4.x is not Shard-Aware 2, Scylla java driver 4.x is Shard-Aware, but the feature does not work behind NAT 3, Scylla java driver 4.x is Shard-Aware and the feature works behind NAT 4, Scylla java driver 4.x must use 19042 to make Shard-Aware feature work 5, Scylla java driver 4.x can use 9042 or 19042, and both ports work for Shard-Aware feature. Can anyone help to clarify those questions? Let’s distinguish 2 terms here: When there is NAT between Scylla and the driver, source port visible by server will be different than the one that the client really uses - so Advanced Shard Awareness won’t work. Drivers that support Advanced Shard Awareness will detect this situation, fall back to port 9042 and build connection port the old way. Shard Awareness will still work in this situation, but fully creating a session will just take longer. If I remember correctly, Java Driver 4.x support Shard Awareness but not Advanced Shard Awareness - @piotr or @Bouncheck please correct me if I’m wrong. So to answer your questions: 1, Scylla Java Driver 4.x is Shard Aware, but does not have Advanced Shard Awareness (for now - because we plan to implement it). 2, Shard Awareness works behing NAT. Advanced Shard Awareness does not, but 4.x does not have it for now anyway. 3, Yes 4, No, Shard Awareness does not depend on the port. Only Advanced Shard Awareness does. 5, Correct. Thank you Lorak, you saved me. Your clarification is what I want. Without your explanation, I don’t know there are two terms for shard-aware, because there are no public docs mentioning it. Thank you very much. Hi Lorak, I have a following question. When there is NAT between Scylla and the driver … Shard Awareness will still work in this situation, but fully creating a session will just take longer Generally how long does this take? Is there any way to know when it’s done (in the code) ? Another question is that, when the driver has not created all connections to all the shards of some node, is the driver already ready to be used to send requests to scylla server? @Huiyong_Wang , for the follow up questions please start a new topic, that way it will be easier for others to find. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-220-2024-03-03/1350 Title: Last week in scylladb.git master (issue #220; 2024-03-03) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 57c408ab5d…94cd235888 range are covered. There were 47 non-merge commits from 9 authors in that period… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-220-2024-03-03/1350 ## Headings Structure: H1: Last week in scylladb.git master (issue #220; 2024-03-03) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #220; 2024-03-03) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 57c408ab5d…94cd235888 range are covered. There were 47 non-merge commits from 9 authors in that period. Some notable commits: Yet another case of repair failures if a table is dropped during repair was fixed. A potential data resurrection problem when cleanup was performed as a side-effect of regular compaction was fixed. A use-after-free in topology coordinator error handling was fixed. A data race in the service that propagates cache hit rate statistics across the cluster was fixed. We now lock raft group 0 while taking a snapshot (to populate another node) to preserve atomicity. The totimestamp() CQL function is now protected against undefined behavior when applied to extreme values. As a side effect of work to reduce test run-time, there is now a config item that can be used to change the internal page size. We now no longer emit a removed node notification to drivers when a node IP address changes. The notification confused drivers, causing then to lose the ability to reconnect. An internal RPC for schema changes, which may only run on shard 0, is now enforced to do so. A deadlock when adding or removing nodes to Raft was fixed. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/scylla-sorts-the-field-in-alphabet-order/1351 Title: Scylla sorts the field in alphabet order - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hey everyone. Is there any way to do not sort the field by alphabet order? Here is my cql CREATE TABLE IF NOT EXISTS auth.users ( email text, password text, refresh_token text, access_token text, user_id uuid, cre… Language: en Canonical URL: https://forum.scylladb.com/t/scylla-sorts-the-field-in-alphabet-order/1351 ## Headings Structure: H1: Scylla sorts the field in alphabet order H3: Related topics ## Main Content: H1: Scylla sorts the field in alphabet order H3: Related topics Hey everyone. Is there any way to do not sort the field by alphabet order? Here is my cql CREATE TABLE IF NOT EXISTS auth.users ( email text, password text, refresh_token text, access_token text, user_id uuid, created_at time, last_login time, PRIMARY KEY (email)); SELECT * FROM auth.users; email | access_token | created_at | last_login | password | refresh_token | user_id -------±-------------±-----------±-----------±---------±--------------±-------- Thanks in advance for any help I’m not sure I understand yoru question, you can determine the order by using the Clustering key. Read more about it in the Basic Data Modeling lesson on ScyllaDB University. Do you mean the columns themselves? Those are sorted indeed by alphabetical order, except for the key columns, those are left in the order provided by the CREATE TABLE statement. There is now way to turn this off. --- ### Page: https://forum.scylladb.com/t/trying-to-setup-a-2-node-multi-dc-cluster-for-the-first-time-seed-node-comes-online-fine-second-node-gets-stuck-repairing-tables-and-constantly-in-a-state-of-uj/1352 Title: Trying to setup a 2 node multi dc cluster for the first time... Seed node comes online fine, second node gets stuck repairing tables and constantly in a state of UJ - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Here are the relevant logs: Mar 03 20:35:18 node02 scylla[6377]: [shard 0:stre] repair - repair[21a0a883-863f-4a32-8a3f-ddd8452c3758]: sync data for keyspace=system_traces, status=started Mar 03 20:35:18 node02 scylla… Language: en Canonical URL: https://forum.scylladb.com/t/trying-to-setup-a-2-node-multi-dc-cluster-for-the-first-time-seed-node-comes-online-fine-second-node-gets-stuck-repairing-tables-and-constantly-in-a-state-of-uj/1352 ## Headings Structure: H1: Trying to setup a 2 node multi dc cluster for the first time... Seed node comes online fine, second node gets stuck repairing tables and constantly in a state of UJ H3: Related topics ## Main Content: H1: Trying to setup a 2 node multi dc cluster for the first time... Seed node comes online fine, second node gets stuck repairing tables and constantly in a state of UJ H3: Related topics Here are the relevant logs: As you can see, node2 is taking over 20 minutes to repair a few small tables (the cluster was only just setup an hour ago and has no data contained therein) and is never coming online. One of the most significant warnings from that long list of logs is this imho: raft_group_registry - (rate limiting dropped 2995 similar messages) Raft server id 545b3f00-e8b3-495b-8ffb-d4e82ce29474 cannot be translated to an IP address. I have no idea what that means unfortunately. I don’t recall setting up any raft related software. Here are my scylla.yaml changes to node1: And my cassandra-rackdc.properties for node1: And here are my scylla.yaml changes to node2: And my cassandra-rackdc.properties for node2: Only other significant change from defaults afaik was adding “–memory 20G --reserve-memory 6G” to SCYLLA_ARGS in /etc/default/scylla-server file. For reference, I am running a dedicated Hetzner EX44 Server in their helsinki datacenter. Any tips or suggestions would be greatly appreciated. I am running Scylla version 5.4.3-0.20240211.cf42ca0c2a65 on Debian 11 that I installed through the web installer. EDIT: Just found some more interesting log lines upon startup of my second datacenter node02: Please open an issue on GitHub (Sign in to GitHub · GitHub) and there post: --- ### Page: https://forum.scylladb.com/t/how-to-parse-the-commitlog/1354 Title: How to parse the commitlog? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Directly view the file CommitLog-2-8901636397.log as garbled code. How to make the file readable by humans? Is there any parsing tool available? Thanks! Language: en Canonical URL: https://forum.scylladb.com/t/how-to-parse-the-commitlog/1354 ## Headings Structure: H1: How to parse the commitlog? H3: Related topics ## Main Content: H1: How to parse the commitlog? H3: Related topics Directly view the file CommitLog-2-8901636397.log as garbled code. How to make the file readable by humans? Is there any parsing tool available? Thanks! Currently there are no end-user tools to parse commit logs. I do plan to add such a tool in the near future, so stay tuned. Out of curiosity, why would you like to dump the content of this commit-log file? There is an indirect way to dump the commit-log: Note that this method will make the data go through insertion to memtable, then flush. Some (or even all) of the data might be compacted away in the process. Thank you, I look forward to it. In order to facilitate the statistics of certain data, I would like to migrate the data to other databases through commitlog. Just an idea. This sounds like a risky reimplementation of CDC. You will need to deduplicate updates from each replica and lose the pre and post-images CDC provides. I do plan to add such a tool in the near future, so stay tuned. Hi, I use the pretty_printer function of frozen_mutation to print the content of commitlog, making it easy for humans to read. In order to be able to call it offline, I added a parsing tool. Now I have a question: how to initialize db and sys_ks? May I ask if you have any suggestions? Thank you. Starting those two services is quite risky in the sense that they might try to create files. Worse still, if you run this tool on a live ScyllaDB node, they might access and mutate the files of the ScyllaDB node. So I think the way to go forward is to cut the dependency between the commitlog replayer and these services, by creating an interface and two implementations. The interface would have all the methods that the commitlog replayer requires from these services, e.g. find_column_family(), extensions(), etc. One implementation would use the db and sys_ks to implement these methods, just like the current code. Another implementation would be in the tools, which would try to do just the minimum amount of implementation possible, to allow the commitlog replayer to replay the commitlog. This might need some experimenting to get it working. --- ### Page: https://forum.scylladb.com/t/are-updates-to-materialized-view-done-synchronously-with-updates-to-base-table/1355 Title: Are updates to materialized view done synchronously with updates to base table? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Are updates to materialized view done synchronously with updates to base table? A write request to base table will update base table and won’t return success to client until MV is also updated, or is update to MV done as… Language: en Canonical URL: https://forum.scylladb.com/t/are-updates-to-materialized-view-done-synchronously-with-updates-to-base-table/1355 ## Headings Structure: H1: Are updates to materialized view done synchronously with updates to base table? H3: Related topics ## Main Content: H1: Are updates to materialized view done synchronously with updates to base table? H3: Related topics Are updates to materialized view done synchronously with updates to base table? A write request to base table will update base table and won’t return success to client until MV is also updated, or is update to MV done async? In general, the answer is no, updates to the views are asynchronous: A write to the base table waits only until CL (consistency level) replicas of the base table are updated, but doesn’t wait for the write to the view table. So for example, if you write to a table with CL=QUORUM and then read from the same table with CL=QUORUM, you are guaranteed to see the value you wrote; But if you read from the view you are not guaranteed to see the new value. The new value will be readable in the view eventually, i.e., eventual consistency, but not immediately. As an extension over Cassandra, Scylla allows marking a materialized view as WITH synchronous_updates = true; (see explanation and syntax in ScyllaDB CQL Extensions | ScyllaDB Docs). When a view has synchronous updates enabled, the write only succeeds once CL copies were written in both base table and view. In this mode, the new data will be immediately readable from the view. The downside of synchronous-updates mode is higher latency for the writes, as well as the risk of unavailability - a write can fail despite being able to reach the desired consistency level in the base table, because of an inability to reach enough view replicas. In some cases, Scylla may do synchronous view updates even without this being explicitly requested - however, the user should not rely on this (and in Cassandra, those updates are always asynchronous). The aforementioned documentation explains: Even in an asynchronous view, some view updates may be done synchronously. This happens when the materialized-view replica is on the same node as the base-table replica. This happens, for example, in tables using vnodes where the base table and the view have the same partition key; But is not the case if the table uses tablets: With tablets, the base and view tablets may migrate to different nodes. In general, users should not, and cannot, rely on these serendipitous synchronous view updates; If synchronous view updates are important, mark the view explicitly with synchronous_updates = true. --- ### Page: https://forum.scylladb.com/t/configurationexception-unrecognized-strategy-option-dc2-scyllau/1358 Title: ConfigurationException: Unrecognized strategy option {DC2} scyllau - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I am trying to do the 101 course but when I tried this *CREATE KEYSPACE scyllaU WITH REPLICATION = {'class' : 'NetworkTopologyStrategy', 'DC1' : 3, 'DC2' : 2};* I got this ConfigurationException: Unrecognized strategy… Language: en Canonical URL: https://forum.scylladb.com/t/configurationexception-unrecognized-strategy-option-dc2-scyllau/1358 ## Headings Structure: H1: ConfigurationException: Unrecognized strategy option {DC2} scyllau H3: Related topics ## Main Content: H1: ConfigurationException: Unrecognized strategy option {DC2} scyllau H3: Related topics I am trying to do the 101 course but when I tried this *CREATE KEYSPACE scyllaU WITH REPLICATION = {'class' : 'NetworkTopologyStrategy', 'DC1' : 3, 'DC2' : 2};* ConfigurationException: Unrecognized strategy option {DC2} passed to org.apache.cassandra.locator.NetworkTopologyStrategy for keyspace scyllau I am on mac with docker ScyllaDB doesn’t recognize DC2 as a datacenter, that is why it complains about it. The DC you use for KEYSPACE replication needs to match the DCs the cluster is aware of from the Snitch, either explicitly with GossipingPropertyFileSnitch or implicitly from Ec2MultiRegionSnitch (for example) . You can use nodetool status to check which DCs the cluster is aware of. --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-2/1359 Title: [RELEASE] ScyllaDB Enterprise 2024.1.2 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2024.1.2, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2024.1 LTS Release. 2024.1.2 patch release includes multiple minor bug fixes… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-2/1359 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2024.1.2 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2024.1.2 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2024.1.2, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2024.1 LTS Release. 2024.1.2 patch release includes multiple minor bug fixes. The following issues are fixed in this release (with an open-source reference, if available): cdc: allow sending writes to the previous generation for some time after switching #7251, #15260. The previous bypass for this issue, retry the request, does not work when using LWT and CDC on the same table. The issue can appear both in CQL CDC and Alternator (Amazon DynamoDB compatible API) Streams. Stability: throughput_limits_test.TestCompactionLimitThroughput.test_can_limit_compaction_throughput is flakey #15721. The root cause is race conditions during server shutdown, between compaction and closing compaction_history table. Stability: query_tombstone_page_limit applies to unpaged queries, while it should not #17241 MV Correctness: in some cases, like one base table with many views, updates to MV might be lost. In one example range tombstones made the issue appears earlier #17117 --- ### Page: https://forum.scylladb.com/t/cassandras-unchecked-tombstone-compaction-not-recognized/1360 Title: Cassandra's unchecked_tombstone_compaction not recognized - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I’m engaging with the community to share an experience and seek insights regarding ScyllaDB’s handling of specific Cassandra’s tombstone compaction option, particularly unchecked_tombstone_compaction. While integrating … Language: en Canonical URL: https://forum.scylladb.com/t/cassandras-unchecked-tombstone-compaction-not-recognized/1360 ## Headings Structure: H1: Cassandra's unchecked_tombstone_compaction not recognized H3: Related topics ## Main Content: H1: Cassandra's unchecked_tombstone_compaction not recognized H4: Add support to compaction properties that relate to tombstone H4: Feedback: ScyllaDB support and compatibility considerations. H3: Related topics I’m engaging with the community to share an experience and seek insights regarding ScyllaDB’s handling of specific Cassandra’s tombstone compaction option, particularly unchecked_tombstone_compaction. While integrating ScyllaDB with a Cassandra-designed plugin, I encountered an issue where the unchecked_tombstone_compaction option was not recognized during a table auto-creation process. This lack of recognition caused the process to fail, despite ScyllaDB’s compatibility with many other tombstone compaction options. Notably, this tombstone compaction is hardcoded in the plugin, which complicates the issue further. To bypass this, I managed a workaround by manually creating the necessary tables without this specific option. While this approach allows me to move forward, it fundamentally deviates from the automated workflow that the plugin is designed to facilitate. This manual intervention not only introduces additional steps but also increases the potential for errors and inconsistencies, impacting the overall efficiency and reliability of the system setup. Why is unchecked_tombstone_compaction specifically not supported or recognized among the other tombstone compaction options that ScyllaDB does accept from Cassandra configurations? Is there a technical limitation or rationale behind this selective recognition? If I understand correctly, this unchecked_tombstone_compaction is a way to bypass tombstone_threshold, starting tombstone compaction based on tombstone_compaction_interval alone. I don’t think there is any technical or product level consideration for us not implementing this (at least not that I’m aware of). Most minor compatibility features like this are developed on-demand and most likely nobody asked for this feature yet. If you would like this feature to be implemented, please try to find an existing issue about this on the ScyllaDB issue tracker and comment on it that it is interesting to you as well, or create a new one if it doesn’t exist. Thank you very much for your insightful response. It’s reassuring to know that the absence of this feature isn’t due to any technical limitations. I came across a closed issue mentioning unchecked_tombstone_compaction , and it left unclear why this option isn’t supported despite indications to the contrary. For additional context, you can review that discussion here: Those are: tombstone_compaction_interval default:864000 (one day) The mini…mum number of seconds after an SSTable is created before Cassandra considers the SSTable for tombstone compaction. Tombstone compaction is triggered if the number of garbage-collectable tombstones in the SSTable is greater than tombstone_threshold. tombstone_threshold default:0.2 The ratio of garbage-collectable tombstones to all contained columns. If the ratio exceeds this limit, Cassandra starts compaction on that table alone, to purge the tombstones. unchecked_tombstone_compaction default:false True allows Cassandra to run tombstone compaction without pre-checking which tables are eligible for this operation. Even without this pre-check, Cassandra checks an SSTable to make sure it is safe to drop tombstones. To provide further clarity and directly engage with this topic, I’d like to share the issue I initiated: When attempting to use the `tables-autocreate` feature specified in an `applicat…ion.conf` file, I encountered a failure due to a compatibility issue between Cassandra and ScyllaDB. The problem arises because the `unchecked_tombstone_compaction` option, used within the compaction strategy configuration, is not recognized by [ScyllaDB](https://opensource.docs.scylladb.com/stable/cql/compaction.html#common-options). The plugin expects this option during the table creation process. The relevant code can be found [here](https://github.com/apache/incubator-pekko-persistence-cassandra/blob/90d80ad20dd72b9eff09ab1998e3354f78227313/core/src/main/scala/org/apache/pekko/persistence/cassandra/compaction/BaseCompactionStrategy.scala#L56). I worked around this by manually creating the necessary tables without using the `unchecked_tombstone_compaction` option. This workaround suggests that broader compatibility might be achievable with some adjustments. Given the growing popularity and adoption of ScyllaDB, it would be nice if future versions of the plugin could consider providing some support for ScyllaDB. Looks like we either forgot about unchecked_tombstone_compaction or didn’t think it is important when doing our initial implementation of tombstone compaction. I found nothing in that tissue suggesting that we made a deliberate decision for not implementing this. Thank you for this clarification. It’s interesting to hear that unchecked_tombstone_compaction might have just slipped through the cracks or wasn’t seen as essential initially. Thank you once again for taking the time to look into this matter and for your ongoing support. --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-2-16/1361 Title: [RELEASE] ScyllaDB 5.2.16 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.2.16, a bugfix release of the ScyllaDB 5.2 stable branch. ScyllaDB Open Source 5.2.16, like all past and future 5.x.y releases, is backward compatible, and supports roll… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-2-16/1361 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.2.16 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.2.16 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.2.16, a bugfix release of the ScyllaDB 5.2 stable branch. ScyllaDB Open Source 5.2.16, like all past and future 5.x.y releases, is backward compatible, and supports rolling upgrades. Note that the latest ScyllaDB Open Source stable release is 5.4 and you are encouraged to upgrade to it. Issue fixed in this release: cdc: allow sending writes to the previous generation for some time after switching #7251, #15260 The previous bypass for this issue, retry the request, does not work when using LWT and CDC on the same table. The issue can appear both in CQL CDC and Alternator (Amazon DynamoDB compatible API) Streams. Stability: streaming might fail when a table is dropped. Streaming can be part of a repair, topology change (adding/removing a node), etc. #15370, #17028, #15598 Stability: in rare cases, nodetool getsstables or calling REST API /column_family/sstables/by_key/ can crash the server #17232 Stability: query_tombstone_page_limit applies to unpaged queries, while it should not #17241 MV Correctness: in some cases, like one base table with many views, updates to MV might be lost. In one example range tombstones made the issue appears earlier #17117 Log: schema::describe: print synchronous_updates only when it’s present in extension map #14924 Build: Regenerate frozen toolchain for gnutls 3.8.3 and clang clang-16.0.6-4. #17285 --- ### Page: https://forum.scylladb.com/t/release-scylla-manager-3-2-6/1362 Title: [RELEASE] Scylla Manager 3.2.6 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces the release of ScyllaDB Manager 3.2.6 production-ready ScyllaDB Manager patch release of the stable ScyllaDB Manager 3.2 branch. ScyllaDB Manager is a centralized cluster administration and r… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-manager-3-2-6/1362 ## Headings Structure: H1: [RELEASE] Scylla Manager 3.2.6 H3: Related topics ## Main Content: H1: [RELEASE] Scylla Manager 3.2.6 H3: Related topics The ScyllaDB team announces the release of ScyllaDB Manager 3.2.6 production-ready ScyllaDB Manager patch release of the stable ScyllaDB Manager 3.2 branch. ScyllaDB Manager is a centralized cluster administration and recurrent tasks automation tool. This release fixes issues in Manager backup, in particular for Alternator, and repair. In this version, we introduce new properties for the cluster. The --force-tls-disabled option compels health checks to establish non-encrypted CQL sessions, even if Scylla is configured to support encrypted sessions. Additionally, the --force-non-ssl-session-port flag instructs Scylla Manager to consistently use native_transport_port instead of native_transport_port_ssl for encrypted CQL sessions (#3679). Cron for tasks scheduling can now be combined with the --start-date flag to specify the exact time and date for the first task execution (#3701, #3700). We have also updated Scylla Manager to handle node IP changes more effectively, particularly in Kubernetes environments (#3707). Support for IMDSv2 has been added. If the region is not defined in scylla-manager-agent.yaml, the region will be determined based on the response from IMDSv2. Previously, only IMDSv1 was supported (#3735). ScyllaDB Enterprise customers are encouraged to upgrade to ScyllaDB Manager 3.2.6 in coordination with the ScyllaDB support team. The new release includes upgrades of both ScyllaDB Manager Server and Agent. Scylla Manager 3.2.6 supports the following Scylla Enterprise releases: And the following Open Source release (limited to 5 nodes see license): You can install and run Scylla Manager on Kubernetes using Scylla Operator. More here. --- ### Page: https://forum.scylladb.com/t/in-what-scenario-will-nodetool-scrub-be-used/1364 Title: In what scenario will `nodetool scrub` be used? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: It can be learned from the document that nodetool scrub can identify and repair damaged SSTables. So, under what circumstances does this tool need to be used in a production environment? Should we use it as a scheduling… Language: en Canonical URL: https://forum.scylladb.com/t/in-what-scenario-will-nodetool-scrub-be-used/1364 ## Headings Structure: H1: In what scenario will `nodetool scrub` be used? H3: Related topics ## Main Content: H1: In what scenario will `nodetool scrub` be used? H3: Related topics It can be learned from the document that nodetool scrub can identify and repair damaged SSTables. So, under what circumstances does this tool need to be used in a production environment? Should we use it as a scheduling task and keep scheduling it in the background, or should we use it when troubleshooting problems? Or is it that we only use it when migrating SSTables? Thanks! Currently, scrub is meant to be used in response to specific problems found in sstables. For the future, we plan to improve it to be a general tool which gets rid of bad sstables (not necessarily repairing them – only when possible) and also scheduled job which regenerates sstables periodically. --- ### Page: https://forum.scylladb.com/t/production-setup-for-single-node-instance/1365 Title: Production setup for single node instance - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I’m sure this is an uncommon request, but what is the production-grade setup for a single node scylladb instance? The use case is high-throughput writing with durability (more than postgres can offer), but not the availa… Language: en Canonical URL: https://forum.scylladb.com/t/production-setup-for-single-node-instance/1365 ## Headings Structure: H1: Production setup for single node instance H3: Related topics ## Main Content: H1: Production setup for single node instance H3: Related topics I’m sure this is an uncommon request, but what is the production-grade setup for a single node scylladb instance? The use case is high-throughput writing with durability (more than postgres can offer), but not the availability requirement of multi-node. For example, do I use SimpleStrategy because there is no need for network topology based strategy, and simple gossip? Also guessing you never need to repair? You won’t get HA, but you already know that. Even with one node, there is no good reason to use SimpleStrategy. SimpleStrategy will make it hard to move to multi-DC in the future, and it needs an advantage. Ok so keep network topology strategy, but what about repair and gossip, are my assumptions correct? Yes, you do not need repair and gossip, but there is no need to disable either or any other service actively. --- ### Page: https://forum.scylladb.com/t/last-month-in-scylla-cluster-tests-git-master-issue-39-2024-03-10/1366 Title: Last month in scylla-cluster-tests.git master (issue #39; 2024-03-10) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last month. Commits in the da5bc7bd…516aa4d9 range are covered. There were 60 non-merge commits from 11 Software Eng… Language: en Canonical URL: https://forum.scylladb.com/t/last-month-in-scylla-cluster-tests-git-master-issue-39-2024-03-10/1366 ## Headings Structure: H1: Last month in scylla-cluster-tests.git master (issue #39; 2024-03-10) H3: Related topics ## Main Content: H1: Last month in scylla-cluster-tests.git master (issue #39; 2024-03-10) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last month. Commits in the da5bc7bd…516aa4d9 range are covered. There were 60 non-merge commits from 11 Software Engineers in Test and 3 Software Engineers in that period. Some notable commits: Readme manual got info about using AWS network configuration refactored recently. Nemesis can be skipped based on Scylla’s open issues, and we have a new command to scan the code for usages of SkipPerIssues and evaluate them if issues are closed and not tagged, so we can consider removing it from code, or tagging as needed. There’s a new script for creating Jenkins pipelines based on directory structure. This is used to re-organize our pipelines, and it supports freestyle jobs and operator job creation. Also pipelines were split to two diffrent folders: oss and enterprise. Since we run into multiple cases on parallel nemesis we introduced lock for target selection so the same node is not used multiple times in the same time. We created a filter for tablets supported nemeses towards scylla version 6.0. To use it, use configurations/tablets_supported_nemeses.yaml in test configuration. Also introduced first tablets specific longevity and multi-dc multi-rack scenarios test. EndOfQuotaNemesis was temporarily disabled due to an issue that won’t be handled meanwhile. Instance types used in AWS tests were updated to the latest ones in most jobs. We added support for streamlining OKTA usage this would check first if with default/used profile we can reach AWS, and if we don’t try to use gimmie-aws-creds tools to create a new profile. New helper class was introduced to report versions of various tools used in SCT test runs in Argus. Currently, scylla’s python-driver version to both log and argus. RestartThenRepairNodeMonkey nemesis was reenabled after a fix in Scylla. New performance test with tablets enabled measuring the latency during grow-shrink case for OSS and Enterprise versions. Disabled later due issues found by it. Fixed weekly triggers use proper CI jobs and it’s parameters for operator tests and added a new weekly trigger for EKS ARM tests to cover ARM support in upcoming scylla-operator-v.1.12 release. Raft feature is enabled by default for Scylla >= 5.5, in other cases it’s based on the consistent-cluster-management flag. New upgrade test to tablets-enabled cluster. The test case starts with scylla 5.4 and during the upgrade process we activate tablets (and raft). K8s tests use the newest scylla-manager version - 3.2.6. Also, we stopped specifying our own Scylla version which gets used as a backend for the scylla-manager and started to rely on the default value. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/how-do-i-model-a-schema-such-that-i-avoid-using-space/1367 Title: How do I model a schema such that I avoid using space - University and Training - ScyllaDB Community NoSQL Forum Meta Description: Say if I have a message schema, & I need a field replyToMessageId which holds the id of the message to which this message instance is supposed to be a reply/comment of, then for multiple messages with this field being nu… Language: en Canonical URL: https://forum.scylladb.com/t/how-do-i-model-a-schema-such-that-i-avoid-using-space/1367 ## Headings Structure: H1: How do I model a schema such that I avoid using space H3: Related topics ## Main Content: H1: How do I model a schema such that I avoid using space H3: Related topics Say if I have a message schema, & I need a field replyToMessageId which holds the id of the message to which this message instance is supposed to be a reply/comment of, then for multiple messages with this field being null, wont it waste bits used to represent null? Is null value stored? & how to model such data better? In ScyllaDB, fields for which you don’t set a value, are simply not stored on disk. Note, that setting a field to null is different than not storing a value at all. A null field still takes some space (it is a cell tombstone). --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-221-2024-03-10/1369 Title: Last week in scylladb.git master (issue #221; 2024-03-10) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 94cd235888…af90910687 range are covered. There were 197 non-merge commits from 26 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-221-2024-03-10/1369 ## Headings Structure: H1: Last week in scylladb.git master (issue #221; 2024-03-10) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #221; 2024-03-10) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 94cd235888…af90910687 range are covered. There were 197 non-merge commits from 26 authors in that period. Some notable commits: Authentication and authorization use system tables in the system_auth keyspace. It now uses a new keyspace, system_auth_v2, which is replicated to all nodes using Raft. This prevents problems due to missing repair, and increases performance, as now the data can always be read locally. Existing clusters will automatically migrate auth data to the new keyspace after an upgrade. Tables that use tablets now maintain their sstables in a specialized data structure that efficiently selects the sstables based on the tablets being accessed, reducing read amplification. Load-and-stream is a mechanism where the database operator can load sstables into any node, and the node then streams the data into nodes which are replicas for that data. This is now supported with tablets. Reading tablet metadata (the system.tablets table) is now careful to avoid stalls. Reshape compaction changes sstables to conform to compaction strategy goals. Reshape compaction now works within tablet boundaries, so it doesn’t violate tablet constraints. More failures are now handled when migrating tablets between nodes. The bundled Python driver was updated to version 3.26.7, in order to fix test stability problems. A recent regression in calculating whether we reached the in-flight hint limit was fixed. There is now more automation for backporting fixes to release branches. The documentation for various scaling procedures was updated to reflect changes for raft-based consistent topology. A deadlock in repair has been fixed. A use-after-free error that manifested during repair with very large partition keys is now fixed. The bundled cqlsh has been updated, with a fix for a COPY TO STDOUT regression. A rare failure when a materialized view update happened to be empty was fixed. The native nodetool command now supports the info subcommand. The native nodetool command now handles ring --resolve-ip correctly. The perf-simple-query benchmark has been enhanced with features around tablets, compaction, and sstable count. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/editor-application-with-a-tree-of-objects-how-to-datamodel/1374 Title: Editor application with a tree of objects - how to datamodel - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/editor-application-with-a-tree-of-objects-how-to-datamodel/1374 ## Headings Structure: H1: Editor application with a tree of objects - how to datamodel H3: Related topics ## Main Content: H1: Editor application with a tree of objects - how to datamodel H3: Related topics Originally from the User Slack @Arthur: I have a question, in a scenario where I am making an editor application, where projects are a tree of objects, and where I should be able to export the whole tree regularly, what would be the best way to model that? My assumption would be that I would need to use a single table for all projects, and to use the project uid as the partition key - that takes care of the requirement of making the whole tree exportable often, since having everything in a project would be in a single partition means a single node can take care of reading it all, bypassing a lot of unnecessary round-trips. I am not sure though if that is not prone to the emergence of hot partitions - is it bad if a single partition becomes too big, since it will make a single node do a lot of work by itself, leaving little resources for other projects it might be responsible for? Then, since I would need to identify each project item, I would use an ULID as a sorting key, since the project might be edited from two clients at once, this allows to have a guarantee of uniqueness and to get the items to be sorted from oldest to newest, which could be useful for some application features, and which i assume would be better for performance, since that means the latest created objects in the tree, which are the most likely to be a lively modified by the user, will be at the bottom of the partition, making it “easy” for Scylla to find? Really unsure about that though. One thing I am really unsure about is that for this model to work I would have to have each field of each object type represented in the table’s rows, which could mean ending up with large rows, which is again bad modelling… Have I gotten things backwards? Should I split up tables per item type instead to keep rows and partitions smaller, despite the fact it will also fragment my data for a single project all across the cluster, knowing that reading the entirety of a project would be a relatively frequent operation? @Felipe_Cardeneti_Mendes: Yes for a single table holding all the projects, but then perhaps break it down into some granularity and incrementally export the tree contents. For example, GET /project/? could retrieve all nested structure (folders?), where . would always be your root and exist. From there you can have projects with a small number of elements, or a lot, but in general bounded (ie: a project with 1000 folders would already be an outlier, but not necessarily an issue). At this point the source client can already mirror the source structure, while you asynchronously fetch inner objects. For example, GET /project/?/identifier , would then sort all inner objects within a specific project and a identifier. The good thing about this approach is that it allows you to asynchronously fetch from multiple sources at once, and you don’t care about the ordering as ultimately each object is unique. Even if the project has thousands of nested structures, you can just run thousands of queries concurrently, and have them spread across to maximize your cluster capacity. We then get to the last part of your question: Identify each project item. The use of a ULID as a sorting key will definitely guarantee uniqueness, and you may control the sorting via the CLUSTERING ORDER BY clause (ideally you’d prefer it to be sorted from newest/oldest?). But then, whenever you traverse the tree you would always read not just the latest objects, but also the inactive ones (this alone brings some questions such as merging, losing updates, etc, but I’ll keep these aside for now). Worse: Unless you keep track of each object in its own separate table, you would get a timestamp-ordered list of changes of mixed objects! Consider a folder with 100 files. You may have 10K changes on top of 99 files, but 1 file was never modified since its creation, and you will only reach it at the top or bottom (depending on your clustering order) of your scan. What would you do with the remainder 9,9K rows you just read? So perhaps an ideal approach is this: Table projects uniquely identify a project, and have a sorting key identifying the structure within that project. PRIMARY KEY (project_id, path) , . always exist and is the parent of all others. Table objects contains all fresh objects within a given path for a specific project. PRIMARY KEY((project_id, path), object_id). Note we do not store the timestamp as part of the key here. This allows a simple SELECT * FROM table WHERE project_id=? AND path=? to always retrieve the latest object_id. Then you can have a 3rd table where you keep track of changes: PRIMARY KEY((project_id, path), object_id, ts). Now you have versioning, but the main “checkout” part of the process is much more lightweight. I may have understood some things wrong on your use case, but hope this helps. --- ### Page: https://forum.scylladb.com/t/docker-on-mac-unable-to-connect-to-scylla-api-server-java-net-connectexception-connection-refused-connection-refused/1375 Title: Docker on Mac: Unable to connect to Scylla API server: java.net.ConnectException: Connection refused (Connection refused) - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/docker-on-mac-unable-to-connect-to-scylla-api-server-java-net-connectexception-connection-refused-connection-refused/1375 ## Headings Structure: H1: Docker on Mac: Unable to connect to Scylla API server: java.net.ConnectException: Connection refused (Connection refused) H3: Related topics ## Main Content: H1: Docker on Mac: Unable to connect to Scylla API server: java.net.ConnectException: Connection refused (Connection refused) H3: Related topics Originally from the User Slack @Joep: Hi all. Trying to follow https://iot.scylladb.com/stable/getting-started.html and when trying to do step 3 of Deploy, running into an issue I don’t know what to do with. Anyone might know what’s up and how to resolve? Getting Started with CarePet: A sample IoT App | ScyllaDB Docs Seems to run now by making this change to the docker-compose. Had to dig through the Scylla logs to find out it was in some stuck state on start-up. Pretty weird. My Docker has 10 Cores available, not sure what Scylla expects by default @Guy: Thanks for reporting, @Attila_Toth. Do you have any ideas? @Attila_Toth: Hi @Joep! Are you using Mac? Also you can share any Docker logs? @Joep: Windows. Just cloned the directory, opened up Docker (default settings), docker-compose, erroring. This doesn’t happen for you locally? @Attila_Toth: This is a known issue. I’m aware of a fix for Linux (increase fs.aio-max-nr), and Mac. Not sure about windows. I see you already came up with a fix yourself, you could also try adding --reactor-backend=epoll to your docker compose file as described here not sure it works on Windows though --- ### Page: https://forum.scylladb.com/t/scylladb-k8-operator-error-unsupported-value-of-nodes-broadcast-address-type-supported-ones-are/1376 Title: ScyllaDB K8 Operator Error: [unsupported value of nodes-broadcast-address-type "", supported ones are: - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-k8-operator-error-unsupported-value-of-nodes-broadcast-address-type-supported-ones-are/1376 ## Headings Structure: H1: ScyllaDB K8 Operator Error: [unsupported value of nodes-broadcast-address-type "", supported ones are: H3: Related topics ## Main Content: H1: ScyllaDB K8 Operator Error: [unsupported value of nodes-broadcast-address-type "", supported ones are: H3: Related topics Originally from the User Slack @Gourav_Upadhyay: hello, how to solve this issue no idea how it occured not able to start please help @Maciej_Zimnoch: probably you’re using latest tag in you Scylla Operator image and versions between operator and Scylla Pod sidecar got out of sync, make sure to use stable version --- ### Page: https://forum.scylladb.com/t/how-many-record-in-system-distributed-cdc-streams-descriptions-v2/1378 Title: How many record in system_distributed.cdc_streams_descriptions_v2? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I use a one dc three node cluster. num_token is 256. smp is 2. this cluster has 768 vnode. I check it by nodetool ring. but: cqlsh> select * from system_distributed.cdc_generation_timestamps ; key | time … Language: en Canonical URL: https://forum.scylladb.com/t/how-many-record-in-system-distributed-cdc-streams-descriptions-v2/1378 ## Headings Structure: H1: How many record in system_distributed.cdc_streams_descriptions_v2? H3: Create a ScyllaDB Cluster - Single Data Center (DC) | ScyllaDB Docs H3: Adding a New Node Into an Existing ScyllaDB Cluster (Out Scale) | ScyllaDB Docs H3: Nodetool checkAndRepairCdcStreams | ScyllaDB Docs H3: Related topics ## Main Content: H1: How many record in system_distributed.cdc_streams_descriptions_v2? H3: Create a ScyllaDB Cluster - Single Data Center (DC) | ScyllaDB Docs H3: Adding a New Node Into an Existing ScyllaDB Cluster (Out Scale) | ScyllaDB Docs H3: Nodetool checkAndRepairCdcStreams | ScyllaDB Docs H3: Related topics I use a one dc three node cluster. num_token is 256. smp is 2. this cluster has 768 vnode. I check it by nodetool ring. but: The last two timestamps, which are less than one second apart, indicate that you did not follow correct node bootstrap procedure. You bootstrapped the second and third node concurrently, which is something that Scylla does not correctly support at the moment (it will only in 6.0 with “Raft based topology”). The documented bootstrap procedure instructs to boot nodes sequentially, i.e. only once a node becomes UN (“UP NORMAL”), it is safe to start booting subsequent node. ScyllaDB is an Apache Cassandra-compatible NoSQL data store that can handle 1 million transactions per second on a single server. ScyllaDB is an Apache Cassandra-compatible NoSQL data store that can handle 1 million transactions per second on a single server. Right now to fix the situation you can use nodetool checkAndRepairCdcStreams to prompt Scylla to create a new CDC generation with the correct number of streams. ScyllaDB is an Apache Cassandra-compatible NoSQL data store that can handle 1 million transactions per second on a single server. --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-4-4/1379 Title: [RELEASE] ScyllaDB 5.4.4 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.4.4, a bugfix release of the ScyllaDB 5.4 stable branch. ScyllaDB Open Source 5.4.4, like all past and future 5.x.y releases, is backward compatible and supports rolling… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-4-4/1379 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.4.4 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.4.4 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.4.4, a bugfix release of the ScyllaDB 5.4 stable branch. ScyllaDB Open Source 5.4.4, like all past and future 5.x.y releases, is backward compatible and supports rolling upgrades. Users are encouraged to upgrade to 5.4.4. Issue fixed in this release: cdc: allow sending writes to the previous generation for some time after switching #7251. The previous bypass for this issue, retry the request, does not work when using LWT and CDC on the same table. The issue can appear both in CQL CDC and Alternator (Amazon DynamoDB compatible API) Streams. Stability: streaming might fail when a table is dropped. Streaming can be part of a repair, topology change (adding/removing a node), etc. #15370, #17028, #15598 Stability: Having an index being dropped during the bootstrap causes a node to fail to start #15598 Stability: throughput_limits_test.TestCompactionLimitThroughput.test_can_limit_compaction_throughput is flakey #15721. The root cause is race conditions during server shutdown, between compaction and closing compaction_history table. Stability: in rare cases, nodetool getsstables or calling REST API /column_family/sstables/by_key/ can crash the server #17232 Stability: query_tombstone_page_limit applies to unpaged queries, while it should not #17241 Stability: rare issue when table drop concurrent to any repair or streaming (so any node ops) #16899, #15425. MV Correctness: in some cases, like one base table with many views, updates to MV might be lost. In one example range tombstones made the issue appears earlier #17117 Build: Regenerate frozen toolchain for gnutls 3.8.3 and clang clang-16.0.6-4. #17285 Build: update Rust dependencies pull#17407. Rust is used to test WASM, an experimental feature in 5.4 --- ### Page: https://forum.scylladb.com/t/event-publishing-outbox-pattern-and-batching/1380 Title: Event Publishing, Outbox Pattern and Batching - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/event-publishing-outbox-pattern-and-batching/1380 ## Headings Structure: H1: Event Publishing, Outbox Pattern and Batching H3: Related topics ## Main Content: H1: Event Publishing, Outbox Pattern and Batching H3: Related topics Originally from the User Slack @Arkam_Fahry: Hello can any some on help me please. I currently have a file upload system which sends events to a queue for many types of events like upload completion and stuff. I would like events to reliably published at least once by persisting them first into scylladb then publishing the event form scylladb. The current system runs mongodb with a outbox collection when a file is uploaded the file metadata object and the upload event is persisted in a multi document transaction. The events are processed every 250ms in micro batches from outbox collection. Is doing a dual table write with a batch operation a viable solution to persist the event and the entity state. @Felipe_Cardeneti_Mendes: I think I addressed a similar question last year, see if it sheds some light (and check the references/history - as they provide additional context) https://forum.scylladb.com/t/are-there-any-plans-for-triggers-procedures-of-any-kind-on-the-scylladb-roadmap/889/4?u=felipemendes ScyllaDB Community NoSQL Forum: Are there any plans for TRIGGERS/PROCEDURES of any kind on the ScyllaDB Roadmap @Arkam_Fahry: Thanks I checked out the answer before posting here. Will using batch operations effect negatively in terms of performance. @Felipe_Cardeneti_Mendes: Well, depends on the kind of batch. But if you’re far from overwhelming the server it works just fine --- ### Page: https://forum.scylladb.com/t/release-scylla-operator-1-11-3/1381 Title: [RELEASE] Scylla Operator 1.11.3 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of Scylla Operator 1.11.3 :rocket:. Scylla Operator 1.11.3 brings a bug fix. As with all of our releases, all API changes are backward compatible. Notable changes … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-operator-1-11-3/1381 ## Headings Structure: H1: [RELEASE] Scylla Operator 1.11.3 H3: Related topics ## Main Content: H1: [RELEASE] Scylla Operator 1.11.3 H3: Related topics The ScyllaDB team is pleased to announce the release of Scylla Operator 1.11.3 . Scylla Operator 1.11.3 brings a bug fix. As with all of our releases, all API changes are backward compatible. For more changes and details check out the GitHub release notes. Upgrade instructions When using kubectl apply, upgrading from v1.10.x or v1.11.x doesn’t require any additional actions, just take the manifest from v1.11.3 tag and substitute for the image. Since helm isn’t capable of handling CRD updates, using helm requires a mandatory manual step for every release. For details, see our upgrade documentation. Best regards, Scylla Operator Team. --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2023-1-6/1383 Title: [RELEASE] ScyllaDB Enterprise 2023.1.6 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2023.1.6, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2023.1 LTS Release. 2023.1.6 patch release includes multiple minor bug fixes… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2023-1-6/1383 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2023.1.6 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2023.1.6 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2023.1.6, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2023.1 LTS Release. 2023.1.6 patch release includes multiple minor bug fixes. ScyllaDB Enterprise’s latest Long Term Support (LTS) is 2024.1. While we continue to support 2023.1, you are encouraged to upgrade to it in coordination with the ScyllaDB Support team. The following issues are fixed in this release (with an open-source reference, if available): cdc: allow sending writes to the previous generation for some time after switching #7251, #15260. The previous bypass for this issue, retry the request, does not work when using LWT and CDC on the same table. The issue can appear both in CQL CDC and Alternator (Amazon DynamoDB compatible API) Streams. Stability: ScyllaDB locks all memory so we don’t experience high latency due to page faults. However, this only applies from the first time the memory is accessed; the first access can still experience stalls, made larger by using transparent huge pages. #8828. Stability: in rare cases, nodetool getsstables or calling REST API /column_family/sstables/by_key/ can crash the server #17232 Stability: query_tombstone_page_limit applies to unpaged queries, while it should not #17241 MV Correctness: in some cases, like one base table with many views, updates to MV might be lost. In one example range tombstones made the issue appears earlier #17117 Log: schema::describe: print synchronous_updates only when it’s present in extension map #14924 Build: Regenerate frozen toolchain for gnutls 3.8.3 and clang clang-16.0.6-4. #17285 --- ### Page: https://forum.scylladb.com/t/java-driver-4-15-01-and-shard-awareness/1385 Title: Java driver 4.15.01 and shard awareness - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello , when testing driver shard awareness, I notice that establishing connection to the 9042 port all the request are redirect to the 9042 and not to the 19042 as I thought it will do. Should I set explicity the 1904… Language: en Canonical URL: https://forum.scylladb.com/t/java-driver-4-15-01-and-shard-awareness/1385 ## Headings Structure: H1: Java driver 4.15.01 and shard awareness H3: Related topics ## Main Content: H1: Java driver 4.15.01 and shard awareness H3: Related topics when testing driver shard awareness, I notice that establishing connection to the 9042 port all the request are redirect to the 9042 and not to the 19042 as I thought it will do. Should I set explicity the 19042? Java driver 4.x does not support advanced shard awareness afaik, meaning that it won’t use port 19042. Shard awareness will still work - the driver may just take longer to connect as it will have to build connection pool the traditional way, by opening connections to “random” shards. See my post here: Questions about scylla java driver 4.x ShardAware feature and NAT issue - #2 by Lorak about my explanation of “shard awareness” and “advanced shard awareness” terms. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-40-2024-03-16/1386 Title: Last week in scylla-cluster-tests.git master (issue #40; 2024-03-16) - Announcements - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from last week. Commits in the 9bf3b22c…90214115 range are covered. There were 16 non-merge commits from 6 authors in that p… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-40-2024-03-16/1386 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #40; 2024-03-16) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #40; 2024-03-16) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from last week. Commits in the 9bf3b22c…90214115 range are covered. There were 16 non-merge commits from 6 authors in that period. Some notable commits: We’ve enhanced SCT usage on MacOS by minimizing sudo requirements for test executions with the docker backend, primarily affecting the monitoring stack setup we run locally. Furthermore, running hydra locally no longer necessitates a dockerhub login, a step mandatory only on Jenkins. Additionally, SCT now also takes charge of cleaning up AWS resources and GCE instances, transferring this responsibility from the cloud-devops repository. This shift aims to leverage SCT’s codebase for better maintainability. Please ensure to use the keep tag (either as metadata or labels on GCE) set to alive or a specific number of hours () if you wish to preserve your resources for an extended period. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-222-2024-03-17/1387 Title: Last week in scylladb.git master (issue #222; 2024-03-17) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the af90910687…7df3acd39c range are covered. There were 127 non-merge commits from 21 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-222-2024-03-17/1387 ## Headings Structure: H1: Last week in scylladb.git master (issue #222; 2024-03-17) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #222; 2024-03-17) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the af90910687…7df3acd39c range are covered. There were 127 non-merge commits from 21 authors in that period. Some notable commits: The bundled Java driver, used with cassandra-stress, now supports tablets. The ownership REST API, used to determine which fraction of the keyspace is owned by which node, now supports tablets. Adjustments were needed since with tablets, individual tables can be assigned to nodes differently. The repair API now supports the ranges option. The repair API now supports the datacenter option when repairing tables with tablets. The repair API now supports the hosts and ignore_nodes options for tables with tablets. The REST API for reporting cache statistics is now more accurate. The materialized view builder was adjusted for concurrent tablet migration. In tables with many small partitions (or many partition tombstones), sstable index pages can contain many entries. They are now destroyed gently to avoid stalls. The native nodetool now implements netstats, tablehistograms, proxyhistograms and status commands. This makes the native nodetool feature complete. Some false-positives were eliminated from the scrub command. All Raft group 0 tables are now made durable under the schema commitlog; previously only some where. We now fail base table writes rather than dropping materialized view updates, to reduce base/view inconsistencies. We now refresh RPC connections to nodes that gained topology information earlier. This avoids failures when data is queried immediately after a node is added. The tablet load balancer will now avoid balancing nodes that are down. Loading speed for tablets was improved by reinstating the sharding metadata stored in the Scylla.db sstable component. The tablets of views and indices associated with a table are now dropped when a table is dropped. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/recovery-after-data-loss-how-do-i-clear-data-for-new-install/1388 Title: Recovery after data loss, how do I clear data for new install? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/recovery-after-data-loss-how-do-i-clear-data-for-new-install/1388 ## Headings Structure: H1: Recovery after data loss, how do I clear data for new install? H3: Related topics ## Main Content: H1: Recovery after data loss, how do I clear data for new install? H3: Related topics Originally from the User Slack @Keith_Mciff: We have a cluster that completely lost all data drives. We have attached new storage and are trying to start the nodes but we are getting the below error when trying to start the first seed node. What file can we delete or command can we run to make this a “clean” first install again? @Felipe_Cardeneti_Mendes: 1. Ensure that the “first seed node” is the node with the “lowest IP address in your seed config. To make it simpler, you can just leave the first node to be the only seed entry and update it later 2. Follow https://opensource.docs.scylladb.com/stable/operating-scylla/procedures/cluster-management/clear-data.html @Keith_Mciff: Ill give it a shot That worked! Thank you! --- ### Page: https://forum.scylladb.com/t/data-modeling-application-with-highly-varying-loads-detecting-large-partitions/1389 Title: Data modeling - application with highly varying loads, detecting large partitions - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/data-modeling-application-with-highly-varying-loads-detecting-large-partitions/1389 ## Headings Structure: H1: Data modeling - application with highly varying loads, detecting large partitions H3: Related topics ## Main Content: H1: Data modeling - application with highly varying loads, detecting large partitions H3: Related topics Originally from the User Slack @Yukkuri: Hi, in case there is no information available to provide partitioning am I understanding correctly that hashing clustering keys and dividing hash space into artificial partition keys is the right approach? In case when load/write amounts can vary drastically from instance of application to instance of application (from one row every few weeks to hundreds of rows per second), dynamic hash space partitioning may also be required, also in my use-case those insert rates can gradually grow for any instance of application over time (but probably never substantially decrease, just keep the level for a while). My current approach is to hold some partitioning generation counters and whenever load on an application instance exceeds certain treshold - introduce new subsequent generation, with more hash space partitions, copying all the data to be redistributed alongside new partitioning scheme, scheduling all write operations while redistribution is in progress for both old and new generations, and after redistribution is complete, old generation partitions can be range-deleted. Basically (domain, generation, hash_bucket), apid where domain, origin of those apids can be derived from apid and apid is hashed into 64-bit value H and hash_bucket is determined using H / (H_max/(2^generation)) My questions are • is periodic querying of system.large_partitions for every node is a reliable/reasonable way to detect partitions growing outside current generation bounds, or could it be preferable to maintain own kind of metric regarding hash bucket fill per domain per generation? • Is there anything to be aware about, maybe other approaches, some potential improvements? @Felipe_Cardeneti_Mendes: yeap, large_partitions is quite effective. Note it is NOT immediate however. large_partitions are populated as part of compactions. large_rows and other large_* tables may be relevant for you. I think the only aspect I don’t understand is why you would copy data around from one place to another. Why don’t you simply keep track of your generations/buckets under a separate table? @Yukkuri: > yeap, large_partitions is quite effective. Note it is NOT immediate however. large_partitions are populated as part of compactions. large_rows and other large_* tables may be relevant for you. Thanks. Why don’t you simply keep track of your generations/buckets under a separate table? That would require keeping track of every apid from past generations there, which could end up being subject to own partitioning pressure; there is no way to know apid’s generation without full-scan of that accounting table. Say there is apid ; and under current generation 0 allowing only one hash bucket (whole hash space) it’s hash_bucket is 0. Now if we increase generation to 1, allowing 2 partitions of hash space, it’s hash_bucket may end up being either 0 or 1 depending on chosen hash function, despite having exact same hash value. To account for this, we must copy a row of (('', 0, 0), '') to (('', 1, $new_hash_bucket), '') so we can continue relying on hash space partitioning to unambiguously map apid hash values to partitioning buckets. That requires extra dance around current/pending generations, to keep both in sync while migration is in progress, but otherwise seem to work. Also one small change from what I described initially: it turned out to be beneficial to store combined generation+hash_bucket value in single column; since it also allows in-queries. In my case both values are fitting nicely in 128-bit uuid. @Felipe_Cardeneti_Mendes: LGTM --- ### Page: https://forum.scylladb.com/t/issue-restoring-tables-from-s3/1394 Title: Issue restoring tables from s3 - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Trying to do a restore-tables with scylla manager from s3. Getting this error: get tombstone_gc of keyspace.table: not found This seems to be happening with all my keyspaces. any thoughts as to why? Scylla OSS 5.2.2 … Language: en Canonical URL: https://forum.scylladb.com/t/issue-restoring-tables-from-s3/1394 ## Headings Structure: H1: Issue restoring tables from s3 H3: Related topics ## Main Content: H1: Issue restoring tables from s3 H3: Related topics Trying to do a restore-tables with scylla manager from s3. Getting this error: get tombstone_gc of keyspace.table: not found This seems to be happening with all my keyspaces. any thoughts as to why? Scylla OSS 5.2.2 Scylla Manager 3.2.6 Can close. I figured out why. The backup was taken from datacenter1 but I was trying to restore into datacenter2. --- ### Page: https://forum.scylladb.com/t/memory-issue-deleting-big-partitions-in-scylladb-with-timewindowcompactionstrategy/1395 Title: Memory issue - deleting big partitions in ScyllaDB with TimeWindowCompactionStrategy - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/memory-issue-deleting-big-partitions-in-scylladb-with-timewindowcompactionstrategy/1395 ## Headings Structure: H1: Memory issue - deleting big partitions in ScyllaDB with TimeWindowCompactionStrategy H3: Related topics ## Main Content: H1: Memory issue - deleting big partitions in ScyllaDB with TimeWindowCompactionStrategy H3: Related topics Originally from the User Slack We try to delete a big partition by just giving partition key, however, no row is deleted since Scylla restart itself in our test env. I tried to delete a partition with around 1 million rows. We have much powerful cluster in our production environments but we try to find the best practices and we come with two solution and I want to discuss with you. 1- We delete the partition chunk by chunk (cluster key consists of just one column with uuid type and we can cut all uuid space into multiple chunks) 2- Divide the only one table into multiple table which contain just one partition and once I want to delete the partition, I just delete the table instead. Solution one seems okey but we found solution 2 much more convenient to implement, however, this time we would create many tables and we thought that it could be overhead for scylla. (we are talking about more than 1000 table) Which solution do you recommend us? How do you delete your big partitions in Scylla. @Felipe_Cardeneti_Mendes: Well, neither 1 nor 2. Scylla shouldn’t restart itself, is it running out of memory? Best is to correct the restart situation and just delete the partition and compact (major) the data away @Jesper_Lundgren: With Cassandra I’ve had to update tombstone_threshold when doing partition deletes as the number of tombstones will be low compared to regular data when doing range deletes. Deletes are not instant, it’s a tombstone write and will eventually delete data in background. Reads won’t return deleted data but reclaiming disk takes time. And it also depends on compaction strategy @golang_learner: @Felipe_Cardeneti_Mendes Hi, actually it does not restart itself but it gives me NoHostAvailable error when I delete the big partition with 1 million rows. @Jesper_Lundgren Thanks jesper. We are using Time-window Compaction Strategy for this table. We do not care about disk claim for now. We just want to delete the big partition in a table which could reach 100 million rows in a safe way. @Felipe_Cardeneti_Mendes: Yeah this shouldn’t happen @golang_learner: Then what should I do? Does the delete operation trigger compaction or something that makes Scylla use a lot memory which causes NoHostAvailable error? I checked that there are 2GB left in the container in which Scylla is running on. Then, can I say that if I allocate more memory such as 8 GB, this NoHostAvailable error should be thrown? @Felipe_Cardeneti_Mendes: Nope, a delete is just a write marker of a tombstone. Compactions are a later process. Is it some kind of LWT statement? Do you have an index/MV involved? These 2 could have a read-before-write which could cause you to bad_alloc. You probably want to trace it to see understand what’s going on. @golang_learner: Our delete statement is like ‘DELETE FROM members where country=‘UK’ and city=‘london’’; I reviewed the logs when we got the error and the errors started with; [shard 1] seastar_memory - oversized allocation: 2052096 bytes. This is non-fatal, but could lead to latency and/or fragmentation issues. [shard 1] storage_proxy - exception during mutation write to 10.18.64.146: std::bad_alloc (std::bad_alloc) We plan to delete the partition chunk by chunk so that we target smaller rows. 'DELETE FROM members where country=‘UK’ and city=‘london’ and member_id >=‘00000000-0000-0000-0000-00000000’ and member_id <=‘3fffffff-ffff-ffff-ffff-ffffffffffff’; 'DELETE FROM members where country=‘UK’ and city=‘london’ and member_id >=‘40000000-0000-0000-0000-00000000’ and member_id >=‘7fffffff-ffff-ffff-ffff-ffffffffffff’; This way we got not error but would you recommend it? We are not sure if it is the best practice. @Felipe_Cardeneti_Mendes: Yeap, bad_alloc is basically an out of memory situation. The question is why… Is this table backed up by a view or an index? A range delete working as opposed to simply a partition delete seems to point to a read-before-write, which would explain why you would see bad_alloc’s and a 2M allocation… The recommendation would be to resolve the OOM situation and just run a normal delete, adding more memory helps, if it still doesn’t address the problem, then maybe opening an issue so we can understand it better. In any case, it is not like range deletes will crash your cluster nor anything, so if you are just comfortable with it, do it and run a major after. But be warned this isn’t a standard approach, you are just working around a problem. @golang_learner: Yes, we defined a materialized view for our main table. Is there any scylla config we can increase so that we do not get this bad_alloc error? Our Scylla version is 4.5.3 open source. @Felipe_Cardeneti_Mendes: not really. maybe one thing to try if you can would be to drop the view, try the delete, and recreate it. Or add more memory. Or do you workaround. @golang_learner: We had deployed scylladb with “–smp 2 --memory 2G --overprovisioned 1 --developer-mode 1” in our test environment. When I re-deployed it with “–memory 14G”, the “seastar_memory - oversized allocation: 2052096 bytes.” warning is logged but there were no bad_alloc error. Can we say that the “seastar_memory - oversized allocation: 2052096 bytes.” warning log is normal for deleting a partition for 1 million row. Since the deletion was success and there was no such error, I think we can say that even though we have seen warning log @Felipe_Cardeneti_Mendes: yeap, given you have a view and a delete to a base table needs to propagate the deletes to the underlying view (which involves a read before write), and given that you have a large partition, a large allocation is expected to some extent see ? Much easier than range deletes That’s how it should be @golang_learner: Our production environments are much more powerful than our test environment and we do not set any memory so ScyllaDB uses all the available memory. We may not expect any bad_alloc error but we are writing a lot of events such as 5 million in a minute. Does the writes consume memory a lot? @Felipe_Cardeneti_Mendes: No, writes are cheap. But you’re right that a delete is effectively a write, thus a delete paired with a view can show off some large allocations, but it shouldn’t get to a point where you bad_alloc if you have sufficient memory. @golang_learner: Okey. We will delete the whole partition with sufficient memory. We would be ready for timeout exceptions (deleting a partition may last more than 30 seconds) with retries in our Kafka consumers. Thanks a lot. I appreciate it --- ### Page: https://forum.scylladb.com/t/regards-incremental-backup-restore-in-open-source/1400 Title: Regards Incremental backup restore in open source - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: when iam restoring incremental backup not getting incremental data. Can you help on this Language: en Canonical URL: https://forum.scylladb.com/t/regards-incremental-backup-restore-in-open-source/1400 ## Headings Structure: H1: Regards Incremental backup restore in open source H3: Related topics ## Main Content: H1: Regards Incremental backup restore in open source H3: Related topics when iam restoring incremental backup not getting incremental data. Can you help on this Please provide information, things like: Iam using Scylla version 5.2 Backup your Data | ScyllaDB Docs For backup iam following this Restore from a Backup and Incremental Backup | ScyllaDB Docs For rstore iam following this If possible, schedule a call with us. Customer is waiting for the update, if we not resolved he will move to the other database. For that we need to clarify him clearly. I use scylla version 5.4.9 I enabled incremental backups and after the flush the backup folders are not getting created in keyspace/table-uuid/ . I followed the same documentation --- ### Page: https://forum.scylladb.com/t/scylladb-cluster-on-aws-with-autoscaling-and-asg/1404 Title: ScyllaDB Cluster on AWS with autoscaling and ASG - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-cluster-on-aws-with-autoscaling-and-asg/1404 ## Headings Structure: H1: ScyllaDB Cluster on AWS with autoscaling and ASG H3: Related topics ## Main Content: H1: ScyllaDB Cluster on AWS with autoscaling and ASG H3: Related topics Originally from the User Slack @Patryk_Kandziora: Hi. I have a question about spinning 3 node scylla cluster on AWS using autoscaling. As for now I spin 1st node with following user_data I use as well lambda and cloudwatch events to create new A record in route53. Thats works fine. Rest of the nodes in the cluster are spin with following user_data JSON where the seeds is the mentioned A record with IP address of the 1 st node. All good so far. Problem: If the 1st node goes down - the ASG (auto-scalling) will spin new EC2 and lambda will update the A record but the node2 and 3 are not checking the A record anymore. The node1 user_data is not pointing to the A record so cannot see the other nodes in the cluster. In this scenario I have 2 clusters now - node1 as stand alone cluster and rest of the nodes as cluster 2. I tried to create the A record for 1st node with 127.0.0.1 IP and then remove it by lambda with 1st run but the node1 is just send errors about problem with connecting to 127.0.0.1 Any ideas how to approach this scenario? @Felipe_Cardeneti_Mendes: You want node1 to point to node2/node3 as seeds so that node1 is able to fetch your cluster’s topology information. That should be sufficient for it to start replacing the previously failed node1 @Patryk_Kandziora: Hi @Felipe_Cardeneti_Mendes - thats the problem here. With ASG I cannot do it because the bootstrap script for 1st node have no one to point to yet. node2 and node3 are not UP yet so there is not IP addresses and with ASG I cannot specify what IP will be there on specific VM - its random based on subnet and subnet mask. Chicken and Egg or catch22 looks like for me @Felipe_Cardeneti_Mendes: Well, that’s the problem with ASG. Why not use our Operator instead? Nonetheless, that’s just one corner case you will have to handle. You could bootstrap the cluster with multiple seed nodes where n1 would be the one with the lowest IP address, then during node1 replacement it would perform a shadow round with all entries as within your n1 script @Patryk_Kandziora: Dont have as much time to go into operator now. All load tests performed on Scylla VM. Didn’t know it will be as big problem really. Will check the alternator - didn’t know about it. Thx @Felipe_Cardeneti_Mendes: ah ignore the above link, that’s was meant to another thread that’s the operator one https://operator.docs.scylladb.com/ @Patryk_Kandziora: Yeah just reading about the alternator and waiting for info about the IP solution No - operator is not an option now - its very late for us at this point. Need to figure out how to approach the ASG - might overwrite the bootstrap script after the initial deployment of node1 and add the seeds option or spin all in the same time and see what happens. Is there retry to read the records provided in seeds when is missing? Option to change that behave or to add delay? @Felipe_Cardeneti_Mendes: yeah, if a seed is unavailable it will eventually retry reading it. keep in mind that any changes to your seed configuration (ie: adding, removing, changing entries) will require a node rolling restart @Patryk_Kandziora: My thinking is - if I add manager to the cluster it will take care of that? Or am I wrong here? @Felipe_Cardeneti_Mendes: it won’t @Patryk_Kandziora: Looks like in long run operator will have to be the option here or further scripting based on cloudwatch events and lambdas. @Felipe_Cardeneti_Mendes you mentioned rolling restart Did you mean I need to trigger on each node sudo systemctl restart scylla-server If so I still can see the old not existing node1 IP addresses as DN from nodetool status It looks like that restarting the scylla-server is not fetching IP addresses from the updated A record. What do I miss here? @Felipe_Cardeneti_Mendes: whenever you update the seed list, yes. Are you using a DNS entry instead of an IP address? There are internal system tablets (eg system.peers) which will hardcode the IP address of a node, so it won’t be reflected in nodetool status until you replace the node with the new one (which then will update the internal tables to reflect the new node) @Patryk_Kandziora: Yes I use DNS entry (aggregated) with IP addresses of all scylla nodes. So each time new instance is in running state the cloudwatch events will trigger lambda and update the IP addresses in that record. @Felipe_Cardeneti_Mendes: that’s fine, just replace the old failed node and then the entry should revert to UN @Patryk_Kandziora: but the new node will have different IP - just restarted the scylla-server service on the new node, updated the scylla.yaml seed_provider: re-pointing it to the aggregated DNS. The node1 is not joining the cluster - neither any update on node2 and node3. hmm…. Yeah so you cannot add new node if there is DN in the cluster. You need to removenode 1st then restart scylla-server on the new node. Regarding the ASG and node01 issue - I finished with overwriting the user_data in the 1st node as additional step and that fix the issue. Still having the issue with removing the down node from the cluster before I can add new node. Its not a brilliant approach but it works. Would be better going with EC2 instead of ASG I think. Question - with replication set to 2 and cluster with 3 nodes - how many nodes can go down to be still operational? I assume just one right? 3 nodes - 1 down 5 nodes - 2 down 7 nodes - 3 down etc. @Felipe_Cardeneti_Mendes: depends on your consistency level. See https://opensource.docs.scylladb.com/stable/cql/consistency-calculator.html Consistency Level Calculator | ScyllaDB Docs --- ### Page: https://forum.scylladb.com/t/removing-the-old-snapshots-of-deboarded-cluster-gracefully-from-s3/1405 Title: Removing the old snapshots of deboarded cluster gracefully from s3 - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Recently we have deboarded one cluster from the scylla manager but the old snapshots are still there in AWS s3 bucket costing us more, Is there any way to gracefully delete all the snapshots stored in s3 and does it af… Language: en Canonical URL: https://forum.scylladb.com/t/removing-the-old-snapshots-of-deboarded-cluster-gracefully-from-s3/1405 ## Headings Structure: H1: Removing the old snapshots of deboarded cluster gracefully from s3 H3: Related topics ## Main Content: H1: Removing the old snapshots of deboarded cluster gracefully from s3 H3: Related topics Recently we have deboarded one cluster from the scylla manager but the old snapshots are still there in AWS s3 bucket costing us more, Is there any way to gracefully delete all the snapshots stored in s3 and does it affect the scylla manager if we delete them manually from s3? this is the structure that we have " backup/meta/cluster/cluster_id/dc/ap-south/node/node_id/ " Specification for backup location defines bucket hierarchy the way that snapshots have different path per cluster Specification | ScyllaDB Docs So, cluster ABC will keep snapshots under: /backup/sst/cluster/ABC/… /backup/meta/cluster/ABC/… /backup/schema/cluster/ABC/… (only if you provided credentials to the cluster) If you don’t need this cluster together with backups anymore, you can just remove these paths on your own. Sctool provides command to list snapshots from given location Backup | ScyllaDB Docs , but looks that we have a small bug here that requires to introduce cluster-id even if you want to list snapshots of all-clusters 'sctool backup list --all-clusters` requires cluster id · Issue #3780 · scylladb/scylla-manager · GitHub . The same statement applies to sctool backup delete , cluster must be managed by current instance of Scylla-Manager Backup | ScyllaDB Docs `sctool backup delete` doesn't work if cluster is removed · Issue #3781 · scylladb/scylla-manager · GitHub Feel free to remove the cluster path in S3 location that consists of /backup/sst/cluster/ABC/… /backup/meta/cluster/ABC/… /backup/schema/cluster/ABC/… --- ### Page: https://forum.scylladb.com/t/error-when-running-scylladb-on-aws-missing-raid-volume-replace-node/1406 Title: Error when running ScyllaDB on AWS - missing RAID volume, replace node - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/error-when-running-scylladb-on-aws-missing-raid-volume-replace-node/1406 ## Headings Structure: H1: Error when running ScyllaDB on AWS - missing RAID volume, replace node H3: Related topics ## Main Content: H1: Error when running ScyllaDB on AWS - missing RAID volume, replace node H3: Related topics Originally from the User Slack @Jasper_Visser: I am running a single self-managed Scylla instance in AWS. After running 1 week without any problems, I got this message from AWS: > EC2 has detected degradation of the underlying hardware hosting your Amazon EC2 instance (instance-ID: i-XX). I needed to stop-and-start my instance. After spinning the instance up again and sshed to the instance, I see it fails to boot because of this error: > Failed mounting RAID volume! > Scylla has aborted startup because of a missing RAID volume. Is there anything I can do to restart scylla with the same dataset as before the restart? @Felipe_Cardeneti_Mendes: Nope, AWS allocated you a new set of local SSDs. You can recreate the array and then start the replace node procedure. @Jasper_Visser: You can recreate the array and then start the replace node procedure. what do you mean by that? So the data is lost? @Felipe_Cardeneti_Mendes: in that particular node, yes. You do have other replicas, no? @Jasper_Visser: No, no other replica’s atm, it’s just for development though @Felipe_Cardeneti_Mendes: > The data on an SSD instance volume persists only for the life of its associated instance. https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ssd-instance-store.html @Jasper_Visser: Ok and do you have any reference to You can recreate the array and then start the replace node procedure.? @Felipe_Cardeneti_Mendes: https://opensource.docs.scylladb.com/stable/operating-scylla/procedures/cluster-management/rebuild-node.html https://opensource.docs.scylladb.com/stable/kb/raid-device.html Depending on your ScyllaDB version you may need to manually delete the systemd .mount resources in your /etc and follow with a systemctl daemon-reload afterwards. Rebuild a Node After Losing the Data Volume | ScyllaDB Docs Recreate RAID devices | ScyllaDB Docs --- ### Page: https://forum.scylladb.com/t/release-scylladb-cloud-21-march-2024/1407 Title: [RELEASE] ScyllaDB Cloud - 21 March 2024 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: We are happy to announce a new update to the ScyllaDB Cloud service. New features and enhancements: When creating a new cluster either via API/Cloud UI, the following default maintenance windows are set (you can modi… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-cloud-21-march-2024/1407 ## Headings Structure: H1: [RELEASE] ScyllaDB Cloud - 21 March 2024 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Cloud - 21 March 2024 H3: Related topics We are happy to announce a new update to the ScyllaDB Cloud service. New features and enhancements: --- ### Page: https://forum.scylladb.com/t/data-modeling-issues/1409 Title: Data modeling issues - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I’m running into issues while modeling the data for an existing application and I’m looking for help and/or additional insights on how to approach my use case. I ran through the related Scylla university topics, but my … Language: en Canonical URL: https://forum.scylladb.com/t/data-modeling-issues/1409 ## Headings Structure: H1: Data modeling issues H3: Related topics ## Main Content: H1: Data modeling issues H4: Cassandra UUID vs TimeUUID benefits and disadvantages H3: Related topics I’m running into issues while modeling the data for an existing application and I’m looking for help and/or additional insights on how to approach my use case. I ran through the related Scylla university topics, but my model seems way different to the heartbeat monitoring use case there and I’m even starting to doubt whether Scylla is a good fit for this project. This is what my domain looks like: Here’s the queries I need to run: (The user count could run up to 50 million.) Do I need to create more than one table to be able to contain the objects with all 3 "ID"s? If so, should I copy the data or just store the ID’s so I can fetch them from the main table? Help would be greatly appreciated You would use create index to create what is a secondary/table index that allows direct lookups. It’s not easy by default to do this, because the point is people need to learn that this is slow (relatively speaking), but when you need to do it you can do it. You don’t need to manually search for and delete old data, just use the default_time_to_live to make the data magically disappear after the maximum age. This saves needing to scan your entire database regularly. For a full table scan, you can do this using the primary key and limit, i.e. Fetch the first 1000 records, then look at the uid of record 999, to do the next query, ie: Assuming of course your querying ordered data. It just depends on what your primary key is and how its sharded. Thanks for sharing these insights! I had no idea Scylla supports indexes, not sure how I missed that in the docs. My post may have not made it entirely clear, but selects on app ID and device ID are also full table scans. So I wonder -especially given your remark about the perf- if an index will be fast enough for that. Do you have any idea if the index solution can cover more than single selects? TLL would indeed be a good option for the cleanup of old data. All three ID properties are GUID’s, so I don’t think it can be considered ordered data for which I can do greater than checks. Additionally, do you think that approach would be faster than the token range scans? Regarding querying rows in an ordered way, I’ve not had to do this very often, but one common solution is to use a timeuuid. It has some downsides, but it is a quick and easy way to make sure the record uuid’s can be scanned easily: With regards to index performance, it’s not too much different to mysql, it simply means a second copy of your data is being stored. It’s not that much different to how mysql handles indexes. The questions you are asking are standard cql (nosql) problems, and have fairly standard answers. Some Youtube video tutorials on how to design a nosql database might be helpful. (Thats how I learn’t all this stuff, back in the day there were some interesting ones, like “how would you model YouTube in Cassandra” that contained a whole lot of helpful insight into how to solve these types of problems in a distributed database) --- ### Page: https://forum.scylladb.com/t/mixed-ttl-values-in-an-update-statement/1410 Title: Mixed TTL values in an update statement? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Recreating my system in scylladb with a little bit more paranoia about keeping personal user data. I cant see that this is possible in the documentation, so I am checking here just in case it is possible. For example, … Language: en Canonical URL: https://forum.scylladb.com/t/mixed-ttl-values-in-an-update-statement/1410 ## Headings Structure: H1: Mixed TTL values in an update statement? H3: Related topics ## Main Content: H1: Mixed TTL values in an update statement? H3: Related topics Recreating my system in scylladb with a little bit more paranoia about keeping personal user data. I cant see that this is possible in the documentation, so I am checking here just in case it is possible. For example, Can we do an insert into a user activity log table, but use different TTL’s for the IP address? I know we could break this up into two different update statements, but that seems silly due to assumedly triggering two round trips on the network and two write activities. It is not possible to specify per-column TTL. They way to achieve this is to use two separate statements. If you don’t want multiple separate round-trips and want to keep the update atomic, you can use a BATCH to bundle the updates into a single one. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-41-2024-03-24/1411 Title: Last week in scylla-cluster-tests.git master (issue #41; 2024-03-24) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 4158cdbb…a452ce97 range are covered. There were 14 non-merge commits from 5 Software Engi… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-41-2024-03-24/1411 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #41; 2024-03-24) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #41; 2024-03-24) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 4158cdbb…a452ce97 range are covered. There were 14 non-merge commits from 5 Software Engineer in Test in that period. Some notable commits: We’ve advanced SCT’s network configuration refactor. Now, AWS availability zones support multiple subnets, identified as primary and secondary, allowing network interfaces to be assigned diversely. All primary subnets in a region connect to one route table, while secondary subnets link to another, facilitating the testing of complex networking setups. K8s functional test results are now sent to Argus and Jenkins in junit.xml format. This is done with help of the new hydra fetch-junit-from-runner command. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/error-on-nodetool-operations-empty-cassandra-directory-installing-scylladb/1412 Title: Error on nodetool operations, empty Cassandra directory, installing ScyllaDB - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/error-on-nodetool-operations-empty-cassandra-directory-installing-scylladb/1412 ## Headings Structure: H1: Error on nodetool operations, empty Cassandra directory, installing ScyllaDB H3: Related topics ## Main Content: H1: Error on nodetool operations, empty Cassandra directory, installing ScyllaDB H3: Related topics Originally from the User Slack @Peeyush: Hello Scylla team, I downloaded scylla version 5.2.5 following the steps listed in the official documentation. I’m unable to perform nodetool operations. Getting the error attached below. And moreover the directory /etc/scylla/cassandra is also empty @syuu1228 I noticed you are the maintainer for scylla-tools-core/files under /etc/scylla/cassandra. Care to comment? @Lubos: FORCE restore from appropriate package when someone manually deletes this dir, it won’t get recreated with apt/deb systems even after reinstall the special option one will do it fwiw - it will be something like likely the options will be different, but you get the gist @Peeyush @Peeyush: Thanks @Lubos Even after uninstalling scylladb, there were some scylla tool that were still left in multiple directories. I deleted all of them and then reinstalled again and got it to work. --- ### Page: https://forum.scylladb.com/t/data-model-with-small-partitions/1413 Title: Data model with small partitions - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/data-model-with-small-partitions/1413 ## Headings Structure: H1: Data model with small partitions H3: Related topics ## Main Content: H1: Data model with small partitions H3: Related topics Originally from the User Slack @Alexandre_Debril: Hello Scylla people We’re designing our data model in such a way that we suspect we will end up with very small partitions (like 1 to 10 rows / partition) spread across nodes. Do you think it can be an issue ? Is there any recommendation we should follow for that kind of data design ? Thanks in advance @Felipe_Cardeneti_Mendes: it primarily depends on the queries. If you frequently need to read 100 rows, you would need to run 10 queries. 1K rows, 100 queries. And so it scales and may add to your client-side tail latency But yea, data distribution wise, this looks good @Alexandre_Debril: ok, great. Thanks for your answer @Felipe_Cardeneti_Mendes --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-223-2024-03-24/1417 Title: Last week in scylladb.git master (issue #223; 2024-03-24) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 7df3acd39c…6bdb456fad range are covered. There were 149 non-merge commits from 24 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-223-2024-03-24/1417 ## Headings Structure: H1: Last week in scylladb.git master (issue #223; 2024-03-24) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #223; 2024-03-24) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 7df3acd39c…6bdb456fad range are covered. There were 149 non-merge commits from 24 authors in that period. Some notable commits: The native nodetool status command now handles joining nodes better. A reactor stall when reading the materializing very large schemas has been fixed. During compaction tombstone garbage collection, we now skip the bloom filter checks if the sstable only contains young data. During compaction tombstone garbage collection, we now consider the memtable only if it contains the key. This prevents old data in memtable from preventing tombstone garbage collection. Tracking of memory within materialized view updates was improved. The replace operation, when running with the same number of nodes as the replication factor (e.g. typically 3), the node replace operation now works correctly. There is now documentation for initiating upgrade into consistent topology management and recovering from a disaster when it is enabled. A crash when a view was created while a TRUNCATE operation is in progress was fixed. The system.token_ring virtual table is now tablet-aware. The native nodetool command now mimics Java-style command line handling for greater compatibility with scripts. The native nodetool command can now query effective ownership on a table basis, useful with tablets. Raft now handles lost quorums more flexibly, allowing callers to decide whether they want to timeout or not. The service level subsystem moved its storage from the system_distributed keyspace to a new table system.service_levels_v2, managed by raft group 0. Data is automatically migrated. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/building-scylladb-on-mageia-9/1418 Title: Building ScyllaDB on Mageia 9 - Database Community - ScyllaDB Community NoSQL Forum Meta Description: Hello there, I hope i’m posting this in the right place. I’m starting my journey in ScyllaDB world by trying to build it under Mageia 9 (mageia.org) for which i’m conributing. I think i’m not close to succeed butt i’… Language: en Canonical URL: https://forum.scylladb.com/t/building-scylladb-on-mageia-9/1418 ## Headings Structure: H1: Building ScyllaDB on Mageia 9 H3: Related topics ## Main Content: H1: Building ScyllaDB on Mageia 9 H3: Related topics I hope i’m posting this in the right place. I’m starting my journey in ScyllaDB world by trying to build it under Mageia 9 (mageia.org) for which i’m conributing. I think i’m not close to succeed butt i’m blocked by a clang error for which i’m getting mad. 1a) Packages that are already in mageia I got there and i thing i forgot nothing : 1b) Packages hat are not already in mageia cargo install cxxbridge-cmd Webassembly : download, build and install Binaryen : download build and install For those 3 i’ll have to package them properly but i want firt to get to the full end before redoing the path cleanly for packaging. So they are build and installed manually and locally in /usr/local/ There it complains about difficulties to build for Debian and for RHEL (which is understandable). Then i tried to build only release : And even only scylla binary ninja build/release/scylla It compiles a lot of thins but in all cases it finishes by failing on the same problem : This “no matching function for call to ‘construct_at’” error drives me crazy. My version of clang is this one : And my verison of ninja is this one : If you could have a look and tell me if i missed something obvious or advice me about how to pass this step that would be really cool. Clang 16 is able to compile this code. I checked and indeed table_info is a simple POD without a constructor. Maybe there was something in C++20 to generate constructors for such PODs (although I couldn’t find anything with a quick search). That said, we have a frozen toolchain (a docker image) in tools/toolchain/dbuild, which is the officially supported way to build scylla binaries. The resulting packages can be installed on any distro. So maybe the best course of action is: This will get you a tarball, with all the binaries and an install.sh script, which installs it. You can wrap this with packaging automation, or you can use the DEB/RPM packages if those work for you (ninja dist-release). Hello, thanks for having taken the time to give a look. If i conclude that my clang 15 is too old am i correct ? For the build process the rule in Mageia is the same as for Debian : packages need to be build on mageia systems with mageia tools and libraries. So if i want mageia users to be able to do a “urpmi scylladb” to install from official mageia mirrors and use it i cannot consider dbuild way, it would never work on the distribution build system If i get over this blocking point, i would then have to work on a way to build also rpms for Mageia and make a pull request to submit it to Scylla community. --- ### Page: https://forum.scylladb.com/t/how-to-ignore-field-in-rust-struct/1419 Title: How to ignore field in rust struct - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I have a UserSession struct, and a User struct. I want to cache a pointer to the user from the UserSession struct to the User struct. To pass them around together, i.e.: #[derive(Debug, FromRow, SerializeRow)] pub stru… Language: en Canonical URL: https://forum.scylladb.com/t/how-to-ignore-field-in-rust-struct/1419 ## Headings Structure: H1: How to ignore field in rust struct H3: Related topics ## Main Content: H1: How to ignore field in rust struct H3: Related topics I have a UserSession struct, and a User struct. I want to cache a pointer to the user from the UserSession struct to the User struct. To pass them around together, i.e.: However, there is something about the User object, the driver doesn’t like, I now get this error in rust: I have spent some time trying to work out how to make this error go away, I notice that there is something called #[scylla(skip)] but it seems not useful in this case. How can I make the driver either ignore that cache info:Option field, or make the info:Option serializable so that the error goes away (even though I don’t need to serialize it). I realize a workaround would be to create two versions of the Session struct, and copy out of that a Session2 struct with the cached field, but why do extra memory copies if we dont need to right? #[scylla(skip)] is (currently) attribute only working with serialization, not deserialization. I don’t think we have any similar attribute for deserialization right now, but there is currently a big deserialization refactor going on. I added a comment to remember about #[scylla(skip)]: Deserialization refactor: macros for the new traits · Issue #962 · scylladb/scylla-rust-driver · GitHub . As for what you can do now: create a separate DBUserSession without this field and impl From for UserSession is one solution, but with a drawback you mentioned. Tthe performance impact should however be negligible, I think worrying about it now is unnecessary. Another solution would be to implement FromRow yourself. Look at the generated implementation for the struct without info field, copy the implementation and add info field initialization to None. Thanks for the info. I was hoping there was an easy answer and I was just missing it in the documentation. I just implemented a workaround for now (defined a new struct to combine pointers to a session and user struct). --- ### Page: https://forum.scylladb.com/t/insert-read-delete-workflow-with-single-rows/1420 Title: INSERT => READ => DELETE workflow with single rows - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/insert-read-delete-workflow-with-single-rows/1420 ## Headings Structure: H1: INSERT => READ => DELETE workflow with single rows H3: Related topics ## Main Content: H1: INSERT => READ => DELETE workflow with single rows H3: Related topics Originally from the User Slack @Dylan_Piette: Hello to the scylla team, I have a question for you In our project we might have a use case with Scylla where we would be forced to do a lot of INSERT => READ => DELETE workflow with single rows I was wondering how would you manage that ? Wouldn’t it be preferable to create another boolean column named “Acknowleged” and have a batch job that will delete them later ? So the new workflow would be INSERT => READ => UPDATE (Acknowleged = true) and later do a batch job consisting of DELETE WHERE Acknowleged = true @dor: It’s hard to say, I’d start with a simplistic approach and just delete the row. Batching can primarily help if you have a window of time where the cluster is not busy and you can even call compaction later. It’s not a must @Dylan_Piette: Ok thanks for your answer, I’ll go with that then I was asking because I saw some posts online that says that cassandra like databases are not meant for these kind of use cases, and as we don’t want to deploy multiple databases system, I was wondering if there was a common strategies with scylla to deal with that @dor: Deletes create tombstones (it basically similar to your ‘update’ and actual delete later Most of the time we handle those tombstones well, so assume that’s the case @Dylan_Piette: Ok great, thank you ! --- ### Page: https://forum.scylladb.com/t/scylla-cluster-membership-issue-after-failed-change/1421 Title: Scylla cluster membership issue after failed change - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Have a little problem. Trying to learn how to deal with and fix clustering issues in scylla. (currently 5.2.2) I created a cluster and then I added 2 nodes from a different datacenter into it (so currently datacenter1 = … Language: en Canonical URL: https://forum.scylladb.com/t/scylla-cluster-membership-issue-after-failed-change/1421 ## Headings Structure: H1: Scylla cluster membership issue after failed change H2: nodetool decommission nodetool: Scylla API server HTTP POST to URL ‘/storage_service/decommission’ failed: std::runtime_error (local node is not a member of the token ring yet) See ‘nodetool help’ or ‘nodetool help ’. H3: Related topics ## Main Content: H1: Scylla cluster membership issue after failed change H2: nodetool decommission nodetool: Scylla API server HTTP POST to URL ‘/storage_service/decommission’ failed: std::runtime_error (local node is not a member of the token ring yet) See ‘nodetool help’ or ‘nodetool help ’. H3: Related topics Have a little problem. Trying to learn how to deal with and fix clustering issues in scylla. (currently 5.2.2) I created a cluster and then I added 2 nodes from a different datacenter into it (so currently datacenter1 = 3 nodes, datacenter2 = 2 nodes, both files use local=true) I made a configuration error in the cassandra-rackdc.properties file when I tried to add the third node to datacenter 2 and added the tag as datacenter1 instead. The node wouldn’t add, just stalled. I tried to readd and it first said that it already existed. After trying to figure this out, i wiped the data and tried again and then it gave me the error that it couldn’t resolve ip addresses for node with id and it gave multiple ID’s that weren’t in the cluster. Now I can’t add any nodes. I created a new node, attempted to add it to the cluster and it just froze. I attempted to reboot and try again and now it’s saying that it already exists in the cluster. However, it doesn’t show up in nodetool status or system.peers. How can I fix this issue so I can add nodes to the cluster? please read the documentation on handling membership change failures, and follow the instructions there: Handling Cluster Membership Change Failures | ScyllaDB Docs Thank you for this. I read through this and at the end it says "If removenode returns an error like: nodetool: Scylla API server HTTP POST to URL ‘/storage_service/remove_node’ failed: std::runtime_error (removenode[12e7e05b-d1ae-4978-b6a6-de0066aa80d8]: Host ID 42405b3b-487e-4759-8590-ddb9bdcebdc5 not found in the cluster) and you’re sure that you’re providing the correct Host ID, it means that the member was already removed and you don’t have to clean up after it. However, Here’s the output from attempts to remove a node. 10.0.137.180 nodetool info ID : dd530cfb-7f4e-42c3-be6a-ce093d263b96 Gossip active : true nodetool: Scylla API server HTTP GET to URL ‘/storage_service/rpc_server’ failed: Not found See ‘nodetool help’ or ‘nodetool help ’. 10.0.130.77 (one of the nodes in the cluster) nodetool describecluster Cluster Information: Name: Veeps Snitch: org.apache.cassandra.locator.GossipingPropertyFileSnitch DynamicEndPointSnitch: disabled Partitioner: org.apache.cassandra.dht.Murmur3Partitioner Schema versions: 6e0ec14b-1d4b-305b-9bf9-d420da03eb45: [10.0.137.180] 7f3ff0e1-96d7-3145-8883-f6125b5522f6: [10.0.130.77, 10.0.137.241] nodetool removenode dd530cfb-7f4e-42c3-be6a-ce093d263b96 nodetool: Scylla API server HTTP POST to URL ‘/storage_service/remove_node’ failed: std::runtime_error (removenode[bf9d0aed-d300-48e0-ad70-8637d11972d6]: Node dd530cfb-7f4e-42c3-be6a-ce093d263b96 not found in the cluster) See ‘nodetool help’ or ‘nodetool help ’. I see you marked my answer as solution, does the problem still persist or have you managed to solve it? I see you tried to run nodetool decommission – on which node, and why? (The error you got is suspicious) Anyway if you’d like us to proceed with investigation then please post the output of: I didn’t mark it as a solution. Someone else did. Status=Up/Down |/ State=Normal/Leaving/Joining/Moving – Address Load Tokens Owns Host ID Rack UN 10.0.137.241 124.76 MB 256 ? 08f06ec2-d35d-444c-afa6-6566b0eada57 rack1 UN 10.0.130.77 123.49 MB 256 ? 0b7f7c40-ba5c-42c1-94f2-207aca6facbe rack1 peer | dc | host_id | load | owns | status | tokens | up ----------------±------------±-------------------------------------±------------±---------±--------±-------±------ 10.250.135.89 | null | null | 2.82419e+06 | null | LEFT | 0 | False 10.250.141.220 | null | null | 3.0423e+06 | null | LEFT | 0 | False 10.0.130.77 | datacenter1 | 0b7f7c40-ba5c-42c1-94f2-207aca6facbe | 1.29489e+08 | 0.500783 | NORMAL | 256 | True 10.0.137.241 | datacenter1 | 08f06ec2-d35d-444c-afa6-6566b0eada57 | 1.30818e+08 | 0.499217 | NORMAL | 256 | True 10.0.137.180 | null | null | | null | UNKNOWN | 0 | True ok so it looks like actually stopping scylla-server on .180 removed it from that list. However what’s your thoughts on the other 2 listed? Those were already offline. ok so it looks like actually stopping scylla-server on .180 removed it from that list. However what’s your thoughts on the other 2 listed? Those were already offline. You mean .89 and .220? They are shown in LEFT status. Expected if you removed them. So now do I understand correctly that you have 2 nodes in the cluster, both in datacenter1/rack1? Sorry, I forgot to ask before – can you also provide output of select * from system.raft_state? cassandra@cqlsh> select * from system.raft_state; group_id | disposition | server_id | can_vote ----------±------------±----------±--------- and yes, 2 nodes, both in datacenter1/rack1. Hm, I was convinced that you have Raft enabled in your cluster (due to the error you said you were getting – that it cannot resolve IP address), but apparently not. Do both nodes return this result when you connect to cqlsh with them? (Empty system.raft_state table) What is your conf/scylla.yaml? Do you have the consistent_cluster_management flag set there? Yeah, that’s part of where I was confused as well. I have consistent_cluster_management set to true on both nodes and has been from day 1. Both nodes return that result. Another thing to mention is this query :select value from system.scylla_local where key = ‘raft_group0_id’; doesn’t return anything either because that key doesn’t exist in scylla_local. That comes from the link you sent me earlier in the post. You haven’t done the “manual Raft recovery procedure”, have you? What is the result of SELECT * FROM system.scylla_local WHERE key = 'group0_upgrade_state'; on each node? Ok, now i feel a little dumb. I went into recovery mode night before last when trying to fix this. I have backed out of this and now this is the result of that query: key | value ----------------------±------------------------- group0_upgrade_state | use_post_raft_procedures WDYM by “backed out”? Have you finished the recovery procedure? If you did anything from the procedure after “enter recovery mode step” (like truncated some Raft tables), then you need to finish it. If you only entered recovery mode, but then set group0_upgrade_state back to use_post_raft_procedures without removing any other data etc., then I guess everything should be fine. So if you did anything else besides just entering recovery mode – you should finish the recovery procedure (starting by going into recovery mode again), and then proceed. Otherwise we should be able to proceed now. If we proceed, then the next step is to check the results of select * from system.raft_state again. Yeah, the only thing I did was enter recovery mode. I didn’t truncate any tables. group_id | disposition | server_id | can_vote --------------------------------------±------------±-------------------------------------±--------- 5d63f9a0-ede9-11ee-8462-8160e0cdd5d0 | CURRENT | 08f06ec2-d35d-444c-afa6-6566b0eada57 | True 5d63f9a0-ede9-11ee-8462-8160e0cdd5d0 | CURRENT | 0b7f7c40-ba5c-42c1-94f2-207aca6facbe | True Is now the result of raft_state So looks like nodetool status is consistent with raft state, there are no “ghost members” anymore. Try to boot the new node again. Make sure you don’t use old work directories from previous boot attempts though – clear the old data from node which failed to boot if you do it on the same machine. If it again gets stuck on “fail to resolve IP” please save the host ID, we’ll have to determine what node this host ID belongs to. Thank you! I will try this now. Ok, did the following (noting it here for posterity) Removed all the data from 10.0.137.180 Ensured cassandra-rackdc.properties was set to datacenter1/rack1 Ensured the seeds were correct in scylla.yaml and all the parameters matched (consistent_cluster_management, etc) Started scylla Now the node has successfully joined datacenter1. I am now going to create 3 new nodes and attempt to join datacenter2 to this cluster. Just make sure you do the joining sequentially (like it’s written in the 5.2 docs.) --- ### Page: https://forum.scylladb.com/t/error-not-enough-memory-cluster-with-different-instance-types/1423 Title: Error, not enough memory - cluster with different instance types - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi team, Iam working with AWS cloud, deployed a cluster. One node configured with like 2CPU and 8GB RAM (m6g.large-instance type) and another node configured with lower 1CPU and 4GB RAM (m6g.medium-instance type). Err… Language: en Canonical URL: https://forum.scylladb.com/t/error-not-enough-memory-cluster-with-different-instance-types/1423 ## Headings Structure: H1: Error, not enough memory - cluster with different instance types H3: Related topics ## Main Content: H1: Error, not enough memory - cluster with different instance types H3: Related topics Hi team, Iam working with AWS cloud, deployed a cluster. One node configured with like 2CPU and 8GB RAM (m6g.large-instance type) and another node configured with lower 1CPU and 4GB RAM (m6g.medium-instance type). Error: After deploying lower configured node that node is not up in the status error it is showing shard distribution error. That screenshots attached below. Team, can you help on this, It will work or not otherwise we need to keep same configuration on both the nodes. please suggest. Thanks, First, running production clusters with different instance types is possible, but: Second, the error code clearly state you do not have enouth RAM per core. You need to either increase the memory, or reduce the number of core. Thanks @tzach. But Iam using for development cluster for data standby. it is ok? Sure, if it’s non-production, and you understand the risk. Almost all the tests we are doing are for clusters with symmetric nodes. Asymmetric nodes with different numbers of cores, should work, but they are less tested. --- ### Page: https://forum.scylladb.com/t/decommissioned-node-still-shows-in-clusterstatus/1424 Title: Decommissioned node still shows in clusterstatus - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: ubuntu@east2-scylla:~$ nodetool removenode dd530cfb-7f4e-42c3-be6a-ce093d263b96 nodetool: Scylla API server HTTP POST to URL ‘/storage_service/remove_node’ failed: std::runtime_error (removenode[be14b77f-ed3f-4f9d-a1ad-3… Language: en Canonical URL: https://forum.scylladb.com/t/decommissioned-node-still-shows-in-clusterstatus/1424 ## Headings Structure: H1: Decommissioned node still shows in clusterstatus H1: Datacenter: datacenter1 H3: Related topics ## Main Content: H1: Decommissioned node still shows in clusterstatus H1: Datacenter: datacenter1 H3: Related topics ubuntu@east2-scylla:~$ nodetool removenode dd530cfb-7f4e-42c3-be6a-ce093d263b96 nodetool: Scylla API server HTTP POST to URL ‘/storage_service/remove_node’ failed: std::runtime_error (removenode[be14b77f-ed3f-4f9d-a1ad-328cd6ce03a2]: Node dd530cfb-7f4e-42c3-be6a-ce093d263b96 not found in the cluster) See ‘nodetool help’ or ‘nodetool help ’. ubuntu@east2-scylla:~$ nodetool describecluster Cluster Information: Name: MyCluster Snitch: org.apache.cassandra.locator.GossipingPropertyFileSnitch DynamicEndPointSnitch: disabled Partitioner: org.apache.cassandra.dht.Murmur3Partitioner Schema versions: 6e0ec14b-1d4b-305b-9bf9-d420da03eb45: [10.0.137.180] 7f3ff0e1-96d7-3145-8883-f6125b5522f6: [10.0.130.77, 10.0.137.241] 10.0.137.180 = dd530cfb-7f4e-42c3-be6a-ce093d263b96 I already ran nodetool decommission on that node what gives? Can you give the output of nodetool status? Status=Up/Down |/ State=Normal/Leaving/Joining/Moving – Address Load Tokens Owns Host ID Rack UN 10.0.137.241 124.9 MB 256 ? 08f06ec2-d35d-444c-afa6-6566b0eada57 rack1 UN 10.0.130.77 123.6 MB 256 ? 0b7f7c40-ba5c-42c1-94f2-207aca6facbe rack1 --- ### Page: https://forum.scylladb.com/t/unavailableexception-not-enough-replicas-available-for-query/1427 Title: UnavailableException: Not enough replicas available for query - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/unavailableexception-not-enough-replicas-available-for-query/1427 ## Headings Structure: H1: UnavailableException: Not enough replicas available for query H3: Related topics ## Main Content: H1: UnavailableException: Not enough replicas available for query H3: Related topics Originally from the User Slack @Praveen: Hi. I have 3 nodes in a data center and one of the node was down and the application got error While checking the cassandra-rackdc.properties I am using only DC and rack. Not using prefer_local. Also I am using GossipingPropertyFileSnitch instead of SimpleSnitch I created the keyspace using below SQL Could you please let me know where is the wrong configuration? @Felipe_Cardeneti_Mendes: Looks ok. Maybe one of the replicas flapped down momentarily and then failed? Btw, prefer NetworkTopologyStrategy, even for single DC. @Praveen: Thanks Felipe. I have tried this option and tried but got the same error. Even I tried to create a new keyspace and still got the same. Am I missing something that still causing this? @Felipe_Cardeneti_Mendes: is this a multi DC deployment by any chance? @Praveen: No. This is not a multi-DC deployment. This is the same configuration on each node except rack changed to scylla-2 and scylla-3 DC is same as DC01 @Felipe_Cardeneti_Mendes: hum, upper case. Well, ensure all nodes are really within the same DC, and that you follow the same case for specifying the DC within your drivers. If this still persists, then trace the failing query, as this should help you understand what’s going on. If it still doesn’t help, open an issue and provide all that info. @Praveen: Sure. Driver might not be updated with the DC info. I will try this again and let you know. Thanks Felipe. --- ### Page: https://forum.scylladb.com/t/using-rest-endpoints-with-scylladb/1428 Title: Using REST endpoints with ScyllaDB - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/using-rest-endpoints-with-scylladb/1428 ## Headings Structure: H1: Using REST endpoints with ScyllaDB H3: Related topics ## Main Content: H1: Using REST endpoints with ScyllaDB H3: Related topics Originally from the User Slack @Sandesh: Hello team, I wanted to know if Scylladb Cloud tables can be queried using REST endpoints? For example, in elasticsearch I use postman to run the queries directly @dor: Not with the CQL api but you can do it with our DynamoDB api @Sandesh: Thanks a lot, sure will check it out --- ### Page: https://forum.scylladb.com/t/forward-pagination-with-mappedasync/1430 Title: Forward pagination with MappedAsync - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Is there a way to implement the following asynchronously with MappedAsyncPagingIterable interface with CompletionStage<MappedAsyncPagingIterable> mappedAsyncPage? I want to implement a linkedin feed feature where first … Language: en Canonical URL: https://forum.scylladb.com/t/forward-pagination-with-mappedasync/1430 ## Headings Structure: H1: Forward pagination with MappedAsync H3: Related topics ## Main Content: H1: Forward pagination with MappedAsync H3: Related topics Is there a way to implement the following asynchronously with MappedAsyncPagingIterable interface with CompletionStage mappedAsyncPage? I want to implement a linkedin feed feature where first request fetches first 5 rows and the following requests will fetch the next 3 rows So If I understood correctly, you would like to make an implementation of AsyncPagingIterable that will have a variable page size set to 5 at first and then to 3 for subsequent pages. While that may be possible to do and make driver use it, I’m not sure if I would be able to provide such implementation. I believe it would be also too complicated to be worth the effort. I’d recommend choosing one page size for the needs of your query and handle the portioning of it on application side and not driver side. For example under the hood, the driver in your application may fetch 10 rows each time, but your application’s front-end will control how many it will show - 5 at first then 3 each time you click “More” button or something like that. --- ### Page: https://forum.scylladb.com/t/scylladb-maximum-node-size-ram-disk-ratio-cpus/1432 Title: ScyllaDB maximum node size, RAM:Disk ratio, CPUs - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-maximum-node-size-ram-disk-ratio-cpus/1432 ## Headings Structure: H1: ScyllaDB maximum node size, RAM:Disk ratio, CPUs H3: Related topics ## Main Content: H1: ScyllaDB maximum node size, RAM:Disk ratio, CPUs H3: Related topics Originally from the User Slack @Nguyen_Huu_Trung_Ha: Dear experts. How big of data a ScyllaDB node can handle? @dor: There is no limit, it depends on the rest of the node resources. In a machine like i3en.metal, we can utilize all of the 60TB It’s good not to exceed a ratio of 1;100 ram:disk @Nguyen_Huu_Trung_Ha: Thank you for your reply @dor. I have another question: what is the recommended data size a node can handle well, including: decommissioning, adding nodes, compaction, cleanup, etc in acceptable time. @dor: It’s primarily a function of your compaction strategy. We recommend having about 30% free space for all of those @Nguyen_Huu_Trung_Ha: For archiving purposes with seldom/rarely access, can we use machines with less CPU, RAM and big storage, for example: 12 CPU cores, 128GB RAM and 20-60TB storage? @dor: The CPUs are less of a problem, it’s the ram:disk ratio, at the end it depends on the access pattern, right now the recommendation is 1:100 and we will work to enlarge it this year to even more than you need --- ### Page: https://forum.scylladb.com/t/why-owns-always-in-the-nodetool-status-command/1434 Title: Why "Owns" always "?" in the nodetool status command - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I am doing all configs correctly and follow the documentation but when I run nodetool status the Owns value is always ?. What can I do about that? Datacenter: 168 =============== Status=Up/Down |/ State=Normal/Leaving/J… Language: en Canonical URL: https://forum.scylladb.com/t/why-owns-always-in-the-nodetool-status-command/1434 ## Headings Structure: H1: Why "Owns" always "?" in the nodetool status command H3: Related topics ## Main Content: H1: Why "Owns" always "?" in the nodetool status command H3: Related topics I am doing all configs correctly and follow the documentation but when I run nodetool status the Owns value is always ?. What can I do about that? To have an actual value for Owns, you need to provide a keyspace parameter for nodetool status, e.g. nodetool status my_keyspace. Thank you. I am just newbie on Scylla. But in cassandra it shows owns only with nodetool status --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-42-2024-03-29/1435 Title: Last week in scylla-cluster-tests.git master (issue #42; 2024-03-29) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 2e49a1e1…a9e63b34 range are covered. There were 16 non-merge commits from 5 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-42-2024-03-29/1435 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #42; 2024-03-29) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #42; 2024-03-29) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 2e49a1e1…a9e63b34 range are covered. There were 16 non-merge commits from 5 authors in that period. Some notable commits: After the networking refactor, we’ve successfully configured multiple network interfaces on DB nodes, loaders, and runners. This enhancement allows for the execution of tests with independent networks, utilizing private IP addresses, marking a significant improvement in our networking setup. We’ve introduced support for the Scylla Alternator feature on Kubernetes. To activate this feature, set k8s_enable_alternator: true in the test configuration. For using the insecure Alternator port, configure alternator_port: 8000. To employ a secure Alternator port, k8s_enable_tls: true and alternator_port: 8043 must be defined. Additionally, YCSB longevity tests have been added to validate this feature. For those running tests locally using Docker, it’s recommended to adopt a dedicated test configuration. This approach minimizes load and adjusts Scylla settings appropriately to prevent overloading the host machine. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/querying-by-non-partition-key-column-creating-an-index/1436 Title: Querying by non partition key column, creating an index - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/querying-by-non-partition-key-column-creating-an-index/1436 ## Headings Structure: H1: Querying by non partition key column, creating an index H3: Related topics ## Main Content: H1: Querying by non partition key column, creating an index H3: Related topics Originally from the User Slack @Benaceur_Ayoub: I have a table “stores” where 'id’ is the partition key, but sometimes I want to query by ‘city’, the problem is each city will contain thousands of stores, so if I create an index on ‘city’ would that would make the partitions large for that index so therefor this is not viable solution. is this correct ? @Felipe_Cardeneti_Mendes: Correct! Such an index would be large and likely very imbalanced @Benaceur_Ayoub: so there is no way I can query by city ? @Felipe_Cardeneti_Mendes: Well, if the ratio of stores per city is up to a few thousands an index will do just fine. If it gets to hundreds of thousands then it wouldn’t. You probably want to make sure you have a StoreByCity kind of table where you add more cardinality with — for example — a zip code range… then you would simply run parallel queries until you walked over all zips for a given city. Or just full table scan with Spark if this is adhoc --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-224-2024-03-31/1437 Title: Last week in scylladb.git master (issue #224; 2024-03-31) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 6bdb456fad…885cb2af07 range are covered. There were 90 non-merge commits from 20 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-224-2024-03-31/1437 ## Headings Structure: H1: Last week in scylladb.git master (issue #224; 2024-03-31) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #224; 2024-03-31) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 6bdb456fad…885cb2af07 range are covered. There were 90 non-merge commits from 20 authors in that period. Some notable commits: In consistent topology mode, it is now possible to change the snitch. Materialized views track the amount of work in progress in order to limit it. Due to a bug, if the base replica had to update both a local view replica and a remote view replica, then only the work to update the local view replica was tracked. This could lead to running out of memory, and is now fixed. In consistent topology more, we now only allow decommission if the node is in NORMAL state. It turned out that some nodetool commands are still not implemented by native nodetool; it now supports getsstables, sstableinfo and checkAndRepairCdcStreams. The CREATE KEYSPACE statement will now warn about unsupported features if tablets are used, allowing the user to disable tablets if they require those features. In consistent topology mode, we now support the initial_token configuration parameter. A bug while copying data from cache during reverse queries was fixed. The unchecked_tombstone_compaction compaction option has been implemented. Setting this option will ignore tombstone_threshold. Alternator, ScyllaDB’s implementation of the DynamoDB API, is now more careful to prevent other queries from stalling when processing large queries. A bug which delayed upgrade to consistent topology until node restart was fixed. The tablet allocator now supports replication factor changes, adding or removing tablets as needed. This is not yet wired to CQL. The native nodetool repair command will now abort on the first failed repair. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/changing-the-number-of-members-with-the-scylladb-operator-expired-certificate/1439 Title: Changing the number of members with the ScyllaDB Operator, expired certificate - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/changing-the-number-of-members-with-the-scylladb-operator-expired-certificate/1439 ## Headings Structure: H1: Changing the number of members with the ScyllaDB Operator, expired certificate H3: Related topics ## Main Content: H1: Changing the number of members with the ScyllaDB Operator, expired certificate H3: Related topics Originally from the User Slack @Nino_Matos: updated: I’m trying to scale my scylladb from 3 members to 4 members and i’m not being able to do it. I’ve tried following the documentation: https://github.com/scylladb/scylla-operator/blob/master/docs/source/generic.md#scale-a-scyllacluster I’ve tried to scale the statefulset and the new node is not able to connect(maybe because there is no service). i’ve edited the members on the scylla cluster itself (scylla.scylladb.com/v1/scyllaclusters) using k9s and nothing… My expectation was that when i change the members of the scylla cluster that would trigger a deployment of the new node. i’ve tried running the kubectl way and got this: @Maciej_Zimnoch: you shouldn’t modify resources managed by Operator on your own, instead modify ScyllaCluster resource. Looks like webhook certificate is expired, new one should be automatically generated and picked up by the webhook. Please provide output of your: @Aleksandar_Kocic: Changing the number of members worked for me. @Nino_Matos: @Maciej_Zimnoch yeah, that’s what i ended up doing. Changing the scyllacluster resource. The problem was that because my certificate was expired i couldn’t change the resource. Once i updated the certificate i was able to change the resource<- @Aleksandar_Kocic Thanks for the suggestions --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-4-5/1440 Title: [RELEASE] ScyllaDB 5.4.5 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.4.5, a bugfix release of the ScyllaDB 5.4 stable branch. ScyllaDB Open Source 5.4.5, like all past and future 5.x.y releases, is backward compatible and supports rolling… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-4-5/1440 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.4.5 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.4.5 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.4.5, a bugfix release of the ScyllaDB 5.4 stable branch. ScyllaDB Open Source 5.4.5, like all past and future 5.x.y releases, is backward compatible and supports rolling upgrades. Users are encouraged to upgrade to 5.4.5. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/release-scylla-manager-3-2-7/1441 Title: [RELEASE] Scylla Manager 3.2.7 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces the release of ScyllaDB Manager 3.2.7 production-ready ScyllaDB Manager patch release of the stable ScyllaDB Manager 3.2 branch. ScyllaDB Manager is a centralized cluster administration and r… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-manager-3-2-7/1441 ## Headings Structure: H1: [RELEASE] Scylla Manager 3.2.7 H3: Related topics ## Main Content: H1: [RELEASE] Scylla Manager 3.2.7 H3: Related topics The ScyllaDB team announces the release of ScyllaDB Manager 3.2.7 production-ready ScyllaDB Manager patch release of the stable ScyllaDB Manager 3.2 branch. ScyllaDB Manager is a centralized cluster administration and recurrent tasks automation tool. In 3.2.7, we updated the grpc package - google.golang.org/grpc - Go Packages dependency to version 1.53.3 #3751. Additionally, we addressed a bug related to improper closure of client connections #3769, improved error logging during task execution failures #3764, and enhanced the output displayed by the sctool tasks command #3750. ScyllaDB Enterprise customers are encouraged to upgrade to ScyllaDB Manager 3.2.7 in coordination with the ScyllaDB support team. The new release includes upgrades of both ScyllaDB Manager Server and Agent. Scylla Manager 3.2.7 supports the following Scylla Enterprise releases: And the following Open Source release (limited to 5 nodes see license): You can install and run Scylla Manager on Kubernetes using Scylla Operator. More here. --- ### Page: https://forum.scylladb.com/t/release-scylla-operator-1-12-0/1446 Title: [RELEASE] Scylla Operator 1.12.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of Scylla Operator 1.12.0. Scylla Operator is an open-source project that helps ScyllaDB Open Source and ScyllaDB Enterprise users run ScyllaDB on Kubernetes. The … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-operator-1-12-0/1446 ## Headings Structure: H1: [RELEASE] Scylla Operator 1.12.0 H1: Notable changes H1: Supported versions H1: Upgrade instructions H2: Getting started with Scylla Operator H2: Related Links H3: Related topics ## Main Content: H1: [RELEASE] Scylla Operator 1.12.0 H1: Notable changes H1: Supported versions H1: Upgrade instructions H2: Getting started with Scylla Operator H2: Related Links H3: Related topics The ScyllaDB team is pleased to announce the release of Scylla Operator 1.12.0. Scylla Operator is an open-source project that helps ScyllaDB Open Source and ScyllaDB Enterprise users run ScyllaDB on Kubernetes. The Scylla Operator manages ScyllaDB clusters deployed to Kubernetes and automates tasks related to operating a ScyllaDB cluster, like installation, vertical and horizontal scaling, as well as rolling upgrades. Scylla Operator 1.12.0 improves stability and brings new features. As with all of our releases, all API changes are backward compatible. Upgrading from v1.11.x with kubectl apply doesn’t require any extra action, just take the manifest from v1.12.0 tag and substitute the released image. Using helm requires a mandatory manual step for every release because helm can’t handle CRDs updates. For details, see our upgrade documentation. We’ll welcome your feedback! Feel free to open an issue or reach out on the #scylla-operator channel in Scylla User Slack. --- ### Page: https://forum.scylladb.com/t/scylladb-on-kubernetes-enabling-authentication/1447 Title: ScyllaDB on Kubernetes, enabling Authentication - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-on-kubernetes-enabling-authentication/1447 ## Headings Structure: H1: ScyllaDB on Kubernetes, enabling Authentication H3: Related topics ## Main Content: H1: ScyllaDB on Kubernetes, enabling Authentication H3: Related topics Originally from the User Slack @Jatt_Singh: Hi Team, I’ve deployed scylla cluster with manager and i want to enable authentication. I’ve followed this doc https://opensource.docs.scylladb.com/stable/operating-scylla/security/authentication.html But i got one issue -> when the ScyllaDB pod restarted changes will be discarded. Is any permanent way to enable authentication? @Felipe_Cardeneti_Mendes: See https://operator.docs.scylladb.com/stable/generic.html#configure-scylla for creating a configmap and merging it with scylla.yaml. You could also have used scyllaArgs within the ScyllaCluster CRD. Deploying Scylla on a Kubernetes Cluster | ScyllaDB Docs @Jatt_Singh: @Felipe_Cardeneti_Mendes Thank you for response. Can you explain how can i pass username, password using scyllaArgs? @Felipe_Cardeneti_Mendes: you mean the default password during cluster creation? Because typically you would simply authenticate using the default cassandra/cassandra and then create users within CQLsh --- ### Page: https://forum.scylladb.com/t/error-when-running-repairs-with-a-mixed-shard-count-cluster/1449 Title: Error when running repairs with a mixed shard-count cluster - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/error-when-running-repairs-with-a-mixed-shard-count-cluster/1449 ## Headings Structure: H1: Error when running repairs with a mixed shard-count cluster H3: Related topics ## Main Content: H1: Error when running repairs with a mixed shard-count cluster H3: Related topics Originally from the User Slack @Gopinath_M: HI, I upgraded my cluster from scylla 4.6.11 to 5.0.13, now while running repairs i am seeing below errors in logs, do we know why we are seeing these errors and fixed in which scylla version ? @dor: hmm, could be the scylla deliberately slow you down to protect against OOM. @Botond_Dénes what do you say? Do you have a large partition or large collection (we have a table for them) @Botond_Dénes: @Gopinath_M do you have a mixed shard-count cluster? These kind of timeouts, during repair are a known problem, when repairing nodes, with different shard count. You can retry the repair, maybe with reduced concurrency. The best workaround is to avoid having nodes with different shard counts in your cluster. Note that even if all your instances are the same, a different --cpuset or --smp command-line argument passed to ScyllaDB can also cause nodes to have different shard counts. @Gopinath_M: we have 3 datacenters. 1 datacenter is in amazon with 32 cpu and 2 other datacenter are local using physical servers having 96 cpu So is there any fix for this or all scylla versions will have this problems in mixed nodes /instances @Botond_Dénes: All current versions suffer from this. We are working on a fix, in the form of tablets, to be released in 6.0 (if everything goes according do plan). This fix involves a complete refactoring of how we replicate data between nodes. Which is to say, it won’t be backported. @Gopinath_M: okay. So if we just have physical servers in all datacenters of same cpu count or just amazon instances with same cpu count in all datacenter, this errors will go away correct? only mixed nodes will cause problems? @Botond_Dénes: Yes, if all nodes have the same shard count, the problem will go away. Note, that having mixed shard count is not a gurantee that this error will appear. It also depends on luck (or the lack of it ). Having large nodes (many CPUs) makes hitting this error more likely. @Gopinath_M: okay Thanks @Botond_Dénes I see 2 more errors, is this also because of the same node configuration mentioned above: but this Errors are seen in 4.6.11 and 5.0.13 as well @Botond_Dénes @Botond_Dénes: Yes, this is probably the same thing. The error is re-reported on different levels. @Gopinath_M: i see, okay thanks very much for your quick response @Botond_Dénes: This is the lowest level error: The keyword here is _streaming_concurrency_sem . This semaphore is only used by repair and streaming and it should never time out. We only have very generous 10min (or 30min) timeout, to break out from situations where there is no progress made. --- ### Page: https://forum.scylladb.com/t/does-scylladb-keyspace-names-are-case-sensitive-or-not/1450 Title: Does ScyllaDb Keyspace names are Case sensitive or not? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: When I am trying to give keyspace name as Mykeyspace1 in scylla.yaml and after restarting the Scylla service when I am executing CQL queries those audit logs are not getting capture in audit log file. But when I am usin… Language: en Canonical URL: https://forum.scylladb.com/t/does-scylladb-keyspace-names-are-case-sensitive-or-not/1450 ## Headings Structure: H1: Does ScyllaDb Keyspace names are Case sensitive or not? H3: Related topics ## Main Content: H1: Does ScyllaDb Keyspace names are Case sensitive or not? H3: Related topics When I am trying to give keyspace name as Mykeyspace1 in scylla.yaml and after restarting the Scylla service when I am executing CQL queries those audit logs are not getting capture in audit log file. But when I am using mykeyspace1 in lower case then everything is perfectly working fine. Cassandra identifiers, such as keyspace, table and column names, are case-insensitive, unless enclosed in double quotation marks. Consider: You can try providing a name enclosed in the double quotation marks in scylla.yaml. Let me know if that worked for you. --- ### Page: https://forum.scylladb.com/t/log-error-of-mutation-write-timeout/1451 Title: Log error of mutation_write_timeout - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: When there is an exception we get from another subsystem shows that there is a mutation_write_timeout exception, but nothing found in Scylla logs while an update_table request failed. Is there any exception catch in ops … Language: en Canonical URL: https://forum.scylladb.com/t/log-error-of-mutation-write-timeout/1451 ## Headings Structure: H1: Log error of mutation_write_timeout H3: Related topics ## Main Content: H1: Log error of mutation_write_timeout H3: Related topics When there is an exception we get from another subsystem shows that there is a mutation_write_timeout exception, but nothing found in Scylla logs while an update_table request failed. Is there any exception catch in ops like update _table, get_item. Timeouts are not logged in general. The main reason is that their frequency can be high and they would generate a lot of noise or huge logs. If there is any light weight methods to log mutation_write_timeout or mutation_read_timeout in info level or error level? You can try enabling debug-level logging for the storage_proxy logger. Note that this will probably result in a lot of log messages. --- ### Page: https://forum.scylladb.com/t/release-scylla-monitoring-stack-4-7-0/1452 Title: [RELEASE] Scylla Monitoring Stack 4.7.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of ScyllaDB Monitoring Stack 4.7.0 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB Enterprise and ScyllaDB Open Source, based on Prometheus … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-monitoring-stack-4-7-0/1452 ## Headings Structure: H1: [RELEASE] Scylla Monitoring Stack 4.7.0 H3: Related topics ## Main Content: H1: [RELEASE] Scylla Monitoring Stack 4.7.0 H3: Related topics The ScyllaDB team is pleased to announce the release of ScyllaDB Monitoring Stack 4.7.0 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB Enterprise and ScyllaDB Open Source, based on Prometheus and Grafana. ScyllaDB Monitoring Stack 4.7.0 supports: Versions updates for ScyllaDB Monitoring Stack 4.7.0 New Information in ScyllaDB Dashboards Detailed Dashboard Change Alternator Dashboard update Overview Dashboard Change --- ### Page: https://forum.scylladb.com/t/support-for-amazon-m7gd-instance-type-aws-graviton3-processor-with-nvme-ssd-storage/1453 Title: Support for Amazon m7gd instance type (AWS Graviton3 Processor with NVMe SSD Storage) - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/support-for-amazon-m7gd-instance-type-aws-graviton3-processor-with-nvme-ssd-storage/1453 ## Headings Structure: H1: Support for Amazon m7gd instance type (AWS Graviton3 Processor with NVMe SSD Storage) H3: Related topics ## Main Content: H1: Support for Amazon m7gd instance type (AWS Graviton3 Processor with NVMe SSD Storage) H3: Related topics Originally from the User Slack @racevedo: Hi everyone! Are there any plans to support aws m7gd instance types out of the box? I see that m6gd is already supported here. It seems like m7gd should provide better performance overall (ref). @dor: Yes but there is no timeline. You can just install the rpm yourself, there isn’t anything special to do for it @racevedo: Sure thing. Thanks! --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-2-17/1454 Title: [RELEASE] ScyllaDB 5.2.17 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.2.17, a bugfix release of the ScyllaDB 5.2 stable branch. ScyllaDB Open Source 5.2.17, like all past and future 5.x.y releases, is backward compatible, and supports roll… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-2-17/1454 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.2.17 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.2.17 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.2.17, a bugfix release of the ScyllaDB 5.2 stable branch. ScyllaDB Open Source 5.2.17, like all past and future 5.x.y releases, is backward compatible, and supports rolling upgrades. Note that the latest ScyllaDB Open Source stable release is 5.4 and you are encouraged to upgrade to it. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2023-1-7/1455 Title: [RELEASE] ScyllaDB Enterprise 2023.1.7 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2023.1.7, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2023.1 LTS Release. 2023.1.7 patch release includes multiple minor bug fixes… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2023-1-7/1455 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2023.1.7 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2023.1.7 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2023.1.7, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2023.1 LTS Release. 2023.1.7 patch release includes multiple minor bug fixes. ScyllaDB Enterprise’s latest Long Term Support (LTS) is 2024.1. While we continue to support 2023.1, you are encouraged to upgrade to it in coordination with the ScyllaDB Support team. The following issues are fixed in this release (with an open-source reference, if available): --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-43-2024-04-05/1456 Title: Last week in scylla-cluster-tests.git master (issue #43; 2024-04-05) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 716e2943…5afc9544 range are covered. There were 4 non-merge commits from 1 authors in that… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-43-2024-04-05/1456 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #43; 2024-04-05) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #43; 2024-04-05) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 716e2943…5afc9544 range are covered. There were 4 non-merge commits from 1 authors in that period. Some notable commits: More fixes to local docker configuration stabilising execution and monitoring. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/release-scylla-operator-v1-11-4/1457 Title: [RELEASE] Scylla Operator v1.11.4 - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of Scylla Operator 1.11.4 :rocket:. Scylla Operator 1.11.4 brings version update of Go and all dependencies. As with all of our releases, all API changes are backwar… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-operator-v1-11-4/1457 ## Headings Structure: H1: [RELEASE] Scylla Operator v1.11.4 H3: Related topics ## Main Content: H1: [RELEASE] Scylla Operator v1.11.4 H3: Related topics The ScyllaDB team is pleased to announce the release of Scylla Operator 1.11.4 . Scylla Operator 1.11.4 brings version update of Go and all dependencies. As with all of our releases, all API changes are backward compatible. For more changes and details check out the GitHub release notes. Upgrade instructions When using kubectl apply, upgrading from v1.10.x or v1.11.3 doesn’t require any additional actions, just take the manifest from v1.11.4 tag and substitute for the image. Since helm isn’t capable of handling CRD updates, using helm requires a mandatory manual step for every release. For details, see our upgrade documentation . Best regards, Scylla Operator Team. --- ### Page: https://forum.scylladb.com/t/release-scylla-operator-v1-12-1/1458 Title: [RELEASE] Scylla Operator v1.12.1 - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of Scylla Operator 1.12.1 :rocket:. Scylla Operator 1.12.1 brings version update of Go and all dependencies. As with all of our releases, all API changes are backwar… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-operator-v1-12-1/1458 ## Headings Structure: H1: [RELEASE] Scylla Operator v1.12.1 H3: Related topics ## Main Content: H1: [RELEASE] Scylla Operator v1.12.1 H3: Related topics The ScyllaDB team is pleased to announce the release of Scylla Operator 1.12.1 . Scylla Operator 1.12.1 brings version update of Go and all dependencies. As with all of our releases, all API changes are backward compatible. For more changes and details check out the GitHub release notes. Upgrade instructions When using kubectl apply, upgrading from v1.11.x or v1.12.0 doesn’t require any additional actions, just take the manifest from v1.12.1 tag and substitute for the image. Since helm isn’t capable of handling CRD updates, using helm requires a mandatory manual step for every release. For details, see our upgrade documentation . Best regards, Scylla Operator Team. --- ### Page: https://forum.scylladb.com/t/timeout-paging-memory-readtimeoutexception-cassandra-timeout-during-read-query-at-consistency/1463 Title: Timeout, paging, memory: ReadTimeoutException: Cassandra timeout during read query at consistency - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/timeout-paging-memory-readtimeoutexception-cassandra-timeout-during-read-query-at-consistency/1463 ## Headings Structure: H1: Timeout, paging, memory: ReadTimeoutException: Cassandra timeout during read query at consistency H3: Related topics ## Main Content: H1: Timeout, paging, memory: ReadTimeoutException: Cassandra timeout during read query at consistency H3: Related topics Originally from the User Slack @Varun_Nagrare: Hi all, recently I’m facing a consistency error while reading data with Spark. I repaired my 3 node cluster with replication factor 3 but still it’s giving the same error. How to avoid this error?! Please help!!! @dor: What’s the error msg? @Varun_Nagrare: @dor Below is the error I’m getting: Why would this occur even after repairing the cluster? @dor: It’s not a consistency error but a timeout. If repair was finished already, could be it brough more sstables and now the nodes need to consolidate it. It depends what the machine is doing now You can run slow query tracing and see how many sstables are read for this query @Varun_Nagrare: Please tell how to do slow query tracing? @dor: it’s a cql feature, described in the docs @Varun_Nagrare: Also when checking the syslog I found out the below: scylla: [shard 0] mutation_partition - Memory usage of unpaged query exceeds soft limit of 1048576 (configured via max_memory_for_unlimited_query_soft_limit) @dor: That can certainly help - unpaged queries aren’t good There is an advisor screen in grafana, check how many queries are unpaged @Varun_Nagrare: I’m reading the data using PySpark. So how to find if a query is unpaged? Since I don’t have grafana installed for monitoring it, is there any alternative for it? @dor: It’s time to start read the docs and and run grafana… @Varun_Nagrare: Yes I know . Just wanted to solve this asap. Anyways thanks for the help. I’ll go through the docs. --- ### Page: https://forum.scylladb.com/t/max-ttl-value-how-is-ttl-applied-after-change-of-value/1464 Title: Max TTL value, how is TTL applied after change of value - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/max-ttl-value-how-is-ttl-applied-after-change-of-value/1464 ## Headings Structure: H1: Max TTL value, how is TTL applied after change of value H3: Related topics ## Main Content: H1: Max TTL value, how is TTL applied after change of value H3: Related topics Originally from the User Slack @jean_carlo: Hello Scylla team, is there any limitation concerning to having a big ttl like 1 year ? @Felipe_Cardeneti_Mendes: Hi jean, nope. Just remember that it will take TTL amount of time before your data expires and disk space can be reclaimed. which is the whole point of TTL after all. @jean_carlo: what about the data present in the data when setting a default ttl non zero? I guess data will not change by itself right ? @Felipe_Cardeneti_Mendes: correct TTL will be applied only to new/updated records after you make the change --- ### Page: https://forum.scylladb.com/t/failed-to-connect-scylladb-from-scala-while-scylladb-nodes-are-running-in-kubernetes/1466 Title: Failed to Connect Scylladb From Scala, while Scylladb Nodes are Running in Kubernetes - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I’m trying to run Scylladb Nodes in k8s and access them from Scala (Phantom Library). Kubernetes Firstly I’ve followed commands from official scylladb documentation to create scylladb cluster. $ minikube start --cpus=… Language: en Canonical URL: https://forum.scylladb.com/t/failed-to-connect-scylladb-from-scala-while-scylladb-nodes-are-running-in-kubernetes/1466 ## Headings Structure: H1: Failed to Connect Scylladb From Scala, while Scylladb Nodes are Running in Kubernetes H3: Related topics ## Main Content: H1: Failed to Connect Scylladb From Scala, while Scylladb Nodes are Running in Kubernetes H3: Related topics I’m trying to run Scylladb Nodes in k8s and access them from Scala (Phantom Library). Firstly I’ve followed commands from official scylladb documentation to create scylladb cluster. After all above commands scylladb cluster got ready, I’ve checked using some commands: After setting above scylladb cluster, now I need to access it via my scala code. Now the problem is whenever I run this, it gives an error: I’ve tried different things in: I’ve tried to give pod IP, service IP, scylladb node IP and service name to connect ContactPoints but nothing worked for me. Guide me how can I connect my above scala service with Scylladb Cluster which is running on Kubernetes. I guess you don’t have connectivity to Kubernetes Pods/Services from machine where you run your Scala code. Either set up this connectivity (check minikube documentation), or run your code as a Pod in your Kube cluster. --- ### Page: https://forum.scylladb.com/t/release-scylladb-cloud-5-april-2024/1467 Title: [RELEASE] ScyllaDB Cloud - 5 April 2024 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: We are happy to announce a new update to the ScyllaDB Cloud service. New features and enhancements: You can edit the cluster’s maintenance windows before launching a new cluster via the Cloud UI. When inviting a user… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-cloud-5-april-2024/1467 ## Headings Structure: H1: [RELEASE] ScyllaDB Cloud - 5 April 2024 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Cloud - 5 April 2024 H3: Related topics We are happy to announce a new update to the ScyllaDB Cloud service. New features and enhancements: --- ### Page: https://forum.scylladb.com/t/is-there-any-problem-with-adding-new-parameters-to-sstable-metadata/1469 Title: Is there any problem with adding new parameters to sstable metadata? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I want to track the tombstone count in each sstable to determine when to execute compaction. Therefore, I added a new parameter, ‘tombstone_count,’ to ‘stats_metadata.’ The newly generated sstable statistics files with t… Language: en Canonical URL: https://forum.scylladb.com/t/is-there-any-problem-with-adding-new-parameters-to-sstable-metadata/1469 ## Headings Structure: H1: Is there any problem with adding new parameters to sstable metadata? H3: Related topics ## Main Content: H1: Is there any problem with adding new parameters to sstable metadata? H3: Related topics I want to track the tombstone count in each sstable to determine when to execute compaction. Therefore, I added a new parameter, ‘tombstone_count,’ to ‘stats_metadata.’ The newly generated sstable statistics files with this parameter have no issues. However, when reading the older sstable statistics files, the value of this parameter appears as a random value. I would like to know if this modification could cause any problems. You need to initialize this new member to 0. Having a random value when not present makes it useless. Note that this count will not be correct when there are TTL’d cells. Also, a tombstone can only be purged under certain conditions (usually when it is 10 days old), which this count won’t be able to account for. It is possible to calculate the exact amount of purgeable tombstone, with a script like this. Run this as: scylla script --script-file=/path/to/purgeable.lua /path/to/sstable. Taking a step back, why do you want to start compactions by hand? Why not enable tombstone compaction instead, so ScyllaDB itself will watch sstables for purgeable tombstone ratio and will kick off compaction when this is above the configured threshold. Thank you. I tried setting this parameter a default value of 0, but it still reads a random value when accessed. Could you please advise on how to enable tombstone compaction for ScyllaDB to monitor the droppable tombstone ratio, and how to configure the threshold for this? See Compaction | ScyllaDB Docs, in particular the options: enabled (enables tombstone compaction), tombstone_threshold and tombstone_compaction_interval. I have got it, thank you. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-225-2024-04-08/1470 Title: Last week in scylladb.git master (issue #225; 2024-04-08) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 885cb2af07b…c01b19fcb3f range are covered. There were 66 non-merge commits from 18 authors in that per… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-225-2024-04-08/1470 ## Headings Structure: H1: Last week in scylladb.git master (issue #225; 2024-04-08) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #225; 2024-04-08) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 885cb2af07b…c01b19fcb3f range are covered. There were 66 non-merge commits from 18 authors in that period. Some notable commits: Bloom filters are used to determine which sstables do not contain a partition key, speeding up reads when sstables can be filtered out. Since they are held in memory, and their size depends on the data (small partitions require larger bloom filters), their memory usage can overwhelm a node. The database will not track the aggregate memory consumption by bloom filters, and drop some filters if memory usage exceeds a fraction of the memory allocated to the shard. Repair history is used to determine which tombstones can be garbage collected, as we only garbage collect tombstones that have been written before the last repair. We now load the repair history table in the background, so we don’t slow node start-up. We now remember which auth tables version are in use, to avoid unnecessary version migration on startup. In maintenance mode, we will skip loading tablet metadata if corrupted, to allow an administrator to fix it. The large cell/row detector will no longer log partition keys as they can be sensitive; they are still written to the system.large_rows (and similar) tables, where access can be restricted. Alternator, ScyllaDB’s implementation of the DynamoDB API, moved back from tablets to vnodes as its default replication method. This is because alternator requires LWT in some configurations, which is not yet supported with tablets. Tablets now use a more compact data structure for the mapping between tablets and compaction groups; this helps conserve memory with extremely large tables. An ALTER TABLE statement that increases the replication factor will now rebuild the table, if it uses tablets. Tracking of memory usage during repair is now more accurate. See you in the next issue of last week in scylladb.git master! An ALTER TABLE statement that increases the replication factor will now rebuild I think you mean the ALTER KEYSPACE statement --- ### Page: https://forum.scylladb.com/t/the-tombston-gc-mode-repair-option-doesnt-remove-tombstones/1471 Title: The tombston_gc={'mode': 'repair'} option doesn't remove tombstones - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I’m using ScyllaDB 5.4.x version. I tested the tombstone_gc={‘mode’: ‘repair’} option and found it doesn’t remove tombstones. I had three ScyllaDB nodes and created keyspace and table like this. CREATE KEYSPACE test W… Language: en Canonical URL: https://forum.scylladb.com/t/the-tombston-gc-mode-repair-option-doesnt-remove-tombstones/1471 ## Headings Structure: H1: The tombston_gc={'mode': 'repair'} option doesn't remove tombstones H3: Related topics ## Main Content: H1: The tombston_gc={'mode': 'repair'} option doesn't remove tombstones H3: Related topics I’m using ScyllaDB 5.4.x version. I tested the tombstone_gc={‘mode’: ‘repair’} option and found it doesn’t remove tombstones. I had three ScyllaDB nodes and created keyspace and table like this. And insert and delete datas to gc_test table. Then I executed nodetool flush, so I could see the tombstone from sstable. Then I executed nodetool repair to all nodes and expected the tombstone to be removed, but it was still there. Why hasn’t the tombstone been removed? You need a compaction after the repair has happened, for the tombstone to be dropped. You can force a major compaction here (nodetool compact) to be able to observe this. I forced major compaction, but the tombstone still remained. Tombstones are eligible for being GC-ed if they are older than repair start timestamp - 1h. So they will disappear if you run repair 1h after deletion happened and then you compact the sstables. This 1h grace period is in order to account for the fact that there could be writes covered by tombstones, which are “stuck” in various pipelines during repair, and we don’t want to drop tombstones before those writes reach replicas and are visible to compaction. Otherwise, those writes, which should be dropped, would be resurrected. --- ### Page: https://forum.scylladb.com/t/some-confusion-about-parameter-murmur3-partitioner-ignore-msb-bits/1472 Title: Some confusion about parameter murmur3 partitioner ignore msb bits - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: We konw that murmur3_partitioner_ignore_msb_bits is used to select which bits are used for sharding. Is its value range between 0 and num_token? Its default value is 12. If we deploy a large-scale cluster (180 nodes), sh… Language: en Canonical URL: https://forum.scylladb.com/t/some-confusion-about-parameter-murmur3-partitioner-ignore-msb-bits/1472 ## Headings Structure: H1: Some confusion about parameter murmur3 partitioner ignore msb bits H3: Related topics ## Main Content: H1: Some confusion about parameter murmur3 partitioner ignore msb bits H3: Related topics We konw that murmur3_partitioner_ignore_msb_bits is used to select which bits are used for sharding. Is its value range between 0 and num_token? Its default value is 12. If we deploy a large-scale cluster (180 nodes), should its value be increased? The range of valid values for this option is [0, 64), i.e. from 0 to the number of bits in the 64 bit token value. This option designates how many bits to ignore from the token’s MSB bits, when calculating the shard that owns this token. This value is completely orthogonal to the number of nodes in the cluster. No need to change it for large deployments. --- ### Page: https://forum.scylladb.com/t/authentication-on-a-cluster-secure-credentials-with-manager/1473 Title: Authentication on a cluster, secure credentials with Manager - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/authentication-on-a-cluster-secure-credentials-with-manager/1473 ## Headings Structure: H1: Authentication on a cluster, secure credentials with Manager H3: Related topics ## Main Content: H1: Authentication on a cluster, secure credentials with Manager H3: Related topics Originally from the User Slack @Jatt_Singh: Hi, if we enable authentication on scylla cluster then we also need to pass credentials in scylla-manager so is there any way to secure this credentials? @Felipe_Cardeneti_Mendes: you can provide a user cert and key and require the server to reject incoming requests without a trusted cert. Keep in mind tho, your manager backend should be just for the manager itself, so hardening it won’t really protect (nor expose) anything --- ### Page: https://forum.scylladb.com/t/error-on-removing-dead-nodes-using-nodetool-removenode/1474 Title: Error on removing dead nodes using nodetool removenode - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/error-on-removing-dead-nodes-using-nodetool-removenode/1474 ## Headings Structure: H1: Error on removing dead nodes using nodetool removenode H3: Related topics ## Main Content: H1: Error on removing dead nodes using nodetool removenode H3: Related topics Originally from the User Slack @Mustafa_Shakir: I have two unreachable dead nodes in my scylla cluster. I want to perform a nodetool removenode on them. but when I execute: nodetool removenode --ignore-dead-nodes ab5438e7-7729-4d14-9e9e-d84459525543 eb4e084e-f61b-4ace-9250-c2c52aec1b13 where ab5438e7-7729-4d14-9e9e-d84459525543 and eb4e084e-f61b-4ace-9250-c2c52aec1b13 are dead unreachable nodes. The command get’s stuck and these are the logs I can see: This error looks similar to this issue: https://github.com/scylladb/scylladb/issues/10291 GitHub: service_level_controller - update_from_distributed_data failed to update configuration · Issue #10291 · scylladb/scylladb If it helps out anyone in the future: I was able to recover my cluster by following this guide https://opensource.docs.scylladb.com/stable/architecture/raft.html#raft-manual-recovery-procedure Raft Consensus Algorithm in ScyllaDB | ScyllaDB Docs --- ### Page: https://forum.scylladb.com/t/a-cql-compatible-percolator-client/1475 Title: A CQL-compatible Percolator client - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hey folks! Just wanted to share a Percolator client that I’ve been working on for CQL, specifically for using with Scylla. I’ve not tested it for 100% transactional security, and there are plenty of optimizations still o… Language: en Canonical URL: https://forum.scylladb.com/t/a-cql-compatible-percolator-client/1475 ## Headings Structure: H1: A CQL-compatible Percolator client H3: Percolators/cql at main · danthegoodman1/Percolators H3: Related topics ## Main Content: H1: A CQL-compatible Percolator client H3: Percolators/cql at main · danthegoodman1/Percolators H3: Related topics Hey folks! Just wanted to share a Percolator client that I’ve been working on for CQL, specifically for using with Scylla. I’ve not tested it for 100% transactional security, and there are plenty of optimizations still on the table, so it’s more of a “toy” implementation (so don’t use it in production!) I think composable transactions are particularly neat (you can see an example in the tests) Contribute to danthegoodman1/Percolators development by creating an account on GitHub. Very interesting, thanks for sharing! Very cool! To validate, is this the project you are referring to: GitHub - percolator/percolator: Semi-supervised learning for peptide identification from shotgun proteomics datasets ? No, it was this: https://www.usenix.org/legacy/event/osdi10/tech/full_papers/Peng.pdf --- ### Page: https://forum.scylladb.com/t/aggregate-shard-level-metrics-to-node-level/1476 Title: Aggregate shard level metrics to node level - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Does there support to use metric_family_config to aggregate shard level metrics to node level by API ? And How ? Because I see here Adding Metrics family config by amnonh · Pull Request #2121 · scylladb/seastar · GitHub … Language: en Canonical URL: https://forum.scylladb.com/t/aggregate-shard-level-metrics-to-node-level/1476 ## Headings Structure: H1: Aggregate shard level metrics to node level H3: Related topics ## Main Content: H1: Aggregate shard level metrics to node level H3: Related topics Does there support to use metric_family_config to aggregate shard level metrics to node level by API ? And How ? Because I see here Adding Metrics family config by amnonh · Pull Request #2121 · scylladb/seastar · GitHub updates such a function It’s still in progress, but we’ll get there --- ### Page: https://forum.scylladb.com/t/using-time-to-live-ttl-with-primary-partition-key-columns/1477 Title: Using Time to Live (TTL) with primary (partition) key columns - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/using-time-to-live-ttl-with-primary-partition-key-columns/1477 ## Headings Structure: H1: Using Time to Live (TTL) with primary (partition) key columns H3: Related topics ## Main Content: H1: Using Time to Live (TTL) with primary (partition) key columns H3: Related topics Originally from the User Slack @Daria_Fedorova: Hi! I am reading on TTL in scylla and one thing seems strange to me If TTL is only defined for non-primary columns , it is not possible to define TTL for a table that has only primary columns ? @Felipe_Cardeneti_Mendes: you can have TTL, but you can’t indeed use the ttl() function on a primary column, yes. @Daria_Fedorova: do you by any chance know the reason for this ? @Felipe_Cardeneti_Mendes: I know that historically Cassandra also lacked support for it - since a decade now https://issues.apache.org/jira/browse/CASSANDRA-9312 It’s a bit of an edge case tho. You can’t really update a primary key, so the TTL will always decrease relative to its writetime() . It is true you can’t use the writetime() function on a key column, but if that’s a necessary use case, you can easily add a non-key column which simply ingests a now() timeuuid. Then you can use either the ttl() function on top of it, or simply deduce from it if you have a fixed TTL For a table that has only primary columns. we insert a row with ttl. now how do we check what is the ttl associated with the row. I am talking about data inserted by this dml INSERT INTO test.table (id, wsid) VALUES (‘test-rand-id’, ‘test-ws-id’) USING TTL 10; and table created using this ddl CREATE TABLE IF NOT EXISTS test.table ( id text, wsid text,PRIMARY KEY ((id, wsid))) WITH bloom_filter_fp_chance = 0.005; You can only use the TTL() function on non-primary key columns. So there’s no easy way to do this as of now. You could try using mutation fragments: But this is more for diagnostics and not for reading the TTL value. --- ### Page: https://forum.scylladb.com/t/95-memory-usage-scylladb-even-when-idle/1479 Title: 95% memory usage scyllaDB even when idle - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I have 3 node scylla cluster with 8 core , 32 Gigs RAM. I am trying to generate data for some testing. Upon running my data generation scripts, memory shot up to 95%. Even after I have stopped the script, memory is sti… Language: en Canonical URL: https://forum.scylladb.com/t/95-memory-usage-scylladb-even-when-idle/1479 ## Headings Structure: H1: 95% memory usage scyllaDB even when idle H3: Related topics ## Main Content: H1: 95% memory usage scyllaDB even when idle H3: Related topics I have 3 node scylla cluster with 8 core , 32 Gigs RAM. I am trying to generate data for some testing. Upon running my data generation scripts, memory shot up to 95%. Even after I have stopped the script, memory is still at 95% even when idle. I am assuming that somewhere data is cached by scylla. How to debug this further and how to bring down the memory? If it is indeed being cached, what is the impact of such high memory usage? TL;DR: it’s a feature, not a bug ScyllaDB is designed to use as much of each node’s resources, such as disk, memory, and network, as possible. For RAM, it will allocate most of it, splitting the space internally between cache, memetable, etc. Cache data will be served Reads faster and reduce storage access. @tzach Thanks for the answer. --- ### Page: https://forum.scylladb.com/t/spark-connector-getting-specific-data-without-full-table-scans/1481 Title: Spark Connector, getting specific data without full table scans - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/spark-connector-getting-specific-data-without-full-table-scans/1481 ## Headings Structure: H1: Spark Connector, getting specific data without full table scans H3: Related topics ## Main Content: H1: Spark Connector, getting specific data without full table scans H3: Related topics Originally from the User Slack @Varun_Nagrare: Hi all, I’m using Spark Cassandra Connector 2.5.2 with Spark 2.4.7. I wanted to know if there is a way to avoid full table scan and get only the required partitions? I have a dataframe with my partition keys and I want to get the data of only those partition keys. I want to know if I can get only that specific data without Spark doing a full table scan as it is also overloading Scylla ultimately giving timeout errors. @dor: There should be a way but I really don’t know spark, we have a 3 blog series about it, you may need to tweak some code @Varun_Nagrare: @dor Can you please provide link to that blog series? I’m trying the tweaks provided on the connector configuration page, but it’s still scanning the whole table and overloading Scylla @dor: https://www.scylladb.com/2018/07/31/spark-scylla/ @Varun_Nagrare: @Botond_Dénes Can you please help? Most of my errors got solved by your guidance. @Botond_Dénes: Unfortunately I also don’t know anything about Spark, @Lubos is our Spark guy, he may be able to help. @Varun_Nagrare: @Lubos Requesting your help @Botond_Dénes Nevermind, I found the solution. I had to filter my partition keys in a for loop and then union the list of dataframes so it only reads the specific partitions into a single dataframe and does not scan the whole table and also Scylla doesn’t overload. This method made my 7+ hours application runtime come down to just 10-15 mins only. --- ### Page: https://forum.scylladb.com/t/using-nodetool-refresh-load-and-stream-to-backup-and-restore-an-entire-cluster/1482 Title: Using nodetool refresh, load and stream to backup and restore an entire cluster - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/using-nodetool-refresh-load-and-stream-to-backup-and-restore-an-entire-cluster/1482 ## Headings Structure: H1: Using nodetool refresh, load and stream to backup and restore an entire cluster H3: Related topics ## Main Content: H1: Using nodetool refresh, load and stream to backup and restore an entire cluster H3: Related topics Originally from the User Slack @Arthur_Li: Hi all, is it possible that I use nodetool snapshot to backup the entire cluster (one node by one node) and restore them on a different cluster with same number of nodes? @Felipe_Cardeneti_Mendes: yes @Arthur_Li: may I ask how can I decide which destination node does the snapshot goes to? is there some calculation that I need to do? @Felipe_Cardeneti_Mendes: the easiest way if you are able to recreate the target cluster is to do a 1:1 source:destination copy by specifying the num_tokens on the target node which matches the source one if you can not do that, then the easiest way is to use load and stream https://opensource.docs.scylladb.com/stable/operating-scylla/nodetool-commands/refresh.html#load-and-stream Nodetool refresh | ScyllaDB Docs @Arthur_Li: will try nodetool refresh by saying num_tokens did you mean the Tokens in the result of nodetool status ? which is default to 256 I guess @Felipe_Cardeneti_Mendes: sorry, no … I meant initial_tokens Run nodetool ring, this will show which tokens are owned by which ones. There will be 256 (the value for num_tokens ) entries for each node transform that into a list, pass as initial_tokens during the bootstrap of the target cluster and done, you have a 1:1 mapping @Arthur_Li: TIL thank you!! @Felipe_Cardeneti_Mendes: > –initial-token arg Used in the single-node-per-token architecture, where a node owns exactly one contiguous range in the ring space. Setting this property overrides num_tokens. > This parameter can be used with num_tokens (vnodes ) in special cases such as Restoring from a snapshot. --- ### Page: https://forum.scylladb.com/t/does-a-cql-query-become-token-aware-depending-on-its-size/1483 Title: Does a CQL query become token-aware depending on its size? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Please provide some background information on how it works. Language: en Canonical URL: https://forum.scylladb.com/t/does-a-cql-query-become-token-aware-depending-on-its-size/1483 ## Headings Structure: H1: Does a CQL query become token-aware depending on its size? H3: Related topics ## Main Content: H1: Does a CQL query become token-aware depending on its size? H3: Related topics Please provide some background information on how it works. In ScyllaDB (and other distributed databases like Cassandra), each node contains only a part of the data. Ideally, a query would reach the node that holds the data (one of the replicas). If this doesn’t happen, the coordinator node (the node that is responsible for communicating with the client for that specific query), will need to send the query internally to a replica, resulting in an extra hop and, thus, higher latency and more resource usage. The size of the query does not cause it to become token-aware (nor does it have any effect on token awareness). Non token aware Queries are queries that did not reach a replica-node, most likely due to not using prepared statements, or not using token aware drivers. Load balancing can also have an effect on token awareness. To test this, you can analyze internal flows within a cluster using Tracing. Also, using ScyllaDB Monitoring you can see if your application is using queries that are not token-aware. Another related concept is Shard Awareness, and the Shard Aware Port which provides an additional performance improvement. --- ### Page: https://forum.scylladb.com/t/how-to-rollback-from-scylldb-to-cassandra/1484 Title: How to rollback from scylldb to cassandra? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Cassandra to Scylla migration is seamless using sstableloader as documented here - Apache Cassandra to Scylla Migration Process | ScyllaDB Docs. But is there any process to roll back from scyllDB to Cassandra? It is me… Language: en Canonical URL: https://forum.scylladb.com/t/how-to-rollback-from-scylldb-to-cassandra/1484 ## Headings Structure: H1: How to rollback from scylldb to cassandra? H3: Related topics ## Main Content: H1: How to rollback from scylldb to cassandra? H3: Related topics Cassandra to Scylla migration is seamless using sstableloader as documented here - Apache Cassandra to Scylla Migration Process | ScyllaDB Docs. But is there any process to roll back from scyllDB to Cassandra? It is mentioned here that Scylla and Cassandra sstables are compatible but when I try to just replace the sstables file, it fails. Also, tried the same scylla to Cassandra migration by just replacing sstables and that too failed. Any leads on how to successfully rollback to Cassandra in case ? It is mentioned here that Scylla and Cassandra sstables are compatible but when I try to just replace the sstables file, it fails. Failed how? What error did you get? What version of ScyllaDB and Apache Cassandra are you using? Also, tried the same scylla to Cassandra migration by just replacing sstables and that too failed. Same questions as above. Another alternative is using SSTableloadr. Since it use CQL, it should work with Apache Cassanra. Cassandra = 4.0 Scylla = 5.1 Cassandra to Scylla migration- Got some corrupted stable format exceptions. Scylla to Cassandra rollback- The actual question was is it actually possible to do so ? --- ### Page: https://forum.scylladb.com/t/timewindow-compaction-strategy-ttl-tombstones-and-window-size/1485 Title: TimeWindow Compaction Strategy, TTL, Tombstones, and window size - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/timewindow-compaction-strategy-ttl-tombstones-and-window-size/1485 ## Headings Structure: H1: TimeWindow Compaction Strategy, TTL, Tombstones, and window size H3: Related topics ## Main Content: H1: TimeWindow Compaction Strategy, TTL, Tombstones, and window size H3: Related topics Originally from the User Slack @Daria_Fedorova: Hi, I have a question about Time Window CompactionStrategy I need to add TTL but I can’t just use 1 TTL value for all records (2 values maybe and a small number of records without TTL ) And I understand that it is not ideal if I leave them all in one table, but shouldn’t tombstone compaction solve the problem of memory release ? This will add strain on disk probably but are there any other problems with multiple TTL in one table ? My only other alternative is joining the the time-series from 3 table on client side (thankfully it is not so difficult) @Felipe_Cardeneti_Mendes: you need to remember that a bucket will only be deleted after all records expire. If you are using 2 different TTLs, then you probably should select a window size that matches your largest TTL. For records with no TTL you shouldn’t be using TWCS, use another strategy instead @Daria_Fedorova: I did not find anything on how to choose window size for TWCS , the one we have I think is a compromise for fast reads for a couple of days , it also maybe was chosen when we had different disks - this memory is lost , maybe 1 week is too small window cassandra doc says something about it but they don’t explain the reasons The larger ttl will be 3 or 5 years , sstables will grow very big , is it not a problem ? if I follow cassandra doc then window size for TTL = 3 years is around 36 days Can you explain what factors to consider when choosing a window ? i think that reads are faster if they work with 1 sstable and for this window == 1 week seems reasonable for us What is the motivation to make it bigger ? @Felipe_Cardeneti_Mendes: You need to remember that SSTables will only be evicted after all data within them expires (hence why you shouldn’t use TWCS if you have records with no TTL). If you have a window of 1 week, but 3 years of TTL, then that will result in >150 windows, which may create heavy memory pressure and slow down your reads/repairs over time. You want to find a good balance between your reads and fewer windows to avoid having too many open files to satisfy a read. ie: consider a 1 hour window and 30 day TTL. Reading all records will require scanning through 720 windows. @Daria_Fedorova: Ok I see ) that is not a big concern for us mostly reads requests a couple of days or month and we split very long reads on client (api) anyway @Felipe_Cardeneti_Mendes: a last thing is that you shouldn’t generally worry much about large SSTables, TWCS compacts everything altogether on a per-window basis to optimize for reads. Plus SSTables are further split by shard. @Daria_Fedorova: I ask a college and he said that they did not want a bigger window because last window uses Size Tiered Compaction and with bigger window when we read range of last records for user it will iterate over many sstables ? and another concern was longer compact ( it affects latency a little now , idk if it will get worse ) @Felipe_Cardeneti_Mendes: all windows use STCS @Felipe_Cardeneti_Mendes: by the way, we are going to have a compaction talk in our upcoming ScyllaDB Summit scylladb.com/summit @Daria_Fedorova: yes, I am planning to tune in --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-3/1486 Title: [RELEASE] ScyllaDB Enterprise 2024.1.3 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2024.1.3, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2024.1 LTS Release. 2024.1.3 patch release includes multiple minor bug fixes… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-3/1486 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2024.1.3 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2024.1.3 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2024.1.3, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2024.1 LTS Release. 2024.1.3 patch release includes multiple minor bug fixes. The following issues are fixed in this release (with an open-source reference, if available): --- ### Page: https://forum.scylladb.com/t/memory-resources-timed-out-dumping-permit-diagnostics/1489 Title: Memory resources: timed out, dumping permit diagnostics - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: HI, I am seeing a lot of these timeouts in my cluster. we have allocated scylla 350gb memory of 376gb. Scylla version 5.2.16. [shard 28] reader_concurrency_semaphore - (rate limiting dropped 3 similar messages) Semaphor… Language: en Canonical URL: https://forum.scylladb.com/t/memory-resources-timed-out-dumping-permit-diagnostics/1489 ## Headings Structure: H1: Memory resources: timed out, dumping permit diagnostics H3: Related topics ## Main Content: H1: Memory resources: timed out, dumping permit diagnostics H3: Related topics HI, I am seeing a lot of these timeouts in my cluster. we have allocated scylla 350gb memory of 376gb. Scylla version 5.2.16. Looks like something is scanning system.local a lot of times. What driver are you using? Are you connecting a lot of new clients? This table is read by drivers upon connecting to the scylla node and we have patched them to use a single partition query, instead of a full scan, because the full scan is much more expensive. If you are using an out-of-date driver, updating it might solve this issue. We are using com.scylladb:java-driver-core-shaded:4.17.0.0 driver. We connect to 250 clients but they connect only once. This is the toppartitions output from one of the nodes. could you check in monitoring, in “Scylla CQL” dashboard we have " Client CQL new connections by Instance" and “Client CQL connections by Instance”. Maybe your driver drops and reconnects for some reason? --- ### Page: https://forum.scylladb.com/t/upgrading-the-scylladb-kubernetes-operator-placement-affinities-issue-with-cleanup-jobs/1490 Title: Upgrading the ScyllaDB Kubernetes Operator, placement affinities issue with cleanup jobs - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/upgrading-the-scylladb-kubernetes-operator-placement-affinities-issue-with-cleanup-jobs/1490 ## Headings Structure: H1: Upgrading the ScyllaDB Kubernetes Operator, placement affinities issue with cleanup jobs H3: Related topics ## Main Content: H1: Upgrading the ScyllaDB Kubernetes Operator, placement affinities issue with cleanup jobs H3: Related topics Originally from the User Slack @Aleksander_Karoński: Hi! I’m trying to upgrade my scylla kubernetes setup and in the process I encountered a problem I deployed scyllaCluster with following placement setup (in a comment below). The setup consists of 3 separate nodes and each node contains one scylla instance. I am separating them by nodeAffinities, podAntiAffinities and tolerations. It worked on older operator version, but now there are some cleanup jobs created with exact same affinities as defined in scyllaCluster(?). Those jobs cannot be scheduled because podAntiAffinity rule prevents it. Do you have any recommendations how to fix such issue? @Maciej_Zimnoch: Instead using requiredDuringSchedulingIgnoredDuringExecution in podAntiAffinityyou can use preferredDuringSchedulingIgnoredDuringExecution which should solve it. @Aleksander_Karoński: Yes, but is there any guarantee that job won’t be scheduled for instance-0 before instance-2 is scheduled? This could lead to 2 instances being placed on the same node (because job have taken a free spot first) @Maciej_Zimnoch: cleanup jobs are created after scaling is finished, so cleanup jobs cannot prevent Scylla Pods from being scheduled @Aleksander_Karoński: Thanks, that’s what I wanted to hear, because I cannot find any docs about it (if such exist, please link here ) @Maciej_Zimnoch: You can read entire proposal for running cleanup here: https://github.com/scylladb/scylla-operator/blob/master/enhancements/proposals/1207-cleanup-after-scaling/README.md --- ### Page: https://forum.scylladb.com/t/toppartition-from-system-auth-keyspace/1492 Title: Toppartition from System_auth keyspace - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: What could be the reason for system_auth:roles to become top hot partition? READS Sampler: Cardinality: ~256 (256 capacity) Top 10 partitions: Partition Count … Language: en Canonical URL: https://forum.scylladb.com/t/toppartition-from-system-auth-keyspace/1492 ## Headings Structure: H1: Toppartition from System_auth keyspace H3: Related topics ## Main Content: H1: Toppartition from System_auth keyspace H3: Related topics What could be the reason for system_auth:roles to become top hot partition? READS Sampler: Cardinality: ~256 (256 capacity) Top 10 partitions: Partition Count +/- (system_auth:roles) user1 19119 85 Is it possible that you create lots of connections instead of reusing them? When we reach to this count ~19119 with system_auth:roles table, cql connection timeouts. How can we identify the maximum no of connections a cluster can support without timeouts? You should be able to observe it in monitoring under “Scylla CQL” dashboard we have " Client CQL new connections by Instance" and “Client CQL connections by Instance”. It’s more about new connection rate than current connection pool, although both have their limits depending on your hardware. Client CQL new connections by Instance I dont see this panel in “Scylla SQL” dashboard. What metrics it checks? Perhaps you don’t have the latest version? It checks scylla_transport_current_connections metric (and it’s rate). Thank you, Added the panel to dashboad. Rate was less than 0 and at some point it goes to 3.7. How to interpret this? I’d say it’s rather low value, by default it would mean connections per second. And how many connections do you have? (remember to make sure it’s aggregated over all nodes and shards) --- ### Page: https://forum.scylladb.com/t/batch-inserts-of-an-entity-using-a-dao/1495 Title: Batch inserts of an @Entity using a @Dao - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: According to the docs, it is possible to declare the return type of an @ Insert to a be a BoundStatement which is “intended for cases where you intend to execute this statement later or in a batch”. This would require, … Language: en Canonical URL: https://forum.scylladb.com/t/batch-inserts-of-an-entity-using-a-dao/1495 ## Headings Structure: H1: Batch inserts of an @Entity using a @Dao H3: Related topics ## Main Content: H1: Batch inserts of an @Entity using a @Dao H3: Related topics According to the docs, it is possible to declare the return type of an @ Insert to a be a BoundStatement which is “intended for cases where you intend to execute this statement later or in a batch”. This would require, however, the consumer of this to get the session object and manually execute the BoundStatements produced which breaks the encapsulation of the @ Dao. Is it possible to create a @ Dao method to perform batch updates? I’m trying to do so with a @ QueryProducer. As part of this, I’m using the generated EntityHelper, however I don’t see a way to create an Insert statement and then bind it to an entity as part of the QueryProducer. My code looks something like this: But this is wrong, it is not actually binding it correctly. Largely, I’m looking for a way to use the EntityHelper to bind the entity to the Insert statement, rather than having to do manually. Is this possible? Am I overlooking something? Hi, I’ve looked at some examples and it seems this is one way to do that (bind the entity to Insert statement): --- ### Page: https://forum.scylladb.com/t/shard-awareness-k8s-environment-and-using-nat/1496 Title: Shard awareness, K8s environment and using NAT - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/shard-awareness-k8s-environment-and-using-nat/1496 ## Headings Structure: H1: Shard awareness, K8s environment and using NAT H3: Related topics ## Main Content: H1: Shard awareness, K8s environment and using NAT H3: Related topics Originally from the User Slack @Nico_Piderman: Can anyone provide any guidance around using advanced shard awareness in a K8s environment, when the clients are in K8s and Scylla is not? The c++ driver complains that it is behind a NAT and disables the feature. @Maciej_Zimnoch: I’m not 100% sure it’s true for C++ driver, but it is for others (Go, Java). I assume it’s the same in C++. Advanced shard awareness is about building connection pool in smarter way. NAT prevents it, because info about to which shard driver attempts to connect to is stored in source port. If NAT interferes, this information is lost, and random shard is assigned. Driver detects that, and disables “smart” connections and falls back to buliding pool using random shard assignment until all shards are covered. Once pool is built, shard awareness still works, as driver is able to pick shard connections from pool. So it’s not as bad as it sounds as it’s only about building connection pool, not about routing requests. The solution is to disable NAT, or live with it. @Nico_Piderman: Ahhh, good point! Thank you for pointing that out --- ### Page: https://forum.scylladb.com/t/gocql-query-paging-issue/1497 Title: Gocql query paging issue - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/gocql-query-paging-issue/1497 ## Headings Structure: H1: Gocql query paging issue H3: Related topics ## Main Content: H1: Gocql query paging issue H3: Related topics Originally from the User Slack @Carlos_V: Is there a reason why a query, using gocql, will return paginated data smaller than PageSize() size, and iter.PageState() for next page will be either empty or using the result as PageState(token) won’t return anything? Querying directly from database using cqlsh, I have a query with 207 rows, but it will only return 205, if I put limit 206 it will return 205 and 30 minutes lates it will return the 206 one. @Marko_Ćorić: All under the same cluster key? And after 30 minutes you see 206th record? @Botond_Dénes: you probably have tons of tombstones @Carlos_V: Yes, after the 30 minutes waiting the --MORE-- the 206 comes back. This is a materializes view. I did run a nodetool flush and repair, no change. @Botond_Dénes: materialized views do have tombstones repair won’t help, try nodetool compact @Carlos_V: Let me try still running compact, in the meantime regarding cluster key question, it a single node scylla instance and the materialized view is this: @Botond_Dénes: this is from the MV definition? @Carlos_V: Yes After 3 hours, node compact has finished, but the same issue happens, seems to be with the same 2 rows. I’m supposed to have 177 rows now. Using limit 175 is fine, with limit 176 then it wait for a long time until it comes back. Funnily enough if I run the same query with COUNT(1) I get 177 qty. @Botond_Dénes: looks like your tombstones could not be purged @Carlos_V: Perhaps I don’t have enough space for that? Use% 77% @Botond_Dénes: tombstones won’t be purged before gc_grace_seconds, this is 10 days by default @Carlos_V: Is there a way to force? the 10 days have passed and the data is surely returning much faster now, only a few seconds now. Count is now 187. If I do LIMIT 187 100 comes then --MORE-- and then another 85 come, and then another --MORE-- and then the 2 come after less than a second. But the go query still won’t bring these two rows. @Botond_Dénes: Maybe there is a problem in how the go driver handles empty pages @Carlos_V: Any change this can be looked at? @Marko_Ćorić: can you share that part of code? @Carlos_V: This is the way I’m fetch data @Marko_Ćorić: Sorry for delay, I had to find our implementation since we rewrite whole production to Rust. This is code that worked for us without single problem for like 2 years --- ### Page: https://forum.scylladb.com/t/last-2-weeks-in-scylla-cluster-tests-git-master-issue-44-2024-04-19/1498 Title: Last 2 weeks in scylla-cluster-tests.git master (issue #44; 2024-04-19) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last 2 weeks. Commits in the b48d6f3f…6c701581 range are covered. There were 12 non-merge commits from 4 Software E… Language: en Canonical URL: https://forum.scylladb.com/t/last-2-weeks-in-scylla-cluster-tests-git-master-issue-44-2024-04-19/1498 ## Headings Structure: H1: Last 2 weeks in scylla-cluster-tests.git master (issue #44; 2024-04-19) H3: Related topics ## Main Content: H1: Last 2 weeks in scylla-cluster-tests.git master (issue #44; 2024-04-19) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last 2 weeks. Commits in the b48d6f3f…6c701581 range are covered. There were 12 non-merge commits from 4 Software Engineers in Test and 1 Software Engineer in that period. Some notable commits: We’ve increased the speed of Scylla cluster setups by parallelizing ‘node_setup’ methods, with only the final bootstrap stage executed serially. This significantly boosts performance in scenarios involving custom disk setups in hybrid configurations. Additionally, datacenter and rack information is now captured in logs and events, aiding analysis in multi DC/rack setups. We resolved several Nemesis failures due to a ‘permission denied’ error by setting the correct ‘user’ attribute in supervisord configs. This marks the completion of our initial validation of Nemesis on the Docker backend, detailed here. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/understanding-timeout-issue-in-logs-count-i-o-and-memory-are-not-saturated/1502 Title: Understanding timeout issue in logs, count (I/O) and memory are not saturated - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/understanding-timeout-issue-in-logs-count-i-o-and-memory-are-not-saturated/1502 ## Headings Structure: H1: Understanding timeout issue in logs, count (I/O) and memory are not saturated H3: Related topics ## Main Content: H1: Understanding timeout issue in logs, count (I/O) and memory are not saturated H3: Related topics Originally from the User Slack Hello All, I am getting similar logs as above frequently in the scylla nodes on production. Can someone help me understand what exactly timed out as per above log. It seems that neither count nor memory threshold reached, still the query timed out. What are permits and how can I increase it so that time outs do not happen. Thanks. @Botond_Dénes: If neither count (I/O) or memory limits is saturated, it will be the third (implicit) one: CPU. Look at your monitoring, you will likely see the shard reporting this hoverig at round 100% CPU usage. @Shantanu_Sachdev: yes, a lot of shards Load frequently peaks to 100%. But the node’s overall CPU never goes above 50%. How can I optimally utilize it ? @Botond_Dénes: You probably have partitions that are more popular then other partitions. The shards hosting these popular partiitons are maxxed out. A scale-out or scale-up can help, if it the problem is multiple reasonably hot partitions – the scale-{out,up} can move these to separate nodes/shards, helping spread the load. If there is just a handful of really hot partitions, I’m afraid you will have to look into changing your data model so that you don’t have such hot partitions. @Shantanu_Sachdev: okay, scale out would become costly for us since we are using i4i.8xlarge nodes. Will inspect more on the hot partitions --- ### Page: https://forum.scylladb.com/t/release-scylladb-monitoring-stack-4-7-1/1503 Title: [RELEASE] ScyllaDB monitoring Stack 4.7.1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce a patch release of ScyllaDB Monitoring Stack 4.7.1 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB Enterprise and ScyllaDB Open Source, based on Prometh… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-monitoring-stack-4-7-1/1503 ## Headings Structure: H1: [RELEASE] ScyllaDB monitoring Stack 4.7.1 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB monitoring Stack 4.7.1 H3: Related topics The ScyllaDB team is pleased to announce a patch release of ScyllaDB Monitoring Stack 4.7.1 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB Enterprise and ScyllaDB Open Source, based on Prometheus and Grafana. ScyllaDB Monitoring Stack 4.7.1 supports: New In ScyllaDB Monitoring 4.7.1 --- ### Page: https://forum.scylladb.com/t/a-number-of-errors-occurring-in-the-java-driver/1505 Title: A number of errors occurring in the java driver - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello, I just want to say thank you very much for your work, scylla is a really cool tool, but here I have a few questions, if you know the answer to them, please let me know. Let me tell you right away that I work on W… Language: en Canonical URL: https://forum.scylladb.com/t/a-number-of-errors-occurring-in-the-java-driver/1505 ## Headings Structure: H1: A number of errors occurring in the java driver H3: Related topics ## Main Content: H1: A number of errors occurring in the java driver H3: Related topics Hello, I just want to say thank you very much for your work, scylla is a really cool tool, but here I have a few questions, if you know the answer to them, please let me know. Let me tell you right away that I work on Windows 11, but on our project we have docker, which runs the container with ScyllaDB Driver version 3.11.5.0 When sending a request to the database, a list of WARNs (about 10 pieces) pops up and reports Advanced shard awareness: requested connection to shard 1, but connected to 0. Is there a NAT between client and server?, I tried for a long time, but it didn’t work I couldn’t solve this problem, I hope you can help me, the request is being executed, but seeing it constantly in the log is quite confusing See my answer here: Questions about scylla java driver 4.x ShardAware feature and NAT issue - #2 by Lorak for a description of shard awareness and advanced shard awareness. Generally this error shouldn’t very be harmful. It affects only session creation, but after it is fully created (which may take a bit more time because of this problem) it will work exactly the same. As to why it may occur - probably Docker on Windows uses some kind of NAT between Scylla and the driver. You may try to run the driver inside Docker too, but I can’t promise it will help. Maybe some Docker support channel would be a better fit for this question? Well, probably the strangest thing is request hashing, I hash the request so that the database is not loaded with prepared requests, but the error that the request needs to be hashed also appears, I have rewritten the code, maybe this is a problem on my part, please let me know Please I’m not sure I get what you are aiming to do. This code fragment: will always call session.prepare which will send a PREPARE request. Which is a bad idea because now you have 2 round trips for each statement you perform - first to prepare unnecessarily and then to execute. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-226-2024-04-21/1506 Title: Last week in scylladb.git master (issue #226; 2024-04-21) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the c01b19fcb3…87b08c957f range are covered. There were 65 non-merge commits from 18 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-226-2024-04-21/1506 ## Headings Structure: H1: Last week in scylladb.git master (issue #226; 2024-04-21) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #226; 2024-04-21) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the c01b19fcb3…87b08c957f range are covered. There were 65 non-merge commits from 18 authors in that period. Some notable commits: A rare case where gossip fails to observe state changes was fixed by adding the appropriate locks. A cluster that is restarted from a cold shutdown is now more robust in its handling of nodes that did not start. Repair now has a more conservative estimate of the number of partitions it repairs, used to generate bloom filters. This solves problems with bloom filter memory usage after repair. The compaction manager will check if a table (or tablet) is eligible for tombstone garbage collection compaction periodically. Since this check is expensive, it is now done only for idle tables/tablets. Active tables/tablets will perform the check as part of regular compaction. The long effort to make the source base compatible with version 10 of the fmt library is concluded. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/how-to-close-read-repair/1507 Title: How to close `read_repair`? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I saw through Alternator query that a certain attribute that should have existed does not exist. But when I query through CQL, it shows that the attribute exists. I suspect it triggered a read repair. Now I would like t… Language: en Canonical URL: https://forum.scylladb.com/t/how-to-close-read-repair/1507 ## Headings Structure: H1: How to close `read_repair`? H3: Related topics ## Main Content: H1: How to close `read_repair`? H3: Related topics I saw through Alternator query that a certain attribute that should have existed does not exist. But when I query through CQL, it shows that the attribute exists. I suspect it triggered a read repair. Now I would like to ask if there is a way to turn off read repair for easier troubleshooting. I see that parameters read_repair_chance and dclocal_read_repair_chance have been hard coded to 0, does this mean that read repair has been turned off? TLDR: If you want to do trouble shooting, you can do reads with CL=ONE, which will never trigger a read-repair, because it only reads from a single replica thus it will not have a chance to notice any differences between the data different replicas have. Read repair cannot be turned off. ScyllaDB will start read-repair every time it notices that replicas that participate in the read have differences in what data they have. Note that whether these differences are noticed by the coordinator or not depends on which replicas are selected for the read. For example, if there are three replicas: A, B and C, with A and B having the same data for some partition and C having different data: a regular CL=QUORUM read which reads from A and B would not trigger read repair, while a read from A and C would. Note that you can force a read repair by reading with CL=ALL, in which case all replicas are read from. I see that parameters read_repair_chance and dclocal_read_repair_chance have been hard coded to 0, does this mean that read repair has been turned off? These two parameter control probabilistic read repair. With read_repair_chance set to 0.5, there is a 50% chance on every read, that the coordinator will also read from all replicas (regardless of which CL the user set) and trigger a read repair if differences are detected. Note that the read will complete as soon as CL replicas responded and then the remaining part will be moved to the background. The dclocal_read_repair_chance works the same way, but only includes extra nodes from the local DC. Setting these to 0 will only disable probabilistic read-repair, but not the regular read repair. There is no way to disable that. Note that these two parameters are deprecated and we are about to remove them. --- ### Page: https://forum.scylladb.com/t/number-of-materialized-views-and-impact-on-performance/1508 Title: Number of Materialized Views and Impact on Performance - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/number-of-materialized-views-and-impact-on-performance/1508 ## Headings Structure: H1: Number of Materialized Views and Impact on Performance H3: Related topics ## Main Content: H1: Number of Materialized Views and Impact on Performance H3: Related topics Originally from the User Slack @Benaceur_Ayoub: Hi, if I have 7 materialized views on one table to support some queries would that hurt performance too much? Also what’s the maximum mv tables you can have per table ? @Felipe_Cardeneti_Mendes: There’s no pre-defined maximum as its down to your own judgement. Considering every view has a distinct key not part of the base table, then every write operation will require * VIEW_COUNT view updates. If you run 1 ops/sec, that’s 8 updates (base table + 7 views). If you write 100K/s, that’s 800K updates. The former is fine. The latter is likely not going to scale well. --- ### Page: https://forum.scylladb.com/t/using-cassandra-drivers-with-scylladb-performance-and-code-changes/1767 Title: Using Cassandra drivers with ScyllaDB, performance and code changes - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/using-cassandra-drivers-with-scylladb-performance-and-code-changes/1767 ## Headings Structure: H1: Using Cassandra drivers with ScyllaDB, performance and code changes H3: Related topics ## Main Content: H1: Using Cassandra drivers with ScyllaDB, performance and code changes H3: Related topics Originally from the User Slack @Dharun: Hi, I have a code written using a datastax driver, can I run the code without changing the driver ? @Felipe_Cardeneti_Mendes: yes, although you will miss our performance improvements, such as shard-awareness and others. You’ll likely also need to tune the driver settings to fully utilize your scylladb cluster capacity Also using scylla drivers should be a drop in replacement, it should be quite straightforward to give a try and see for yourself --- ### Page: https://forum.scylladb.com/t/need-help-in-explaination-on-rlatencyp95-metrics-exposed-by-scylla/1778 Title: Need help in explaination on rlatencyp95 metrics exposed by scylla - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: We are facing latency issues and while debugging through grafana rlatencyp95, it is exposing on instance,shard level and cluster and DC level. When the issue occurs, rlatencyp95 for instance level still seems in microse… Language: en Canonical URL: https://forum.scylladb.com/t/need-help-in-explaination-on-rlatencyp95-metrics-exposed-by-scylla/1778 ## Headings Structure: H1: Need help in explaination on rlatencyp95 metrics exposed by scylla H3: Nodetool toppartitions | ScyllaDB Docs H3: Hot Partition, Large Partition, Single Node Check H3: Related topics ## Main Content: H1: Need help in explaination on rlatencyp95 metrics exposed by scylla H3: Nodetool toppartitions | ScyllaDB Docs H3: Hot Partition, Large Partition, Single Node Check H3: Related topics We are facing latency issues and while debugging through grafana rlatencyp95, it is exposing on instance,shard level and cluster and DC level. When the issue occurs, rlatencyp95 for instance level still seems in microseconds but on cluster level it seems spiking in seconds. Whats the diff in metrics on instance and cluster level ?Below is the query which i used to debug the latency issue. I wanted a breakdown at instance level to identify if there is any particular instance causing this latency issue. avg(rlatencyp99{by=“cluster”, instance, cluster=~“scyl-test”, scheduling_group_name!=“streaming”} > 0) by (cluster, instance) But after applying this query, I observed that the latency shown at instance level is still in microseconds while the cluster level metrics still shows 30 sec as latency. Can we get any official documentation for understanding the scope. ? above snapshot graph has 2 scylla nodes Hi! Hot partition or other imbalance is the first thing to check, please experiment with: ScyllaDB is an Apache Cassandra-compatible NoSQL data store that can handle 1 million transactions per second on a single server. 8 min to completeIn the case of replica imbalance, the common issues are hot partitions, large partitions, and issues related to the specific node. The node-specific issues include CPU, I/O, and OS related issues. The lesson explains these issues and... and in general better to debug using scylla monitoring because we have a lot of specialized graphs for such things (e.g. it’s better to view the whole cluster on a per shard (vcpu) level). Marcin, the problem over here is that this is happening to all nodes in same project (GCP ) irrespective of cluster. So we feel its not workload issue. So any network related issue can cause this ? --- ### Page: https://forum.scylladb.com/t/running-scylladb-and-spark-on-the-same-node-like-cassandra-spark/1811 Title: Running ScyllaDB and Spark on the same node, (like Cassandra + Spark) - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/running-scylladb-and-spark-on-the-same-node-like-cassandra-spark/1811 ## Headings Structure: H1: Running ScyllaDB and Spark on the same node, (like Cassandra + Spark) H3: Related topics ## Main Content: H1: Running ScyllaDB and Spark on the same node, (like Cassandra + Spark) H3: Related topics Originally from the User Slack @Dharun: Hi, is it okay to run scylladb with spark on the same node like how datastax has (Cassandra + spark) ? If yes, will it impact any performance ? @Felipe_Cardeneti_Mendes @Felipe_Cardeneti_Mendes: tldr; it is possible, as long as you tune both services properly. The easiest thing to do is to simply run them separated. Long answer: Similarly as it happens with Cassandra, the main reason NOT to run both colocated into the same machine is to avoid Spark from consuming DB resources. With ScyllaDB, you want to ensure that your spark workers are pinned to CPUs not in use by the database to avoid performance problems. @Dharun: is there any documentation for tuning, we have dse setup, just thinking for similar setup with open source spark & scylladb @Felipe_Cardeneti_Mendes: https://opensource.docs.scylladb.com/stable/getting-started/scylla-in-a-shared-environment.html ScyllaDB in a Shared Environment | ScyllaDB Docs this is ScyllaDB specific tuning. Check Spark documentation on how to pin them to specific cores @Dharun: @Felipe_Cardeneti_Mendes I see the following is mentioned but under /etc/sysconfig/scylla-server where do i set the args ? SCYLLA_ARGS ? @Felipe_Cardeneti_Mendes: or you can run scylla_memory_setup --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-4-6/1812 Title: [RELEASE] ScyllaDB 5.4.6 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.4.6, a bugfix release of the ScyllaDB 5.4 stable branch. ScyllaDB Open Source 5.4.6, like all past and future 5.x.y releases, is backward compatible and supports rolling… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-4-6/1812 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.4.6 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.4.6 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.4.6, a bugfix release of the ScyllaDB 5.4 stable branch. ScyllaDB Open Source 5.4.6, like all past and future 5.x.y releases, is backward compatible and supports rolling upgrades. Users are encouraged to upgrade to 5.4.6. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/change-cl-at-cqlsh-login-time/1819 Title: Change CL at CQLSH login time - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: How can I change CL when login to CQLSH, getting following error message="Operation timed out for system_auth.roles - received only xx responses from xx CL=QUORUM." . I want to change CL=ONE at login time. Language: en Canonical URL: https://forum.scylladb.com/t/change-cl-at-cqlsh-login-time/1819 ## Headings Structure: H1: Change CL at CQLSH login time H3: Related topics ## Main Content: H1: Change CL at CQLSH login time H3: Related topics How can I change CL when login to CQLSH, getting following error message="Operation timed out for system_auth.roles - received only xx responses from xx CL=QUORUM." . I want to change CL=ONE at login time. I’m not sure I understand what you mean by changing the consistency level at login. You can see a hands-on example of how to change the CL using CQLSH in High Availability ScyllaDB University lesson. There is no way to change this, this is an internal query done by ScyllaDB and the CL it uses is not tunable. It is kind of tunable , as non ‘cassandra’ user would use local one cl. More about it here Creating a Custom Superuser | ScyllaDB Docs It’s an old quirk which soon will be fixed. What I noticed even java driver uses the same CL when it tries to connect. We are facing issue with connection timeout due to CL=quorum. If all quorum nodes does not respond connection timeouts. Is there any tunable parameter in java driver. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-45-2024-04-26/1837 Title: Last week in scylla-cluster-tests.git master (issue #45; 2024-04-26) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 7380b6bf…e8cdc0ba range are covered . There were 15 non-merge commits from 3 Software Engi… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-45-2024-04-26/1837 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #45; 2024-04-26) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #45; 2024-04-26) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 7380b6bf…e8cdc0ba range are covered . There were 15 non-merge commits from 3 Software Engineers in Test and 1 Software Engineer in that period. Some notable commits: The simulated_racks SCT config option, which allows configuration of different Scylla DB racks within a single availability zone, is now complemented by a new simulated_regions option. This new feature enables the simulation of multiple Scylla DB regions with nodes provisioned within the same regional boundary. Additionally, we’ve enhanced our logging capabilities by incorporating hostname information into RemoteCmdRunner output, making it easier to analyze test results. Due to the EOL status of Centos-7, we’ve transitioned to Ubuntu for loader and monitor nodes. Testing of tablets at scale (covering multiple keyspaces and tables) is now integrated into CI. The new ScyllaMetricsController class simplifies the disabling of specific Scylla metrics collection when needed. Manager version comparison now accounts for snapshot versions as well. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/ephemeral-storage-persistency-availability-and-data-recovery-on-nvme-ssd-ebs/1859 Title: Ephemeral storage, persistency, availability and data recovery (on nvme, ssd, ebs) - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/ephemeral-storage-persistency-availability-and-data-recovery-on-nvme-ssd-ebs/1859 ## Headings Structure: H1: Ephemeral storage, persistency, availability and data recovery (on nvme, ssd, ebs) H3: Related topics ## Main Content: H1: Ephemeral storage, persistency, availability and data recovery (on nvme, ssd, ebs) H3: Related topics Originally from the User Slack @Pranit_Chugh: Hi Community, I am trying to setup scylla 3 node cluster on production on i4i instance. I am a little apprehensive about using ephemeral storage as a primary disk. I have read @Felipe_Cardeneti_Mendes answer to this here : https://www.scylladb.com/2023/07/17/top-mistakes-with-scylladb-storage/ : Locally-attached disk mythbusting But I’m still not sure if a node goes down or is stopped and started, other than backup through replication what are my options for data recovery? Also I want to know would it be suggested to use an ebs disk attached with i4i for backup. If yes how will that work? Read and write through nvme ssd and backup on ebs or is it not possible/recommended. If attaching ebs with nvme disk wouldn’t be utilising i4i’s cabability, then what would be the next best bet for instance type for scylla with persistant storage? @Botond_Dénes: I don’t about AWS best practices, but if you have a ScyllaDB cluster with 3 noes and you create your keyspace with RF=3, you read/write with CL=QUORUM, your data should be safe, even if a node goes down. Don’t forget to repair regularly. @Pranit_Chugh: @Botond_Dénes thanks for taking out time to answer the question. But this doesn’t give me full clarity about data backup. With ebs disks even if all nodes go down I can replace nodes attach ebs and am good to go. But with ephemeral storages not sure what the recommended path for backup is. @Botond_Dénes: You can use ScyllaManager to do backups. It is free with a small number of nodes AFAIK. @Stewart: @Pranit_Chugh scylladb is architecturally recommended to use a local ssd instance, and reliability is ensured by backup and replication factors. If you are interested in using EBS and local ssd simultaneously for greater reliability, check out this related article. https://discord.com/blog/how-discord-supercharges-network-disks-for-extreme-low-latency This article details the work done on discord to raid local ssd and EBS to increase reliability. How Discord Supercharges Network Disks for Extreme Low Latency @Pranit_Chugh: Thanks this is a beautifully written article and tells exactly what I wanted to know. @Stewart: Additionally, we are currently using the i4i type. This type is very fast and can achieve amazing scylladb performance (i3en) compared to the i3en type. Unfortunately, it has smaller disk than the i3en type, but it is recommended for very high performance instances. We are also working on how to RAID EBS and local SSDs on Kubernetes based on that article. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-227-2024-04-28/1861 Title: Last week in scylladb.git master (issue #227; 2024-04-28) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 87b08c957f…d8313dda43 range are covered. There were 91 non-merge commits from 19 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-227-2024-04-28/1861 ## Headings Structure: H1: Last week in scylladb.git master (issue #227; 2024-04-28) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #227; 2024-04-28) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 87b08c957f…d8313dda43 range are covered. There were 91 non-merge commits from 19 authors in that period. Some notable commits: Probabilistic read repair, long deprecated, has been removed. The native nodetool implementation will now be used for every invocation of nodetool. The Java-based nodetool is still available as a fallback. Consistent topology changes (using Raft) is no longer experimental and is the default for new clusters. Write failures during replace-node-with-same-address procedure with consistent topology enabled, due to the new node being advertised as alive too late, are fixed. Off-strategy compaction is used to being sstables into compliance with the compaction strategy invariants after a node operation such as rebuild. It is now enabled for tablets. The large partition system tables now count range tombstones as regular rows, so that large partitions with mainly range tombstones are reported. When dropping bloom filters in order to reclaim memory, we could have left a bloom filter component on disk after its sstable was compacted. This is now fixed. The tombstone_gc table configuration species how to garbage collect tombstones. It is now set to repair for tables using tablets, indicating that tombstones can be garbage collected as soon as a full repair has been run. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/compaction-throughput-mb-per-sec-default-value/1870 Title: Compaction_throughput mb_per_sec default value - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: What was the default value of compaction_throughput_mb_per_sec in Scylla 4.6 and did this default value change to 0 from Scylla 5.1 ? Language: en Canonical URL: https://forum.scylladb.com/t/compaction-throughput-mb-per-sec-default-value/1870 ## Headings Structure: H1: Compaction_throughput mb_per_sec default value H3: Related topics ## Main Content: H1: Compaction_throughput mb_per_sec default value H3: Related topics What was the default value of compaction_throughput_mb_per_sec in Scylla 4.6 and did this default value change to 0 from Scylla 5.1 ? This configuration item is not used by ScyllaDB, it exists only for backward compatibility with Apache Cassandra, its value is ignored. ScyllaDB uses the user-space scheduling provided by seastar (the framework it is built on) and has controllers for compaction CPU and IO, making sure it uses the appropriate amount of these resources, based on the current compaction backlog. See this blog post for more details: Taming the Beast: How ScyllaDB Leverages Control Theory to Keep Compactions Under Control - ScyllaDB. --- ### Page: https://forum.scylladb.com/t/installation-through-rpms/1874 Title: Installation through rpms - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi Team, I am trying to install scylladb through rpms on RHEL, I am getting below error. Please letus know how to download below dependent rpms. Is there any blog to install through rpm. Thanks Language: en Canonical URL: https://forum.scylladb.com/t/installation-through-rpms/1874 ## Headings Structure: H1: Installation through rpms H3: Related topics ## Main Content: H1: Installation through rpms H3: Related topics I am trying to install scylladb through rpms on RHEL, I am getting below error. Please letus know how to download below dependent rpms. Is there any blog to install through rpm. Hey, you should be able to do it following the steps here: ScyllaDB | Get Started with ScyllaDB. The example I posted is for Open Source on Debian 11, so please remember to change it to the enterprise version. Here’s also a good reference - ScyllaDB Unified Installer (relocatable executable) | ScyllaDB Docs I have installed with that approach on RHEL 9.3 version, I am getting below error. Seems like your OS is not registered, try this How to fix RED HAT Error — This system is not registered with an entitlement server. You can use subscription-manager to register. - Amit Kumar Gupta - Medium Hi @Piotr_Smaron - unable to register in rhel suscription manager,verification email not receiving from RHEL. I have downloaded from below location https://downloads.scylladb.com/ I have tried tried with rpms, but i am getting below error.from which location i need to download below dependencies.Is there any doc how to install through rpm based. Error: Problem: cannot install the best candidate for the job --- ### Page: https://forum.scylladb.com/t/live-migration-from-dynamodb-to-scylladb-migrator-and-different-options/1897 Title: Live migration from DynamoDB to ScyllaDB, Migrator and different options - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/live-migration-from-dynamodb-to-scylladb-migrator-and-different-options/1897 ## Headings Structure: H1: Live migration from DynamoDB to ScyllaDB, Migrator and different options H3: DynamoDB – How to Move Out? H3: Inside a DynamoDB to ScyllaDB Migration H3: Related topics ## Main Content: H1: Live migration from DynamoDB to ScyllaDB, Migrator and different options H3: DynamoDB – How to Move Out? H3: Inside a DynamoDB to ScyllaDB Migration H3: Related topics Originally from the User Slack @Stewart: Is there a good way to do a live migration from dynamodb to scylladb? @Felipe_Cardeneti_Mendes: Alternator or CQL? For the former you should be able to use our migrator if your source has a fixed schema For the latter, you’ll probably want something along the lines of https://github.com/fee-mendes/scylladb-alternator-lambdastreams/ and replace the alternator logic with CQL. We are in process of publishing an article discussing these. @Stewart: we are trying to migrate with CQL . @Felipe_Cardeneti_Mendes does scylla spark migrator don’t support DDB to CQL ? @Felipe_Cardeneti_Mendes: it doesnt (yet) unfortunately @Stewart: so sad When do you think this feature will be released? @Felipe_Cardeneti_Mendes: I don’t think I have a good estimate. the Migrator for DynamoDB today relies on an deprecated & no longer supported library which we likely should replace, and only after plan to extend its functionality @Stewart: If so, what is the best way to live migrate to ddb -> scylladb as currently recommended by scylladb? @Felipe_Cardeneti_Mendes: Export a DDB S3 backup, load it to ScyllaDB, capture and replay events from ddb streams via lambda, kinesis or other method you prefer. pretty much a write heavy CQL code, at most a few deletes once you get to the streams replay part @Stewart: Is there a blog post or article that explains the process well? I am currently working on moving all of our ddb’s to scylladb and would like to do a POC. @Felipe_Cardeneti_Mendes: the GH repo above contains most of it, but it starts with a DDB->Alternator migration. Changing to CQL would require updating the s3Restore.py and the lambda function to use CQL instead, but the idea is pretty much there. and yes, there’s an article in the works as I mentioned last week. Probably should go out next week or so and this is how you export a S3 backup https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/S3DataExport_Requesting.html Requesting a table export in DynamoDB - Amazon DynamoDB @Stewart: I haven’t worked with ddb -> scylladb migration yet, so I’m a little confused. Can I check out the relevant documentation first and then ask some more questions? @Felipe_Cardeneti_Mendes If so, is it possible to use spark migrator to live migrate the table using alternator and then read/write the created table based on cql? @Felipe_Cardeneti_Mendes: unfortunately not, the non-key fields within alternator are stored in a blob map which is unintelligible for CQL clients. I will share you the docs we are currently working on today - ping me if I somehow forget lol @Stewart: ok thanks ! @Felipe_Cardeneti_Mendes any updates? @Felipe_Cardeneti_Mendes: I sent you an email yesterday following your Slack registered address @Stewart: thanks!! @Felipe_Cardeneti_Mendes I found a very interesting way to do a dynamodb -> scylladb live migration. If it works, I’d be happy to share it with you. It’s currently in PoC. @Felipe_Cardeneti_Mendes: certainly @Stewart: I just successfully implemented the code and PoC for dynamodb live migration, and wanted to share it briefly. First of all, the architecture is as follows flow is dynamodb -> dynamodb cdc source connector -> kafka -> scylladb sink connector (with ddb cdc transform) -> scylladb With this flow, data flows and data is loaded from ddb to scylladb in real time. If you are curious about my implementation, I can schedule a meeting for you. @Felipe_Cardeneti_Mendes GitHub: GitHub - fetch-rewards/kafka-connect-dynamodb: A Kafka Connect Source Connector for DynamoDB @Felipe_Cardeneti_Mendes: Oh, that’s a great idea, which effectively solves the DDB->CQL problem. How have you accomplished the bulk loading part? This is typically a requirement for a full migration (ie: export from DDB->ScyllaDB, and only after carry out your steps for syncing changes) @Stewart: if you see that repo, they support initial snapshot for connector. initial sync - automatically detects and if needed performs initial(existing) data replication before tracking changes from the DynamoDB table stream Synced(Source) DynamoDB table unit capacity must be large enough to ensure INIT_SYNC to be finished in around 16 hours. Otherwise there is a risk INIT_SYNC being restarted just as soon as it’s finished because DynamoDB Streams store change events only for 24 hours. INIT_SYNC can be skipped with init.sync.skip=true configuration @Felipe_Cardeneti_Mendes: Indeed. Super cool This is a great find actually! I probably should play with it, but are you interested in sharing your walkthrough and history with the community? We can definitely meet @Stewart: maybe I could write a blog about it later. currently I just finished testing with sample table, and need to do real mirgration with it. @Felipe_Cardeneti_Mendes: Sounds like a plan! Let us know how it goes @Stewart: I’ll upload the sample code to github after testing the migration, but if you want to test it before then, I can send you the source code. I ran a test yesterday with a real table base, and it worked fine. However, I realized that for struct and list types in dynamodb, there might be a problem. For nested data types like struct in struct and list in struct in dynamodb, the scylladb sink connector could not create a table normally. In this case, I think I have to develop an application that consumes dynamodb cdc topic and writes to scylladb by myself, or trasnform the cdc topic to KSQLDB and make it into a format that can be used in scylladb, but I wonder how to handle this case in the case of dynamodb alternator. @Felipe_Cardeneti_Mendes: Well, it makes sense as it can be defined to a udt, or frozen/unfrozen collection, nested structs can be considered an edge case (though unfortunately not that uncommon in ddb) but for writing directly to dynamodb alternator it should be a no-brainer. @Stewart: Surprisingly, there are quite a few cases where we use the nested type. In addition, I need to do a full scan of both dynamodb and scylladb to compare the data between them to see if it was migrated properly, which is also not easy. @Felipe_Cardeneti_Mendes: There’s also the even more edgy case where a list/map may contain values of different types. Whereas in CQL a list/map can be only of a specific type. @Stewart: yes right Rather, I’ve been running scylladb for about 2~3 years so I’m pretty familiar with it, but I’m not incredibly knowledgeable about dynamodb. I just found out yesterday that dynamodb supports this type of thing for the first time, because scylladb obviously doesn’t support it. lol @Felipe_Cardeneti_Mendes: Well, we DO support it (via Alternator). The way it is accomplished is that the payload gets serialized to a blob, and then we don’t need to care about its type, as long as the input is valid. But CQL has this restriction and switching protocols will inevitably require you some app changes when using such types. And these are likely app-specific, which makes it almost impossible to automate it. @Stewart: So currently my single message transformer (smt) supports boolean, integer, long, list, map, and I was going to support nested map and list types as well, but I stopped because I think the PR needs to be merged to use nested types. https://github.com/scylladb/kafka-connect-scylladb/pull/71 GitHub: Add complex types support and tests by Bouncheck · Pull Request #71 · scylladb/kafka-connect-scylladb @Felipe_Cardeneti_Mendes: A cool project which would probably be worth spending some hours on for fun would be to allow a CQL driver to manipulate Alternator tables. This would allow shard-awareness, prepared statements and all other goodies from the CQL protocol. But the penalty would be the serialization/deserialization being deferred to the client-side. Whether it is worth it or not depends on the application @Stewart: In preparation for this migration, I also wanted to check what tables are created internally when using Alternator, but I found a cool way to migrate to cdc and didn’t check it out. @Felipe_Cardeneti_Mendes: You may bump the PR comments, it has no progress for almost 2 years now As an update to this chat, we recently published the following articles on DynamoDB to ScyllaDB migrations: What does a DynamoDB migration really look like? Should you dual-write? Are there any tools to help? A detailed walk-through of an end-to-end DynamoDB to ScyllaDB migration. Further, the ScyllaDB Migrator received considerable improvements for DynamoDB migrations, such as: Note that there’s still no one-solution-fits-all for a DynamoDB->CQL transition (the main inquiry in this post), users are encouraged to report back their solutions/results as well as interest in accomplishing said task. --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-2-19/1905 Title: [RELEASE] ScyllaDB Enterprise 2022.2.19 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2022.2.19, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.2 Feature Release. Note the latest ScyllaDB Enterprise release is 2024… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2022-2-19/1905 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2022.2.19 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2022.2.19 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2022.2.19, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2022.2 Feature Release. Note the latest ScyllaDB Enterprise release is 2024.1 LTS, and you are encouraged to upgrade to it in coordination with the ScyllaDB Support team. Get ScyllaDB Enterprise 2022.2.19 (customers only, or 30-day evaluation) Upgrade from 2022.1.x to 2022.2.y Upgrade from ScyllaDB Enterprise 2022.2.x to 2024.1.y Upgrade from ScyllaDB Open Source 5.1 to ScyllaDB 2022.2 The following issues are fixed in this release (with an open-source reference, if available): MV Correctness: in some cases, like one base table with many views, updates to MV might be lost. In one example range tombstones made the issue appear earlier #17117. Stability: wrong range used for partition estimation for mixed-shard repair, when different nodes have different shared counts, resulting in larger than needed filters #17863. --- ### Page: https://forum.scylladb.com/t/schemaless-json-in-a-single-column/1918 Title: Schemaless JSON in a single column - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I understand that Scylla has support for storing JSON data and getting data in JSON form but from what I can see this is only schemafull. Every field in the JSON object will map to a column. Is it possible to store schem… Language: en Canonical URL: https://forum.scylladb.com/t/schemaless-json-in-a-single-column/1918 ## Headings Structure: H1: Schemaless JSON in a single column H3: Related topics ## Main Content: H1: Schemaless JSON in a single column H3: Related topics I understand that Scylla has support for storing JSON data and getting data in JSON form but from what I can see this is only schemafull. Every field in the JSON object will map to a column. Is it possible to store schemaless JSON in a single columns? for e.g store {a:“B”} in column3 It is, but then you will not be able to update a subset of the stored JSON document. how is it possible? I only need to insert said data Just create a text or blob column in your schema and dump the JSON in it. You will have to serialize/unserialize on the client side. That’s what I’ve been doing uptill now. Not exactly the solution I was hoping for. Hope the team adds support for pure schemaless JSON data soon It is, but then you will not be able to update a subset of the stored JSON document. That’s true. If you store the entire document as JSON, updating a specific subset can become more challenging. --- ### Page: https://forum.scylladb.com/t/superuser-account-in-scylladb-should-the-default-cassandra-user-be-dropped/1921 Title: Superuser account in ScyllaDB, should the default Cassandra user be dropped? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/superuser-account-in-scylladb-should-the-default-cassandra-user-be-dropped/1921 ## Headings Structure: H1: Superuser account in ScyllaDB, should the default Cassandra user be dropped? H3: Related topics ## Main Content: H1: Superuser account in ScyllaDB, should the default Cassandra user be dropped? H3: Related topics Originally from the User Slack @ishaan: Hi everyone ,we are currently working with Scylla and have a question about managing user accounts. After creating another superuser in Scylla, is it considered a best practice to drop the default Cassandra user, or should we keep it and simply downgrade it to a non-superuser role and revoke all permissions? I’m trying to understand the implications and best practices around this decision. Any insights or recommendations would be greatly appreciated. @Greg_Day: yes it’s best practice to drop the default user https://opensource.docs.scylladb.com/stable/operating-scylla/security/create-superuser.html Creating a Custom Superuser | ScyllaDB Docs --- ### Page: https://forum.scylladb.com/t/support-of-ip-range-as-primary-key/1927 Title: Support of ip range as primary key - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, Community. I’m currently working on storing a table with IP ranges as the partition key (ip_from and ip_to declared as inet type). Since IP ranges may overlap, I’ve added an additional unique clustering key row_num,… Language: en Canonical URL: https://forum.scylladb.com/t/support-of-ip-range-as-primary-key/1927 ## Headings Structure: H1: Support of ip range as primary key H3: Related topics ## Main Content: H1: Support of ip range as primary key H3: Related topics Hi, Community. I’m currently working on storing a table with IP ranges as the partition key (ip_from and ip_to declared as inet type). Since IP ranges may overlap, I’ve added an additional unique clustering key row_num, to ensure that I can always select the same row from a set of ranges. The create table statement looks like this: create table my_table ( ip_from inet, ip_to inet, row_num bigint, …, primary key ((ip_from, ip_to), row_num) ) I’m executing queries to find a row with an IP range that includes a provided IP address: select * from my_table where ip_from <= ? and ip_to >= ? limit 1 allow filtering In some cases, I manage to find the row I’m looking for (for example, any IP in the range 2607:fb90:e200:8000:: to 2607:fb90:e200:80ff:ffff:ffff:ffff:ffff), but in other cases, nothing is found (for example, any IP in the range 2603:2000:: to 2603:203f:ffff:ffff:ffff:ffff:ffff:ffff). Could anyone explain this behavior? (I understand that this might not be the most appropriate task to solve using Scylla, but perhaps there’s a way to tackle it somehow) UPD: I found workaround hack but I don’t think that it is best solution. I converted all ip ranges to CIDR format and used 2 parts of it as 2 primary keys. For example, ip range 2607:fb90:e200:8000:: - 2607:fb90:e200:80ff:ffff:ffff:ffff:ffff can be presented as 2607:fb90:e200:8000::/56 - 2607:fb90:e200:8000:: as first key and 56 as second key. It is good, if all ranges has the same ip prefix in CIDR, but in my case this is not the true. I have to iterate with passed ip to find needed ip_prefix from 128 to 1 for IPv6 and from 32 to 1 for IPv4. It looks like the best solution is to use vector search mechanism like Cassandra supports now, but as I know Scylla don’t support it yet. Can you post an example that fails? There is no fail or error. Just nothing found while I know that the record exists. If I execute query like: and pass 2607:fb90:e200:8000:: for ip_from and 2607:fb90:e200:80ff:ffff:ffff:ffff:ffff for ip_to then I will find the record with this range. But any other query that is looking for range by specific ip, that this range includes, like: will return nothing as soon as the table size becomes more than 100 000 records. Please post a script that creates such a table, populates it, and runs the failing query. Best in a new issue. --- ### Page: https://forum.scylladb.com/t/release-introducing-database-level-encryption-for-aws-clusters-1-may-2024/1928 Title: [RELEASE] Introducing database-level encryption for AWS clusters - 1 May 2024 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: ScyllaDB Cloud now offers sensitive data protection by enabling database-level encryption for AWS based database clusters. This encryption is on top of the existing storage-based encryption.When creating a new cluster, u… Language: en Canonical URL: https://forum.scylladb.com/t/release-introducing-database-level-encryption-for-aws-clusters-1-may-2024/1928 ## Headings Structure: H1: [RELEASE] Introducing database-level encryption for AWS clusters - 1 May 2024 H3: Related topics ## Main Content: H1: [RELEASE] Introducing database-level encryption for AWS clusters - 1 May 2024 H3: Related topics ScyllaDB Cloud now offers sensitive data protection by enabling database-level encryption for AWS based database clusters. This encryption is on top of the existing storage-based encryption.When creating a new cluster, users can enable database-level encryption by either creating a key via their own AWS KMS account (available in our Premium tier offering), or by using a ScyllaDB-managed key. For more information, see our documentation. Amazing …i didn’t expect this… --- ### Page: https://forum.scylladb.com/t/issue-when-trying-to-install-an-older-version-to-test-the-upgrade-process/1930 Title: Issue when trying to install an older version to test the upgrade process - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/issue-when-trying-to-install-an-older-version-to-test-the-upgrade-process/1930 ## Headings Structure: H1: Issue when trying to install an older version to test the upgrade process H3: Related topics ## Main Content: H1: Issue when trying to install an older version to test the upgrade process H3: Related topics Originally from the User Slack @Kristjan_Janežič: Hello everyone. As per the instructions I am trying to install an older version of Scylla (so we can test the upgrade process). I have added older repo in apt list and updated. I can see the packages with “ap-cache madison” or “apt list --all-versions”, but I cannot install them with “apt install” as the packages are not found. How can I install the older version? Thank you in advance @Felipe_Cardeneti_Mendes: Working just fine here though you should help yourself and don’t use a rc @Kristjan_Janežič: hey, thank you for the reply. I need to use rc because i need to test upgrade from that version . Could you share your apt sources? or the keyring command you used? I have also tried with these instructions, with the same result. @Felipe_Cardeneti_Mendes: > deb [arch=amd64,arm64 signed-by=/etc/apt/keyrings/scylladb.gpg] https://downloads.scylladb.com/downloads/scylla/deb/debian-ubuntu/scylladb-5.4 stable main > deb [arch=amd64,arm64 signed-by=/etc/apt/keyrings/scylladb.gpg] https://downloads.scylladb.com/downloads/scylla/deb/debian-ubuntu/scylladb-5.0 stable main history snippet: > 10 gpg --homedir /tmp --no-default-keyring --keyring /etc/apt/keyrings/scylladb.gpg --keyserver hkp://keyserver.ubuntu.com:80 --recv-keys d0a112e067426ab2 > 11 apt update > 12 apt-cache madison scylla > 13 apt install scylla=5.0~rc8-0.20220612.f28542a71-1 > 14 apt install scylla{,-server,-jmx,-tools,-tools-core,-kernel-conf,-node-exporter,-conf,-python3}=5.0~rc8-0.20220612.f28542a71-1 --- ### Page: https://forum.scylladb.com/t/applying-configurations-after-provisioning-a-scyllacluster-using-scylla-operator/1951 Title: Applying configurations after provisioning a ScyllaCluster using Scylla-operator - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/applying-configurations-after-provisioning-a-scyllacluster-using-scylla-operator/1951 ## Headings Structure: H1: Applying configurations after provisioning a ScyllaCluster using Scylla-operator H3: Related topics ## Main Content: H1: Applying configurations after provisioning a ScyllaCluster using Scylla-operator H3: Related topics Originally from the User Slack @Greg_Day: Hello, I am looking for some guidance on the proper ways to apply configurations after provisioning a ScyllaCluster manifest via scylla-operator. If I want to add or remove racks, do I simply need to reconfigure the manifest and scylla-operator will handle it? Also is there a recommended way to change immutable fields, like exposeOptions, or does that require a complete re-provisioning? additionally, I do not see a way to use preallocated IP addresses for GKE internal loadbalancers: I tried using the annotation on the CRD, but it turns out this is wrong and does not apply to internal passthrough loadbalancers, only Ingresses. @Maciej_Zimnoch: > If I want to add or remove racks, do I simply need to reconfigure the manifest and scylla-operator will handle it? Simply change ScyllaCluster spec and scylla-operator will take care of reconciling it. Also is there a recommended way to change immutable fields, like exposeOptions, or does that require a complete re-provisioning Immutable fields are immutable. You have to recreate the cluster. If you don’t delete PVC’s (they survive ScyllaCluster deletion), you won’t loose any data, there will be a downtime though. > I do not see a way to use preallocated IP addresses for GKE internal loadbalancers: Annotation you mentioned only works on Ingresses. To expose ScyllaCluster via Ingress there’s exposeOptions.cql.ingress but it requires SNI proxy. Seems like we are missing a field to support preallocated LB address. You may submit a feature request in scylla-operator repo. @Greg_Day: > Immutable fields are immutable. yes, sorry for the confusing language, I was just looking for tips on that recreation process, but I can handle that no problem. > Seems like we are missing a field to support preallocated LB address. You may submit a feature request in scylla-operator repo. sounds good, I might even take a look and see if I could add it and submit a PR --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-228-2024-05-05/1952 Title: Last week in scylladb.git master (issue #228; 2024-05-05) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the d8313dda43…53b98a8610 range are covered. There were 65 non-merge commits from 10 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-228-2024-05-05/1952 ## Headings Structure: H1: Last week in scylladb.git master (issue #228; 2024-05-05) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #228; 2024-05-05) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the d8313dda43…53b98a8610 range are covered. There were 65 non-merge commits from 10 authors in that period. Some notable commits: The ‘ALTER TABLE … DROP column’ statement gained support of the USING TIMESTAMP clause. This is useful to restore tables that had columns dropped in the past and then re-created while avoiding data resurrection. The nodetool command now understands the resetlocalschema subcommand. The rpm and deb packages will no longer pull in the scylla-tools and scylla-jmx subpackages by default. They can still be installed directly. The SELECT statement now chooses a consistent plan when multiple secondary indexes are available, preventing inconsistency when a paged query hops to a different coordinator. Handling of very large mutations has been improved, reducing stalls while managing cluster metadata (e.g. topology and tablets). See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/how-to-connect-to-scylladb-monitor-installed-on-docker-on-wsl2/1954 Title: How to connect to ScyllaDB Monitor installed on docker on WSL2 - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello, I have installed the ScyllaDB Monitor Stack on Windows WSL2 (Ubuntu), and I have started the monitor successfully accessing localhost like below (run from wsl2 ubuntu): hzb@DESKTOP-V2CLPP0:~/scylla-monitoring$ ./… Language: en Canonical URL: https://forum.scylladb.com/t/how-to-connect-to-scylladb-monitor-installed-on-docker-on-wsl2/1954 ## Headings Structure: H1: How to connect to ScyllaDB Monitor installed on docker on WSL2 H3: Related topics ## Main Content: H1: How to connect to ScyllaDB Monitor installed on docker on WSL2 H3: Related topics Hello, I have installed the ScyllaDB Monitor Stack on Windows WSL2 (Ubuntu), and I have started the monitor successfully accessing localhost like below (run from wsl2 ubuntu): It show the start is successful, however I can’t access localhost:3000, neither on my PC browser or firefox browser installed in WSL2 ubuntu. Here is my scylla_servers.yml: Does ScyllaDB Monitor Stack support WSL2 install. If so what I have done wrong? I can’t find more help from the docuementation. If it helps, here is the status of ScyllaDB (in another container on docker): Problem solved after I use docker network instead of localhost by changing scylla_servers.yml to: and running the Monitor Stack using : ./start-all.sh -d prometheus-data Now I can access the monitor on localhost. Looks like it runs pretty well on WSL2 --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-2-18/1956 Title: [RELEASE] ScyllaDB 5.2.18 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.2.18, a bugfix release of the ScyllaDB 5.2 stable branch. ScyllaDB Open Source 5.2.18, like all past and future 5.x.y releases, is backward compatible, and supports roll… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-2-18/1956 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.2.18 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.2.18 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.2.18, a bugfix release of the ScyllaDB 5.2 stable branch. ScyllaDB Open Source 5.2.18, like all past and future 5.x.y releases, is backward compatible, and supports rolling upgrades. Note that the latest ScyllaDB Open Source stable release is 5.4 and you are encouraged to upgrade to it. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/removing-a-node-from-a-cluster-without-replicating-the-data/1963 Title: Removing a node from a cluster without replicating the data - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/removing-a-node-from-a-cluster-without-replicating-the-data/1963 ## Headings Structure: H1: Removing a node from a cluster without replicating the data H3: Related topics ## Main Content: H1: Removing a node from a cluster without replicating the data H3: Related topics Originally from the User Slack @Rinson_John: Hi, Do we have any alternative for nodetool assassinate in ScyllaDB? I would like to remove a node from the cluster without re-replicating any data. Using scylla 4.2.1 version. @Morten_Bo: New user so not an expert - but wouldn’t that mean you get a replication factor issue? Since if you take out a node, whatever data that node was carrying must either be replicated elsewhere (to keep whatever replication level you have) or you will have data which is NOT replicated as per your replication setup. @Felipe_Cardeneti_Mendes: Your understanding is correct @Morten_Bo. @Rinson_John you can use nodetool removenode force after a removenode, and this should work. But it goes without saying that your deployment is ancient, so you should upgrade. --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-4/1964 Title: [RELEASE] ScyllaDB Enterprise 2024.1.4 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2024.1.4, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2024.1 LTS Release. 2024.1.4 patch release includes multiple minor bug fixes… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-4/1964 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2024.1.4 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2024.1.4 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2024.1.4, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2024.1 LTS Release. 2024.1.4 patch release includes multiple minor bug fixes. The following issues are fixed in this release (with an open-source reference, if available): --- ### Page: https://forum.scylladb.com/t/nodetool-refresh-to-migrate-data/1965 Title: Nodetool Refresh to Migrate Data - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, I’m looking into options to migrate Scylla data in a few different ways. We have multiple clusters with customer data split across them. We have a few customers that were initially split (for downstream “ease” to set… Language: en Canonical URL: https://forum.scylladb.com/t/nodetool-refresh-to-migrate-data/1965 ## Headings Structure: H1: Nodetool Refresh to Migrate Data H3: Related topics ## Main Content: H1: Nodetool Refresh to Migrate Data H3: Related topics Hi, I’m looking into options to migrate Scylla data in a few different ways. We have multiple clusters with customer data split across them. We have a few customers that were initially split (for downstream “ease” to setup), but we’re looking to merge them now. I think nodetool refresh is the way we want to go with this, but we have a few scenarios that we will run into. This seems fairly cut and dry. Should we still run load and stream, or is it redundant in this case? Will we need to do a nodetool cleanup if not? Load and stream should handle the cluster chnage as far as I can tell, but should we just copy the SSTables Node1 > Node1, or would it make more sense to split it (Node1 > 50% to Node1/ 50% to Node2) to distribute the load on the cluster and speed things up? Or would there be issues with splitting the SSTables like this? I assume we would still need a refresh if we are just renaming a keyspace (but not the tables), or is there some better way to handle that? load-and-stream is preferable as it doesn’t require you to copy the data everywhere, and doesn’t require nodetool cleanup afterwards. Thanks for the reply Avi. Though I still have some use questions: It’s better to run work on all nodes. There’s no fast alternative to rename a keyspace. --- ### Page: https://forum.scylladb.com/t/adding-a-column/1967 Title: Adding a column - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: after adding a column to scylla table , we are getting this error in app gocql: not enough columns to scan into: have 7 want 8 Language: en Canonical URL: https://forum.scylladb.com/t/adding-a-column/1967 ## Headings Structure: H1: Adding a column H3: Related topics ## Main Content: H1: Adding a column H4: ScyllaDB with golang driver H3: Related topics after adding a column to scylla table , we are getting this error in app gocql: not enough columns to scan into: have 7 want 8 not enough columns to scan into Which gocql are you using? A similar issue was sloved many years ago I tried the gocassa driver(https://github.com/megamsys/gocassa) built on top of … github.com/gocql/gocql with scylladb. **Cassandra 2.1.8** In cassandra a query is performed on inside the code `github.com/gocql/gocql` for this query. SELECT data_center, rack, host_id, tokens, release_version FROM system.local WHERE key='local' I received the results ``` Enter code here...&gocql.Iter{err:error(nil), pos:0, meta:gocql.resultMetadata{flags:1, pagingState:[]uint8(nil), columns:[]gocql.ColumnInfo{gocql.ColumnInfo{Keyspace:"system", Table:"local", Name:"data_center", TypeInfo:gocql.NativeType{proto:0x3, typ:13, custom:""}}, gocql.ColumnInfo{Keyspace:"system", Table:"local", Name:"rack", TypeInfo:gocql.NativeType{proto:0x3, typ:13, custom:""}}, gocql.ColumnInfo{Keyspace:"system", Table:"local", Name:"host_id", TypeInfo:gocql.NativeType{proto:0x3, typ:12, custom:""}}, gocql.ColumnInfo{Keyspace:"system", Table:"local", Name:"tokens", TypeInfo:gocql.CollectionType{NativeType:gocql.NativeType{proto:0x3, typ:34, custom:""}, Key:gocql.TypeInfo(nil), Elem:gocql.NativeType{proto:0x3, typ:13, custom:""}}}, gocql.ColumnInfo{Keyspace:"system", Table:"local", Name:"release_version", TypeInfo:gocql.NativeType{proto:0x3, typ:13, custom:""}}}, colCount:5, actualColCount:5}, numRows:1, next:(*gocql.nextIter)(nil), host:(*gocql.HostInfo)(nil), framer:(*gocql.framer)(0xc8200d80b0), once:sync.Once{m:sync.Mutex{state:0, sema:0x0}, done:0x0}} ``` and it works **In ScyllaDB 0.17** The same query returns back with no info from `github.com/gocql/gocql` ``` &gocql.Iter{err:error(nil), pos:0, meta:gocql.resultMetadata{flags:4, pagingState:[]uint8(nil), columns:[]gocql.ColumnInfo(nil), colCount:0, actualColCount:0}, numRows:1, next:(*gocql.nextIter)(nil), host:(*gocql.HostInfo)(nil), framer:(*gocql.framer)(0xc820094370), once:sync.Once{m:sync.Mutex{state:0, sema:0x0}, done:0x0}} ``` `unable to fetch host info for 127.0.0.1: gocql: not enough columns to scan into: have 5 want 0` I tried using `cqlsh to scylladb` and ran the same query, but it works. Can somebody throw light on what i can do in `github.com/gocql/gocql` to fix this problem ? --- ### Page: https://forum.scylladb.com/t/storing-data-using-key-value-where-value-is-json-avro-options-for-generating-the-timeuuid/1971 Title: Storing data using Key Value where value is Json/Avro, options for generating the timeuuid - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/storing-data-using-key-value-where-value-is-json-avro-options-for-generating-the-timeuuid/1971 ## Headings Structure: H1: Storing data using Key Value where value is Json/Avro, options for generating the timeuuid H3: Related topics ## Main Content: H1: Storing data using Key Value where value is Json/Avro, options for generating the timeuuid H3: Related topics Originally from the User Slack We are looking to use ScyllaDB for storing different types of data in a kind of KeyValue setup where the Value of the Data is JSON/Avro or some other format which can hold multiple records and columns within a single column. We utilize a PrimaryKey that looks kind of like this: ((datasetid, inserteddate),inserteddatetime,timeuuid) Since in theory different pieces of data can be submitted for a specific datasetid, inserteddate and inserteddatetime I want to utilize the built in currentTimeUUID() function in my inserts on the table, are there any caveats to using the system-generated timeUUID function (as well as currenttimestamp()) when inserting data? @Felipe_Cardeneti_Mendes: insertdate is good for bucketing. But insertdatetime + timeuuid is redundant. You shouldn’t have any collisions just with timeuuid , and you can filter all records ingested within a given timeuuid window with the maxTimeuuid() and minTimeuuid() functions @Morten_Bo: I am aware of the redundancy but it is just really nice to be able to visually deduce the timestamp without resorting to using maxtimeuuid() and mintimeuuid() functions. I might remove inserteddatetime from the PK and simply have it as a regular column on the table @Felipe_Cardeneti_Mendes: that would be toTimestamp() or toUnixTimestamp(), not max/min. If your use case is append-only, then you can also rely on writetime() on a regular column to retrieve its epoch. Plenty of options to choose from. My use case is append-only. Basically any reading of the data will typically be getting all data for a given dataset in a given timeframe. So can span multiple days and be everything within a day or just a subset of the day. I might play around with different formats. --- ### Page: https://forum.scylladb.com/t/high-compaction-at-peak/1973 Title: High Compaction at peak - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, I want to know how to adjust compaction_static_shares parameter according to the load. We are experiencing high compaction running at peak read and latency increases to 5s. Help us to understand the above parameter. … Language: en Canonical URL: https://forum.scylladb.com/t/high-compaction-at-peak/1973 ## Headings Structure: H1: High Compaction at peak H3: Related topics ## Main Content: H1: High Compaction at peak H3: Related topics I want to know how to adjust compaction_static_shares parameter according to the load. We are experiencing high compaction running at peak read and latency increases to 5s. Help us to understand the above parameter. usually, we set compaction_static_shares to 100 to minimize impact. 1k is the maximum value. do you have monitoring set up? latency increases to that only during compaction? --- ### Page: https://forum.scylladb.com/t/materialized-views-and-indexing-filtering-columns-by-range-allow-filtering/1974 Title: Materialized Views and Indexing, filtering columns by range, ALLOW FILTERING - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/materialized-views-and-indexing-filtering-columns-by-range-allow-filtering/1974 ## Headings Structure: H1: Materialized Views and Indexing, filtering columns by range, ALLOW FILTERING H3: Related topics ## Main Content: H1: Materialized Views and Indexing, filtering columns by range, ALLOW FILTERING H3: Related topics Originally from the User Slack @Artem_Golovko: Hello everyone. I’m a new guy using scylla and just play around watching the scylla university courses. Would love if someone could help me with my confusion. I have a simple table: I want to answer on the queries by receiveTime or by sourceTime, so for that I have created a local index. If I describe the keyspace the index under the hood uses the following MV By receiveTime obvious works, no questions. By sourceTime not working and requires filtering, but I can’t understand why it is so Despite that I can query index directly and then use IN clause (do things manually that scylla should do for me) Would love if you can clarify that behavior. Thanks! @Nik: Hey Artem as per my understanding of MV they are created internally by Scylla to optimise reads and get over the problem of out of sync during updation so as stated in this case a table can have multiple MV for faster reads ,schema alteration doesn’t happen on the base table rather a new MV is created for specific use case . @Artem_Golovko: @Nik According to the documentation https://opensource.docs.scylladb.com/stable/using-scylla/local-secondary-indexes.html scylla should do 2 requests for me. First, request the index view to know primary keys of the base table and second request the base table with found primary keys. But it’s not working when I’m filtering indexed column by range, but it should work, because in that case indexed column is a clustering key @Felipe_Cardeneti_Mendes: you may be interested in https://github.com/scylladb/scylladb/issues/5547 GitHub: Local indexes could allow range queries on indexed elements · Issue #5547 · scylladb/scylladb @Artem_Golovko: @Felipe_Cardeneti_Mendes thanks! --- ### Page: https://forum.scylladb.com/t/loadbalancer-per-node-when-deploying-with-the-kubernetes-operator/1985 Title: LoadBalancer per node when deploying with the Kubernetes Operator - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/loadbalancer-per-node-when-deploying-with-the-kubernetes-operator/1985 ## Headings Structure: H1: LoadBalancer per node when deploying with the Kubernetes Operator H3: Related topics ## Main Content: H1: LoadBalancer per node when deploying with the Kubernetes Operator H3: Related topics Originally from the User Slack @Christopher_Wong: I’m trying to expose a 3-node scylla cluster created with the scylla-operator. When the LoadBalancer service gets created, I see a LoadBalancer per node rather than the client service having a load balancer pointed at all 3 nodes. Is this the expected behavior? Using these exposeOptions @Maciej_Zimnoch: yes, nodes have their own LB service to allow drivers to connect to particular node. This is required to have node-awareness and shard-awareness, as traffic load balancing happens on the driver side, not server side. @Guilherme_Nogueira: We just rely on the LB service mechanism to allow the external connectivity into the pods. As Maciej mentioned, the driver themselves are responsible to load-balance the request across any nodes in the cluster @Christopher_Wong: gotcha - I’m seeing some weird behavior where the nodes just continuously log TLS handshake errors when I set the client type to ServiceLoadBalancerIngress. The cluster starts, but no nodes are available to take requests setting clients.type back to CLusterIP resolves the issue but doesn’t allow for an externally exposed load balancer @Greg_Day: @Christopher_Wong I don’t know much but you might just need to set the correct annotations (or specify a custom loadBalancerClass) depending on your environment. there are examples in the docs for EKS and GKE https://operator.docs.scylladb.com/stable/exposing.html#loadbalancer-type that’s all I know! Exposing ScyllaCluster | ScyllaDB Docs @Maciej_Zimnoch: LoadBalancer Service without special annotation on cloud providers results in ScyllaCluster being exposed to Internet, where various scanners or LB healthchecks are trying to connect to different ports. These usually are visible in logs. If you want to expose it only within cloud VPC, consider using either PodIP or LB with special annotation. Refer to cloud documentation regarding that. --- ### Page: https://forum.scylladb.com/t/background-reclaim/1986 Title: Background Reclaim - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Background reclaim - Time used by Instance and Time spent in task quota violation by instance are high in few nodes. Is that means those nodes has bad memory or what could be the reason? Language: en Canonical URL: https://forum.scylladb.com/t/background-reclaim/1986 ## Headings Structure: H1: Background Reclaim H3: Related topics ## Main Content: H1: Background Reclaim H3: Related topics Background reclaim - Time used by Instance and Time spent in task quota violation by instance are high in few nodes. Is that means those nodes has bad memory or what could be the reason? This means that you have high memory pressure. ScyllaDB tries to maintain some reserve memory. When this runs low (below some threshold), the background reclaim kicks in to free some reclaimable memory. If you have high CPU usage in this scheduling group, it means all easily reclaimable memory was already reclaimed and the reclaimer resorted to expensive LSA memory compaction. Check your memory metrics, you probably have high non-LSA memory consumption. A typical cause of this is high bloom-filter memory usage. --- ### Page: https://forum.scylladb.com/t/release-scylladb-rust-driver-0-13-0/1987 Title: [RELEASE] ScyllaDB Rust Driver 0.13.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce ScyllaDB Rust Driver 0.13.0, an asynchronous CQL driver for Rust, optimized for Scylla, but also compatible with Apache Cassandra! Some interesting statistics: over 1,615k dow… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-rust-driver-0-13-0/1987 ## Headings Structure: H1: [RELEASE] ScyllaDB Rust Driver 0.13.0 H2: Changes H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Rust Driver 0.13.0 H2: Changes H3: Related topics The ScyllaDB team is pleased to announce ScyllaDB Rust Driver 0.13.0, an asynchronous CQL driver for Rust, optimized for Scylla, but also compatible with Apache Cassandra! Some interesting statistics: API cleanups / breaking changes: New features / enhancements: CI / developer tool improvements: Congrats to all contributors and thanks everyone for using our driver! The source code of the driver can be found here: The official crates.io registry entry is here: Thank you for your attention, please do not hesitate to contact us if you have any questions, issues, feature requests, or are simply interested in our driver! Contributors since the last release: --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2023-1-8/1988 Title: [RELEASE] ScyllaDB Enterprise 2023.1.8 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2023.1.8, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2023.1 LTS Release. 2023.1.8 patch release includes multiple minor bug fixes… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2023-1-8/1988 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2023.1.8 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2023.1.8 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2023.1.8, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2023.1 LTS Release. 2023.1.8 patch release includes multiple minor bug fixes. ScyllaDB Enterprise’s latest Long Term Support (LTS) is 2024.1. While we continue to support 2023.1, you are encouraged to upgrade to it in coordination with the ScyllaDB Support team. The following issues are fixed in this release (with an open-source reference, if available): --- ### Page: https://forum.scylladb.com/t/last-2-weeks-in-scylla-cluster-tests-git-master-issue-46-2024-05-10/1989 Title: Last 2 weeks in scylla-cluster-tests.git master (issue #46; 2024-05-10) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last 2 weeks. Commits in the fdef448a…fa9af3cf range are covered. There were 44 non-merge commits from 10 Software … Language: en Canonical URL: https://forum.scylladb.com/t/last-2-weeks-in-scylla-cluster-tests-git-master-issue-46-2024-05-10/1989 ## Headings Structure: H1: Last 2 weeks in scylla-cluster-tests.git master (issue #46; 2024-05-10) H3: Related topics ## Main Content: H1: Last 2 weeks in scylla-cluster-tests.git master (issue #46; 2024-05-10) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last 2 weeks. Commits in the fdef448a…fa9af3cf range are covered. There were 44 non-merge commits from 10 Software Engineers in Test and 2 Software Engineers in that period. Some notable commits: Longevity tests can now be executed on Jenkins using a docker backend, which allows for testing in AWS environments similar to local setups. These tests are run on sct-runner instances equipped with NVMe disks. The cluster provisioning process was sped up by avoiding re-execution of the machine configuration script if it has already been run, reducing the risk of bugs related to syslog-ng reconfiguration. A new check ensures that non-voter nodes do not remain in the cluster with consistent-topology changes. Nodetool scrub in validate mode is now utilized to verify the Scylla cluster post-longevity test. In the event of corruption, affected sstables are quarantined and uploaded to S3, and a corresponding error event is logged. This feature was introduced in multidc tier1 tests and will be expanded to additional tests. The monitoring instance now uses a prepared image (version 4.6.2) to expedite provisioning on AWS and GCE, accompanied by updated documentation on image updates. The Latte loader image size was significantly reduced from 580MB to 35MB, and its version updated, now used to test custom workflows. Tests now have the ability to reserve AWS instances for their duration, which helps mitigate capacity issues common in performance tests using placement groups. It also includes fallback options to another AZ to increase test run chances. Artifact tests now include running scylla-doctor and validating the JSON output file. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/metric-does-not-match/1990 Title: Metric does not match - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Following metrics values does not match, please explain what it means? scylla_storage_proxy_coordinator_foreground_reads scylla_storage_proxy_coordinator_read_latency_count Language: en Canonical URL: https://forum.scylladb.com/t/metric-does-not-match/1990 ## Headings Structure: H1: Metric does not match H3: Related topics ## Main Content: H1: Metric does not match H3: Related topics Following metrics values does not match, please explain what it means? scylla_storage_proxy_coordinator_foreground_reads scylla_storage_proxy_coordinator_read_latency_count Does not match what? ? Anyhow if you install scylla monitoring ( ScyllaDB Monitoring Stack | ScyllaDB Docs ), in dashboard of monitoring you can hover your mouse above (i) icon to get explanation of metrics, below are: scylla_storage_proxy_coordinator_foreground_reads scylla_storage_proxy_coordinator_read_latency_count you can get other help for metrics by doing on Scylla node: curl http://:9180/metrics --- ### Page: https://forum.scylladb.com/t/creating-a-superuser-with-custom-credentials-salting-method-and-password-hash-generation/1999 Title: Creating a Superuser with custom credentials, salting method and password hash generation - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/creating-a-superuser-with-custom-credentials-salting-method-and-password-hash-generation/1999 ## Headings Structure: H1: Creating a Superuser with custom credentials, salting method and password hash generation H3: Related topics ## Main Content: H1: Creating a Superuser with custom credentials, salting method and password hash generation H3: Related topics Originally from the User Slack @Greg_Day: Hello, hopefully this is a quick one for anyone who knows. On this documentation page https://opensource.docs.scylladb.com/stable/operating-scylla/security/create-superuser.html#setting-custom-superuser-credentials-in-scylla-yaml this configuration value is described: What method of salting is used? Can I simply do mkpasswd '$secure-password' and use that output here? Thanks. Creating a Custom Superuser | ScyllaDB Docs that did not work, I still had to manually create the alternate superuser. What am I doing wrong? @LikDan: idk how this works on Linux, but on MacOS to generate password’s hash you can use openssl passwd -6 password12345678 where password12345678 is your password than you will have output similar to this $6$/PRr56xlMXctw5B5$kwxT3g4MTekVc.0o46izORDbR2DsCaztI2L0kxs6mKaHsnNCns8vsrYVzcz2iRphNewLco8Iydg.ZuiHmw4Ii/ and this is you have to paste in scylla.yml file p.s. if I am not mistaken scylla use SHA-512 algorithm @Greg_Day: Thanks! yeah openssl passwd -6 claims it uses: > -6 Use the SHA256 / SHA512 based algorithms defined by Ulrich Drepper. See . with mkpasswd that’s mkpasswd -m sha512crypt "yourpasswordhere" --- ### Page: https://forum.scylladb.com/t/scylladb-labs-training-event-may-2024/2000 Title: ScyllaDB Labs Training Event - May 2024 - Announcements - ScyllaDB Community NoSQL Forum Meta Description: On the 29th of May, we’ll host ScyllaDB Labs Building High-Performance Apps, a 2-hour virtual hands-on training event. It’s an interactive hands-on workshop, where you’ll learn how to get started with ScyllaDB and see… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-labs-training-event-may-2024/2000 ## Headings Structure: H1: ScyllaDB Labs Training Event - May 2024 H3: Related topics ## Main Content: H1: ScyllaDB Labs Training Event - May 2024 H3: Related topics On the 29th of May, we’ll host ScyllaDB Labs Building High-Performance Apps, a 2-hour virtual hands-on training event. It’s an interactive hands-on workshop, where you’ll learn how to get started with ScyllaDB and see some NoSQL development in action. Some of the topics that we will cover are: Save your spot here . Hope to see you there! This is happening this week, and (free) registration is still open, see you there! --- ### Page: https://forum.scylladb.com/t/release-scylladb-monitoring-stack-4-7-2/2001 Title: [RELEASE] ScyllaDB Monitoring Stack 4.7.2 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce a patch release of ScyllaDB Monitoring Stack 4.7.2 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB Enterprise and ScyllaDB Open Source, based on Prometh… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-monitoring-stack-4-7-2/2001 ## Headings Structure: H1: [RELEASE] ScyllaDB Monitoring Stack 4.7.2 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Monitoring Stack 4.7.2 H3: Related topics The ScyllaDB team is pleased to announce a patch release of ScyllaDB Monitoring Stack 4.7.2 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB Enterprise and ScyllaDB Open Source, based on Prometheus and Grafana. ScyllaDB Monitoring Stack 4.7.2 supports: --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-229-2024-05-12/2002 Title: Last week in scylladb.git master (issue #229; 2024-05-12) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 53b98a8610…2ad13e5d76 range are covered. There were 109 non-merge commits from 23 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-229-2024-05-12/2002 ## Headings Structure: H1: Last week in scylladb.git master (issue #229; 2024-05-12) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #229; 2024-05-12) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 53b98a8610…2ad13e5d76 range are covered. There were 109 non-merge commits from 23 authors in that period. Some notable commits: INSERT JSON now understands more complex data types when used as keys. The ScyllaDB source code now bundles the abseil library (again), as new versions of abseil cannot be used from distribution packages. The hinted handoff code now uses host IDs instead of IP addresses. This conforms with consistent topology, which uses host IDs throughout. A bug related to the cluster’s view of the cluster size during upgrade to consistent topology has been fixed. Change Data Capture (CDC) maintains a history of the cluster topology to allow clients to fetch older change records. This history is now pruned earlier to reclaim space. We will avoid running migrations from auth tables if we’re already using Raft based auth tables. During compaction, all cell timestamps are relative to a base timestamp. We now exclude fully expired sstables from calculation of the base timestamp, in order to get a more accurate base. The failure detector’s ping timeout is now tunable. The system.large_partitions table now includes a column for the number of deleted rows. Tables that use tablets will now reject lightweight transactions (LWTs) as they are still not implemented. The REST API now understands the special table names used to back alternator local secondary indexes. ScyllaDB caches a queries that have been paused by the paging mechanism so the next page can be fetched more efficiently. It will now drop cached queries when the tablet they refer to has been migrated. Recently, ScyllaDB started to drop Bloom filters when they use up too much space. It will now reload them when space is available again, for example due to compaction. ScyllaDB estimates a compaction’s partition count in order to correctly size the bloom filter. It will now improve the estimate for garbage collection sstables. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/mysql-to-nosql-data-retention-similar-to-aws-s3-retention-and-tiered-storage-similar-to-elastic-search-data-tier-s3/2004 Title: MySQL to NoSQL, Data retention (similar to AWS S3 Retention) and tiered storage (similar to Elastic Search Data Tier, S3) - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/mysql-to-nosql-data-retention-similar-to-aws-s3-retention-and-tiered-storage-similar-to-elastic-search-data-tier-s3/2004 ## Headings Structure: H1: MySQL to NoSQL, Data retention (similar to AWS S3 Retention) and tiered storage (similar to Elastic Search Data Tier, S3) H3: Related topics ## Main Content: H1: MySQL to NoSQL, Data retention (similar to AWS S3 Retention) and tiered storage (similar to Elastic Search Data Tier, S3) H3: Related topics Originally from the User Slack @Diego_Mendes: Hi There, I am currently investigating alternatives to replace one of our MySQL databases currently used as a metadata server. At this moment we have tables crossing the 10TB size and we notice soon or later we might have to reconsider our current setup. My team is investigating between refactoring our current database to use partitioning and gain a few more years or move to a NoSQL solution, due to our current use not being pure relational, we are tempted to move away from MySQL if the effort is the same. We came up with ScyllaDB as an option, but we have 2 requirements to satisfy the new solution: Am I correct to say these features are not supported nor will be in the short term? Are there any tools for Scylla today that can handle these use cases? @Felipe_Cardeneti_Mendes: First one is simple, TTL. Tiered Storage is on the roadmap. Thought there is nothing that prevents you from manually offloading data to tiered storage. ie: TTL automatically handles expiration while data is already in another media @Diego_Mendes: For the tiered storage I was looking for something similar to elastic search where we can keep the data in the server but with a different storage like HDD or s3. --- ### Page: https://forum.scylladb.com/t/can-i-restore-a-keyspace-with-a-different-schema/2005 Title: Can I restore a keyspace with a different schema? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I had to DROP a keyspace. Now I want to restore it from a backup - I want the data back, but I want to configure different schema options. Can I do it upon restore? Or do I have to restore the keyspace first and ALTER i… Language: en Canonical URL: https://forum.scylladb.com/t/can-i-restore-a-keyspace-with-a-different-schema/2005 ## Headings Structure: H1: Can I restore a keyspace with a different schema? H3: Related topics ## Main Content: H1: Can I restore a keyspace with a different schema? H3: Related topics I had to DROP a keyspace. Now I want to restore it from a backup - I want the data back, but I want to configure different schema options. Can I do it upon restore? Or do I have to restore the keyspace first and ALTER it later? Using Create+Alter or just Create with the new option, doesn’t make a difference. What makes a difference is what options you want to change. You can change the replication options, but note that when restoring the data, you have to take the changes into account, you will probably have to use nodetool refresh --load-and-stream. --- ### Page: https://forum.scylladb.com/t/hi-scylladb-community-my-name-is-ali-i-m-happy-to-be-here-i-don-t-use-scylladb-but-i-m-curious-to-learn-more/2009 Title: Hi ScyllaDB community. My name is Ali. I’m happy to be here. I don’t use ScyllaDB but I’m curious to learn more - Database Community - ScyllaDB Community NoSQL Forum Meta Description: Hi ScyllaDB community. My name is Ali. I’m happy to be here. I don’t use ScyllaDB but I’m curious to learn more. Language: en Canonical URL: https://forum.scylladb.com/t/hi-scylladb-community-my-name-is-ali-i-m-happy-to-be-here-i-don-t-use-scylladb-but-i-m-curious-to-learn-more/2009 ## Headings Structure: H1: Hi ScyllaDB community. My name is Ali. I’m happy to be here. I don’t use ScyllaDB but I’m curious to learn more H3: Related topics ## Main Content: H1: Hi ScyllaDB community. My name is Ali. I’m happy to be here. I don’t use ScyllaDB but I’m curious to learn more H3: Related topics Hi ScyllaDB community. My name is Ali. I’m happy to be here. I don’t use ScyllaDB but I’m curious to learn more. --- ### Page: https://forum.scylladb.com/t/is-there-anything-bad-when-executes-nodetool-repair-whose-host-ip-doesnt-match-token-range/2019 Title: Is there anything bad when executes nodetool repair whose host ip doesn't match token range? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: In order to make it specific enough, here set an example, a cluster with nodes 127.0.0.1, 127.0.0.2, 127.0.0.3 , and consistency level is quorum. Here is the process: firstly, get the token range by executing nodetool d… Language: en Canonical URL: https://forum.scylladb.com/t/is-there-anything-bad-when-executes-nodetool-repair-whose-host-ip-doesnt-match-token-range/2019 ## Headings Structure: H1: Is there anything bad when executes nodetool repair whose host ip doesn't match token range? H3: Related topics ## Main Content: H1: Is there anything bad when executes nodetool repair whose host ip doesn't match token range? H3: Related topics In order to make it specific enough, here set an example, a cluster with nodes 127.0.0.1, 127.0.0.2, 127.0.0.3 , and consistency level is quorum. Here is the process: firstly, get the token range by executing nodetool describering , then get: Finally, logs print repair successfully. Is there anything bad when executes nodetool repair whose host ip doesn’t match token range? I don’t think there’s any problem. --- ### Page: https://forum.scylladb.com/t/whats-the-sql-transaction-equivalent-in-nosql-select-values-in-batch-and-consistency/2020 Title: What's the SQL Transaction equivalent in NoSQL, select values in batch and consistency - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/whats-the-sql-transaction-equivalent-in-nosql-select-values-in-batch-and-consistency/2020 ## Headings Structure: H1: What's the SQL Transaction equivalent in NoSQL, select values in batch and consistency H3: Related topics ## Main Content: H1: What's the SQL Transaction equivalent in NoSQL, select values in batch and consistency H3: Related topics Originally from the User Slack @TN: what would be the equivalent of this SQL scenario in scylla? I haven’t found any information on how to select values in a batch statement, so since they can change in the meantime (after step 2), you’d need to use counters to ensure that columns are updated correctly but that means that you’d have to update all such types to counter types that take up 64 bits. Is that the correct way? And if the value doesn’t exist then I think I could just create some transaction that uses IF NOT EXISTS and first check whether the counter update was successful and if not then use the insert? @Felipe_Cardeneti_Mendes: --- ### Page: https://forum.scylladb.com/t/what-is-tombstone-compaction-interval-mean/2028 Title: What is tombstone_compaction_interval mean? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I read the doc troubleshooting/pointless-compactions here is my table schema ALTER TABLE foo.bar WITH compaction = {'class' : 'LeveledCompactionStrategy','tombstone_threshold' : 0.1,'sstable_size_in_mb' : 160, 'tombs… Language: en Canonical URL: https://forum.scylladb.com/t/what-is-tombstone-compaction-interval-mean/2028 ## Headings Structure: H1: What is tombstone_compaction_interval mean? H3: Related topics ## Main Content: H1: What is tombstone_compaction_interval mean? H3: Related topics I read the doc troubleshooting/pointless-compactions here is my table schema I expected scylla will check sstable’s Estimated droppable tombstones every 60s if Estimated droppable tombstones > 0.1 then will exec minor compact for this table but in my test after I delete data scylla don’t exec minor compact or should I write same primary key, when scylla flush to sstable then will trigger minor compact ? tombstone_compaction_interval is a lower-bound on the earliest time a new tombstone compaction can be started. Meaning that if an sstable was compacted at time X, the earliest time it will be considered for tombstone compaction again is X + tombstone_compaction_interval. This is not a guarantee that sstables will be considered for compaction immediately after tombstone_compaction_interval time has elapsed. I think in your test, you might have too little data or too little tombstones. It is hard to say why tombstone compaction is not triggering, without more details. @denesb I want to understand the trigger mechanism for compaction. I tested it by setting tombstone_compaction_interval to 120, which is 2 minutes, but I’ve observed that sometimes the compaction doesn’t happen at fixed intervals. What is the logic in the source code for this? How often does it scan? --- ### Page: https://forum.scylladb.com/t/running-scylladb-on-docker-with-ipv6/2030 Title: Running ScyllaDB on Docker with IPv6 - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/running-scylladb-on-docker-with-ipv6/2030 ## Headings Structure: H1: Running ScyllaDB on Docker with IPv6 H3: Related topics ## Main Content: H1: Running ScyllaDB on Docker with IPv6 H3: Related topics Originally from the User Slack @Nathan_Tamez: Hey I’m trying to run Scylla locally via docker with this image scylladb/scylla and it looks as though Scylla is just failing to start as it keeps trying to restart and i keep getting this log i’m on Mac OS with a M1 Pro chip @Felipe_Cardeneti_Mendes: Hi, can you share the output of hostname -i in your container? I think that must be it: > if self._listenAddress is None: > self._listenAddress = subprocess.check_output([‘hostname’, ‘-i’]).decode(‘ascii’).strip() I use a Mac myself but I don’t get an IPV6 address like you do, did you manually enabled it? @Nathan_Tamez: it was IPv6 haha, i needed to enable it for a test with some IPv6 for some other project. @Felipe_Cardeneti_Mendes: ok, as expected. https://github.com/scylladb/scylladb/issues/16502 GitHub: docker node crashes with error: too many positional options have been specified on the command line · Issue #16502 · scylladb/scylladb @Nathan_Tamez: thanks for the help --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-47-2024-05-17/2032 Title: Last week in scylla-cluster-tests.git master (issue #47; 2024-05-17) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the d3021bbd…819ad1e7 range are covered. There were 15 non-merge commits from 5 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-47-2024-05-17/2032 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #47; 2024-05-17) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #47; 2024-05-17) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the d3021bbd…819ad1e7 range are covered. There were 15 non-merge commits from 5 authors in that period. Some notable commits: To speed up test setup, we now configure and start the DB cluster, loaders, monitors, and Oracle cluster in parallel. This change has reduced setup time by approximately 20 minutes for a 6-node DB cluster. We added a bisect decorator that allows bisection functionality for selected SCT tests. This feature is still experimental, as it has only been tested in one performance scenario, and requires some preparation of tests to work fully. More details can be found in the documentation. It is limited to published binaries in ScyllaDB downloads. Since cassandra-stress is not included in the latest Scylla Docker images, loaders will use version 5.4.6. To further speed up SCT test setup, monitors and loaders will no longer configure memory swap space. Running a multi-DC cluster test from a local machine has been fixed. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/high-read-failures/2036 Title: High read failures - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: How to trace the fail reads, like select query? What keys are failing? Language: en Canonical URL: https://forum.scylladb.com/t/high-read-failures/2036 ## Headings Structure: H1: High read failures H3: Related topics ## Main Content: H1: High read failures H3: Related topics How to trace the fail reads, like select query? What keys are failing? I’m not sure what you mean by “What keys are failing”. There are different options for Tracing in ScyllaDB, read more about it here. Need to trace which key’s or queries are failing. --- ### Page: https://forum.scylladb.com/t/what-is-the-process-for-upgrading-scylladb-and-the-os-centos-to-rocky-linux/2040 Title: What is the process for upgrading ScyllaDB and the OS (CentOS to Rocky Linux)? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/what-is-the-process-for-upgrading-scylladb-and-the-os-centos-to-rocky-linux/2040 ## Headings Structure: H1: What is the process for upgrading ScyllaDB and the OS (CentOS to Rocky Linux)? H3: Related topics ## Main Content: H1: What is the process for upgrading ScyllaDB and the OS (CentOS to Rocky Linux)? H3: Related topics Originally from the User Slack @Akki: My current Scylla Cluster is 5.2 running on CentOS 7. Scylla 5.4 does not support C7, what steps I need to take to smoothly transition from C7 to Rocky? @Felipe_Cardeneti_Mendes: Standard practice is to add new nodes and then decommission the old ones. Of course other ways exist. @Akki: Adding new nodes is not an option. I was thinking for removing one node at a time rebuilding it with new RockyOS and Scylla5.4 and adding it back to the cluster. But my question is, as the rolling upgrade continues the cluster will have hosts with different OS. Will it be an issue? @Felipe_Cardeneti_Mendes: You can’t bootstrap a node in a diff version as it introduces newer features. You can take it down and upgrade it/move around data then start This is unsafe of course as if another node fails in between you lost quorum @Akki: What would be the best way this can be handled Centos7 to RockyOS? (rolling preferred) @Felipe_Cardeneti_Mendes: You should be ok with just postponing the upgrade for later. Upgrade first (remove, upgrade, add) to same version, then just upgrade ScyllaDB once complete @Akki: kool, so do we update OS on seeds first? @Felipe_Cardeneti_Mendes: Doesn’t matter much @Akki: Sounds good, will keep posted --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-230-2024-05-19/2041 Title: Last week in scylladb.git master (issue #230; 2024-05-19) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 2ad13e5d76…c93a7d2664 range are covered. There were 96 non-merge commits from 23 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-230-2024-05-19/2041 ## Headings Structure: H1: Last week in scylladb.git master (issue #230; 2024-05-19) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #230; 2024-05-19) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 2ad13e5d76…c93a7d2664 range are covered. There were 96 non-merge commits from 23 authors in that period. Some notable commits: Alternator, ScyllaDB’s implementation of the DynamoDB API, now includes an end-to-end benchmark (including the HTTP server). With tablets, each tablet replica’s sstables are isolated in a storage group. We now dynamically allocate these storage groups only when a node actually hosts a particular tablet. This saves memory and allows for large tables. The container image no longer contains the JMX server and cassandra-stress. The bundled cqlsh version was updated to v6.0.18. The iotune utility is used to measure a disk’s performance when installing ScyllaDB. It now works when executed on machines with a very large core count. A performance problem with range scans on tablets was fixed. There are now some metrics reported per table. The metrics are not reported per shard to avoid a combinatorial explosion. The DESCRIBE TABLES command now omits the special CDC tables. These aren’t needed to restore the schema since the base CREATE TABLE command will recreate them. ScyllaDB disseminates the materialized view update backlog in order to control update rates. A bug that prevented this in some circumstances was fixed. The nodetool ring command now supports tablets. Materialized views now throttle updates better, when an update generates a large amount of materialized view writes. A hang on some CQL queries with redundant constraints was fixed. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/how-to-migrate-data-between-twc-tables-to-change-the-ttl/2043 Title: How to migrate data between TWC tables to change the TTL - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/how-to-migrate-data-between-twc-tables-to-change-the-ttl/2043 ## Headings Structure: H1: How to migrate data between TWC tables to change the TTL H3: Related topics ## Main Content: H1: How to migrate data between TWC tables to change the TTL H3: Related topics Originally from the User Slack @Daria_Fedorova: Hi I am planning to migrate some data between TWC tables (to change TTL of this data) and I am afraid my data will become commingled. The doc https://opensource.docs.scylladb.com/stable/kb/compaction.html seems to say that there is no way to have a table available for real time write during migration of historical data ? The plan was 1 full scan and filter needed data from old table and write it to the new table with TTL=new_table_ttl - record_age (moved data up to timeuuid T0) 2 start writing in new table from timeuuid T1 3 migrate historical data missed from T0 to T1 But third step would involve writing old and new data at the same time. And this is not allowed ? Compaction | ScyllaDB Docs @avi: If you use the USING TIMESTAMP option to write historical data, it will get sorted into the right buckets. Best is to use the original timestamps. --- ### Page: https://forum.scylladb.com/t/release-scylla-manager-3-2-8/2044 Title: [RELEASE] Scylla Manager 3.2.8 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces the release of ScyllaDB Manager 3.2.8 production-ready ScyllaDB Manager patch release of the stable ScyllaDB Manager 3.2 branch. ScyllaDB Manager is a centralized cluster administration and r… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-manager-3-2-8/2044 ## Headings Structure: H1: [RELEASE] Scylla Manager 3.2.8 H3: Related topics ## Main Content: H1: [RELEASE] Scylla Manager 3.2.8 H3: Related topics The ScyllaDB team announces the release of ScyllaDB Manager 3.2.8 production-ready ScyllaDB Manager patch release of the stable ScyllaDB Manager 3.2 branch. ScyllaDB Manager is a centralized cluster administration and recurrent tasks automation tool. In version 3.2.8, we improved the reliability of health check tasks by making them fully independent from side effects unrelated to the primary purpose of the health check (#3767 #3847). We also included a small table optimization feature, available via the Scylla API (>= 5.5 Open Source, >= 2024.1.5 Enterprise), in the repair process (#3842). In addition to these fixes, this patch release prepares Scylla Manager to start supporting tablet-enabled environments (#3753 #3759 #3773). Post-restore repair is now executed with maximum intensity and parallelism (#3832). ScyllaDB Enterprise customers are encouraged to upgrade to ScyllaDB Manager 3.2.8 in coordination with the ScyllaDB support team. The new release includes upgrades of both ScyllaDB Manager Server and Agent. Scylla Manager 3.2.8 supports the following Scylla Enterprise releases: And the following Open Source release (limited to 5 nodes see license): You can install and run Scylla Manager on Kubernetes using Scylla Operator. More here. --- ### Page: https://forum.scylladb.com/t/loss-of-availability-and-timeout-errors-kubernetes-nodes-de-scheduled/2047 Title: Loss of Availability and Timeout errors, Kubernetes nodes de-scheduled - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/loss-of-availability-and-timeout-errors-kubernetes-nodes-de-scheduled/2047 ## Headings Structure: H1: Loss of Availability and Timeout errors, Kubernetes nodes de-scheduled H3: Related topics ## Main Content: H1: Loss of Availability and Timeout errors, Kubernetes nodes de-scheduled H3: Related topics Originally from the User Slack @ahmed_grati: Hello, We’re running ScyllaDB with 1 data center and 3 nodes across 3 AZs with a replication factor of 2. These nodes are spot instances. The version of Scylla is 5.2.9 We started experiencing many timeout errors when one of the nodes gets de-scheduled because the underlying Kubernetes nodes are de-scheduled also. It takes around 10 minutes to re-schedule another node (both Kubernetes and Scylla nodes). After the Scylla node was re-scheduled, we started seeing this kind of timeout errors: It should be noted that this happens only when we use LWT. My assumption is that since the replication factor is 2, that means that every request should be sent to all of the other nodes in the data center (since we have only 3 nodes), and since the node is de-scheduled these transactions would be put in a hinted handoff. And when the node gets re-scheduled it needs to handle both live and hinted transactions which can increase the size of the queue and cause a timeout issue for old requests. Again this is just an assumption and I’m reaching out to you searching for an explanation and a solution for that since this is a blocker for us to use ScyllaDB in our production. @avi: RF=3 is needed for LWT (and spot instances are very dangerous) @ahmed_grati: @avi Can you elaborate more on why spot instances are very dangerous for Scylla? (I saw some blog posts on ScyllaDB for people deploying it on Spot instances) @avi: They’re dangerous because AWS will take them away @ahmed_grati: Yes I know, but how does that influence Scylla? and is that related to the issue that we faced? @avi: Ah it’s only dangerous if you use instances with local storage @ahmed_grati: Nope, we’re using EBS volumes @avi: EBS is okay (but can lose availability) @ahmed_grati: are timeouts related to spot instaces? or you have another explanation? @avi: If you lose quorum you’ll get timeouts @ahmed_grati: Thanks, @avi for being responsive. Just one last question, Is there any way/solution to run Scylla on the spot instance without having timeouts? From what I understood, running on spot would decrease the availability and it would also engender timeouts. Please correct me if I’m wrong. @avi: Loss of availability = timeouts @ahmed_grati: @avi more on this, I’m still seeing the timeout error even though no node has been de-scheduled and the load is 20%. Any clue? @avi: Check the advanced dashboard in metrics to see if I/O or CPU is overloaded --- ### Page: https://forum.scylladb.com/t/node-stuck-in-joining-mode-while-bootstrapping-a-4-6-5-scylla-node/2049 Title: Node stuck in joining mode while bootstrapping a 4.6.5 Scylla node - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, We are trying to bootstrap a node in to the 4.6.5 Scylla cluster but looks like node is stuck in joining mode and allowed_repair_based_node_ops={replace} option is picked up by Scylla even when we disable or enable … Language: en Canonical URL: https://forum.scylladb.com/t/node-stuck-in-joining-mode-while-bootstrapping-a-4-6-5-scylla-node/2049 ## Headings Structure: H1: Node stuck in joining mode while bootstrapping a 4.6.5 Scylla node H3: Related topics ## Main Content: H1: Node stuck in joining mode while bootstrapping a 4.6.5 Scylla node H3: Related topics We are trying to bootstrap a node in to the 4.6.5 Scylla cluster but looks like node is stuck in joining mode and allowed_repair_based_node_ops={replace} option is picked up by Scylla even when we disable or enable RBNO. Logs shows compaction of shard 0 system column families and there are numerous compactions of system schema tables in a loop and node never joins the cluster. Can anyone please throw some light or any inputs on how to fix this issue? Thanks & Regards, Satish Hey Satish, did you figure this out? If not, please share the logs and any other details you have. --- ### Page: https://forum.scylladb.com/t/no-host-available/2053 Title: No Host Available - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Am getting following error, when I run queries using cqlsh. Please help me to resolve this error. 'NoHostAvailable - ('Unable to connect to any servers', {'10.xx.xx.xx:9042': error(24, "Tried connecting to [('10.xx.xx.x… Language: en Canonical URL: https://forum.scylladb.com/t/no-host-available/2053 ## Headings Structure: H1: No Host Available H3: Related topics ## Main Content: H1: No Host Available H3: Related topics Am getting following error, when I run queries using cqlsh. Please help me to resolve this error. 'NoHostAvailable - ('Unable to connect to any servers', {'10.xx.xx.xx:9042': error(24, "Tried connecting to [('10.xx.xx.xx', 9042)]. Last error: Too many open files")})'] (will try again later attempt 3 of 5) But I don’t see Too many open files error in logs. Try increasing the open-file limit on your server. --- ### Page: https://forum.scylladb.com/t/unknown-schema-in-advance-using-json/2054 Title: Unknown Schema in advance, using JSON - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/unknown-schema-in-advance-using-json/2054 ## Headings Structure: H1: Unknown Schema in advance, using JSON H3: Related topics ## Main Content: H1: Unknown Schema in advance, using JSON H3: Related topics Originally from the User Slack @Dhruv_garg: Is there a way to store JSON in one of the columns of scylladb table? we want to store some arbitrary data for which we won’t know the schema before hand @avi: Use a text column @Dhruv_garg: is storing in a blob also a good idea? @avi: It’s equivalent --- ### Page: https://forum.scylladb.com/t/release-scylla-operator-v1-12-2/2057 Title: [RELEASE] Scylla Operator v1.12.2 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of Scylla Operator 1.12.2 Scylla Operator 1.12.2 brings version update of Go and all dependencies. As with all of our releases, all API changes are backward compatib… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-operator-v1-12-2/2057 ## Headings Structure: H1: [RELEASE] Scylla Operator v1.12.2 H3: Related topics ## Main Content: H1: [RELEASE] Scylla Operator v1.12.2 H3: Related topics The ScyllaDB team is pleased to announce the release of Scylla Operator 1.12.2 Scylla Operator 1.12.2 brings version update of Go and all dependencies. As with all of our releases, all API changes are backward compatible. For more changes and details check out the GitHub release notes. Upgrade instructions Upgrading from v1.11.x or v1.12.1 with kubectl apply doesn’t require any extra action, just take the manifest from v1.12.2 tag and substitute the released image. Using helm requires a mandatory manual step for every release because helm can’t handle CRDs updates. For details, see our upgrade documentation. Regards, Scylla Operator Team --- ### Page: https://forum.scylladb.com/t/release-scylla-operator-v1-11-5/2058 Title: [RELEASE] Scylla Operator v1.11.5 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of Scylla Operator 1.11.5 Scylla Operator 1.11.5 brings version update of Go and all dependencies. As with all of our releases, all API changes are backward compatib… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-operator-v1-11-5/2058 ## Headings Structure: H1: [RELEASE] Scylla Operator v1.11.5 H3: Related topics ## Main Content: H1: [RELEASE] Scylla Operator v1.11.5 H3: Related topics The ScyllaDB team is pleased to announce the release of Scylla Operator 1.11.5 Scylla Operator 1.11.5 brings version update of Go and all dependencies. As with all of our releases, all API changes are backward compatible. For more changes and details check out the GitHub release notes. Upgrade instructions Upgrading from v1.10.x or v1.11.4 with kubectl apply doesn’t require any extra action, just take the manifest from v1.11.5 tag and substitute the released image. Using helm requires a mandatory manual step for every release because helm can’t handle CRDs updates. For details, see our upgrade documentation. Regards, Scylla Operator Team --- ### Page: https://forum.scylladb.com/t/modeling-transactional-behavior-modifying-two-tables/2059 Title: Modeling "transactional" behavior, modifying two tables - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello! Evaluating ScyallDB for PosstgreSQL replacement and I need help in modeling of “transactional” scenario in Scylladb. I have one table, kind of “resources”, and second table which holds information about user pai… Language: en Canonical URL: https://forum.scylladb.com/t/modeling-transactional-behavior-modifying-two-tables/2059 ## Headings Structure: H1: Modeling "transactional" behavior, modifying two tables H3: Related topics ## Main Content: H1: Modeling "transactional" behavior, modifying two tables H3: Related topics Evaluating ScyallDB for PosstgreSQL replacement and I need help in modeling of “transactional” scenario in Scylladb. I have one table, kind of “resources”, and second table which holds information about user paid quota, so user can create and store only N resources in resource table. I’m trying to model scenario when quota value incrementing/decrementing when resource created/deleted, but it is kind of untrivial. For now I came up with: So, is it idempotent workflow? Am I missing something? Is it correct usage of lwt? For now it seems to me it is ok, but I’m not sure, please advise. #transaction #atomic lwt acid Your flow looks reasonable, but you may want to have a history table to address some potential issues you’ve already outlined. See, for example, the transfers table in Getting the Most out of Lightweight Transactions in ScyllaDB - ScyllaDB . Another idea would be to know who locked the resource_quota table, as in case the process fails in the middle, you know the authoritative process able to resume the operation (if that’s something you want), rather than waiting for the TTL to expire and failing everything until then. One potential optimization is whether you really need resource_id to be part of the partition key on the resource table, and whether updated_at can be a non-key column. If you don’t, then you maybe don’t need used_value, and all you would have to do would be to simply acquire a resource_quota lock. Scan the resource’s partition in question (user_id, resource_type) and from there infer the user_id’s consumption. As you acquired the lock, other instances of your application know it can read the table as needed, but it can’t issue a read-before-write because there may be another update in progress. As you satisfy the constraints, then simply batch all updates to resource and release the lock from resource_quota. In the end it would end up with something like (note we eliminated one LWT write): Thus, after (2) above, you are ok with responding the client - you don’t need to wait for unlocking as you ain’t updating anything else (ie: used_value) in the resource_quota table. This also assumes that upserting values is ok for you use case, which AFAICT it looks doable. --- ### Page: https://forum.scylladb.com/t/data-modeling-question-choosing-the-partition-key-clustering-key-indexing-and-performance-impact/2060 Title: Data Modeling question, choosing the Partition Key, Clustering Key, Indexing and performance impact - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/data-modeling-question-choosing-the-partition-key-clustering-key-indexing-and-performance-impact/2060 ## Headings Structure: H1: Data Modeling question, choosing the Partition Key, Clustering Key, Indexing and performance impact H3: Related topics ## Main Content: H1: Data Modeling question, choosing the Partition Key, Clustering Key, Indexing and performance impact H3: Related topics Originally from the User Slack @Mosca_Careca: This might be a really easy and dumb question, but it’s my first time modeling with CQL so please keep that in mind Let’s say I have a Table about Books. Each book has certain attributes and one of those is the price. I want to be able to efficiently query by price. If I was querying by something like tags, I was taught that I could create a secondary table like “books_by_tag” where the PK is the tag name, instead of the book ID and each tag has the books associated with it, some sort of FK in SQL. The thing is, I cannot do that using a float. How would I go about querying by price? Is the only way available to use an index on the price attribute? If so, wouldn’t that be “bad” because it’s potentially scanning a large portion of the dataset? @Felipe_Cardeneti_Mendes: Normally price would be a clustering key of some partition grouping all entries you want to retrieve (ie: tag in your example). However, if you want to retrieve all items matching a specific price (this is a bit weird - I don’t see people or websites where you enter “give me all books which match $1 USD”), then an index is the way. @Mosca_Careca: So you’re saying usually it would be something like: where I could then query like so: ? Yeah sorry, index was not what I meant to say. Like you said, it makes no sense to use an index here, because I want price ranges and not exact pricing. @Felipe_Cardeneti_Mendes: > PRIMARY KEY (book_id, price) No, this isn’t good, because a book_id uniquely identify a book. And thus it can have only one price at a time. You probably want to follow the books_by_tag approach, where you have: And this will allow you to retrieve books from within a category where the price matches a range: SELECT book_id FROM books_by_category WHERE category_id = ? AND price >= ? AND price <= ?; @Mosca_Careca: Ok makes sense. In my previous approach I would be creating one partition per book right? I understand the approach you’re taking, I just don’t understand why that’s needed. I can see that adding a category to encompass larger groups of books is good. Could I add an “all” category and query from there? I also have a question about how much redundancy should be present. Say I have the books table and the books_by_tag table. Is it a good idea to have all the info of a book both in the books table and in the books_by_tag table or is it enough to have just the book_id of a book in the books_by_tag table? @Felipe_Cardeneti_Mendes: Actually, the PRIMARY KEY would be PRIMARY KEY(category_id, book_price, book_id) – as there can be books falling under the same price. But this creates the need of delete/insert books whenever a price changes. So there are some tradeoffs even with that approach. Another way would be to simply get rid of the book_price and do an ALLOW FILTERING instead if the group is small (as it should be). > Could I add an “all” category and query from there? You can, but the efficiency of it will depend on how many books to scan. Why not break all into smaller queries per group? For your last question, you probably want the minimum data you need in auxiliary tables. For example, you don’t need to hold the book description on books_by_tag, because ultimately you would present that data to the end user only after he clicked on the book. However, you may want to present the path for a thumbnail, as the usual case for retrieving a price range involves listing You can also make it as small as possible and for every book_id matched you simply read from the main table, there’s room for you to play on what works best for you @Mosca_Careca: Uhm but then, on your approach you’re assuming that the categories are finite, which is not the case for user input categories (like in my case, I have user input tags). I still don’t totally understand PRIMARY KEYS in CQL. I have to admit that I didn’t understand much of this last approach you said. I’ll have to investigate further and get back to it later. > You can, but the efficiency of it will depend on how many books to scan. Why not break all into smaller queries per group? I have just shy of 1M entries in that table. Wouldn’t breaking into smaller groups force me to execute multiple queries, one per tag or category and then unite the result sets? Is that more efficient than a big query? > You can also make it as small as possible and for every book_id matched you simply read from the main table, there’s room for you to play on what works best for you A gain could be made by saving the need to have this secondary query, but then the schema wouldn’t be as “universal” @Felipe_Cardeneti_Mendes: ah 1M should be ok. --- ### Page: https://forum.scylladb.com/t/steps-to-have-the-scylladb-on-plain-ec2-right-from-the-download-installation-configuration-and-best-practices/2061 Title: Steps to have the scyllaDB on plain ec2 right from the download, installation,configuration and best practices - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello, i want to lauch scylladb on plain ec2 right from the download, installation,configuration and best practices.suppose i have 3 instance 172.0.0.1,172.0.0.2,172.0.0.3. give me proper scylla.yaml file or cassandra-ra… Language: en Canonical URL: https://forum.scylladb.com/t/steps-to-have-the-scylladb-on-plain-ec2-right-from-the-download-installation-configuration-and-best-practices/2061 ## Headings Structure: H1: Steps to have the scyllaDB on plain ec2 right from the download, installation,configuration and best practices H3: Related topics ## Main Content: H1: Steps to have the scyllaDB on plain ec2 right from the download, installation,configuration and best practices H3: Related topics Hello, i want to lauch scylladb on plain ec2 right from the download, installation,configuration and best practices.suppose i have 3 instance 172.0.0.1,172.0.0.2,172.0.0.3. give me proper scylla.yaml file or cassandra-rackdc.properties file with proper configuration . thank you --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-48-2024-05-25/2064 Title: Last week in scylla-cluster-tests.git master (issue #48; 2024-05-25) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 3b8467a1…fe35c3ad range are covered. There were 24 non-merge commits from 6 authors in tha… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-48-2024-05-25/2064 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #48; 2024-05-25) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #48; 2024-05-25) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 3b8467a1…fe35c3ad range are covered. There were 24 non-merge commits from 6 authors in that period. Some notable commits: We removed deprecated system_init property in Distro class, cleaning up the codebase. Fixed issue with handling ipv6 addresses across codebase enabling IPv6 tests. SCT sends JUnit results as xml files to Argus. We no longer need to update major distros images (e.g. AMI’s on AWS) manually. Latest of specific types are automatically resolved. Ubuntu 24.04 is now supported as the basis for Scylla images. Manager jenkins pipeline uses provision stage speeding up the test setup time and becoming in compliance to all other tests. Verification if Scylla’s node is up is done with higher frequency, shortening 6-node cluster preparation by around 4 minutes (for recent ScyllaDB versions). Now test setup time is bound by the time it takes to configure monitor and loader nodes. Adapted submitting parsed scylla version to argus to include build_version field used by Argus. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/how-do-i-change-the-ttl-of-an-existing-table-and-delete-old-data/2066 Title: How do I change the TTL of an existing Table and delete old data? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/how-do-i-change-the-ttl-of-an-existing-table-and-delete-old-data/2066 ## Headings Structure: H1: How do I change the TTL of an existing Table and delete old data? H3: Related topics ## Main Content: H1: How do I change the TTL of an existing Table and delete old data? H3: Related topics Originally from the User Slack @Guillaume_Lorin: Hello everyone! I have a table setup with TWCS compaction and a table default TTL of 1 year. I want to reduce the TTL on this table to 1 month, and right away delete all the previously-inserted-data that is older than that to save disk space. So I understand from the documentation that data written before the change is not affected by the new TTL. And running DELETE is not recommended with TWCS. So I was wondering if there is a recommended strategy to delete it safely? Or should I just start a new table ? @Felipe_Cardeneti_Mendes: You can also change the compaction strategy temporarily. But given the problem at hand I guess I would just create a new table and switch past your TTL period --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-231-2024-05-26/2067 Title: Last week in scylladb.git master (issue #231; 2024-05-26) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the c93a7d2664…de798775fd range are covered. There were 131 non-merge commits from 18 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-231-2024-05-26/2067 ## Headings Structure: H1: Last week in scylladb.git master (issue #231; 2024-05-26) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #231; 2024-05-26) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the c93a7d2664…de798775fd range are covered. There were 131 non-merge commits from 18 authors in that period. Some notable commits: Tablet replicas can now be migrated within nodes, from shard to shard, not just from one node to another. This is important when a small table grows and needs to use all vcpus of a node. DESCRIBE statements will now sort CREATE TYPE statements in dependency order, so that if type A is used in type B, CREATE TYPE A will appear before CREATE TYPE B. This is useful for restoring schemas. ScyllaDB maintains two commitlogs, one for data and one for metadata. The metadata commitlog is now performance-isolated from the data commitlog. This maintains responsiveness in a heavily loaded cluster. The version number was bumped from 5.5 to 6.0, indicating the next release will be called 6.0. It was then bumped further to 6.1, indicating that 6.0 was branched. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/batch-performance-considerations-when-grouping-queries-on-the-client-side/2076 Title: Batch performance considerations, when grouping queries on the client side - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/batch-performance-considerations-when-grouping-queries-on-the-client-side/2076 ## Headings Structure: H1: Batch performance considerations, when grouping queries on the client side H3: Related topics ## Main Content: H1: Batch performance considerations, when grouping queries on the client side H3: Related topics Originally from the User Slack @Artem_Golovko: Hello everyone. I would like to ask about batch performance consideration. Are there any performance improvements from grouping several INSERT statements with the same PK and same table together with UNLOGGED batch on the client side? @dor: Probably, since it will go to the same replicas anyway. Write performance in general is very good, not sure you need to over think it @Artem_Golovko: @dor if you will have the following structure table1: ((key), f1, f2), value table2: ((key), f2, f1), value so absolutely the same tables, but first one allow you to get value based on the f1, while the second one allow you to get value based on the f2. And you need 2 inserts for every record. My idea was to merge these two tables into the one table: common_table: ((key), type, f1, f2), value now I can use a batch insert: BEGIN BATCH INSERT key, false, f1, f2, value INSERT key, true, f2, f1, value END BATCH @dor: Oh, two separate tables, that wouldn’t improve performance. Each table has its own set of sstables ,caches, etc --- ### Page: https://forum.scylladb.com/t/scylladb-process-is-more-than-100-utilized-even-it-is-ideal/2079 Title: Scylladb process is more than 100% utilized even it is ideal - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello, We are facing high CPU utilization by Scylladb process even if it is ideal. And when we put a small load CPU jump to 400% and more. Is there a way we can trace or enable advance logging to get this issue resolve… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-process-is-more-than-100-utilized-even-it-is-ideal/2079 ## Headings Structure: H1: Scylladb process is more than 100% utilized even it is ideal H3: Related topics ## Main Content: H1: Scylladb process is more than 100% utilized even it is ideal H3: Related topics We are facing high CPU utilization by Scylladb process even if it is ideal. And when we put a small load CPU jump to 400% and more. Is there a way we can trace or enable advance logging to get this issue resolve. Where do you see this high CPU utilization? In top or similar? Note that checking ScyllaDB’s resource utilization in external tools is misleading. These external tools will report high CPU and memory utilization, even when the ScyllaDB is lightly loaded. To get a real picture of how much is ScyllaDB loaded, setup monitoring and check the load and memory metrics there. Thanks for your reply. I am also surprise with the output from top command. Load average is 1.07, 1.65, 1.67 but in CPU% it says 146.5. Memory utilization is 47GB used as we configured it 48GB max. I will try to setup monitoring. If you can suggest anything will be appreciated. top - 11:37:01 up 2 days, 23:23, 1 user, load average: 1.07, 1.65, 1.67 Tasks: 406 total, 1 running, 405 sleeping, 0 stopped, 0 zombie %Cpu(s): 2.7 us, 1.1 sy, 0.0 ni, 96.2 id, 0.0 wa, 0.0 hi, 0.0 si, 0.0 st MiB Mem : 64087.2 total, 10160.3 free, 48765.5 used, 5161.3 buff/cache MiB Swap: 32735.0 total, 32735.0 free, 0.0 used. 13244.8 avail Mem 10251 scylla 20 0 16.0t 46.7g 109392 S 146.5 74.6 9721:10 scylla I will try to setup monitoring. If you can suggest anything will be appreciated. We have our own monitoring stack (based on widely used open-source software), see ScyllaDB Monitoring Stack | ScyllaDB Docs for more information. Thank you Botond, I will try ScyllaDB Monitoring Stack I am also surprise with the output from top command. Load average is 1.07, 1.65, 1.67 but in CPU% it says 146.5. Memory utilization is 47GB used as we configured it 48GB max. I will try to setup monitoring. If you can suggest anything will be appreciated. top - 11:37:01 up 2 days, 23:23, 1 user, load average: 1.07, 1.65, 1.67 Tasks: 406 total, 1 running, 405 sleeping, 0 stopped, 0 zombie %Cpu(s): 2.7 us, 1.1 sy, 0.0 ni, 96.2 id, 0.0 wa, 0.0 hi, 0.0 si, 0.0 st MiB Mem : 64087.2 total, 10160.3 free, 48765.5 used, 5161.3 buff/cache MiB Swap: 32735.0 total, 32735.0 free, 0.0 used. 13244.8 avail Mem It does seem unusual to have such a high CPU utilization reported by top while the load average remains relatively low. --- ### Page: https://forum.scylladb.com/t/should-i-use-list-float-or-blob-as-a-data-type-for-data-that-doesnt-mutate-for-reducing-storage-cost/2080 Title: Should I use list or blob as a data type for data that doesn't mutate, for reducing storage cost? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/should-i-use-list-float-or-blob-as-a-data-type-for-data-that-doesnt-mutate-for-reducing-storage-cost/2080 ## Headings Structure: H1: Should I use list or blob as a data type for data that doesn't mutate, for reducing storage cost? H3: Related topics ## Main Content: H1: Should I use list or blob as a data type for data that doesn't mutate, for reducing storage cost? H3: Related topics Originally from the User Slack @Evaldas_Buinauskas: Hello! I would like to use scylla as an intermediate store in our data engineering pipeline where we need to keep search index up to date, very similar project to https://doordash.engineering/2021/07/14/open-source-search-indexing/ One thing that we need to store is clip image embeddings which are represented as a collection of floats. Should I store them as a list or as a blob type? I’ll never need to mutate the list and would like to save on performance and reduce storage cost. I am also fine handling the float ↔ byte conversion at the code level. @Botond_Dénes: If you don’t need to ever mutate this list after writing it, I recommend either frozen> or blob. Frozen collections are much more lightweight at the storage level, they are essentially treated as a blob, with some added convenience for users. @Evaldas_Buinauskas: Would I still be able to set it to null when I need to delete it? @Evaldas_Buinauskas: cool, let me try that, thanks! --- ### Page: https://forum.scylladb.com/t/authenticator-ldap-and-user-password-together/2089 Title: Authenticator LDAP and user/password together - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Good afternoon, does anyone know if there is a way to use LDAP authentication and user/password in the Enterprise version? Only one or the other works, not both together authenticator: PasswordAuthenticator authentica… Language: en Canonical URL: https://forum.scylladb.com/t/authenticator-ldap-and-user-password-together/2089 ## Headings Structure: H1: Authenticator LDAP and user/password together H3: Related topics ## Main Content: H1: Authenticator LDAP and user/password together H3: Related topics Good afternoon, does anyone know if there is a way to use LDAP authentication and user/password in the Enterprise version? Only one or the other works, not both together authenticator: PasswordAuthenticator authenticator: com.scylladb.auth.SaslauthdAuthenticator I tested it like this and it didn’t work either authenticator: PasswordAuthenticator, com.scylladb.auth.SaslauthdAuthenticator Hello! We don’t have such possibility yet, there is a request for implementing it: [RFE] Support Multiple Authenticators · Issue #18232 · scylladb/scylladb · GitHub but it’s not planned yet. --- ### Page: https://forum.scylladb.com/t/scylladb-tokio-parallel-tasks-concurrency-and-optimizing-the-performance/2094 Title: ScyllaDB, Tokio, parallel tasks, concurrency, and optimizing the performance - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-tokio-parallel-tasks-concurrency-and-optimizing-the-performance/2094 ## Headings Structure: H1: ScyllaDB, Tokio, parallel tasks, concurrency, and optimizing the performance H3: Related topics ## Main Content: H1: ScyllaDB, Tokio, parallel tasks, concurrency, and optimizing the performance H3: Related topics Originally from the User Slack @Jimmy: hi there i saw the presentation of storing telemetry in syclladb, (https://www.youtube.com/watch?v=wZ9xc5LnsB80 storing 80k entries in about a second is very impressive, my question is are 80k tokio tasks built right? each task handling 1 insertion, is this generates overhead despite being massive parallel? would not make sense to batch the insertion a bit like assigning 100 (or X number) entries to 1 tasks or that is not needed at all? ty how u optimize the calls to the db, cc: @Felipe_Cardeneti_Mendes @Felipe_Cardeneti_Mendes: No, there aren’t 80k concurrent tasks. There’s a limit imposed by a semaphore. In any case, it’s been quite a while I don’t touch that code, happy to see people find it useful still. To your question, tokio is very lightweight on CPU, so handling several tasks in parallel isn’t typically a problem. What you want to avoid, however, is to end up with high concurrency on the database - where many queries are overwhelming what a single shard can handle. You could assign a pool of workers and submit requests to then, but the code in question was meant to be easy for users to understand. Batching in this specific case could work as we are ingesting several rows per device. You shouldn’t be batching across different partitions though, as it is less efficient. --- ### Page: https://forum.scylladb.com/t/why-flush-hints-when-gc-mode-is-repair/2100 Title: Why flush_hints when gc_mode is repair? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I saw in repair.cc that flush_hints are required when tombstone_gc_mode is repair. When writing fails, request data will be recorded in the hints directory. Does flush_hints flush these records to disk? Language: en Canonical URL: https://forum.scylladb.com/t/why-flush-hints-when-gc-mode-is-repair/2100 ## Headings Structure: H1: Why flush_hints when gc_mode is repair? H3: Related topics ## Main Content: H1: Why flush_hints when gc_mode is repair? H3: Related topics I saw in repair.cc that flush_hints are required when tombstone_gc_mode is repair. When writing fails, request data will be recorded in the hints directory. Does flush_hints flush these records to disk? Flushing hints and batchlog when repairing with tombstone_gc is required to prevent any data in either causing data resurrection. With tombstone_gc, any tombstone that was written before the last repair, can be garbage-collected, so it is important to include any data that is in the batchlog or in hints in the repair, so any tombstone that might shadow them can take effect. Specifically, to ensure that no data resurrection occurs in which scenario? I found that not performing the flush_hints operation has no effect, as shown in the following figure. Because in order to clear the tombstone, all replicas must participate in the repair. So even without flush_hints, data that fails to be written will still be repaired, including tombstones. Hints can contain data which is deleted by a tombstone. Repair runs, the tombstone is GC’d, then hints are replayed and data is resurrected. This is why hints have to be replayed before repair. If coordinator (A) fails to write a=1 to node B, then A records the hints. After a period of time, there is a request to delete a=1 and write it. If the deletion fails, hints will also be remembered. And then the unified flush_hints have no effect. If the write is successful, it means B is alive, and the hints will also be sent over. If the write is successful, it means B is alive, and the hints will also be sent over. This won’t have any impact, right? So why do we have to flushvints? Flushing hints is a safety mechanism, in case some hints linger on the nodes. Normally, there wouldn’t be any, this is to cover corner cases. this is to cover corner cases. I can’t figure it out no matter what. Can you give me a simple example? Thank you. In other words, what are the consequences of not executing flush_hints? How is data resurrected? I can’t figure it out no matter what. Can you give me a simple example? Thank you. In other words, what are the consequences of not executing flush_hints? How is data resurrected? Taking your example from: If coordinator (A) fails to write a=1 to node B, then A records the hints. After a period of time, there is a request to delete a=1 and write it. If the deletion fails, hints will also be remembered. And then the unified flush_hints have no effect. If the write is successful, it means B is alive, and the hints will also be sent over. If the write is successful, it means B is alive, and the hints will also be sent over. This won’t have any impact, right? So why do we have to flushvints? Node (A) may miss the delete a. It may not notice (B) becoming online. Sending the hints may fail. These are unlikely but it is very hard to prove they cannot happen. So we usually program defensively and want to make sure there are no hints left on any node after a repair, so data resurrection is off the table. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-49-2024-05-31/2105 Title: Last week in scylla-cluster-tests.git master (issue #49; 2024-05-31) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 9c59cc58…da5f0d43 range are covered. There were 29 non-merge commits from 6 authors in tha… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-49-2024-05-31/2105 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #49; 2024-05-31) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #49; 2024-05-31) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 9c59cc58…da5f0d43 range are covered. There were 29 non-merge commits from 6 authors in that period. Some notable commits: SCT logs are now printing hostnames for each RemoteCommandRunner output line, which helps differentiate and filter log lines for specific nodes. We introduced Amazon Linux 2023 tests as Amazon Linux 2 will be deprecated next year, so we are beginning support for the next version. Another speed improvement for database node setup was achieved by moving syslogng-exporter installation to cloud-init. Enabled testing of Kafka connectors with a Kafka cluster setup on sct-runner instance using Docker Compose. With a Kafka consumer thread, the test reads data written by the connector and validates the received information, starting with row count. Initial pipelines for the basic case were added, using Docker and AWS backends. A small improvement for the apt packages installer to use a built-in feature to wait for locks. When topology operations are aborted by SCT request, the severity of expected raft_topology error messages are decreased to a warning. The large partition longevity test stress commands were refactored to simplify and behave correctly when triggered with a custom stress duration. A new longevity test with RF=1 was added to cover issues observed in the field. Rolling restart now waits for all nodes to be in UN state before proceeding to the next node. Continued adaptation of Mergify to automate the backporting process within SCT. Report screenshots are now created with Grafana Image Renderer, greatly improving the speed (from 9-10 minutes to 2.5 minutes). Lots of code could be dropped and selenium removed from the Hydra image. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/why-bootstrap-on-new-server-with-existing-data-folder/2106 Title: Why bootstrap on new server with existing data folder? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, I replaced the hardware of a scylla server yesterday. The complete data folder was backed up and put on the new server. Nevertheless scylla was complaining about failed bootstrap because the IP already exists. Do… Language: en Canonical URL: https://forum.scylladb.com/t/why-bootstrap-on-new-server-with-existing-data-folder/2106 ## Headings Structure: H1: Why bootstrap on new server with existing data folder? H3: Related topics ## Main Content: H1: Why bootstrap on new server with existing data folder? H3: Related topics Hi, I replaced the hardware of a scylla server yesterday. The complete data folder was backed up and put on the new server. Nevertheless scylla was complaining about failed bootstrap because the IP already exists. Does scylla read out some kind of hardware-id to detect such situations? I fail to understand why scylla did not just start up - like in a regular restart. How did it know its a new server? If you haven’t replaced the node by another node with the same IP, and you have backed up and restored all keyspaces, including the system keyspace, the node should have indeed be able to rejoin the cluster, as the node’s host_id is kept in system.local. We’ll probably need to look at the logs and try to figure out what’s the reason for that. Best if you could open an issue on github to collect the logs. If you haven’t replaced the node by another node with the same IP The only thing I can imagine is that the old node was started back up after the snapshot was taken. Additionally the new node was started with scylla 5.2, while before it was 5.1. I would expect that should make no difference. There is not enough information in the log excerpt about the node’s host_id. We’s also need nodetool gossipinfo and/or logs from the other nodes to see what they think this endpoint’s HOST_ID is. @kbr anything else? Does scylla read out some kind of hardware-id to detect such situations? I fail to understand why scylla did not just start up - like in a regular restart. How did it know its a new server? Maybe it tried to use a different data directory due to a different scylla.yaml configuration file. Did you copy the config file as well? Also, there are multiple data directories and a commitlog directory. Make sure you copied all of them. The fact that you upgraded Scylla version might also be problematic. Do one thing at a time – either upgrade version or hardware, not both at the same time. If you want to upgrade from 5.1 to 5.2 then follow the documented rolling upgrade procedure, keeping the hardware static. All of that said, I haven’t found whatever you’re doing as a documented procedure in our docs, so you’re doing something that is not generally supported by Scylla, so we don’t test it, don’t be surprised if it doesn’t work. --- ### Page: https://forum.scylladb.com/t/data-modeling-question-with-a-high-consistency-requirement-using-lwt-mv-si-or-frozen-maps/2108 Title: Data Modeling question with a high consistency requirement, using LWT, MV, SI or frozen maps? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/data-modeling-question-with-a-high-consistency-requirement-using-lwt-mv-si-or-frozen-maps/2108 ## Headings Structure: H1: Data Modeling question with a high consistency requirement, using LWT, MV, SI or frozen maps? H3: Related topics ## Main Content: H1: Data Modeling question with a high consistency requirement, using LWT, MV, SI or frozen maps? H3: Related topics Originally from the User Slack @XH_L: Hello everyone! I am considering using ScyllaDB as the only database, but I have encountered some consistency issues. Could anyone point me in the right direction or provide a better solution? Thanks! For instance, I want to make email and phone be unique, similar to the UNIQUE in SQL. I noticed that ScyllaDB has a LightWeight Transaction (LWT) IF NOT EXISTS. Based on this, I have made some attempts: • Adding email and phone to the PRIMARY KEY, then using INSERT ... IF NOT EXISTS ◦ It checks the combination of the entire primary key ◦ This conflicts with my desired query to find the password using email or phone • Creating two new tables which use email or phone as the PRIMARY KEY, then performing BATCH INSERT ... IF NOT EXISTS on both tables atomicity ◦ Queries need access multi tables ◦ The primary key could lead to unbalanced distribution ◦ BATCH cannot perform transactional insertions on different tables, and separating the insertions could lead to new consistency issues By the way, I noticed that cassandra plan to support ACID, do we have similar plans? https://thenewstack.io/acid-transactions-change-the-game-for-cassandra-developers/ The New Stack: ACID Transactions Change the Game for Cassandra Developers @Felipe_Cardeneti_Mendes: Based on your description it seems like you want to support retrieving a password for either email or phone inputs Why don’t you use a view or an index then? If the input is an email you read from the main table. If it’s a phone you read from the view . Both should return the same password From your proposed data model your partition is ID, but if you already know the id early in that stage then this question is irrelevant. Perhaps you should find the email/phone first and discover the ID to support other functions. @XH_L: Thanks for your reply! Yes, mv or si can implement my queries, but my main concern is how to ensure email and phone are globally unique and keep my queries work. This is my final attempt, It work, but probably not a good practice: There are still some problems with applcation level atomic insert. It is possible that the phone insertion is successful but the email insertion fails. At this time, it needs to be deleted the phone. I think it is too complicated and it is estimated that it will not have good distribution and performance. @Felipe_Cardeneti_Mendes: Maybe a frozen map then? Email and phone are authentication methods, and you perhaps may want to support others in the future. A method maps to an id. It is unique and all under a single table. What Id do would be to require gating of these methods. One must confirm their identity to each one. At the same time, they can register with either one— or even both. But we require just one method initially, and they are free to add others (and confirm they own it) later on @XH_L: I’ll look into forzen map, thanks for the suggestion! --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-232-2024-06-02/2111 Title: Last week in scylladb.git master (issue #232; 2024-06-02) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the de798775fd…5b4e688668 range are covered. There were 113 non-merge commits from 16 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-232-2024-06-02/2111 ## Headings Structure: H1: Last week in scylladb.git master (issue #232; 2024-06-02) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #232; 2024-06-02) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the de798775fd…5b4e688668 range are covered. There were 113 non-merge commits from 16 authors in that period. Some notable commits: Previously, we moved the tables backing authentication to be under Raft management, and moved them from the system_auth keyspace to the system_auth_v2 keyspace. To simplify things further, the authentication tables are moved to the system keyspace. It is now possible to change the replication factor of keyspaces using tablets. The tablets feature has been marked as generally available, and removed from the list of experimental features. It is now available by default. Note it still has limitations in interacting with other features. The memory allowance for bloom filters has been increased from 10% to 20%, recognizing that some workloads need more memory for bloom filters. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-6-0-rc1/2112 Title: [RELEASE] ScyllaDB 6.0 RC1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce ScyllaDB Open Source 6.0 RC1, the first Release Candidate for the ScyllaDB Open Source 6.0 major release. ScyllaDB 6.0 introduces two major features which change the way ScyllaDB… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-6-0-rc1/2112 ## Headings Structure: H1: [RELEASE] ScyllaDB 6.0 RC1 H2: New features H3: Tablets H3: Strongly Consistent Topology Updates H3: Strongly Consistent Auth Updates H3: Strongly Consistent Service Levels H3: Describe Schema with Internals H3: Native Nodetool H3: Removing the JMX Server H3: Maintenance Mode H3: Maintenance Socket H3: Deployment H2: Improvements H3: Bloom Filters H3: Stability and performance H3: Guardrails H3: Replication strategies Guardrails can be used to deny SimpleStratagy strategy, for example, which is not recommended for production. H3: CQL H3: Alternator H3: Images and packaging H3: Tooling and REST API H3: Configuration H3: Tracing H3: Monitoring H3: Deprecated and removed features H3: Bug fixes and stability H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 6.0 RC1 H2: New features H3: Tablets H4: Using Tablets H4: Procedures H4: Monitor Tablets H4: Driver Support H3: Strongly Consistent Topology Updates H3: Strongly Consistent Auth Updates H3: Strongly Consistent Service Levels H3: Describe Schema with Internals H3: Native Nodetool H3: Removing the JMX Server H3: Maintenance Mode H3: Maintenance Socket H3: Deployment H2: Improvements H3: Bloom Filters H3: Stability and performance H4: Compaction Related H4: Commitlog Related H4: Cluster Operation Related H4: Materialized Views Related H4: Performance Related H4: All kinds of edge cases H3: Guardrails H3: Replication strategies Guardrails can be used to deny SimpleStratagy strategy, for example, which is not recommended for production. H3: CQL H3: Alternator H3: Images and packaging H3: Tooling and REST API H3: Configuration H3: Tracing H3: Monitoring H3: Deprecated and removed features H3: Bug fixes and stability H3: Related topics The ScyllaDB team is pleased to announce ScyllaDB Open Source 6.0 RC1, the first Release Candidate for the ScyllaDB Open Source 6.0 major release. ScyllaDB 6.0 introduces two major features which change the way ScyllaDB works: In addition, ScyllaDB 6.0 includes many other improvements in functionality, stability, UX and performance. We encourage you to run ScyllaDB 6.0 release candidates on your test environments; this will help ensure that an upgrade to ScyllaDB 6.0 General Availability will proceed smoothly with your workload. Only the last two minor releases of the ScyllaDB Open Source project are supported. Once ScyllaDB Open Source 6.0 is officially released, only ScyllaDB Open Source 6.0 and ScyllaDB 5.4 will be supported, and ScyllaDB 5.2 will be retired. Get ScyllaDB Open Source 6.0 as binary packages (RPM/DEB), AWS AMI, GCP Image and Docker image Upgrade from ScyllaDB 5.4 to ScyllaDB 6.0 In this release, ScyllaDB enabled Tablets, a new data distribution algorithm to replace the legacy vNodes approach inherited from Apache Cassandra. While the vNodes approach statically distributes all tables across all nodes and shards based on the token ring, the Tablets approach dynamically distributes each table to a subset of nodes and shards based on its size. In the future, distribution will use CPU, OPS, and other information to further optimize the distribution. In particular, Tablets provide the following: Read more about Tablets here. Tablets are enabled by default for new clusters. You can set the initial number of Tablets per table using the “initial” parameter: CREATE KEYSPACE … WITH TABLETS = { ‘initial’: 1 } Or disable Tablets for the keyspace with the “‘enabled’ parameter: CREATE KEYSPACE … WITH TABLETS = { ‘enabled’: false } (See CREATE KEYSPACE docs) All tables created in this Keyspace will use Tablets by default. In this release, you can not use the following features with Tablets: If you are planning to use one of these features, disable Tablets when creating the Keyspace. Note: you can not ALTER an existing Keyspace to switch between Tablets and vNode based table and back. We will remove these restrictions in upcoming patches and minor releases. With Tablets, the Replication Factor (RF) cannot be updated to a value higher than the number of nodes per Data Center (DC). This feature protects the Admin from setting an impossible-to-support RF. This affects the following operations: Node Decommission / Remove Starting from 6.0, you cannot decommission or remove a node if the resulting number of nodes would be smaller than the largest non-zero replication factor (for any keyspace) in this DC 1 DC, 5 nodes, a KS with RF=5 The decommission request will fail CREATE and ALERT a Keyspace You can create a KS with an RF that is greater than the number of nodes, but you cannot create a Table in this KS until you add nodes to match the RF. You can not alter the RF of a KS with Tablets, to be less than the available nodes. To Monitor Tablets in real time, upgrade ScyllaDB Monitoring Stack to release 4.7, and use the new dynamic Tablet panels, below. The Following Drivers support Tablets Legacy ScyllaDB and Apache Cassandra drivers will continue to work with ScyllaDB but will be less efficient when working with tablet-based Keyspaces. With Raft-managed topology enabled, all topology operations are internally sequenced consistently. A centralized coordination process ensures that topology metadata is synchronized across the nodes on each step of a topology change procedure. This makes topology updates fast and safe, as the cluster administrator can trigger many topology operations concurrently, and the coordination process will safely drive all of them to completion. For example, multiple nodes can be bootstrapped concurrently, which couldn’t be done with the previous gossip-based topology. Strongly Consistent Topology Updates is now the default for new clusters, and should be enabled after upgrade for existing clusters. System-auth-2 is a reimplementation of the Authentication and Authorization systems in a strongly consistent way on top of the Raft sub-system. This means that Role-Based Access Control (RBAC) commands like create role or grant permission are safe to run in parallel without a risk of getting out of sync with themselves and other metadata operations, like schema changes. As a result, there is no need to update system_auth RF or run repair when adding a DataCenter. Service Levels allow you to define attributes like timeout per workload. Service levels are now strongly consistent using Raft, like Schema, Topology and Auth. Until this release, CQL DESCRIBE SCHEMA was not sufficient to do a full schema restore from backup. For example, it lacks information about dropped columns. In 6.0, the DESC SCHEMA WITH INTERNALS command provides more information, streamlining the restore process. The nodetool utility provides simple command-line interface operations and attributes. ScyllaDB inherited the Java based nodetool from Apache Cassandra. In this release, the Java implementation was replaced with a backward-compatible native nodetool. The native nodetool works much faster. Unlike the Java version ,the native nodetool is part of the ScyllaDB repo, and allows easier and faster updates. With the Native Nodetool (above), the JMX server has become redundant and will no longer be part of the default ScyllaDB Installation or image. If you are using the JMX server directly (not via nodetool): Related issues: #15588 #18566 #18472 #18566 Maintenance mode is a new mode in which the node does not communicate with clients or other nodes and only listens to the local maintenance socket and the REST API. It can be used to fix damaged nodes – for example, by using nodetool compact or nodetool scrub. In maintenance mode, ScyllaDB skips loading tablet metadata if it is corrupted to allow an administrator to fix it. The Maintenance Socket provides a new way to interact with ScyllaDB from within the node it runs on. It is mainly for debugging. You can use CQLSh with the Maintenance Socket as described in the Maintenance Socket docs. #16172 The following is a list of improvements and bug fixes included in the release, grouped by domain. Bloom filters are used to determine which SStables do not contain a partition key, speeding up reads when SStables can be filtered out. Since the Bloom filters are held in memory, and their size depends on the data (small partitions require larger Bloom filters), there is a tradeoff between allocated memory, and the risk of OOM and the filter efficiency. The following improvements were made to Bloom filters in this release: Topology changes, Repairs, etc Guardrails is a framework to protect ScyllaDB users and admins from common mistakes and pitfalls. In this release the following Guardrails are added: Two new configurations replication_strategy_warn_list replication_strategy_fail_list Replace the old restrict_replication_simplestrategy and give more granularity for DB Admin to warn, or block non production strategies. Scylla REST API is now documented! (beta) The sstable validation tools, scylla sstable validate-checksum and scylla sstable validate, now returns output in json format. Admin API: a new API for asynchronous compaction: /tasks/compaction/keyspace_compaction/{keyspace} Similar to the existing synchronos storage_service API. Scylla Monitoring Stack released 4.7 and later supports ScyllaDB 6.0. See metrics update between 5.4 and 6.0 here, as well as the new, beta, metrics reference here. More monitoring related updates: This is a mirror for CASSANDRA-13910. See git log for a full (long) list of fixed issues. --- ### Page: https://forum.scylladb.com/t/running-on-multiple-data-centers-in-different-locations-latency-and-performance-impact-and-async-replication/2114 Title: Running on multiple data centers in different locations, latency (and performance) impact and async replication - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/running-on-multiple-data-centers-in-different-locations-latency-and-performance-impact-and-async-replication/2114 ## Headings Structure: H1: Running on multiple data centers in different locations, latency (and performance) impact and async replication H3: Related topics ## Main Content: H1: Running on multiple data centers in different locations, latency (and performance) impact and async replication H3: Related topics Originally from the User Slack @TO: Hi all, need some thoughts or recommendations. I want to run a three node of ScyllaDB cluster with each note in different data centres (geolocations). The data centres are about 30ms to 90ms apart. Is this good, or will this set-up be potentially problematic? @Felipe_Cardeneti_Mendes: Well, it is not common as you can imagine. I guess it is up to you to say whether this will be problematic or not, as you seem to already know beforehand the latency penalty for quorum and the risk in lower consistency levels. @TO: Thanks @Felipe_Cardeneti_Mendes, I am entirely new to ScyllaDB. In this case, how can the risks be managed in terms of quorum and consistency? Does Scylla support async replication? Also, what would have been the best practise for multi datacentres (geolocations)? @Felipe_Cardeneti_Mendes: ScyllaDB can be configured to fulfill the topology you want. For example, you could have a multi-region cluster where each node is logically placed in a different region. This means that quorum queries will need to traverse regions and will impact latency. To avoid a latency impact, you could reduce the consistency level to ONE (or local_one), but the failure of a single node could cause data loss. So you need to understand these trade-offs for this sort of topology. We do support async replication, and the best practice for multi-region clusters is to have at least 3 replicas in a primary region, potentially the same in other regions, but depending on your usage can be refined. @TO: Thanks @Felipe_Cardeneti_Mendes for the insight. --- ### Page: https://forum.scylladb.com/t/backup-and-restore-to-local-storage/2117 Title: Backup and Restore to Local storage - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: A few questions surrounding backups with ScyllaDB. is there any way to configure scylla manager to back up to local storage? (Or a network share). All of the documentation seem to point to an S3 bucket or azure/gcp ba… Language: en Canonical URL: https://forum.scylladb.com/t/backup-and-restore-to-local-storage/2117 ## Headings Structure: H1: Backup and Restore to Local storage H3: Related topics ## Main Content: H1: Backup and Restore to Local storage H3: Related topics A few questions surrounding backups with ScyllaDB. is there any way to configure scylla manager to back up to local storage? (Or a network share). All of the documentation seem to point to an S3 bucket or azure/gcp based options. If you have a 3 node cluster with RF3 do you still need to backup and restore all nodes via snapshot? or Can you use a snapshot from one node since it has all data on it (due to rf3). is there any way to configure scylla manager to back up to local storage? (Or a network share). All of the documentation seem to point to an S3 bucket or azure/gcp based options. Yes, you can configure the backup on Minio. See this Setup S3 compatible storage | ScyllaDB Docs. You need to configure Minio the way that it stores to local filesystem then. If you have a 3 node cluster with RF3 do you still need to backup and restore all nodes via snapshot? or Can you use a snapshot from one node since it has all data on it (due to rf3). The is no option to choose which nodes should participate in the backup (if it’s done via scylla-manager). But you are right, if the keyspace is RF3 and the cluster is of 3 nodes size, then it should be enough to take the snapshot from a single node. --- ### Page: https://forum.scylladb.com/t/select-with-empty-output-error/2118 Title: Select * with empty output error - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi everyone, I have question about a “select” query. We are a big data company writing over 300K to our cluster every second. When we run a select query with “consistency 1” from the logs table we see the empty output… Language: en Canonical URL: https://forum.scylladb.com/t/select-with-empty-output-error/2118 ## Headings Structure: H1: Select * with empty output error H3: Related topics ## Main Content: H1: Select * with empty output error H3: Related topics I have question about a “select” query. We are a big data company writing over 300K to our cluster every second. When we run a select query with “consistency 1” from the logs table we see the empty output. However, when I change the consistency to “all” then I receive the data. What would be the reason for this? QUERY: select * from (table name) Furthermore, this usually happens when we select one of the logs from the logs table, process it and delete it after that and selecting another one. I look forward for your help. When using “consistency 1”, data will be served by a single replica. If this replica happens to miss the data you want to query, then you will get empty result. When using “consistency all”, all replicas participate in the read and if at least one of them has the data, you will get it in the response. Furthermore, ScyllaDB will detect that there was a difference between the data different replicas provided and this will trigger a “read repair”, which will make sure all replicas hold the same data. After such a query, subsequent “consistency 1” requests should also return the data reliably. I recommend running periodic repair on the cluster. Furthermore, if you want to reliably get back the data you wrote, both your reads and your writes should use “consistency quorum” at least. Anything less and your queries will be prone to differences like you describe above. I recommend running periodic repair on the cluster Check out Scylla Manager, a tool that do just that. --- ### Page: https://forum.scylladb.com/t/process-for-making-configuration-changes-on-a-running-cluster-with-no-downtime-is-nodetool-drain-required/2120 Title: Process for making configuration changes on a running cluster, with no downtime. Is nodetool drain required? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/process-for-making-configuration-changes-on-a-running-cluster-with-no-downtime-is-nodetool-drain-required/2120 ## Headings Structure: H1: Process for making configuration changes on a running cluster, with no downtime. Is nodetool drain required? H3: Related topics ## Main Content: H1: Process for making configuration changes on a running cluster, with no downtime. Is nodetool drain required? H3: Related topics Originally from the User Slack @Bohdan_Smal: Hello. I wanted to clarify a few basic questions. In the Production Readiness Guidelines, it’s stated that “All configuration settings for all nodes in the same cluster should be identical or coherent.” However, I couldn’t find additional description of parameters indicating whether they are coherent or not. There’s only information about whether liveness updates are supported or not. For example, on the production cluster with 6 nodes (Ec2Snitch, 2 nodes in each zone), I need to enable additional metrics by adding the “enable_keyspace_column_family_metrics parameter” to the configuration. Most likely, there shouldn’t be any issues with this setting, but I still wanted to clarify. Additionally, I wanted to ask if it’s sufficient to simply sequentially execute “sudo systemctl start scylla-server.service” on each node after adding this parameter to the configuration, or if it would be more reliable to first execute “nodetool drain” and then restart the scylladb service. These questions may seem trivial, but I wanted to clarify with you. Thank you. Production Readiness Guidelines | ScyllaDB Docs Configuration Parameters | ScyllaDB Docs @Felipe_Cardeneti_Mendes: The warning idea is to prevent configuration drifts. Like having enable_keyspace_column_family_metrics enabled in just one node and not across other nodes in the cluster. Of course, you can test specific settings under a single node before deploying it globally. But it’s a good idea to always employ some way (such as ansible), to ensure that the configuration is consistent across the board eventually. It is a good practice to drain the node first, then restart. @Bohdan_Smal: @Felipe_Cardeneti_Mendes Please, could you provide more details on what you meant? The cluster is under load, so to avoid downtime, I’ll sequentially update the configuration on each node, drain the node, and restart the Scylla service to bring the node back into the cluster. Therefore, for a certain period, the configuration on the nodes will differ, but approximately within 10 minutes, when I execute my commands on each node in sequence, the configuration will be the same everywhere. My main goal is to avoid downtime or any errors on the ScyllaDB side, as requests will continue to be processed during this time. @Felipe_Cardeneti_Mendes: Your flow is correct. --- ### Page: https://forum.scylladb.com/t/operation-timeout-on-delete/2121 Title: Operation timeout on Delete - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/operation-timeout-on-delete/2121 ## Headings Structure: H1: Operation timeout on Delete H3: Related topics ## Main Content: H1: Operation timeout on Delete H3: Related topics Originally from the User Slack @Marko_Ćorić: Ok, quick question. Any chance we can increase timeout on DELETE request WITH TIMEOUT 500s is not available here. We are using Rust driver scylla = "0.12.0". This is error message: @Piotr_Smaroń: have you tried setting requests timeout in the RUST driver? From cqlsh you would do something like: tried to add it to execution profile, but no success. We are using caching session, maybe that’s problem? just give up after 2 sec @Piotr_Smaroń: 2s looks suspicious, can you run this query with tracing enabled? @Marko_Ćorić: yup, lemme get logs @Piotr_Smaroń: oh, I rather meant to enable query tracing - https://opensource.docs.scylladb.com/stable/using-scylla/tracing.html - by running but on 2nd thought, IDK if that’d yield any info, since the query fails/times out Tracing | ScyllaDB Docs @Michał_Chojnowski: There are two kinds of a timeout: • server-side timeout, which happens when replica nodes don’t respond to the coordinator in time • client-side timeout, which happens when the coordinator doesn’t respond to the client in time What you see here is a server timeout — the coordinator reported that only one replica (probably itself) has responded within the configured time, but at least two were needed to satisfy the quorum. .request_timeout in the Rust driver only sets the client-side timeout. If you make it larger, the client will only wait longer for the coordinator — but the coordinator will still just give up after 2 seconds. To change the coordinator timeout, you can either add USING TIMEOUT xyz in the query, or you can change it globally in the server config. (E.g. via the write_request_timeout_in_ms entry in scylla.yaml. It is set to write_request_timeout_in_ms: 2000 by default, hence 2 seconds. @Marko_Ćorić: can this help maybe? should I go for “one by one” delete? Get data with HistoryCollector and then remove them in loop? @Michał_Chojnowski I can’t add USING TIMEOUT in delete query, tried that first @Michał_Chojnowski: > should I go for “one by one” delete? No, a range delete should usually perform better than thousands of single-row deletes. Though it depends on the details. What Scylla version are you using? Lingering tombstones have been a major source of performance problems for a long time, but there have been improvements over time. @Marko_Ćorić: That’s true, but can’t go further. 5.2.13-0.20240103.c57a0a7a46c6 that’s version we are using atm atm I’m using manual remove when tombstone “hit”, apply query to record day-by-day (run query from DataGrip), but that takes time … also, sometimes rollback happens, so removed records are not removed. @Michał_Chojnowski: You can do USING TIMEOUT, but the syntax is that USING has to be before the WHERE. I.e.: delete from ks.t using timeout 500s where key = x will work @Marko_Ćorić: hmmm, ok lemme try @Michał_Chojnowski: I just checked that USING TIMESTAMP works. But increasing the timeout is only going to hide the underlying performance problem. (It’s hard to say what it is without the details). A delete shouldn’t take 2s in the first place. (Unless there is a materialized view involved, because then a single statement might have to delete any number of partitions). @Marko_Ćorić: jesus, this is solution. Thanks a lot m8. yes, USING TIMEOUT before WHERE do the job I tried at the end of query, following documentation for SELECT we have materialized view, bit it’s not used in this query @Michał_Chojnowski: If an increased timeout is a good enough solution for you, then good. But if the real problem (which causes things to happen slow enough that they exceed the timeout) is about having too many tombstones, then you might have to do something to deal with the tombstones. Unfortunately, isn’t that very straightforward — you’d have to run repair, adjust tombstone GC settings, flush and compact, and then maybe even clean the cache, to flush all old tombstones out. I see that https://opensource.docs.scylladb.com/stable/kb/tombstones-flush.html has some info about that. (But it seems somewhat outdated, because since Scylla 5.0 you’d probably want to use ALTER TABLE WITH tombstone_gc = {'mode':'repair'} ; before the repair instead of altering gc_grace_seconds after the repair). Fiddling with tombstone GC settings requires you to know what you’re doing though, otherwise you might accidentally undelete data. But again, it’s hard to know what the performance problem is without taking a close look. @Marko_Ćorić: We just had to remove old data, so that’s good. That will help us with large partitions @Michał_Chojnowski: > I tried at the end of query, following documentation for SELECT Can’t blame you for that; it’s weird that SELECT and DELETE have a different ordering between WHERE and USING. I wonder how the syntax ended up like that. I guess the author of USING TIMEOUT just didn’t notice the inconsistency. Maybe the grammar should be relaxed in a future release so that either ordering works. --- ### Page: https://forum.scylladb.com/t/scylla-manager-3-2-8-cql-ssl-timeout/2123 Title: Scylla Manager 3.2.8 CQL SSL Timeout - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello, We are currently using ScyllaDB with Scylla Manager Open Source, and I’ve encountered an issue. When running the sctool status command, the CQL returns a TIMEOUT SSL error. I couldn’t find specific information i… Language: en Canonical URL: https://forum.scylladb.com/t/scylla-manager-3-2-8-cql-ssl-timeout/2123 ## Headings Structure: H1: Scylla Manager 3.2.8 CQL SSL Timeout H3: Related topics ## Main Content: H1: Scylla Manager 3.2.8 CQL SSL Timeout H3: Related topics We are currently using ScyllaDB with Scylla Manager Open Source, and I’ve encountered an issue. When running the sctool status command, the CQL returns a TIMEOUT SSL error. I couldn’t find specific information in the documentation regarding the correct configuration for SSL keys. While our application is configured to use SSL correctly with the necessary certificates, my attempts to apply the same configuration to the manager have been unsuccessful. Could you please provide guidance on how to properly configure SSL for Scylla Manager? 3 Machines (for the Scylla Cluster) + 1 machine (for the Scylla Manager + scylla) Everything runs on baremetal. My last(current) attempt for the /etc/scylla-manager/scylla-manager.yaml on the manager machine. Note: The scylla at 127.0.0.1 is a standalone node reserved for the manager. Here is how I fixed my issue: The SSL configuration in /etc/scylla-manager/scylla-manager.yaml The reason I have to do this is that the cluster uses the port 9042 for SSL/TLS References in scylladb documentation (I cannot post links…): --- ### Page: https://forum.scylladb.com/t/can-we-expand-multiple-nodes-in-multi-dc-cluster/2125 Title: Can we expand multiple nodes in multi-DC cluster? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Now I have a cluster with three Data Centers and I need to expand this cluster. Can I expand one node in each Data Center at the same time? Additionally, I should set the parameter ‘consistent_rangemovement’ to ‘false’ t… Language: en Canonical URL: https://forum.scylladb.com/t/can-we-expand-multiple-nodes-in-multi-dc-cluster/2125 ## Headings Structure: H1: Can we expand multiple nodes in multi-DC cluster? H3: Related topics ## Main Content: H1: Can we expand multiple nodes in multi-DC cluster? H3: Related topics Now I have a cluster with three Data Centers and I need to expand this cluster. Can I expand one node in each Data Center at the same time? Additionally, I should set the parameter ‘consistent_rangemovement’ to ‘false’ to expand multiple nodes simultaneously. I’m not sure whether this operation will have negative effects on the cluster. No, only one node can be added at a time. In ScyllaDB 6.0 with tablets, you can add nodes simultaenously. --- ### Page: https://forum.scylladb.com/t/getting-error-in-lab-getting-started-with-scylladb-cloud/2126 Title: Getting Error in Lab: Getting Started with ScyllaDB Cloud - University and Training - ScyllaDB Community NoSQL Forum Meta Description: I am running the docker command as defined in the courseware but I am getting an error: (I have masked the password and IP address) ❯ docker run -it --rm --entrypoint cqlsh scylladb/scylla-cqlsh -u scylla -p XXXXX XXX.X… Language: en Canonical URL: https://forum.scylladb.com/t/getting-error-in-lab-getting-started-with-scylladb-cloud/2126 ## Headings Structure: H1: Getting Error in Lab: Getting Started with ScyllaDB Cloud H3: Related topics ## Main Content: H1: Getting Error in Lab: Getting Started with ScyllaDB Cloud H3: Related topics I am running the docker command as defined in the courseware but I am getting an error: (I have masked the password and IP address) My IP4 address is in the “Allowed IPs List.” Also, I would like to note in the course that two quiz questions topics were not covered during the video. I opened a support ticket first and then discovered this support forum. It’s a long shot but just wondering if you are using any VPN? If, yes, try to stop the VPN add your new public IP to the “Allowed List” and retry. Hi Michael, did you manage to solve the issue? With the cloud connection no but a work around with a local docker container got me through the exercise. Could you share your workaround? I’d like to see if/how we can improve the instructions for this Lab exercise. It was simply to complete the exercise from a docker container. I went to docker hub, pulled down a copy, ran it and completed the exercise. I basically did exactly what the very next exercise did. --- ### Page: https://forum.scylladb.com/t/bloom-filters-memory-reclamation-and-affects-on-ram-and-disk-i-o/2127 Title: Bloom Filters, memory reclamation, and affects on RAM and Disk I/O - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/bloom-filters-memory-reclamation-and-affects-on-ram-and-disk-i-o/2127 ## Headings Structure: H1: Bloom Filters, memory reclamation, and affects on RAM and Disk I/O H3: Related topics ## Main Content: H1: Bloom Filters, memory reclamation, and affects on RAM and Disk I/O H3: Related topics Originally from the User Slack @Skunnyk: Hi I’m testing scylladb 5.4.6 which should solve lots of our LSA/memory problems due to excessive bloom filter size (thank you @Botond_Dénes for the help ), and I see that https://github.com/scylladb/scylladb/commit/1aedc7372de857bb80ef54aa0b87b21a217d82a3 is also present in the release. Non-LSA Ram usage on cluster upgraded to 5.4.6 (from 5.4.x) is lower, and I try to understand this new “feature”. Does this means that “big” bloom filter are now reclaimed if they exceed 10% of the shard memory ? "Ratio of available memory for all in-memory components of SSTables in a shard beyond which the memory will be reclaimed from components until it falls back under the threshold. Currently, this limit is only enforced for bloom filters." . Does this means that we can expect a bit more disk i/o because theses blooms filters are dropped ? GitHub: Merge ‘[Backport 5.4] : Track and limit memory used by bloom filters’… · scylladb/scylladb@1aedc73 @avi: Yes, they get thrown out, you’ll get more disk I/O you can adapt by increasing the ratio (will eat into your cache, so more I/O from that), or increasing the false positive chance (more false positives, so more I/O), or installing more RAM (less I/O) @Skunnyk: Yes, that was my understanding Thank you ! Also, we upgraded to 5.4.6, and our bad_alloc and non-lsa problems are now gone --- ### Page: https://forum.scylladb.com/t/probabilistic-data-hyperloglog/2129 Title: Probabilistic Data: HyperLogLog - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Does it exist HyperLogLog data structure for distinct counting like in Redis or Aerospike? Language: en Canonical URL: https://forum.scylladb.com/t/probabilistic-data-hyperloglog/2129 ## Headings Structure: H1: Probabilistic Data: HyperLogLog H3: Related topics ## Main Content: H1: Probabilistic Data: HyperLogLog H3: Related topics Does it exist HyperLogLog data structure for distinct counting like in Redis or Aerospike? Did you find a solution for this? It seems that Spark supports this, you can try to look at connecting Spark with ScyllaDB. No, but it’s an interesting idea! I suggest opening an issue for it, or commenting on Implement custom merging algorithms · Issue #1321 · scylladb/scylladb · GitHub --- ### Page: https://forum.scylladb.com/t/adding-a-node-in-a-running-production-cluster/2136 Title: Adding a node in a running Production Cluster - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: We have scylladb multidc cluster setup with 3 nodes in dc1 for writes and 6 in dc2 for reads Currently dc1 nodes have 16cpu with 80% utilisation We want to add 3 more nodes in dc1 for smooth write operation in future a… Language: en Canonical URL: https://forum.scylladb.com/t/adding-a-node-in-a-running-production-cluster/2136 ## Headings Structure: H1: Adding a node in a running Production Cluster H3: Related topics ## Main Content: H1: Adding a node in a running Production Cluster H3: Related topics We have scylladb multidc cluster setup with 3 nodes in dc1 for writes and 6 in dc2 for reads Currently dc1 nodes have 16cpu with 80% utilisation We want to add 3 more nodes in dc1 for smooth write operation in future as app traffic is growing Please suggest correct method to add nodes in cluster with minimum downtime Following above link of ScyllaDB node addition its not clear when to alter system_auth keyspaces , either before starting new node scylla service or after Note: We have 1.8 tb of data present on each node Scylla version : 5.4.1 --- ### Page: https://forum.scylladb.com/t/release-scylladb-6-0-rc2/2137 Title: [RELEASE] ScyllaDB 6.0 RC2 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The Scylla team is pleased to announce ScyllaDB Open Source 6.0 RC2, the second Release Candidate for the Scylla Open Source 6.0 minor release. Only the last two minor releases of the ScyllaDB Open Source project are su… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-6-0-rc2/2137 ## Headings Structure: H1: [RELEASE] ScyllaDB 6.0 RC2 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 6.0 RC2 H3: Related topics The Scylla team is pleased to announce ScyllaDB Open Source 6.0 RC2, the second Release Candidate for the Scylla Open Source 6.0 minor release. Only the last two minor releases of the ScyllaDB Open Source project are supported. Once Scylla Open Source 6.0 is officially released, ScyllaDB Open Source 6.0 and 5.4 will be supported, and ScyllaDB 5.2 will be retired. For a complete description of ScyllaDB 6.0 see ScyllaDB 6.0 RC1. Get ScyllaDB Open Source 6.0 (under “More Versions” for each distro) Updates and bug fixes since 6.0 RC1 (not including tests and docs updates) mv: handle different ERMs for base and view tables. The result might be a crash when MV updates concurrent with a decommission #17786, #18709 Upgrade: Schema mismatch after rollback during rolling upgrade from 5.4 to 6.0 #18098 --- ### Page: https://forum.scylladb.com/t/hosting-and-operating-scylladb-vs-cassandra-ease-of-use-compaction-garbage-collection-and-complexity/2139 Title: Hosting and operating ScyllaDB vs. Cassandra, ease of use, Compaction, Garbage Collection and complexity - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/hosting-and-operating-scylladb-vs-cassandra-ease-of-use-compaction-garbage-collection-and-complexity/2139 ## Headings Structure: H1: Hosting and operating ScyllaDB vs. Cassandra, ease of use, Compaction, Garbage Collection and complexity H3: Related topics ## Main Content: H1: Hosting and operating ScyllaDB vs. Cassandra, ease of use, Compaction, Garbage Collection and complexity H3: Related topics Originally from the User Slack @sandr8: Is hosting and operating open-source Scylla as stressful as it is to host and operate Cassandra? Or are many of the culprits taken care of in a better way in Scylla? Any personal experience? @dor: It’s less. Compaction is a solved problem. There is no GC, etc. The operational complexity as a whole is reduced. Still, you do need to be an expert for an always-on distributed DB @sandr8: doesn’t the new compaction strategy require enterprise licensing? Also: wouldn’t there always be scary situations related to tombstones and nodes being down? @dor: ICS does require an enterprise license but you can use LCS, TWCS and STCS for free and compaction works better than cassandra there too When nodes are down, it’s not easy but it still supposed to be better than C*. Especially now that we got Raft’s consistent topology in place @sandr8: :gratitude_thank_you: --- ### Page: https://forum.scylladb.com/t/release-scylladb-6-0-rc3/2141 Title: [RELEASE] ScyllaDB 6.0 RC3 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The Scylla team is pleased to announce ScyllaDB Open Source 6.0 RC3, the 3rd Release Candidate for the Scylla Open Source 6.0 minor release. Only the last two minor releases of the ScyllaDB Open Source project are suppo… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-6-0-rc3/2141 ## Headings Structure: H1: [RELEASE] ScyllaDB 6.0 RC3 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 6.0 RC3 H3: Related topics The Scylla team is pleased to announce ScyllaDB Open Source 6.0 RC3, the 3rd Release Candidate for the Scylla Open Source 6.0 minor release. Only the last two minor releases of the ScyllaDB Open Source project are supported. Once Scylla Open Source 6.0 is officially released, ScyllaDB Open Source 6.0 and 5.4 will be supported, and ScyllaDB 5.2 will be retired. For a complete description of ScyllaDB 6.0 see ScyllaDB 6.0 RC1. Get ScyllaDB Open Source 6.0 (under “More Versions” for each distro) Updates and bug fixes since 6.0 RC2 (not including tests and docs updates) --- ### Page: https://forum.scylladb.com/t/error-on-drop-role-ldap/2142 Title: Error on drop role ldap - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Has anyone seen this error when deleting an ldap role? request resulted in cassandra_error, stream 123, code 256, message [Cannot delete passwords with SaslauthdAuthenticator] Language: en Canonical URL: https://forum.scylladb.com/t/error-on-drop-role-ldap/2142 ## Headings Structure: H1: Error on drop role ldap H3: Related topics ## Main Content: H1: Error on drop role ldap H3: Related topics Has anyone seen this error when deleting an ldap role? request resulted in cassandra_error, stream 123, code 256, message [Cannot delete passwords with SaslauthdAuthenticator] I think it will always happen as part of role deletion is password deletion. But since there should be nothing to delete maybe it’s a bug to throw an exception here and instead should be a nop. --- ### Page: https://forum.scylladb.com/t/release-scylladb-6-0/2143 Title: [RELEASE] ScyllaDB 6.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce ScyllaDB Open Source 6.0, a production-ready major release. ScyllaDB 6.0 introduces two major features which change the way ScyllaDB works: Tablets, a dynamic way to distribute… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-6-0/2143 ## Headings Structure: H1: [RELEASE] ScyllaDB 6.0 H2: Related Links H2: New features H3: Tablets H3: Strongly Consistent Topology Updates H3: Strongly Consistent Auth Updates H3: Strongly Consistent Service Levels H3: Describe Schema with Internals H3: Native Nodetool H3: Removing the JMX Server H3: Maintenance Mode H3: Maintenance Socket H3: Deployment H2: Improvements H3: Bloom Filters H3: Stability and performance H3: Guardrails H3: Replication strategies Guardrails can be used to deny SimpleStratagy strategy, for example, which is not recommended for production. H3: CQL H3: Alternator H3: Images and packaging H3: Tooling and REST API H3: Configuration H3: Tracing H3: Monitoring H3: Deprecated and removed features H3: Bug fixes and stability H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 6.0 H2: Related Links H2: New features H3: Tablets H4: Using Tablets H4: Procedures H4: Monitor Tablets H4: Driver Support H3: Strongly Consistent Topology Updates H3: Strongly Consistent Auth Updates H3: Strongly Consistent Service Levels H3: Describe Schema with Internals H3: Native Nodetool H3: Removing the JMX Server H3: Maintenance Mode H3: Maintenance Socket H3: Deployment H2: Improvements H3: Bloom Filters H3: Stability and performance H4: Compaction Related H4: Commitlog Related H4: Cluster Operation Related H4: Materialized Views Related H4: Performance Related H4: All kinds of edge cases H3: Guardrails H3: Replication strategies Guardrails can be used to deny SimpleStratagy strategy, for example, which is not recommended for production. H3: CQL H3: Alternator H3: Images and packaging H3: Tooling and REST API H3: Configuration H3: Tracing H3: Monitoring H3: Deprecated and removed features H3: Bug fixes and stability H3: Related topics The ScyllaDB team is pleased to announce ScyllaDB Open Source 6.0, a production-ready major release. ScyllaDB 6.0 introduces two major features which change the way ScyllaDB works: In addition, ScyllaDB 6.0 includes many other improvements in functionality, stability, UX and performance. Only the latest two minor releases of the ScyllaDB Open Source project are supported. With this release, only ScyllaDB Open Source 6.0 and 5.4 are supported. Users running earlier releases are encouraged to upgrade to one of these two releases. Get ScyllaDB Open Source 6.0 as binary packages (RPM/DEB), AWS AMI, GCP Image and Docker image Upgrade from ScyllaDB 5.4 to ScyllaDB 6.0 In this release, ScyllaDB enabled Tablets, a new data distribution algorithm to replace the legacy vNodes approach inherited from Apache Cassandra. While the vNodes approach statically distributes all tables across all nodes and shards based on the token ring, the Tablets approach dynamically distributes each table to a subset of nodes and shards based on its size. In the future, distribution will use CPU, OPS, and other information to further optimize the distribution. In particular, Tablets provide the following: Read more about Tablets here. Tablets are enabled by default for new clusters. You can set the initial number of Tablets per table using the “initial” parameter: CREATE KEYSPACE … WITH TABLETS = { ‘initial’: 1 } Or disable Tablets for the keyspace with the “‘enabled’ parameter: CREATE KEYSPACE … WITH TABLETS = { ‘enabled’: false } (See CREATE KEYSPACE docs) All tables created in this Keyspace will use Tablets by default. In this release, you can not use the following features with Tablets: If you are planning to use one of these features, disable Tablets when creating the Keyspace. Note: you can not ALTER an existing Keyspace to switch between Tablets and vNode based table and back. We will remove these restrictions in upcoming patches and minor releases. With Tablets, the Replication Factor (RF) cannot be updated to a value higher than the number of nodes per Data Center (DC). This feature protects the Admin from setting an impossible-to-support RF. This affects the following operations: Node Decommission / Remove Starting from 6.0, you cannot decommission or remove a node if the resulting number of nodes would be smaller than the largest non-zero replication factor (for any keyspace) in this DC 1 DC, 5 nodes, a KS with RF=5 The decommission request will fail The Replication Factor (RF) of Keyspaces must be less than or equal to the number of available nodes per Data Center (DC) Once a tablets-enabled Keyspace has tables, you can not ALTER its Replication Factor to be greater than the number of available nodes per DC. If you create such a Keyspace, you won’t be able to create Tables until you fix the RF or add more nodes. To Monitor Tablets in real time, upgrade ScyllaDB Monitoring Stack to release 4.7, and use the new dynamic Tablet panels, below. The Following Drivers support Tablets Legacy ScyllaDB and Apache Cassandra drivers will continue to work with ScyllaDB but will be less efficient when working with tablet-based Keyspaces. With Raft-managed topology enabled, all topology operations are internally sequenced consistently. A centralized coordination process ensures that topology metadata is synchronized across the nodes on each step of a topology change procedure. This makes topology updates fast and safe, as the cluster administrator can trigger many topology operations concurrently, and the coordination process will safely drive all of them to completion. For example, multiple nodes can be bootstrapped concurrently, which couldn’t be done with the previous gossip-based topology. Strongly Consistent Topology Updates is now the default for new clusters, and should be enabled after upgrade for existing clusters. System-auth-2 is a reimplementation of the Authentication and Authorization systems in a strongly consistent way on top of the Raft sub-system. This means that Role-Based Access Control (RBAC) commands like create role or grant permission are safe to run in parallel without a risk of getting out of sync with themselves and other metadata operations, like schema changes. As a result, there is no need to update system_auth RF or run repair when adding a DataCenter. Service Levels allow you to define attributes like timeout per workload. Service levels are now strongly consistent using Raft, like Schema, Topology and Auth. Until this release, CQL DESCRIBE SCHEMA was not sufficient to do a full schema restore from backup. For example, it lacks information about dropped columns. In 6.0, the DESC SCHEMA WITH INTERNALS command provides more information, streamlining the restore process. The nodetool utility provides simple command-line interface operations and attributes. ScyllaDB inherited the Java based nodetool from Apache Cassandra. In this release, the Java implementation was replaced with a backward-compatible native nodetool. The native nodetool works much faster. Unlike the Java version ,the native nodetool is part of the ScyllaDB repo, and allows easier and faster updates. With the Native Nodetool (above), the JMX server has become redundant and will no longer be part of the default ScyllaDB Installation or image. If you are using the JMX server directly (not via nodetool): Related issues: #15588 #18566 #18472 #18566 Maintenance mode is a new mode in which the node does not communicate with clients or other nodes and only listens to the local maintenance socket and the REST API. It can be used to fix damaged nodes – for example, by using nodetool compact or nodetool scrub. In maintenance mode, ScyllaDB skips loading tablet metadata if it is corrupted to allow an administrator to fix it. #5489 The Maintenance Socket provides a new way to interact with ScyllaDB from within the node it runs on. It is mainly for debugging. You can use CQLSh with the Maintenance Socket as described in the Maintenance Socket docs. #16172 Ubuntu 24.04 is now supported. RHEL / CentOS 7 support is deprecated. The setup utility now works with disks that do not have UUIDs, such as those in some virtualized environments #13803 The scylladb-kernel-conf package tunes the Linux kernel scheduler via sysfs to improve latency. These tunings were lost in Linux 5.13+ due to kernel changes. They are now restored. #16077 Docker: can not connect to Scylla 5.4 with CQLSh without providing host IP #16329 On Ubuntu, the installer now handles conflicts between a system process updating apt metadata and the installer itself.#16537 The following is a list of improvements and bug fixes included in the release, grouped by domain. Bloom filters are used to determine which SStables do not contain a partition key, speeding up reads when SStables can be filtered out. Since the Bloom filters are held in memory, and their size depends on the data (small partitions require larger Bloom filters), there is a tradeoff between allocated memory, and the risk of OOM and the filter efficiency. The following improvements were made to Bloom filters in this release: Topology changes, Repairs, etc Guardrails is a framework to protect ScyllaDB users and admins from common mistakes and pitfalls. In this release the following Guardrails are added: Two new configurations Replace the old restrict_replication_simplestrategy and give more granularity for DB Admin to warn, or block non production strategies. Scylla REST API is now documented! (beta) The sstable validation tools, scylla sstable validate-checksum and scylla sstable validate, now returns output in json format. Admin API: a new API for asynchronous compaction: /tasks/compaction/keyspace_compaction/{keyspace} Similar to the existing synchronos storage_service API. #15092 Bundled tools now use the Seastar epoll reactor backend rather than linux-aio; this reduces the risk of startup failures. The bundled cqlsh has been updated , with a fix for a COPY TO STDOUT regression. #17451 The REST API for reporting cache statistics is now more accurate. #9418 Some false-positives were eliminated from the scrub command. #16326 The iotune utility is used to measure a disk’s performance when installing ScyllaDB. It now works when executed on machines with a very large core count (208 cores )#18546 Scylla Monitoring Stack released 4.7 and later supports ScyllaDB 6.0. See metrics update between 5.4 and 6.0 here, as well as the new, beta, metrics reference here. More monitoring related updates: See git log for a full (long) list of fixed issues. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-50-2024-06-07/2144 Title: Last week in scylla-cluster-tests.git master (issue #50; 2024-06-07) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the e943b9b9…f99d2881 range are covered. There were 36 non-merge commits from 10 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-50-2024-06-07/2144 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #50; 2024-06-07) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #50; 2024-06-07) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the e943b9b9…f99d2881 range are covered. There were 36 non-merge commits from 10 authors in that period. Some notable commits: Added wait for cloud-init before starting node configuration, as it now takes longer and caused issues when the provision step was not used. This allowed us to remove the hack of disabling SSH at the start and end of the cloud-init script. Manager version is now collected in Argus. Dependabot will track Python driver releases and create PRs when a new version is available, e.g., bumping scylla-driver from 3.25.11 to 3.26.8. PR backport labels are now required. Further backport automation has been added by automatically adding backport-done. New, short installation and sanity tests of Manager are now available on different OS distributions. The append_scylla_yaml attribute, previously defined in test configurations as a multiline string representation and not composable from multiple configs, is now changed to a dict to allow for native merging by anyconfig. See the hydra upload command, which allows users to specify a test ID and file path to upload an arbitrary file to S3 and link it in Argus. Files ending in .png or .jpg are also added to the test run screenshot collection inside Argus. Latte tool was updated to support tablets. See new functions and properties that allow users to get supported and enabled Scylla features on a cluster. These functions are a better choice for checking new features like tablets and raft consistent topology changes than the current method based on versions and the scylla.yaml file. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/scylladb-returning-a-smaller-page-than-expected-to-maintain-low-latency-handling-of-tombstones/2145 Title: ScyllaDB returning a smaller page than expected to maintain low latency, handling of tombstones - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-returning-a-smaller-page-than-expected-to-maintain-low-latency-handling-of-tombstones/2145 ## Headings Structure: H1: ScyllaDB returning a smaller page than expected to maintain low latency, handling of tombstones H3: Related topics ## Main Content: H1: ScyllaDB returning a smaller page than expected to maintain low latency, handling of tombstones H3: Related topics Originally from the User Slack @Carlos_V: Is there a reason why a query, using gocql, will return paginated data smaller than PageSize() size, and iter.PageState() for next page will be either empty or using the result as PageState(token) won’t return anything? Querying directly from database using cqlsh, I have a query with 207 rows, but it will only return 205, if I put limit 206 it will return 205 and 30 minutes lates it will return the 206 one. @Marko_Ćorić: All under the same cluster key? And after 30 minutes you see 206th record? @Botond_Dénes: you probably have tons of tombstones @Carlos_V: Yes, after the 30 minutes waiting the --MORE-- the 206 comes back. This is a materializes view. I did run a nodetool flush and repair, no change. @Botond_Dénes: materialized views do have tombstones repair won’t help, try nodetool compact @Carlos_V: Let me try still running compact, in the meantime regarding cluster key question, it a single node scylla instance and the materialized view is this: @Botond_Dénes: this is from the MV definition? @Carlos_V: Yes After 3 hours, node compact has finished, but the same issue happens, seems to be with the same 2 rows. I’m supposed to have 177 rows now. Using limit 175 is fine, with limit 176 then it wait for a long time until it comes back. Funnily enough if I run the same query with COUNT(1) I get 177 qty. @Botond_Dénes: looks like your tombstones could not be purged @Carlos_V: Perhaps I don’t have enough space for that? Use% 77% @Botond_Dénes: tombstones won’t be purged before gc_grace_seconds, this is 10 days by default @Carlos_V: Is there a way to force? the 10 days have passed and the data is surely returning much faster now, only a few seconds now. Count is now 187. If I do LIMIT 187 100 comes then --MORE-- and then another 85 come, and then another --MORE-- and then the 2 come after less than a second. But the go query still won’t bring these two rows. @Botond_Dénes: Maybe there is a problem in how the go driver handles empty pages @Carlos_V: Any change this can be looked at? @Marko_Ćorić: can you share that part of code? @Carlos_V: This is the way I’m fetch data @Marko_Ćorić: Sorry for delay, I had to find our implementation since we rewrite whole production to Rust. This is code that worked for us without single problem for like 2 years @Carlos_V: Thanks @Marko_Ćorić, just managed to get back working on this now. It’s a lesson learned from scylla. So the way it was happening here is, as an example: @avi: It’s expected. ScyllaDB will return a smaller page sometimes in order to maintain low latency @Carlos_V: OK, makes sense. The smaller make sense for me, the 0 page is more concerning I suppose. @avi: 0 pages are sent when there is a long series of tombstones. Older versions tried to fetch at least one live row, but if you have a series of a million tombstones, a timeout is inevitable. @Carlos_V: Thanks Avi. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-233-2024-06-09/2146 Title: Last week in scylladb.git master (issue #233; 2024-06-09) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 5b4e688668…6e3b997e04 range are covered. There were 111 non-merge commits from 18 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-233-2024-06-09/2146 ## Headings Structure: H1: Last week in scylladb.git master (issue #233; 2024-06-09) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #233; 2024-06-09) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 5b4e688668…6e3b997e04 range are covered. There were 111 non-merge commits from 18 authors in that period. Some notable commits: In Alternator, ScyllaDB’s implementation of the DynamoDB API, we now yield when generating large JSON responses to avoid inducing latency for concurrent requests. The tablet load balancer now uses randomization to select candidates for tablet replica migration to avoid tablets for a particular table clustering in some nodes or shards, which can cause CPU imbalance. The primary replica algorithm for tablets has been adjusted so that nodetool repair -pr balances repair work more evenly across the cluster. Materialized view flow control is based on adjusting the write rate based on backlog. The backlog calculations are now more accurate. The task manager provides observability to the operator about internal operations. It is now careful to conserve memory by not keeping state for completed tasks. Hints were recently changed to be stored per host ID rather than per IP address. However, hints can still be stored with IP-based directories from older cluster. A bug draining those hints when a node leaves the cluster was fixed. The bundled cqlsh version was updated to 6.0.20. Repair and tablet migration are now serialized. Schema-modifying statements (DDL) and authentication were moved to rely on Raft separately. A DDL statement that grants permissions (for example, CREATE TABLE) will now execute in a single transaction. This prevents failures from leaving only part of the operation committed. The schema and topology coordinator now runs in the [gossip scheduling group](group0, topology coordinator: run group0 and the topology coordinator… · scylladb/scylladb@34cf5c8 · GitHub Author: Gleb Natapov gleb@scylladb.com); previously it was run in the streaming scheduling group. This prevents operations like repair and streaming from competing with metadata management. A bug in write back-pressure management, that could trigger an out-of-memory condition, was fixed. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-2-19/2148 Title: [RELEASE] ScyllaDB 5.2.19 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.2.19, a bugfix release of the ScyllaDB 5.2 stable branch. ScyllaDB Open Source 5.2.19, like all past and future 5.x.y releases, is backward compatible, and supports roll… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-2-19/2148 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.2.19 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.2.19 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.2.19, a bugfix release of the ScyllaDB 5.2 stable branch. ScyllaDB Open Source 5.2.19, like all past and future 5.x.y releases, is backward compatible, and supports rolling upgrades. Note that the latest ScyllaDB Open Source stable release is 6.0, and you are encouraged to upgrade to it. This is Scylla 5.2’s last patch release. Scylla 5.2 is no longer supported. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/slow-queries-when-indexing-bypass-cache-and-non-expired-tombstones/2150 Title: Slow queries when indexing, BYPASS CACHE and non-expired tombstones - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/slow-queries-when-indexing-bypass-cache-and-non-expired-tombstones/2150 ## Headings Structure: H1: Slow queries when indexing, BYPASS CACHE and non-expired tombstones H3: Related topics ## Main Content: H1: Slow queries when indexing, BYPASS CACHE and non-expired tombstones H3: Related topics Originally from the User Slack @Dylan_Piette: Hello to the scylla team ! I’m trying to understand why some of my queries are slow and how I can change that Basically I created an index on a table and then I directly query this index to get the distinct partition key (scheduled_windows) of the index And it takes a long time even if there’s no data in it I understand that it might be related to tombstones So I changed the gc_grace_periods of the table and the materialized view created by the index I then ran nodetool compact and expected all the tombstones to disappear and then to get better performance but still I get horrible performances for an empty table Here is a tracing session of the query Could someone help me understand what’s going wrong ? And for information I have only one node and it’s on a big machine with an nvme disk @Felipe_Cardeneti_Mendes: shard 3 was picked as the query coordinator. You can see the source_elapsed columns increases considerably as you hit ranges with dead rows. First one is here: > Page stats: 0 partition(s), 0 static row(s) (0 live, 0 dead), 530275 clustering row(s) (0 live, 530275 dead) and 0 range tombstone(s) [shard 3] | 2024-05-02 15:26:24.799904 | 172.16.0.4 | 927112 | 172.16.0.4 Second one is: > Page stats: 0 partition(s), 0 static row(s) (0 live, 0 dead), 355826 clustering row(s) (0 live, 355826 dead) and 0 range tombstone(s) [shard 3] | 2024-05-02 15:26:25.426583 | 172.16.0.4 | 1553790 | 172.16.0.4 And the time it took to fully run the query: > Done processing - preparing a result [shard 3] | 2024-05-02 15:26:25.427603 | 172.16.0.4 | 1554810 | 172.16.0.4 So you’re running a fullscan, this full scan takes a while to populate a page worth of data, eventually hits ranges still with tombstones, and processing time increases. You may see if you get better result with BYPASS CACHE but, in general you would want a restriction while querying to avoid the full scan all the time @Piette_Dylan: Thanks for your answer What do you mean by putting a restriction ? Because I need to do a full scan if I want to do get all the different partition keys no ? And how do you explain that changing the gc_grace_periods of the table and the view, compacting them still doesn’t make the situation better ? @Felipe_Cardeneti_Mendes: If you frequently want to retrieve all distinct keys, then do a efficient token scan instead and append BYPASS CACHE to it. This will break down the scan to smaller queries which are going to be picked up by different shards and nodes, rather than a single query landing in an individual shard having to carry on the entire work. > And how do you explain that changing the gc_grace_periods of the table and the view, compacting them still doesn’t make the situation better ? I don’t know how the situation was before, so I can’t say whether the situation improved or not. One possibility could be that you had data in memtables shadowing SSTables, which would be true if you use USING TIMESTAMP for inserts. Another possibility could be simply that your gc_grace wasn’t low enough to expire the tombstones. In any case, your trace clearly shows that scans went all to cache, maybe you’ve hit https://github.com/scylladb/scylladb/issues/6033 - a way to check is to just BYPASS CACHE or restart the node the empty the cache. @Piette_Dylan: It is indeed that, the tombstones are not evicted from the cache The bypass cache works great, thanks ! If I understood correctly that problem would be solved if I upgrade to 5.4 ? @Felipe_Cardeneti_Mendes: Things should be improved, yes - But maybe not the way you’d expect. In particular, expired row and range tombstones are evicted on access from the cache. It means that if your cache accumulated many of them, an initial scan will still be penalized, but subsequent ones won’t. Cell tombstones and partition tombstones aren’t addressed yet - the bottom of the issue explains the rationale around those. Most likely partition tombstones are just fine, but cell tombstones aren’t. But all this may be orthogonal to your situation actually. The thing with full scans is that it will require scanning your entire data set, populating the cache causing potentially eviction of important rows, and so on. So for full scans specifically, it is generally just better to simply BYPASS CACHE, unless you know for certain your cache can hold your entire data set or that evictions won’t bottleneck other maybe important queries to frequently accessed partitions. There’s also the more common scenario of non-expired tombstones. It is important to try to understand why many ended up accumulating after all, maybe tuning compaction settings. @avi: Note: ScyllaDB is particularly bad with empty or almost-empty tables, since it can’t amortize the large number of vnodes and shards over a large number of rows. This is expected to improve with tablets. --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-4-7/2153 Title: [RELEASE] ScyllaDB 5.4.7 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.4.7, a bugfix release of the ScyllaDB 5.4 stable branch. ScyllaDB Open Source 5.4.7, like all past and future 5.x.y releases, is backward compatible and supports rolling… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-4-7/2153 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.4.7 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.4.7 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.4.7, a bugfix release of the ScyllaDB 5.4 stable branch. ScyllaDB Open Source 5.4.7, like all past and future 5.x.y releases, is backward compatible and supports rolling upgrades. Note that the latest ScyllaDB Open Source stable release is 6.0 and you are encouraged to upgrade to it. Issue fixed in this release: Tracing: the ability of figuring which shard owns a partition from system.large_partition by analyzing a corresponding sstable id was broken when moving to UUID based SSTable numbering #18381 Instead, one can use scylla sstable shard-of to extract shards which own the specified SSTables. #16343 Stability: Bad estimation for temporary sstables (containing GCed data only) causes, in the worst case, of 2x space amplification for filters #18283 Stability: clustering_range intersection() can cause an infinite loop #18688 Image: while dropping openssh-server from the ScyllaDB image, other system packages used by ScyllaDB scripts were also dropped.#17787 The regression was introduced in 5.4.6. Config: default task_ttl_in_seconds is 0, but scylla.yaml changes the value to 10. #16714 Stability: direct failure detector: make ping timeout configurable and increase the default, from 300 to 600 ms #16607. Failure detector is an internal inter-node mechanism. Stability: fromJson() or INSERT JSON fails to set a map #18477 Stability: handle_paxos_accept() fails to record a trace message when done with handling #18725 Stability: mutation_fragment_stream_validating_filter doesn’t respect validation_level::none #18662. The issue might cause false-positive validation errors during repair/streaming, leading to aborting the operation. Stability: Materialized View (MV): memory used by a view updates batch is incorrectly divided #17854 Stability: MV semaphore units are not kept alive until a view update finishes #17890 Tooling: native sstable validation yields false positive errors (limited to MX format) #16326 Stability: Potential use-after-move in streaming’s stream_result_future (courtesy of clang-tidy) #18332 Stability: Reclaimed bloom filter is left back in disk when the SSTable is deleted #18398 repair: add control for repair percentage for partition count estimation #18615. The new parameter is repair_partition_count_estimation_ratio. The default 10% has not changed. Stability: Scylla crash when reading from mutation fragments with token() filtering #18637 Performance: The algorithm for picking an index for a request is not always deterministic #7969 Stability: Use-after-move in sstable writer (courtesy of clang-tidy) #18323 Stability: Use-after-move in thrift/handler.cc:make_non_overlapping_ranges() (courtesy of clang-tidy) #18356 Stability: utils::chunked_vector fill constructor is exception unsafe #18635 Stability: Wrong exception is printed in build step exception handling #18423 --- ### Page: https://forum.scylladb.com/t/facing-issue-on-nodetool-rebuild-keeps-getting-stuck-for-a-long-time/2156 Title: Facing Issue on nodetool rebuild : keeps getting stuck for a long time - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I am trying to add a new Data Center to an existing 3 node scylla cluster. The older DC and the new DC are at different location (connected via IPsec). I have successfully joined the new DC and is UN. The problem arise… Language: en Canonical URL: https://forum.scylladb.com/t/facing-issue-on-nodetool-rebuild-keeps-getting-stuck-for-a-long-time/2156 ## Headings Structure: H1: Facing Issue on nodetool rebuild : keeps getting stuck for a long time H1: I am trying to add a new Data Center to an existing 3 node scylla cluster. H1: Followed the steps exactly described on Adding a New Data Center Into an Existing ScyllaDB Cluster H3: Related topics ## Main Content: H1: Facing Issue on nodetool rebuild : keeps getting stuck for a long time H1: I am trying to add a new Data Center to an existing 3 node scylla cluster. H1: Followed the steps exactly described on Adding a New Data Center Into an Existing ScyllaDB Cluster H3: Related topics The older DC and the new DC are at different location (connected via IPsec). I have successfully joined the new DC and is UN. The problem arises while running nodetool rebuild command on the new node. It runs smoothly for a while but gets stuck at a certain point and the process stops suddenly. What could be the possible cause for this issue and how can i fix it ? sudo nodetool status displays all nodes Up and Normal Version : 5.4.4-0.20240228.58a1be93b212 Logs from journalctl : get_row_diff_with_rpc_stream_handler from=192.168.10.11, repair_meta_id=6060: seastar::nested_exception: seastar::rpc::clos> [shard 10:stre] rpc - client 192.168.10.11:56965: server connection dropped: recv: Connection timed out [shard 2:stre] rpc - client 192.168.10.11:60587: server connection dropped: recv: Connection timed out [shard 10:stre] rpc - client 192.168.10.11:63025: server connection dropped: recv: Connection timed out [shard 2:stre] rpc - client 192.168.10.11:65042: server connection dropped: recv: Connection timed out [shard 4:stre] repair - Failed to process get_row_diff_with_rpc_stream_handler from=192.168.10.11, repair_meta_id=6027: seastar::nested_exception: seastar::rpc::clos> [shard 0:stre] repair - Failed to process get_row_diff_with_rpc_stream_handler from=192.168.10.11, repair_meta_id=6008: seastar::nested_exception: seastar::rpc::clos> [shard 9:stre] repair - Failed to process get_row_diff_with_rpc_stream_handler from=192.168.10.11, repair_meta_id=6009: seastar::nested_exception: seastar::rpc::clos> [shard 1:stre] repair - Failed to process get_row_diff_with_rpc_stream_handler from=192.168.10.11, repair_meta_id=6002: seastar::nested_exception: seastar::rpc::clos> [shard 3:stre] repair - Failed to process get_row_diff_with_rpc_stream_handler from=192.168.10.11, repair_meta_id=6035: seastar::nested_exception: seastar::rpc::clos> [shard 14:stre] repair - Failed to process get_row_diff_with_rpc_stream_handler from=192.168.10.11, repair_meta_id=6025: seastar::nested_exception: seastar::rpc::clos> [shard 8:stre] rpc - client 192.168.10.11:51413: server connection dropped: recv: Connection timed out [shard 6:stre] rpc - client 192.168.10.11:49731: server connection dropped: recv: Connection timed out [shard 9:stre] rpc - client 192.168.10.11:57324: server connection dropped: recv: Connection timed out [shard 9:stre] repair - Failed to process get_row_diff_with_rpc_stream_handler from=192.168.10.11, repair_meta_id=6037: seastar::nested_exception: seastar::rpc::clos> [shard 14:stre] rpc - client 192.168.10.11:59324: server connection dropped: recv: Connection timed out [shard 7:stre] rpc - client 192.168.10.11:49627: server connection dropped: recv: Connection timed out [shard 13:stre] rpc - client 192.168.10.11:59263: server connection dropped: recv: Connection timed out [shard 3:stre] repair - Failed to process get_row_diff_with_rpc_stream_handler from=192.168.10.11, repair_meta_id=6052: seastar::nested_exception: seastar::rpc::clos> [shard 14:stre] rpc - client 192.168.10.11:55484: server connection dropped: recv: Connection reset by peer [shard 3:stre] rpc - client 192.168.10.11:55953: server connection dropped: recv: Connection reset by peer [shard 3:stre] rpc - client 192.168.10.11:49698: server connection dropped: recv: Connection reset by peer [shard 5:stre] rpc - client 192.168.10.11:54500: server connection dropped: recv: Connection reset by peer [shard 11:stre] rpc - client 192.168.10.11:53981: server connection dropped: recv: Connection reset by peer [shard 7:stre] rpc - client 192.168.10.11:59392: server connection dropped: recv: Connection reset by peer [shard 13:stre] rpc - client 192.168.10.11:51928: server connection dropped: recv: Connection reset by peer [shard 13:stre] rpc - client 192.168.10.11:60148: server connection dropped: recv: Connection reset by peer [shard 0:stre] rpc - client 192.168.10.11:57195: server connection dropped: recv: Connection reset by peer [shard 2:stre] rpc - client 192.168.10.11:64142: server connection dropped: recv: Connection reset by peer [shard 5:stre] rpc - client 192.168.10.11:63395: server connection dropped: recv: Connection reset by peer [shard 4:stre] rpc - client 192.168.10.11:57259: server connection dropped: recv: Connection reset by peer [shard 3:stre] rpc - client 192.168.10.11:52008: server connection dropped: recv: Connection reset by peer [shard 0:stre] gossip - failure_detector_loop: Send echo to node 192.168.10.11, status = failed: seastar::rpc::closed_error (connection is closed) [shard 0:goss] gossip - Fail to send EchoMessage to 192.168.10.11: seastar::rpc::closed_error (connection is closed) [shard 0:stre] gossip - failure_detector_loop: Send echo to node 192.168.10.11, status = failed: seastar::rpc::timeout_error (rpc call timed out) s get_row_diff_with_rpc_stream_handler from=192.168.10.11, repair_meta_id=6027: se Looks like the logs are truncated on the right edge. Please provide complete logs. Also look at the logs on the other nodes from the same time range. --- ### Page: https://forum.scylladb.com/t/compaction-strategy-best-for-deletion-of-a-large-number-of-records/2161 Title: Compaction strategy best for deletion of a large number of records - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello! I wanted to ask for some suggestions on what compaction strategy to use for my use case. I’m using ScyllaDB to store the data generated by running in parallel multiple scripts. These scripts can finish running in… Language: en Canonical URL: https://forum.scylladb.com/t/compaction-strategy-best-for-deletion-of-a-large-number-of-records/2161 ## Headings Structure: H1: Compaction strategy best for deletion of a large number of records H3: Related topics ## Main Content: H1: Compaction strategy best for deletion of a large number of records H3: Related topics Hello! I wanted to ask for some suggestions on what compaction strategy to use for my use case. I’m using ScyllaDB to store the data generated by running in parallel multiple scripts. These scripts can finish running independently and run for different periods of time. Because of that, I’ve added a default TTL value of 1 day so that the data that’s generated by a script that runs for multiple days and that’s older than 1 day is automatically deleted. At the same time, once the scripts finish running, I have some logic that deletes everything related to the specific run with something like DELETE FROM table WHERE run_id = "specific_run". My question is what kind of compaction strategy do you think would best fit this use case? Up until now I’ve been using TWCS along with these settings: In the beginning I started with a 1 day time window, but I felt that the compaction was not triggered frequently enough to properly get rid of all of the tombstones created by the DELETE part at the end of the scripts. Now I read that apparently TWCS is not that good for manual deletes and I’m not sure anymore on what to use. Can you please help me out with a suggestion? Tombstones are challenging with all compaction strategies, there is no “best one” for delete heavy workloads. Therefore I would recommend going with the generalist STCS, as TWCS will handle explicit deletions very poorly and LCS is only adequate for mostly read workloads. --- ### Page: https://forum.scylladb.com/t/connection-failing-how-to-correctly-configure-encryption-ssl-tls/2163 Title: Connection failing, how to correctly configure encryption - SSL/TLS? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/connection-failing-how-to-correctly-configure-encryption-ssl-tls/2163 ## Headings Structure: H1: Connection failing, how to correctly configure encryption - SSL/TLS? H3: Related topics ## Main Content: H1: Connection failing, how to correctly configure encryption - SSL/TLS? H3: Related topics Originally from the User Slack @Omkar_Jadhav: Hello, I am trying to setup a single scylla node on my local system using docker-compose. This is my basic setup I am trying to use the same client code that we use in production to connect to this node. The client is configured to use SSL/TLS. As a result the connection is failing. How do I configure the scylla node in docker to use SSL/TLS? @Felipe_Cardeneti_Mendes: You can pass the --*client-encryption*-options command line argument to the container. However, given that you will need to copy the certificate, and keys anywhere, might be easier to simply append these arguments to scylla.yaml directly. See: https://opensource.docs.scylladb.com/stable/operating-scylla/security/client-node-encryption.html Encryption: Data in Transit Client to Node | ScyllaDB Docs @Omkar_Jadhav: I see, will try this. Thank you --- ### Page: https://forum.scylladb.com/t/error-i-n-h-c-c-decompressionexception-crc-value-mismatch/2164 Title: Error : i.n.h.c.c. DecompressionException: CRC value mismatch - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, We are using Scylldb 5.4.6 and while processing our request getting following error in our getling test tool. i.n.h.c.c. DecompressionException: CRC value mismatch. Expected 1 (0.23%) : 3151494194, Got: 3450875255 … Language: en Canonical URL: https://forum.scylladb.com/t/error-i-n-h-c-c-decompressionexception-crc-value-mismatch/2164 ## Headings Structure: H1: Error : i.n.h.c.c. DecompressionException: CRC value mismatch H3: Related topics ## Main Content: H1: Error : i.n.h.c.c. DecompressionException: CRC value mismatch H3: Related topics We are using Scylldb 5.4.6 and while processing our request getting following error in our getling test tool. i.n.h.c.c. DecompressionException: CRC value mismatch. Expected 1 (0.23%) : 3151494194, Got: 3450875255 Anyone had this issue before ? if yes what could be the reason for this. We need more details here. Where is this error raised from? From your tool, the driver or the DB (I don’t recognize this error)? What is this calculated on? What does “i.n.h.c.c.” mean? --- ### Page: https://forum.scylladb.com/t/multiple-scylla-cluster-manager-pair-deployment/2167 Title: Multiple scylla cluster-manager pair deployment - Kubernetes Operator - ScyllaDB Community NoSQL Forum Meta Description: Hi, I tried to deploy a scylla cluster - manager pair in k8s cluster working as expected, but on deploying another pair of it. Manager Controller for the first pair registers the second cluster in the first manager, and … Language: en Canonical URL: https://forum.scylladb.com/t/multiple-scylla-cluster-manager-pair-deployment/2167 ## Headings Structure: H1: Multiple scylla cluster-manager pair deployment H3: Related topics ## Main Content: H1: Multiple scylla cluster-manager pair deployment H3: Related topics Hi, I tried to deploy a scylla cluster - manager pair in k8s cluster working as expected, but on deploying another pair of it. Manager Controller for the first pair registers the second cluster in the first manager, and second manager controller also registers the second cluster in second manager. Both the pairs are in different namespaces, I would like to ask if it is possible to deploy two or multiple pairs of one-one mapped scylla cluster and manager?, I was looking through the controller code and found that scylla clusters are discovered across all the namespace, is it possible to control it and limit it to a specific namespace, or some specific annotations? No, it’s not possible with current design. Current integration assumes Scylla Manager is global and it registers all discovered ScyllaClusters. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-51-2024-06-14/2170 Title: Last week in scylla-cluster-tests.git master (issue #51; 2024-06-14) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the ea6a460f…ef159f51 range are covered. There were 13 non-merge commits from 5 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-51-2024-06-14/2170 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #51; 2024-06-14) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #51; 2024-06-14) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the ea6a460f…ef159f51 range are covered. There were 13 non-merge commits from 5 authors in that period. Some notable commits: Monitoring was updated to version 4.7. SCT Events containing backtraces are tested against an issue-keyword map. If a given keyword is found in this map, the corresponding issue is added to the event message, easing investigations. Currently, the REACTOR_STALLED event is supported, but it’s easily expandable. Base K8S version can be specified in Jenkins job parameters. It works with EKS and GKE, but local Kind is not supported. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/how-can-i-speed-up-scylladb-node-start-up-time-for-use-with-unit-tests-specifically-testcontainers/2171 Title: How can I speed up ScyllaDB node start up time for use with unit tests (specifically Testcontainers)? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/how-can-i-speed-up-scylladb-node-start-up-time-for-use-with-unit-tests-specifically-testcontainers/2171 ## Headings Structure: H1: How can I speed up ScyllaDB node start up time for use with unit tests (specifically Testcontainers)? H3: Related topics ## Main Content: H1: How can I speed up ScyllaDB node start up time for use with unit tests (specifically Testcontainers)? H3: Related topics Originally from the User Slack @Stuart: Does anyone know if you can get speed up scylla container start up time? I’m looking to use testcontainers to power some unit tests and waiting for the scylla image to start can take up to 30 seconds @Felipe_Cardeneti_Mendes: docker run --rm --disable-version-check --skip-wait-for-gossip-to-settle 0 This should be instant for single node (well, for multi-node as well, but then you probably want gossip to settle) @Stuart: perfect this cut down startup time by 25 seconds --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-5/2172 Title: [RELEASE] ScyllaDB Enterprise 2024.1.5 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2024.1.5, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2024.1 LTS Release. Improvements in this release: Alternator Performance i… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-5/2172 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2024.1.5 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2024.1.5 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2024.1.5, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2024.1 LTS Release. Improvements in this release: Varies optimization improved the throughput of the Alternator by up to 35%, depending on the workload. This release adds Workload Prioritization to Alternator, ScyllaDB’s Amazon DynamoDB-compatible API. To use Workload Prioritization, one needs to enable alternator_enforce_authorization in the configuration. Read more about this feature here Repair: Introduce small table optimization #15974 #16011 On a cluster with multiple datacenters, with high latency between them, repairing small tables, like system_auth can take unreasonable time. The root cause is token range based repair, which is useful for large tables, generating a huge overhead when repairing small tables. This optimization fixes this issue by avoiding token range based repair for small tables. The following issues are fixed in this release (with an open-source reference, if available): Varies optimization improved the throughput of the Alternator by up to 35%, depending on the workload. Nice! Is this also available in the open-source version? Are there any benchmark writeups? The improvement is limited to Enterprise, similar to the Enterprise only improvement for CQL in 2024.1.0 We will share more information on Benchmark shortly. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-234-2024-06-16/2173 Title: Last week in scylladb.git master (issue #234; 2024-06-16) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 6e3b997e04…bbb424a757 range are covered. There were 108 non-merge commits from 19 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-234-2024-06-16/2173 ## Headings Structure: H1: Last week in scylladb.git master (issue #234; 2024-06-16) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #234; 2024-06-16) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 6e3b997e04…bbb424a757 range are covered. There were 108 non-merge commits from 19 authors in that period. Some notable commits: Alterator, ScyllaDB’s implementation of the DynamoDB API, runs TTL expiration as an internal background process. This expiration process is run in the maintenance scheduling group to prevent it from dominating the user workload; however this isolation was broken when making calls to replica nodes. TTL expiration is now properly isolated in the maintenance group. A bug when updating materialized views on a newly-migrated tablet was fixed. The scylla sstable command can now recover the schema from the sstable itself, if it is new enough (ma format or later)`. Materialized views perform flow control by measuring a backlog and keeping it under control. The propagation of backlog to all shards in a node has been fixed. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/starting-the-scylladb-manager-agent-during-bootstrap-of-a-node-configuration-issue/2174 Title: Starting the ScyllaDB Manager agent during bootstrap of a node, configuration issue - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/starting-the-scylladb-manager-agent-during-bootstrap-of-a-node-configuration-issue/2174 ## Headings Structure: H1: Starting the ScyllaDB Manager agent during bootstrap of a node, configuration issue H3: Related topics ## Main Content: H1: Starting the ScyllaDB Manager agent during bootstrap of a node, configuration issue H3: Related topics Originally from the User Slack @Patryk_Kandziora: Hi. I use scyllaDB on AWS using the marketplace images in auto-scalling group. All works good (to some extend ;)) The problem is with Scylla Manager and agent to be installed during the boot time. My bootstrap script is as the following: The above is part of the script in Json format. Then we have base64 encoded post_configuration_script which looks as the following: The scyllaDB will fail to start with the last line of the script sudo systemctl start scylla-manager-agent If I remove the last line - scyllaDB will start without issue. Error: The interesting part is that when I SSH to the node and trigger the command sudo systemctl start scylla-manager-agent manually - all works fine. Any idea why is that and how to start the agent during the bootstrap of scyllaDB node in the cluster? @Felipe_Cardeneti_Mendes: If I remember it correctly, the agent depends on the database. So when you start the agent the database is still configuring itself, and calls the its own dependency, which is the scylla-machine-image service, but this one is already running. It’s best to check whether the server is up and running, then start the agent. One example can be doing something like https://github.com/fee-mendes/rust-driver-example/blob/main/docker-compose/check_cluster_healthy GitHub: rust-driver-example/docker-compose/check_cluster_healthy at main · fee-mendes/rust-driver-example @Patryk_Kandziora: Thx @Felipe_Cardeneti_Mendes - I used something like this: It doesnt work though. It looks like the script is executed before the scyllaDB service actually starts. If we take into consideration there is 600 sec timeout - then thats the problem here - I think. Yeah the service fails to start: @Felipe_Cardeneti_Mendes: Oh ok, just now I realized these are the AMI parms. lol Sorry Yeah, it can’t be in post_configuration_script indeed. The naming is correct, it will run after the configuration, but before the service is started. So you need this script to be run somewhere else, maybe as cloud-init? @Patryk_Kandziora: Yeah the problem is that if you pass the json config to bootstrap wrapper there is no way to pass anything else - at least I have not found solution for that. That is why I thought there is the post_configuration_script argument to overcome such limitation. Maybe I miss something here? @Felipe_Cardeneti_Mendes: I think if you really want something in the AMI that is more like a “post_start_script”, then the best place would be to check in https://github.com/scylladb/scylla-machine-image I don’t see a way out other than maybe forking the process (&) - IIUC it should just let the “post config” move forward. Or maybe just set: start_scylla_on_first_boot: false , past the configuration do something like: Reboot and everyone should live happily thereafter @Patryk_Kandziora: Thanks @Felipe_Cardeneti_Mendes - all good and works. Working bootstrap script if anyone will face the same issue with scyllaDB on AWS looks like this: Docs which helps: • https://opensource.docs.scylladb.com/stable/getting-started/install-scylla/launch-on-aws.html • https://github.com/scylladb/scylla-machine-image • https://cloudinit.readthedocs.io/en/latest/reference/examples.html#run-commands-on-first-boot Thanks again @Felipe_Cardeneti_Mendes Launch ScyllaDB on AWS | ScyllaDB Docs GitHub: GitHub - scylladb/scylla-machine-image --- ### Page: https://forum.scylladb.com/t/how-to-ensure-the-atomicity-of-data-mutation-and-cdc-log/2175 Title: How to ensure the atomicity of data mutation and CDC log? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: hi there, i’m using the cdc feature to replicate data from scylla to some other database (version 5.2). Can scylla ensure the atomicity of the data mutation and cdc_log mutation? In other word, can i assume that if the r… Language: en Canonical URL: https://forum.scylladb.com/t/how-to-ensure-the-atomicity-of-data-mutation-and-cdc-log/2175 ## Headings Structure: H1: How to ensure the atomicity of data mutation and CDC log? H3: Related topics ## Main Content: H1: How to ensure the atomicity of data mutation and CDC log? H3: Related topics hi there, i’m using the cdc feature to replicate data from scylla to some other database (version 5.2). Can scylla ensure the atomicity of the data mutation and cdc_log mutation? In other word, can i assume that if the response for a wirite operation for a cdc-enabled table is success, then the base-table data and cdc-log data are both visiable on scylla cluster? I also have the same question. If cdc_log cannot guarantee success, it will also have an impact on data migration due to cdc_log data loss. CDC has a built-in delay to account for out-of-order writes, so while the CDC entry will be on the cluster, it will still take time for the CDC library to see it. --- ### Page: https://forum.scylladb.com/t/release-database-level-encryption-using-scylladb-managed-key-now-enabled-by-default-17-jun-2024/2176 Title: [RELEASE] Database-Level Encryption using ScyllaDB-managed key now enabled by default - 17 Jun 2024 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: We are happy to announce that database-level encryption using ScyllaDB-managed keys is now enabled by default in ScyllaDB Cloud, at no extra cost. This enhancement ensures that your sensitive data is protected even more … Language: en Canonical URL: https://forum.scylladb.com/t/release-database-level-encryption-using-scylladb-managed-key-now-enabled-by-default-17-jun-2024/2176 ## Headings Structure: H1: [RELEASE] Database-Level Encryption using ScyllaDB-managed key now enabled by default - 17 Jun 2024 H3: Related topics ## Main Content: H1: [RELEASE] Database-Level Encryption using ScyllaDB-managed key now enabled by default - 17 Jun 2024 H3: Related topics We are happy to announce that database-level encryption using ScyllaDB-managed keys is now enabled by default in ScyllaDB Cloud, at no extra cost. This enhancement ensures that your sensitive data is protected even more robustly, providing an additional layer of security on top of the existing storage-based encryption. For more details on this feature, please refer to our documentation. --- ### Page: https://forum.scylladb.com/t/creating-an-index-on-user-defined-types/2178 Title: Creating an index on User Defined Types - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/creating-an-index-on-user-defined-types/2178 ## Headings Structure: H1: Creating an index on User Defined Types H3: Related topics ## Main Content: H1: Creating an index on User Defined Types H3: Related topics Originally from the User Slack @Dhruv_garg: can we create index on user defined types? e.g. CREATE INDEX ON nudge.rewards_coupon_code (client_id, coupon.code); @Felipe_Cardeneti_Mendes: Well, not for specific fields yet. You can index the entire UDT column @Dhruv_garg: so after indexing whole column, I can search on it’s specific fields as well? like where coupon.code = 'xyz', will this work? @Felipe_Cardeneti_Mendes: Nope, the entire column gets indexed not individual fields atm @Dhruv_garg: how is that helpful? what kind of queries can I run using that? @Felipe_Cardeneti_Mendes: Of course 1:1 queries where you require an exact match on the full UDT. You can open an issue to enable indexing individual fields. IIUC it should be simple with the existing infrastructure to index collection columns, but we’ll likely also discuss compatibility wrt https://issues.apache.org/jira/browse/CASSANDRA-6382 which is not yet solved by cassandra as well @Dhruv_garg: got it, thanks --- ### Page: https://forum.scylladb.com/t/scylladb-university-live-june-18-2024/2179 Title: ScyllaDB University LIVE - June 18, 2024 - University and Training - ScyllaDB Community NoSQL Forum Meta Description: Join us for a half day of (free) instructor-led training. We’re going to have two parallel tracks with many interesting topics, also touching on the new 6.0 release. At the end of the day, we will have an expert panel w… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-university-live-june-18-2024/2179 ## Headings Structure: H1: ScyllaDB University LIVE - June 18, 2024 H3: Related topics ## Main Content: H1: ScyllaDB University LIVE - June 18, 2024 H3: Related topics Join us for a half day of (free) instructor-led training. We’re going to have two parallel tracks with many interesting topics, also touching on the new 6.0 release. At the end of the day, we will have an expert panel with @avikivity, our CTO, and some other experts, who will be taking your questions. Save your spot here, hope to see you there! --- ### Page: https://forum.scylladb.com/t/i-attempted-to-set-up-a-scylladb-using-three-nodes-but-its-not-working-as-expected/2185 Title: I attempted to set up a ScyllaDB using three nodes, but it's not working as expected - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I am trying to set up a Scylla cluster on three separate servers, following this course: courses/s905-scylla-university-live-2024-essentials-track/lessons/scylla-essentials/topic/consistency-level-demo-part-1 Server 1(… Language: en Canonical URL: https://forum.scylladb.com/t/i-attempted-to-set-up-a-scylladb-using-three-nodes-but-its-not-working-as-expected/2185 ## Headings Structure: H1: I attempted to set up a ScyllaDB using three nodes, but it's not working as expected H1: Server 1(10.0.0.81) H1: Server 2(10.0.0.147) H1: Server3(10.0.0.18) H1: Server1 H1: Server2 H1: Server 3 H3: Related topics ## Main Content: H1: I attempted to set up a ScyllaDB using three nodes, but it's not working as expected H1: Server 1(10.0.0.81) H1: Server 2(10.0.0.147) H1: Server3(10.0.0.18) H1: Server1 H1: Server2 H1: Server 3 H3: Related topics I am trying to set up a Scylla cluster on three separate servers, following this course: courses/s905-scylla-university-live-2024-essentials-track/lessons/scylla-essentials/topic/consistency-level-demo-part-1 However, I encountered the following results and here are the logs for each node. Could you help me identify the problem? Usually it takes a minute or so for the the nodes to connect etc. Did you try waiting a bit and running the nodetool status command again? I have set up the cluster and waited for 3 hours, but the status remains the same. The logs of each node indicate that the connections are established correctly, but only one node is visible in the status The issue is that you’re doing this on separate standalone docker nodes. To get this to work, you’ll need to export the ports used by ScyllaDB. At a minimum, you should add -p 7000:7000 to allow the nodes to talk to each other. ScyllaDB’s ports are listed at its Administration Guide, which I can’t link to. My current compose file for this doing docker testing with separate standalone docker hosts looks like: I don’t claim this is a best practice setup, but it ended up working for testing. I ended up using host mode networking rather than port forwarding. The hostname/extra_hosts combo is necessary because of how scylla decides what IPs to listen on, it needs to be able to resolve its hostname to the external IP address of the docker host its running on, not the internal docker assigned IP. The port section is unnecessary in host mode, it’s there for documentation. --- ### Page: https://forum.scylladb.com/t/large-io-starvation-with-low-query-io-delay/2186 Title: Large io starvation with low query IO delay - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, in our load testing we are currently seeing a large Disk query starvation time (up to 30ms) and low disk query IO queue delay (<1ms). This combination is not mentioned in the monitoring dashboards. What does this i… Language: en Canonical URL: https://forum.scylladb.com/t/large-io-starvation-with-low-query-io-delay/2186 ## Headings Structure: H1: Large io starvation with low query IO delay H3: Related topics ## Main Content: H1: Large io starvation with low query IO delay H3: Related topics in our load testing we are currently seeing a large Disk query starvation time (up to 30ms) and low disk query IO queue delay (<1ms). This combination is not mentioned in the monitoring dashboards. What does this indicate? Hard to tell without additional info. Can you please follow How to Report a ScyllaDB Problem | ScyllaDB Docs and share more info? Hi, sure, we can get more stats. I just did not want to overload you with specifics. I was more interested in understanding the different metrics. Here is our understanding of the metrics: Is this explanation correct? Its worth noting, that we have kernel block cache activated, as we run on HDDs here. --- ### Page: https://forum.scylladb.com/t/maximum-blob-binary-large-object-size-limit-and-recommendations/2187 Title: Maximum Blob (Binary Large OBject) size limit and recommendations - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/maximum-blob-binary-large-object-size-limit-and-recommendations/2187 ## Headings Structure: H1: Maximum Blob (Binary Large OBject) size limit and recommendations H3: Related topics ## Main Content: H1: Maximum Blob (Binary Large OBject) size limit and recommendations H3: Related topics Originally from the User Slack @Kishore: Hello, What is the maximum blob size in Scylla? @dor: There is no limit. We don’t recommend more than several MBs. @avi: Scylla will reject blobs larger than half the commitlog segment size, so 16MB @Kishore: okay, thank you. --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2023-1-9/2189 Title: [RELEASE] ScyllaDB Enterprise 2023.1.9 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2023.1.9, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2023.1 LTS Release. 2023.1.9 patch release includes multiple minor bug fixes… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2023-1-9/2189 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2023.1.9 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2023.1.9 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2023.1.9, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2023.1 LTS Release. 2023.1.9 patch release includes multiple minor bug fixes. ScyllaDB Enterprise’s latest Long Term Support (LTS) is 2024.1. While we continue to support 2023.1, you are encouraged to upgrade to it in coordination with the ScyllaDB Support team. The following issues are fixed in this release (with an open-source reference, if available): --- ### Page: https://forum.scylladb.com/t/mutli-table-batches-atomicity-and-isolation/2190 Title: Mutli table Batches Atomicity and Isolation - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/mutli-table-batches-atomicity-and-isolation/2190 ## Headings Structure: H1: Mutli table Batches Atomicity and Isolation H3: Related topics ## Main Content: H1: Mutli table Batches Atomicity and Isolation H3: Related topics Originally from the User Slack @scyllero: Hi, new user here. Are scylladb multi table batch updates atomic and isolated within the same node? @Botond_Dénes: No, multi-table batches are never atomic and are not isolated. Only batches that affect the same partition of the same table are atomic. @scyllero: “… dml statements to achieve atomicity and isolation when targeting a single partition or only atomicity when targeting multiple partitions” I got that from Cassandra docs for batch command. So I thought atomicity is always guaranteed (through logged), but atomicity AND isolation only if the batch comprises a single partition of a single table. Can you confirm your second answer please, that changes a lot for me @Botond_Dénes: Multi-partition batches (from the same table) are only atomic if they happen to all be owned by the same replica. And this is very hard to guarantee and thus it is best to consider it non-atomic. @scyllero: Ok, so basically there is no way that I can create , let say and order along their items, ATOMICALLY. I mean, this wouldn’t work atomically: begin batch insert into orders… (order_id = 1) insert into order_items… (order_id=1) apply batch @Botond_Dénes: It depends, if all the content of the batch refers to a single partition, then it is atomic. Otherwise, you may observe a state, where some of the statements already applied but others didn’t yet. You can de-normalize your data-model and have everything you want to update in a single table and partition. Then you can make the updates atomic. @scyllero: Im not getting why it dependes. In that example we are talking about two partitions from two different tables partition order_id=1 for orders AND partition order_id=1 for order_partition @Botond_Dénes: Yes, I noticed that after I typed my answer. In that case indeed there is no way to make it atomic. --- ### Page: https://forum.scylladb.com/t/release-scylla-operator-1-13-0/2191 Title: [RELEASE] Scylla Operator 1.13.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of Scylla Operator 1.13.0. Scylla Operator is an open-source project that helps ScyllaDB Open Source and ScyllaDB Enterprise users run ScyllaDB on Kubernetes. The Scy… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-operator-1-13-0/2191 ## Headings Structure: H1: [RELEASE] Scylla Operator 1.13.0 H1: Notable changes H1: Supported versions H1: Upgrade instructions H2: Getting started with Scylla Operator H2: Related Links H3: Related topics ## Main Content: H1: [RELEASE] Scylla Operator 1.13.0 H1: Notable changes H1: Supported versions H1: Upgrade instructions H2: Getting started with Scylla Operator H2: Related Links H3: Related topics The ScyllaDB team is pleased to announce the release of Scylla Operator 1.13.0. Scylla Operator is an open-source project that helps ScyllaDB Open Source and ScyllaDB Enterprise users run ScyllaDB on Kubernetes. The Scylla Operator manages ScyllaDB clusters deployed to Kubernetes and automates tasks related to operating a ScyllaDB cluster, like installation, vertical and horizontal scaling, as well as rolling upgrades. Scylla Operator 1.13.0 improves stability and brings new features. As with all of our releases, all API changes are backward compatible. For more changes and details check out the GitHub release notes. Upgrading from v1.12.x with kubectl apply doesn’t require any extra action, just take the manifest from v1.13.0 tag and substitute the released image. Using helm requires a mandatory manual step for every release because helm can’t handle CRDs updates. For details, see our upgrade documentation. We’ll welcome your feedback! Feel free to open an issue or reach out on the #scylla-operator channel in Scylla User Slack. --- ### Page: https://forum.scylladb.com/t/why-is-only-the-update-item-latency-relatively-high/2192 Title: Why is only the update_item latency relatively high? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: During the normal pressure testing process, there is a high latency when only the update_item operation occurs. Why does this phenomenon occur? The CPU load and TPS are not high. Language: en Canonical URL: https://forum.scylladb.com/t/why-is-only-the-update-item-latency-relatively-high/2192 ## Headings Structure: H1: Why is only the update_item latency relatively high? H3: Related topics ## Main Content: H1: Why is only the update_item latency relatively high? H3: Related topics During the normal pressure testing process, there is a high latency when only the update_item operation occurs. Why does this phenomenon occur? The CPU load and TPS are not high. The update operation is an LWT request. And there are some timeout return messages, but the actual request time is within 200ms. (These are known phenomena). Update_items will heavily update the same partitionion. Will this behavior cause an increase in latency? Yes, concurrent updates to the same partition will increase latency. Make sure to use the latest version of the drivers, which mitigate this somewhat. I found that when there is put traffic, the update latency will increase, and stopping will decrease. What is the situation? I am using Raid5, is it okay with this? --- ### Page: https://forum.scylladb.com/t/tried-to-execute-unprepared-query-error/2194 Title: Tried to execute unprepared query ERROR - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, I am facing an issue with scyllaDB open source hosted on a linux server. I am using java-driver-core 4.17.0.1 to connect to the database. The DB was running fine for a few days, however now I am frequently getting a… Language: en Canonical URL: https://forum.scylladb.com/t/tried-to-execute-unprepared-query-error/2194 ## Headings Structure: H1: Tried to execute unprepared query ERROR H3: Related topics ## Main Content: H1: Tried to execute unprepared query ERROR H3: Related topics Hi, I am facing an issue with scyllaDB open source hosted on a linux server. I am using java-driver-core 4.17.0.1 to connect to the database. The DB was running fine for a few days, however now I am frequently getting an error: How can I resolve this issue? Has something significant happened to the cluster in that time? Have you added/removed nodes or performed any unusual operations? Regardless of that, you can try setting in your configuration advanced.prepared-statements.prepared-cache.weak-values to false. For reference, here’s manual page on configuration: ht tps://java-driver.docs.scylladb.com/stable/manual/core/configuration/ This will make driver use strong values for prepared statements cache instead of weak values. This may increase memory usage of your application so keep that in mind. No I haven’t added any nodes. There is only a single scylla-server running on an ubuntu server. I tried setting prepared_statements_cache_size_mb to a higher value, but the behaviour is still the same. Isn’t the query supposed to be reprepared if its not found in the cache automatically? How can I prevent this from happening without increasing memory usage? Yes, the query will be reprepared as long as the driver has the data. There are 2 things you should check for: If it turns out that 2. is the cause then your memory usage will probably still rise a little, but there’s no helping it. To be clear that would be normal memory usage. Your current usage is lower because you are discarding something needed (assuming 2. is the cause) --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-52-2024-06-21/2197 Title: Last week in scylla-cluster-tests.git master (issue #52; 2024-06-21) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the ba94f20a…335a3018 range are covered. There were 30 non-merge commits from 7 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-52-2024-06-21/2197 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #52; 2024-06-21) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #52; 2024-06-21) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the ba94f20a…335a3018 range are covered. There were 30 non-merge commits from 7 authors in that period. Some notable commits: We introduced the parallel_node_operations config parameter to start DB nodes in parallel (since 6.0). This speeds up cluster provisioning and allows further parallel node operations in SCT. Temporarily enabled for a small subset of tests, it will soon be enabled by default. Previous ScyllaDB versions will start sequentially regardless of the setting. Disabled apt daily triggers at an early stage in cloud-init. Improved labelling of monitoring targets to be DC-aware, enabling node filtering with the ‘dc’ filter. With centos8 deprecated, we introduced the exact same test on top of centos9. Updated client/server certificates management, so each test generates and distributes its own individual certificates for each node, allowing mutual TLS and hostname validation. Committing to the SCT repository is now faster thanks to the replacement of pylint with Ruff. Ruff, a Rust-based linter for Python, works incredibly fast and helps address performance issues in SCT CI. Added the replication_factors property to the ReplicationStrategy class to obtain the RF for each DC used in the keyspace. This can be used in various situations, such as filtering out keyspaces with RF one to avoid using them with the toggle_table_gc_mode nemesis, as it is not supported in that case. It has become increasingly useful to build and test with SCT custom ScyllaDB images. We added the possibility to bring your own ScyllaDB for longevity and upgrades with unmerged changes to the ScyllaDB repo. To use this, provide a path to the Jenkins job that builds the image in the byo_job_path parameter. More details will be showcased soon in the R&D meeting. As the tablets feature is now enabled by default, the tablets-enabled tests Jenkins folder has been renamed to tablets-disabled. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/reshaping-details-in-scylladb/2199 Title: Reshaping Details in ScyllaDB - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Whenver our scylla-server restarts, the reshaping processes take place which takes many hours. Have a few questions regarding the same - How can we see overall progress for the reshaping? How much has completed and how… Language: en Canonical URL: https://forum.scylladb.com/t/reshaping-details-in-scylladb/2199 ## Headings Structure: H1: Reshaping Details in ScyllaDB H3: Related topics ## Main Content: H1: Reshaping Details in ScyllaDB H3: Related topics Whenver our scylla-server restarts, the reshaping processes take place which takes many hours. Have a few questions regarding the same - Looks like your compaction is not in a healthy state. Do you have static shares set for compaction? You cannot abort or skip reshape, but I think you can monitor the progress using the /task_manager REST endpoints. --- ### Page: https://forum.scylladb.com/t/how-to-implement-full-text-search-with-scylladb/2200 Title: How to implement full text search with ScyllaDB? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/how-to-implement-full-text-search-with-scylladb/2200 ## Headings Structure: H1: How to implement full text search with ScyllaDB? H3: Related topics ## Main Content: H1: How to implement full text search with ScyllaDB? H3: Related topics Originally from the User Slack @Mosca_Careca: Is full text search supported in scyllaDB or is a secondary service like ElasticSearch still needed? @Felipe_Cardeneti_Mendes: You can probably mimic it if the query in question always specify the partition key, but if it doesn’t then a secondary service is the way to go @scyllero: @Felipe_Cardeneti_Mendes and how would you handle that? Let say you have a chat app: you write the message to Scylla first, and then to elastic . Ok, but how do you do that reliably? @Felipe_Cardeneti_Mendes: If you are using ScyllaDB as your source of truth, in that case you would deltas via CDC. https://github.com/scylladb/scylla-cdc-source-connector/blob/master/QUICKSTART-ELASTICSEARCH-INTEGRATION.md GitHub: scylla-cdc-source-connector/QUICKSTART-ELASTICSEARCH-INTEGRATION.md at master · scylladb/scylla-cdc-source-connector @Mosca_Careca: Does ScyllaDB even have any type of Keyword-based search over text fields? Let’s say I want my users to be able to give me a query so that I can retrieve products whose description fits that query. Am I just forced to use a secondary service or is there any other way of solving this problem? Technically I could have a description_tokens table and implement my own full-text search, but I’d rather not reinvent the wheel here @scyllero: Go for a secondary service. Do not waste your time. Refer to that last reply from @Felipe_Cardeneti_Mendes The flow would be your app emitting products - scylladb - cdc source connector - kafka - sink connector (or custom consumer) - mongodb (if text queries are simple) or elasticsearch @Mosca_Careca: Yeah sadly, that’ll have to be done. I feel like that many things running on my machine will nuke it, but we’ll see. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-235-2024-06-23/2202 Title: Last week in scylladb.git master (issue #235; 2024-06-23) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the bbb424a757…d8009ed843 range are covered. There were 90 non-merge commits from 19 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-235-2024-06-23/2202 ## Headings Structure: H1: Last week in scylladb.git master (issue #235; 2024-06-23) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #235; 2024-06-23) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the bbb424a757…d8009ed843 range are covered. There were 90 non-merge commits from 19 authors in that period. Some notable commits: A DDL statement CREATE … IF NOT EXISTS will now tell the driver a schema change occurred even if it did not make any changes (because the table or keyspace already existed). This helps tools like cassandra-stress that create keyspaces from multiple processes; before the change such a tool could miss the keyspace creation. Support for the Thrift protocol, deprecated since ScyllaDB 5.2, has been removed. There is now configuration for the number of concurrent reads allowed for maintenance operations (e.g. repair). This helps repair in some situations where different nodes have different shard counts. Internal access to the roles table was optimized. Off-strategy compaction is used to make sstables conform to the compaction strategy after an operation such as repair. Off-strategy compaction for TWCS will now have less space overhead. Statements such as SELECT count(*) use an internal map-reduce service to parallelize the query. We no longer do so for single-partition queries as they don’t benefit from it. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-6-0-1/2203 Title: [RELEASE] ScyllaDB 6.0.1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 6.0.1, a bugfix patch release of the ScyllaDB 6.0 stable branch. ScyllaDB Open Source 6.0.1, like all past and future 6.x.y releases, is backward compatible and supports r… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-6-0-1/2203 ## Headings Structure: H1: [RELEASE] ScyllaDB 6.0.1 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 6.0.1 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 6.0.1, a bugfix patch release of the ScyllaDB 6.0 stable branch. ScyllaDB Open Source 6.0.1, like all past and future 6.x.y releases, is backward compatible and supports rolling upgrades. Users are encouraged to upgrade to 6.0.1. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/scylladb-listens-on-all-interfaces/2204 Title: Scylladb listens on all interfaces - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello, I am here to find out and if it would be possible to have connectivity with clients on all the interfaces I have on the server and connection between the various nodes only with internal IP. Thank you all Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-listens-on-all-interfaces/2204 ## Headings Structure: H1: Scylladb listens on all interfaces H3: Related topics ## Main Content: H1: Scylladb listens on all interfaces H3: Related topics Hello, I am here to find out and if it would be possible to have connectivity with clients on all the interfaces I have on the server and connection between the various nodes only with internal IP. Thank you all Perhaps Configure Scylla Networking with Multiple NIC/IP Combinations | ScyllaDB Docs describes your case? --- ### Page: https://forum.scylladb.com/t/using-a-clustering-key-impact-on-performance-data-distribution-and-partition-size/2205 Title: Using a clustering key, impact on performance, data distribution and partition size - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/using-a-clustering-key-impact-on-performance-data-distribution-and-partition-size/2205 ## Headings Structure: H1: Using a clustering key, impact on performance, data distribution and partition size H3: Related topics ## Main Content: H1: Using a clustering key, impact on performance, data distribution and partition size H3: Related topics Originally from the User Slack @Bohdan_Smal: Hello everyone, I have a somewhat general question. Could you please advise on the potential risks of not using a clustering key? Currently, we have a table where the primary key is a combination of brand and client_id, and the clustering key is transaction_id. We have observed that we could achieve better performance and more even data distribution if we don’t use a clustering key. Instead, we would set the primary key as transaction_id, brand, and client_id. This way, each transaction becomes a separate partition, ensuring no imbalance in partition distribution across nodes, even if some clients have more transactions. However, we are not fully aware of the potential risks associated with having a large number of partitions. Can anyone explain the possible risks of this approach? Thank you! @Karol_Baryła: I’m not sure what are the performance implications of large number of partitions. What comes to my mind is usability. With such a schema you can’t e.g: • Select all transactions for a given user without using ALLOW FILTERING and making the query much slower this way. • Use LWT / Batches with LWT to atomically update several transactions for a given user. @Bohdan_Smal: Got it, thank you for your response. Have a great day! @avi: In fact having more and smaller partitions is better than having fewer and larger partitions. So if you don’t need a clustering key for sorting and grouping, don’t use it. @Bohdan_Smal: thank you) --- ### Page: https://forum.scylladb.com/t/throughput-and-latency-impact-when-doing-a-rolling-restart-in-a-3-node-cluster/2213 Title: Throughput and latency impact when doing a rolling restart in a 3 node cluster - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/throughput-and-latency-impact-when-doing-a-rolling-restart-in-a-3-node-cluster/2213 ## Headings Structure: H1: Throughput and latency impact when doing a rolling restart in a 3 node cluster H3: Related topics ## Main Content: H1: Throughput and latency impact when doing a rolling restart in a 3 node cluster H3: Related topics Originally from the User Slack I have a 3 node cluster with and RF of 3 and network topology strategy. The clients have a CL of Local Quorum. When I did a rolling restart of the nodes 1 at a time I saw reads and writes dramatically drop off. My understanding was that if local quorum was in place with this that reads would still be maintained? So I’m not sure what I misunderstood or what setting I got wrong? @Felipe_Cardeneti_Mendes: Well yes but you lose 1/3 of your capacity as you shutdown a replica. If the remaining capacity is unable to supply increased demand your latency increases. Sometimes it may also happen that you could have restarted too fast and all nodes came up with cold caches, thus reads became more expensive as they needed to go to disk instead of fetching from the database cache. These are some examples, you probably want to look into metrics and find out what really happened @J: Ok thanks. I noticed it as soon as the first node came down. So I don’t think it’s that. In terms of capacity we were only doing 80krps across the cluster when it came down. These nodes are in K8s running on i4i.8xlarge each pod has 12 cpu and 32Gi RAM. Is that not enough capacity? Is there any other setting I should check? I did during this move the CL to one and it continued fine during the rest of the rolling restarts. But I do see in the logs both our clients saying cannot achieve CL LOCAL_ONE requires 1, alive 0. Though at no point was more than one node down @Felipe_Cardeneti_Mendes: From your description seems like enough capacity. I doubt CPU was a bottleneck, but check on the advanced dashboard if anything pops up wrt disks or CPU @J: Ok thanks, let me look into that --- ### Page: https://forum.scylladb.com/t/is-it-possible-to-have-read-only-access-in-scylladb-or-apache-cassandra/2217 Title: Is it possible to have read-only access in ScyllaDB (or Apache Cassandra)? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/is-it-possible-to-have-read-only-access-in-scylladb-or-apache-cassandra/2217 ## Headings Structure: H1: Is it possible to have read-only access in ScyllaDB (or Apache Cassandra)? H3: Related topics ## Main Content: H1: Is it possible to have read-only access in ScyllaDB (or Apache Cassandra)? H3: Related topics Originally from the User Slack @Shubham_Jain: Hi guys, Is it possible to run read-only nodes in Scylla ? @Piotr_Smaroń: I think if you enable authentication and grant read-only permissions to users, then yes @Shubham_Jain: Yeah, I thought the same Thanks @Piotr_Smaroń @Piotr_Smaroń: BTW. here’s a related doc - https://opensource.docs.scylladb.com/stable/operating-scylla/security/authorization.html Grant Authorization CQL Reference | ScyllaDB Docs --- ### Page: https://forum.scylladb.com/t/scylla-write-failure/2219 Title: Scylla Write failure - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi All, I am getting below exception while trying to load data in Scylla DB using PySpark Application (running in AWS EMR) - Traceback (most recent call last): File “/mnt/tmp/aip-workflows/scylla-load/src/s3-to-scylla… Language: en Canonical URL: https://forum.scylladb.com/t/scylla-write-failure/2219 ## Headings Structure: H1: Scylla Write failure H3: Related topics ## Main Content: H1: Scylla Write failure H3: Related topics I am getting below exception while trying to load data in Scylla DB using PySpark Application (running in AWS EMR) - Traceback (most recent call last): File “/mnt/tmp/aip-workflows/scylla-load/src/s3-to-scylla.py”, line 215, in source_json.write.format(cassandra_write_format).mode(‘append’).options( File “/usr/lib/spark/python/lib/pyspark.zip/pyspark/sql/readwriter.py”, line 1461, in save File “/usr/lib/spark/python/lib/py4j-0.10.9.7-src.zip/py4j/java_gateway.py”, line 1322, in call File “/usr/lib/spark/python/lib/pyspark.zip/pyspark/errors/exceptions/captured.py”, line 179, in deco File “/usr/lib/spark/python/lib/py4j-0.10.9.7-src.zip/py4j/protocol.py”, line 326, in get_return_value py4j.protocol.Py4JJavaError: An error occurred while calling o115.save. : org.apache.spark.SparkException: Job aborted due to stage failure: Authorized committer (attemptNumber=0, stage=2, partition=2133) failed; but task commit success, data duplication may happen. reason=ExceptionFailure(java.io.IOException,Failed to write statements to oeidp_artemis.json_content. The latest exception was Cassandra failure during write query at consistency ALL (2 responses were required but only 1 replica responded, 1 failed) Please check the executor logs for more exceptions and information ,[Ljava.lang.StackTraceElement;@29602c51,java.io.IOException: Failed to write statements to oeidp_artemis.json_content. The latest exception was Cassandra failure during write query at consistency ALL (2 responses were required but only 1 replica responded, 1 failed) Please check the executor logs for more exceptions and information latest exception was Cassandra failure during write query at consistency ALL (2 responses were required but only 1 replica responded, 1 failed) Please check the executor logs for more exceptions and information latest exception was Cassandra failure during write query at consistency ALL (2 responses were required but only 1 replica responded, 1 failed) Please check the executor logs for more exceptions and information ,Some(org.apache.spark.ThrowableSerializationWrapper@7c4eed1c),Vector(AccumulableInfo(83,None,Some(64134),None,false,true,None), AccumulableInfo(85,None,Some(0),None,false,true,None), AccumulableInfo(86,None,Some(10),None,false,true,None), AccumulableInfo(112,None,Some(150024615),None,false,true,None), AccumulableInfo(113,None,Some(19624),None,false,true,None)),Vector(LongAccumulator(id: 83, name: Some(internal.metrics.executorRunTime), value: 64134), LongAccumulator(id: 85, name: Some(internal.metrics.resultSize), value: 0), LongAccumulator(id: 86, name: Some(internal.metrics.jvmGCTime), value: 10), LongAccumulator(id: 112, name: Some(internal.metrics.input.bytesRead), value: 150024615), LongAccumulator(id: 113, name: Some(internal.metrics.input.recordsRead), value: 19624)),WrappedArray(3078441872, 162501632, 0, 0, 35523537, 0, 35523537, 0, 290824738, 0, 0, 0, 0, 0, 0, 0, 21, 1468, 8, 1076, 2544)) at org.apache.spark.scheduler.DAGScheduler.failJobAndIndependentStages(DAGScheduler.scala:3067) at org.apache.spark.scheduler.DAGScheduler.$anonfun$abortStage$2(DAGScheduler.scala:3003) at org.apache.spark.scheduler.DAGScheduler.$anonfun$abortStage$2$adapted(DAGScheduler.scala:3002) at scala.collection.mutable.ResizableArray.foreach(ResizableArray.scala:62) at scala.collection.mutable.ResizableArray.foreach$(ResizableArray.scala:55) at scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:49) at org.apache.spark.scheduler.DAGScheduler.abortStage(DAGScheduler.scala:3002) at org.apache.spark.scheduler.DAGScheduler.$anonfun$handleStageFailed$1(DAGScheduler.scala:1311) at org.apache.spark.scheduler.DAGScheduler.$anonfun$handleStageFailed$1$adapted(DAGScheduler.scala:1311) at scala.Option.foreach(Option.scala:407) at org.apache.spark.scheduler.DAGScheduler.handleStageFailed(DAGScheduler.scala:1311) at org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.doOnReceive(DAGScheduler.scala:3268) at org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.onReceive(DAGScheduler.scala:3205) at org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.onReceive(DAGScheduler.scala:3194) at org.apache.spark.util.EventLoop$$anon$1.run(EventLoop.scala:49) at org.apache.spark.scheduler.DAGScheduler.runJob(DAGScheduler.scala:1041) at org.apache.spark.SparkContext.runJob(SparkContext.scala:2406) at org.apache.spark.sql.execution.datasources.v2.V2TableWriteExec.writeWithV2(WriteToDataSourceV2Exec.scala:385) at org.apache.spark.sql.execution.datasources.v2.V2TableWriteExec.writeWithV2$(WriteToDataSourceV2Exec.scala:359) at org.apache.spark.sql.execution.datasources.v2.AppendDataExec.writeWithV2(WriteToDataSourceV2Exec.scala:225) at org.apache.spark.sql.execution.datasources.v2.V2ExistingTableWriteExec.run(WriteToDataSourceV2Exec.scala:337) at org.apache.spark.sql.execution.datasources.v2.V2ExistingTableWriteExec.run$(WriteToDataSourceV2Exec.scala:336) at org.apache.spark.sql.execution.datasources.v2.AppendDataExec.run(WriteToDataSourceV2Exec.scala:225) at org.apache.spark.sql.execution.datasources.v2.V2CommandExec.result$lzycompute(V2CommandExec.scala:43) at org.apache.spark.sql.execution.datasources.v2.V2CommandExec.result(V2CommandExec.scala:43) at org.apache.spark.sql.execution.datasources.v2.V2CommandExec.executeCollect(V2CommandExec.scala:49) at org.apache.spark.sql.execution.QueryExecution$$anonfun$eagerlyExecuteCommands$1.$anonfun$applyOrElse$1(QueryExecution.scala:113) at org.apache.spark.sql.catalyst.QueryPlanningTracker$.withTracker(QueryPlanningTracker.scala:108) at org.apache.spark.sql.execution.SQLExecution$.withTracker(SQLExecution.scala:255) at org.apache.spark.sql.execution.SQLExecution$.executeQuery$1(SQLExecution.scala:129) at org.apache.spark.sql.execution.SQLExecution$.$anonfun$withNewExecutionId$9(SQLExecution.scala:165) at org.apache.spark.sql.catalyst.QueryPlanningTracker$.withTracker(QueryPlanningTracker.scala:108) at org.apache.spark.sql.execution.SQLExecution$.withTracker(SQLExecution.scala:255) at org.apache.spark.sql.execution.SQLExecution$.$anonfun$withNewExecutionId$8(SQLExecution.scala:165) at org.apache.spark.sql.execution.SQLExecution$.withSQLConfPropagated(SQLExecution.scala:276) at org.apache.spark.sql.execution.SQLExecution$.$anonfun$withNewExecutionId$1(SQLExecution.scala:164) at org.apache.spark.sql.SparkSession.withActive(SparkSession.scala:900) at org.apache.spark.sql.execution.SQLExecution$.withNewExecutionId(SQLExecution.scala:70) at org.apache.spark.sql.execution.QueryExecution$$anonfun$eagerlyExecuteCommands$1.applyOrElse(QueryExecution.scala:110) at org.apache.spark.sql.execution.QueryExecution$$anonfun$eagerlyExecuteCommands$1.applyOrElse(QueryExecution.scala:101) at org.apache.spark.sql.catalyst.trees.TreeNode.$anonfun$transformDownWithPruning$1(TreeNode.scala:503) at org.apache.spark.sql.catalyst.trees.CurrentOrigin$.withOrigin(origin.scala:76) at org.apache.spark.sql.catalyst.trees.TreeNode.transformDownWithPruning(TreeNode.scala:503) at org.apache.spark.sql.catalyst.plans.logical.LogicalPlan.org$apache$spark$sql$catalyst$plans$logical$AnalysisHelper$$super$transformDownWithPruning(LogicalPlan.scala:33) at org.apache.spark.sql.catalyst.plans.logical.AnalysisHelper.transformDownWithPruning(AnalysisHelper.scala:267) at org.apache.spark.sql.catalyst.plans.logical.AnalysisHelper.transformDownWithPruning$(AnalysisHelper.scala:263) at org.apache.spark.sql.catalyst.plans.logical.LogicalPlan.transformDownWithPruning(LogicalPlan.scala:33) at org.apache.spark.sql.catalyst.plans.logical.LogicalPlan.transformDownWithPruning(LogicalPlan.scala:33) at org.apache.spark.sql.catalyst.trees.TreeNode.transformDown(TreeNode.scala:479) at org.apache.spark.sql.execution.QueryExecution.eagerlyExecuteCommands(QueryExecution.scala:101) at org.apache.spark.sql.execution.QueryExecution.commandExecuted$lzycompute(QueryExecution.scala:88) at org.apache.spark.sql.execution.QueryExecution.commandExecuted(QueryExecution.scala:86) at org.apache.spark.sql.execution.QueryExecution.assertCommandExecuted(QueryExecution.scala:151) at org.apache.spark.sql.DataFrameWriter.runCommand(DataFrameWriter.scala:859) at org.apache.spark.sql.DataFrameWriter.saveInternal(DataFrameWriter.scala:312) at org.apache.spark.sql.DataFrameWriter.save(DataFrameWriter.scala:248) at java.base/jdk.internal.reflect.NativeMethodAccessorImpl.invoke0(Native Method) at java.base/jdk.internal.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:77) at java.base/jdk.internal.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43) at java.base/java.lang.reflect.Method.invoke(Method.java:568) at py4j.reflection.MethodInvoker.invoke(MethodInvoker.java:244) at py4j.reflection.ReflectionEngine.invoke(ReflectionEngine.java:374) at py4j.Gateway.invoke(Gateway.java:282) at py4j.commands.AbstractCommand.invokeMethod(AbstractCommand.java:132) at py4j.commands.CallCommand.execute(CallCommand.java:79) at py4j.ClientServerConnection.waitForCommands(ClientServerConnection.java:182) at py4j.ClientServerConnection.run(ClientServerConnection.java:106) at java.base/java.lang.Thread.run(Thread.java:840) Also I am unable to find any error log path for Scylla DB installed in CentOS system. Can you please let me know possible solution for this issue? Though other table loads are running fine, facing issue for one of the table only. ScyllaDB log is available in journalctl Cassandra failure during write query at consistency ALL (2 responses were required but only 1 replica responded, 1 failed) Seems to be an issue related to replication factor and consistency. Try to run it again with setting consistency as 1, it should work. To debug further, I would need more information regarding your cluster ( number of nodes, replication factor of keyspace, consistency set at ingestion if any?) Seems to be an issue related to replication factor and consistency. Try to run it again with setting consistency as 1, it should work. To debug further, I would need more information regarding your cluster ( number of nodes, replication factor of keyspace, consistency set at ingestion if any?) How many nodes in your cluster? --- ### Page: https://forum.scylladb.com/t/release-scylla-manager-3-3-0/2231 Title: [RELEASE] Scylla Manager 3.3.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces the release of ScyllaDB Manager 3.3 production-ready ScyllaDB Manager minor release of the stable ScyllaDB Manager 3.3 branch. ScyllaDB Manager is a centralized cluster administration and rec… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-manager-3-3-0/2231 ## Headings Structure: H1: [RELEASE] Scylla Manager 3.3.0 H3: Related topics ## Main Content: H1: [RELEASE] Scylla Manager 3.3.0 H3: Related topics The ScyllaDB team announces the release of ScyllaDB Manager 3.3 production-ready ScyllaDB Manager minor release of the stable ScyllaDB Manager 3.3 branch. ScyllaDB Manager is a centralized cluster administration and recurrent tasks automation tool. In version 3.3, we ensure that ScyllaDB Manager works seamlessly with ScyllaDB servers supporting the Raft Consensus Algorithm and Data Distribution with Tablets. Tablets and Raft will be available in the upcoming ScyllaDB Enterprise Feature release, and in ScyllaDB Open Source 6.0. Enabling Raft has impacted how ScyllaDB Manager restores the schema on servers with this feature enabled (issues #3868, #3887). Support for Data Distribution with Tablets is ensured by addressing the following issue #3516. ScyllaDB Enterprise customers are encouraged to upgrade to ScyllaDB Manager 3.3 in coordination with the ScyllaDB support team. The new release includes upgrades of both ScyllaDB Manager Server and Agent. Scylla Manager 3.3 supports the following Scylla Enterprise releases: And the following Open Source release (limited to 5 nodes see license): You can install and run Scylla Manager on Kubernetes using Scylla Operator. More here. --- ### Page: https://forum.scylladb.com/t/how-to-sync-from-cassandra-scylladb-to-elasticsearch-without-using-cdc/2232 Title: How to sync from cassandra/scylladb to elasticsearch without using CDC? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/how-to-sync-from-cassandra-scylladb-to-elasticsearch-without-using-cdc/2232 ## Headings Structure: H1: How to sync from cassandra/scylladb to elasticsearch without using CDC? H3: Related topics ## Main Content: H1: How to sync from cassandra/scylladb to elasticsearch without using CDC? H3: Related topics Originally from the User Slack @scyllero: Hi everyone, Lets say that for whatever reason you’re not allowed to use cdc. What would you use to sync from cassandra/scylladb to elasticsearch? I’m not asking for a tutorial, just some keywords to look up to I was thinking some kind of ES: write to amqp or kafka, then have 2 consumers write to cassandra and elasticsearch concurrently. This could lead to problems. Are there better approaches? @Felipe_Cardeneti_Mendes: well the kafka one would be my pick. Unless you want to reinvent the wheel and defer it to a single consumer to coordinate updating both. Maybe listen to yourself pattern, smth like that. @scyllero: I will look into that pattern. never heard of it --- ### Page: https://forum.scylladb.com/t/connect-with-django/2233 Title: Connect with django - Database Community - ScyllaDB Community NoSQL Forum Meta Description: docker+ django + scyalla Language: en Canonical URL: https://forum.scylladb.com/t/connect-with-django/2233 ## Headings Structure: H1: Connect with django H3: GitHub - r4fek/django-cassandra-engine: Django Cassandra Engine - the... H3: Related topics ## Main Content: H1: Connect with django H3: GitHub - r4fek/django-cassandra-engine: Django Cassandra Engine - the... H3: Related topics docker+ django + scyalla you might want to try out this repo: Django Cassandra Engine - the Cassandra backend for Django - r4fek/django-cassandra-engine I think it should work exactly the same ontop of scylla --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-53-2024-06-28/2235 Title: Last week in scylla-cluster-tests.git master (issue #53; 2024-06-28) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 9b2a1cd9…1932795e range are covered . There were 32 non-merge commits from 8 authors in t… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-53-2024-06-28/2235 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #53; 2024-06-28) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #53; 2024-06-28) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 9b2a1cd9…1932795e range are covered . There were 32 non-merge commits from 8 authors in that period. Some notable commits: SCT now supports zstd compressed coredumps. Fixed problems with GitHub request quota limits by caching issue info in S3. This cache is used in the SkipPerIssue feature, populated several times a day, and can be updated manually from the GitHub Actions UI. Reworked monitoring node setup to support Rocky, so it can be used in manager installation tests for Rocky distro. create_keyspace Tester method now allows specifying tablet configuration when creating a keyspace, enabling disabling tablets or setting the initial number of tablets. Added a new Scylla Manager test performing sanity checks for clusters consisting of vnodes and tablets simultaneously. SCT now supports adding and decommissioning multiple nodes in parallel, useful in scenarios like grow-shrink cluster nemesis. This can be enabled by setting parallel_node_operations and specifying nemesis_add_node_cnt greater than 1. Fixed the issue of broken coredumps by creating a hard link as soon as a coredump is detected before compression and uploading. Bumped scylla-driver from 3.26.8 to 3.26.9. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-236-2024-06-30/2238 Title: Last week in scylladb.git master (issue #236; 2024-06-30) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the d8009ed843…d034cde01f range are covered. There were 75 non-merge commits from 19 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-236-2024-06-30/2238 ## Headings Structure: H1: Last week in scylladb.git master (issue #236; 2024-06-30) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #236; 2024-06-30) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the d8009ed843…d034cde01f range are covered. There were 75 non-merge commits from 19 authors in that period. Some notable commits: During compaction, we estimate the required bloom filter size and allocates it. Now, if after compaction it turns out we over-estimated the bloom filter size, we will rebuild the filter. This conserves memory since bloom filters are always held in RAM. Batches (as generated by the BATCH statement) are held in the system.batchlog table. As this table can accumulate a lot of tombstones, we now take steps to ensure these tombstones are purged eagerly. This is important for repair, which replays the batch log. When a node is started, it will now make a best-effort attempt to notify other nodes that it is up. This speeds up rolling restart, as we don’t have to wait for nodes to notice the node is up via pings. The bundled cqlsh package now includes a Python driver that is compiled for the target architecture, making it faster. Tracing of speculative retries is improved. ALTER KEYSPACE will now refuse to switch a keyspace from tablets to vnodes. The build toolchain is now based on Fedora 40; this moves the compiler from clang 16 to clang 18. Due to compiler immaturity, we previously had to restrict optimization on the aarch64 platform. As the compiler bugs have been fixed, these restrictions are now removed. The source language was updated from C++20 to C++23. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-4-8/2239 Title: [RELEASE] ScyllaDB 5.4.8 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.4.8, a bugfix release of the ScyllaDB 5.4 stable branch. ScyllaDB Open Source 5.4.8, like all past and future 5.x.y releases, is backward compatible and supports rolling… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-4-8/2239 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.4.8 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.4.8 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.4.8, a bugfix release of the ScyllaDB 5.4 stable branch. ScyllaDB Open Source 5.4.8, like all past and future 5.x.y releases, is backward compatible and supports rolling upgrades. Note that the latest ScyllaDB Open Source stable release is 6.0 and you are encouraged to upgrade to it. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-6/2249 Title: [RELEASE] ScyllaDB Enterprise 2024.1.6 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2024.1.6, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2024.1 LTS Release. Related Links Get ScyllaDB Enterprise 2024.1 (customer… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-6/2249 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2024.1.6 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2024.1.6 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2024.1.6, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2024.1 LTS Release. The following issues are fixed in this release (with an open-source reference, if available): --- ### Page: https://forum.scylladb.com/t/scylladb-connectivity-with-managedkafka-gcp/2251 Title: Scylladb-connectivity-with-managedkafka-GCP - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi i want to connect my scylladb with my managed kafka service in GCP without using confluent cloud ?, is that possible to connect Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-connectivity-with-managedkafka-gcp/2251 ## Headings Structure: H1: Scylladb-connectivity-with-managedkafka-GCP H3: Related topics ## Main Content: H1: Scylladb-connectivity-with-managedkafka-GCP H3: Related topics Hi i want to connect my scylladb with my managed kafka service in GCP without using confluent cloud ?, is that possible to connect If you’re using GCPs Kafka Connect offering then according to the public documentation I’ve found it is not possible. Per the limitations listed at Kafka Connect overview  |  Google Cloud Managed Service for Apache Kafka : The service does not support uploading custom connector plugins to your Kafka Connect cluster. Which is unfortunate. However if you could setup your own Kafka Connect cluster and configure it so that it works with brokers used by your Managed Kafka Service then it should be doable. In such case you could run any custom connector on your Connect cluster, which includes GitHub - scylladb/scylla-cdc-source-connector: A Kafka source connector capturing Scylla CDC changes and GitHub - scylladb/kafka-connect-scylladb: Kafka Connect Scylladb Sink One is a CDC connector capable of reading CDC log tables and streaming them to Kafka and another one is a connector reading from Kafka topics and inserting messages into ScyllaDB tables. Before trying to set anything up make sure that you are not hindered by their limitations. Not every complex data type or mode is currently supported. --- ### Page: https://forum.scylladb.com/t/scylla-manager-backup-failure/2255 Title: Scylla-manager backup failure - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello All, A backup to azure from scylladb manager fails when we run a dry run backup using “sctool backup -c Cluster -L azure:scylladb-backup --dry-run”. But when we run the check location on each scylla cluster node … Language: en Canonical URL: https://forum.scylladb.com/t/scylla-manager-backup-failure/2255 ## Headings Structure: H1: Scylla-manager backup failure H3: Setup Azure Blob Storage | ScyllaDB Docs H3: Related topics ## Main Content: H1: Scylla-manager backup failure H3: Setup Azure Blob Storage | ScyllaDB Docs H3: Related topics A backup to azure from scylladb manager fails when we run a dry run backup using “sctool backup -c Cluster -L azure:scylladb-backup --dry-run”. But when we run the check location on each scylla cluster node with scylla-manager-agent check-location --location azure:scylladb-backup it wont return any error. The same thing happens for scylla-manager-agent check-location --debug --location azure:scylla-db what we have done so far; check location on all cluster nodes using the check location command with debug and they all returned without any error! the check location with debug on all cluster nodes returned successfully writing a temp file to the azure location and deleting them. The issue I am having now is when the command from scylladb manager is executed “sctool backup -c Cluster -L azure:scylladb-backup --dry-run”, we get the error below. "we get the error below: Jun 13 08:55:52 scylladb scylla-manager[1957]: {“L”:“ERROR”,“T”:“2024-06-13T08:55:52.685Z”,“N”:“backup”,“M”:“Failed to access location from node”,“node”:“XXX.XX.XX.XX”,“location”:“azure:scylladb-backup”,“error”:“17X.XX.XX.XX: giving up after 2 attempts: after 30s: context deadline” Kindly help with steps or ways to troubleshoot and resolve this issue. scylla-manager-agent reads the configuration directly from /etc/scylla-manager-agent/scylla-manager-agent.yaml when it’s executed. Therefore, any changes you make to the configuration file are automatically reflected when you call the scylla-manager-agent CLI. The Scylla Manager server performs location checks through the agent running as a service on the node. Whenever you make changes to scylla-manager-agent.yaml, you must restart the service to apply the recent configuration changes. Have you updated the location in the .yaml file but haven’t restarted the service yet? Please try restarting the scylla-manager-agent.service on all nodes. ScyllaDB is an Apache Cassandra-compatible NoSQL data store that can handle 1 million transactions per second on a single server. Hello @Karol_Kokoszka , Thank you for the feedback. Yes, the agent service has been restarted on all nodes but it still fails to backup the file to Azure and we still get the following errors in the logs as seen below. Aug 29 10:20:14 scylladb-u scylla-manager[306555]: {“L”:“INFO”,“T”:“2024-08-29T10:20:14.933+0100”,“N”:“cluster.client”,“M”:“HTTP retry backoff”,“operation”:“OperationsCheckPermissions”,“wait”:“1s”,“error”:“after 30s: context deadline exceeded”,“_trace_id”:“1fFp5ltHQnmQd2F38Sy8fg”} Aug 29 10:20:14 scylladb-u scylla-manager[306555]: message repeated 2 times: [ {“L”:“INFO”,“T”:“2024-08-29T10:20:14.933+0100”,“N”:“cluster.client”,“M”:“HTTP retry backoff”,“operation”:“OperationsCheckPermissions”,“wait”:“1s”,“error”:“after 30s: context deadline exceeded”,“_trace_id”:“1fFp5ltHQnmQd2F38Sy8fg”}] Aug 29 10:20:45 scylladb-u scylla-manager[306555]: {“L”:“INFO”,“T”:“2024-08-29T10:20:45.936+0100”,“N”:“backup”,“M”:“Location check FAILED”,“host”:“1XX.XX.XX.XXX”,“location”:“azure:scyllatestbackup”,“error”:“giving up after 2 attempts: after 30s: context deadline exceeded”,“_trace_id”:“1fFp5ltHQnmQd2F38Sy8fg”} Aug 29 10:20:45 scylladb-u scylla-manager[306555]: {“L”:“ERROR”,“T”:“2024-08-29T10:20:45.936+0100”,“N”:“backup”,“M”:"Failed to access location from node. What else can we check and try, could it be an issue with the Azure location/access? 29T10:20:14.933+0100”,“N”:“cluster.client”,“M”:“HTTP retry backoff”,“operation”:“OperationsCheckPermissions”,“wait”:“1s”,“error”:“after 30s: context deadline exceeded”,“_trace_id”:“1fFp5ltHQnmQd2F38Sy8fg”} @Mike The error above suggests that scylla-manager-server cannot access agent’s from the nodes. Can you call sctool status and paste the output ? Status | ScyllaDB Docs Hello @Karol_Kokoszka , kindly see the output for the sctool status below ±—±---------±---------±--------------±-----------±-----±--------±-------±------±-------------------------------------+ | | CQL | REST | Address | Uptime | CPUs | Memory | Scylla | Agent | Host ID | ±—±---------±---------±--------------±-----------±-----±--------±-------±------±-------------------------------------+ | UN | UP (0ms) | UP (1ms) | 1XX.XX.XX.1XX | 116h19m19s | 8 | 62.797G | 5.2.6 | 3.3.0 | b8f1375x-8855-4f02-913e-b093fcaff1c3 | | UN | UP (0ms) | UP (0ms) | 1XX.XX.XX.2XX | 572h41m1s | 8 | 62.788G | 5.2.6 | 3.3.0 | ad57011x-6597-427a-9f42-15390fa7c379 | | UN | UP (0ms) | UP (0ms) | 1XX.XX.XX.3XX | 1315h1m7s | 8 | 62.788G | 5.2.6 | 3.3.0 | b3577f6x-dc40-4330-bb7a-0612724f68db | ±—±---------±---------±--------------±-----------±-----±--------±-------±------±-------------------------------------+ Location check FAILED @Mike Is there a chance to see the scylla-manager-agent logs as well ? I have no clue yet, what may be a reason. Call to manager-agent times out definitely. You can change the log-level in scylla-manager-agent.yaml as well to see bit more of the details, by providing this config value (top level) to the file: Hello @Karol_Kokoszka, Thank for the help, I appreciate. We finally resolved the issue(this was due to permissions on the cloud). Can you please share why our backup which is now running could be slow? is there a parameter(s) in the scyllamanager yaml file or metrics we could monitor to help increase the backup speed? See details of the running status of our cluster backup to Azure. It has been running for more than 3 days now!!! scylladbmanager:~# sctool task progress -c Cluster backup/e42a9107-1cc5-423b-8c9a-e5e30b0bd118 Command “progress” is deprecated, use sctool backup|repair progress instead. Run: 373c1d8b-6c82-11ef-b5c7-005056a70a02 Status: RUNNING (uploading data) Start time: 06 Sep 24 19:00:00 UTC Duration: 86h57m26s Progress: 32% Snapshot Tag: sm_20240906190002UTC Datacenters: ±------------±---------±-------±--------±-------------±-------+ | Host | Progress | Size | Success | Deduplicated | Failed | ±------------±---------±-------±--------±-------------±-------+ | 1XX.38.1.3 | 18% | 6.059T | 1.091T | 0 | 0 | | 1XX.38.1.4 | 60% | 6.470T | 3.885T | 0 | 0 | | 1XX.38.1.8 | 18% | 6.327T | 1.154T | 0 | 0 | ±------------±---------±-------±--------±-------------±-------+ --- ### Page: https://forum.scylladb.com/t/implementation-and-usage-of-paxos-vs-raft-in-scylla/2256 Title: Implementation and usage of Paxos vs Raft in Scylla - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Recently i completed Data modeling course from Scylla University and learned that Initial decision was to use RAFT in scylla and later Paxos protocol was used. I saw a latest update in scylla open source 6.0 that scylla… Language: en Canonical URL: https://forum.scylladb.com/t/implementation-and-usage-of-paxos-vs-raft-in-scylla/2256 ## Headings Structure: H1: Implementation and usage of Paxos vs Raft in Scylla H3: Related topics ## Main Content: H1: Implementation and usage of Paxos vs Raft in Scylla H3: Related topics Recently i completed Data modeling course from Scylla University and learned that Initial decision was to use RAFT in scylla and later Paxos protocol was used. I saw a latest update in scylla open source 6.0 that scylla uses RAFt in its new Tablet based architecture. while i am not a Database expert, but I am trying to seek clarity and understand if these protocols have different use case in Scylla or Paxos has been replaced with RAFT. Any insights would be helpful to get clarity. Paxos is used to implement LWT queries in ScyllaDB. RAFT is used to implement strongly consistent metadata storage: so far this includes schema, topology, tablets metadata and some other internal system tables. Going forward we will use RAFT for more and more things, one day perhaps we will have a RAFT-based LWT implementation too. Thank you for the clarification. --- ### Page: https://forum.scylladb.com/t/openssh-vulnerability-cve-2024-6387-mitigation/2257 Title: OpenSSH Vulnerability CVE-2024-6387 Mitigation - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Issue Summary A vulnerability has been identified in OpenSSH that incorrectly handles signal management. This flaw, referenced as CVE-2024-6387, could allow a remote attacker to bypass authentication and gain unauthorize… Language: en Canonical URL: https://forum.scylladb.com/t/openssh-vulnerability-cve-2024-6387-mitigation/2257 ## Headings Structure: H1: OpenSSH Vulnerability CVE-2024-6387 Mitigation H3: Related topics ## Main Content: H1: OpenSSH Vulnerability CVE-2024-6387 Mitigation H4: Issue Summary H4: Mitigation Steps H3: Related topics A vulnerability has been identified in OpenSSH that incorrectly handles signal management. This flaw, referenced as CVE-2024-6387, could allow a remote attacker to bypass authentication and gain unauthorized system access. To mitigate this issue immediately, set LoginGraceTime to 0 in /etc/ssh/sshd_config. Actions Taken by ScyllaDB ScyllaDB Cloud Users: ScyllaDB Enterprise and Open Source Users: Please address this issue promptly to ensure your systems remain secure. If you have any questions or need further assistance, please contact our support team. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-54-2024-07-05/2260 Title: Last week in scylla-cluster-tests.git master (issue #54; 2024-07-05) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the e3bdc706…927c2429 range are covered. There were 38 non-merge commits from 8 authors in tha… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-54-2024-07-05/2260 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #54; 2024-07-05) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #54; 2024-07-05) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the e3bdc706…927c2429 range are covered. There were 38 non-merge commits from 8 authors in that period. Some notable commits: New weekly trigger with tablets disabled tests. Performance regression tests now disable tablets by default; we’ll select specific cases to use tablets in the upcoming weeks. Until we have an official release of the cassandra-stress Docker image separate from Scylla, we’ll use Scylla 6.0, which contains drivers with tablet support. cql-stress also received a version update and we added auth support for it in SCT. The Kafka connector also gained support for auth, as our integration tests use auth by default. Not only disrupt_nodetool_cleanup but all nemeses will now run nodetool cleanup in parallel on all nodes. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/failed-calling-webhook-webhook-scylla-scylladb-com/2262 Title: Failed calling webhook "webhook.scylla.scylladb.com" - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello, I’m fairly new to Scylla. I’m trying to install it on an EKS cluster so my devs can load data into it. I’ve installed the cert-manager and scylla operator with helm and they’re running without any issues. When i … Language: en Canonical URL: https://forum.scylladb.com/t/failed-calling-webhook-webhook-scylla-scylladb-com/2262 ## Headings Structure: H1: Failed calling webhook "webhook.scylla.scylladb.com" H3: Related topics ## Main Content: H1: Failed calling webhook "webhook.scylla.scylladb.com" H3: Related topics I’m fairly new to Scylla. I’m trying to install it on an EKS cluster so my devs can load data into it. I’ve installed the cert-manager and scylla operator with helm and they’re running without any issues. When i try to install Scylla, i’m getting following error. My EKS cluster is private but it has a NAT gateway and has access to the internet. My EKS node group allows Ingress from EKS cluster on tcp ports 443,9443,6443,4443,10250,8443. Can someone please advise the issue here? Probably you’re missing a rule allowing traffic from Kubernetes master nodes to worker nodes on 443 port. --- ### Page: https://forum.scylladb.com/t/issue-in-creating-mat-view-with-aggregations-in-view-itself/2264 Title: Issue in creating Mat view with aggregations in view itself - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I was trying to create mat views on over table with some aggregations in views its self. For referance you cna refer following dummy query: create materialized view test.dummy_vw as select user_id, sum(add_cash) from … Language: en Canonical URL: https://forum.scylladb.com/t/issue-in-creating-mat-view-with-aggregations-in-view-itself/2264 ## Headings Structure: H1: Issue in creating Mat view with aggregations in view itself H3: Related topics ## Main Content: H1: Issue in creating Mat view with aggregations in view itself H3: Related topics I was trying to create mat views on over table with some aggregations in views its self. For referance you cna refer following dummy query: Following statement didn’t gets executed. Does Mat views in scylla Didn’t support aggregation in view itself. Your statement has a syntax error: it’s PRIMARY KEY, not primary_key in the last line. Other than that you’re right that MVs don’t support aggregates, per doc: Materialized Views | ScyllaDB Docs If you rewrite your query, you’ll get an explicit error message: thanks for this info! --- ### Page: https://forum.scylladb.com/t/scylladb-hangs-during-startup/2267 Title: Scylladb hangs during startup - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: hello Everyone, I am new in Community and in Scylladb, i was trying to deploy scylladb cluster with 9 node in 3 datacenter with 3 nodes in each datacenter. There are site to site tunnels between datacenters which means … Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-hangs-during-startup/2267 ## Headings Structure: H1: Scylladb hangs during startup H3: Related topics ## Main Content: H1: Scylladb hangs during startup H3: Related topics hello Everyone, I am new in Community and in Scylladb, i was trying to deploy scylladb cluster with 9 node in 3 datacenter with 3 nodes in each datacenter. There are site to site tunnels between datacenters which means nodes communicates through private ips insde tunnels. so my main question is, i am trying to deploy scylla cluster using ansible while roles and the process is hanging on the stage of starting CQL, this is the message: TASK [ansible-scylla-node : Wait for CQL port on 172.16.16.101] on my control host and in the logs of that host i get this: **ansible-wait_for Invoked with port=9042 host=172.16.16.101 timeout=25200 connect_timeout=5 delay=0 active_connection_states=['ESTABLISHED', 'FIN_WAIT1', 'FIN_WAIT2', 'SYN_RECV', 'SYN_SENT', 'TIME_WAIT'] state=started sleep=1 path=None search_regex=None exclude_hosts=None msg=None** and it has been like this for hours. any one has idea why ? thank you in advance for your help. i am trying to deploy scylla cluster using ansible while roles and the process is hanging on the stage of starting CQL, this is the message: TASK [ansible-scylla-node : Wait for CQL port on 172.16.16.101] on my control host and in the logs of that host i get this: **ansible-wait_for Invoked with port=9042 host=172.16.16.101 timeout=25200 connect_time It’s really hard to understand, please format the message so it’s more readable. i few words, i am deploying scylladb cluster with ansible, and the process got stuck on cassandra trying to start on a new node for several hours: TASK [ansible-scylla-node : Wait for CQL port on 172.16.16.101]. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-237-2024-07-07/2268 Title: Last week in scylladb.git master (issue #237; 2024-07-07) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the d034cde01f…407274e828 range are covered. There were 67 non-merge commits from 15 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-237-2024-07-07/2268 ## Headings Structure: H1: Last week in scylladb.git master (issue #237; 2024-07-07) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #237; 2024-07-07) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the d034cde01f…407274e828 range are covered. There were 67 non-merge commits from 15 authors in that period. Some notable commits: A problem with a node forgetting its own IP address was fixed. If the CPU is busy processing a query, ScyllaDB will let that query complete before starting another one, since several queries using the CPU concurrently on a shard will make all of them slower. There is now an option to allow CPU concurrency on queries for workloads where this helps. A write to a n new replica of base table will no longer update any materialized views while it is still joining, since that work will be cancelled later (in some cases) or be unnecessary (in others). Materialized views calculate a backlog to decide on whether throttling of base table writes is requires. This calculation is now more accurate. Support for Debian 10 was removed, since that distribution has reached end-of-life. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/counter-table-vs-select-count-for-partition-row-count/2271 Title: Counter Table vs SELECT COUNT for partition row count - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I need to get the number of rows in a partition. Right now, I’m planning to just store a counter table and increment counters when new rows are made in the corresponding partition. But it made me wonder if SELECT COUNT (… Language: en Canonical URL: https://forum.scylladb.com/t/counter-table-vs-select-count-for-partition-row-count/2271 ## Headings Structure: H1: Counter Table vs SELECT COUNT for partition row count H3: Related topics ## Main Content: H1: Counter Table vs SELECT COUNT for partition row count H3: Related topics I need to get the number of rows in a partition. Right now, I’m planning to just store a counter table and increment counters when new rows are made in the corresponding partition. But it made me wonder if SELECT COUNT (1) FROM table WHERE partition_key = ? would be optimized in this case (given that partitions are distributed), allowing me to avoid storing redundant data (and avoid some writes, but that doesn’t matter). Does Scylla store the number of rows per partition? And would SELECT COUNT be efficient in this case? Taking a gander at the source code for [cql3::functions::aggregate_fcts::make_count_function] (and countRows), it doesn’t look like it handles partitions in any special way (unless it is handled elsewhere). I guess I’ll go with Counters for now. Indeed, ScyllaDB does not maintain a row count for partitions internally, instead COUNT() will invoke the regular COUNT code-path which will request all rows from the replicas and count them on the coordinator. Normally, this is a very expensive operation, because a full scan on a table is expensive. However, when invoked on a single partition, it should not be particularly expensive. I would recommend you to test the performance and impact of this query and decide based on the results whether you really need a row-count column. --- ### Page: https://forum.scylladb.com/t/running-scylladb-on-aws-how-to-use-ebs-volumes-ephemeral-storage-reboot-or-restart/2276 Title: Running ScyllaDB on AWS, how to use EBS volumes, ephemeral storage reboot or restart - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/running-scylladb-on-aws-how-to-use-ebs-volumes-ephemeral-storage-reboot-or-restart/2276 ## Headings Structure: H1: Running ScyllaDB on AWS, how to use EBS volumes, ephemeral storage reboot or restart H3: Related topics ## Main Content: H1: Running ScyllaDB on AWS, how to use EBS volumes, ephemeral storage reboot or restart H3: Related topics Originally from the User Slack @채수상_(soo): hello. I’m trying to use scylladb for the first time… I ran scylladb using the combination of ami image and i3.xlarge as instructed in the scylla document. Afterwards, I stopped the instance and started it again, but the instance storage was not mounted and scylla-server was not running. @fodo: seems like more of aws questions rather than scylla :simple_smile: @채수상_(soo): Yes, that’s right. It seems like a question related to Linux and AWS. However, I used ami produced by scylla, and it seems that ephemral storage is automatically mounted when booting and DB is installed on that storage. I wonder if I can simply mount the storage in fstab, reboot, and start scylla-server, or if I need to make additional settings. @fodo: IIRC scylladb ami will mount it for you automatically https://github.com/scylladb/scylla-machine-image/blob/next/common/scylla_create_devices If your concern is about data loss, it’s data will be lost after restart since it’s “i type” which means its data layer uses local ssd. it’s not ebs so when you need to restart scylladb node, make sure that you have replication factor > 1 for your data to replicate over other nodes @avi: AWS i3 have ephemeral storage, do not stop and start them. You can reboot them. btw, use i4i, they’re much better than i3 @채수상_(soo): Thanks, I’ll try using i4i. Want to understand something here. It seems the recommendation is to not stop/start the instance. What I am not understanding is how to ensure data retention in if the instance goes down outside our control. Is that just not possible with a replication factor of 1? --- ### Page: https://forum.scylladb.com/t/number-of-connections-and-cpus-vcpus-whnen-workign-with-k8s/2278 Title: Number of connections and CPUs/VCPUs whnen workign with K8s - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/number-of-connections-and-cpus-vcpus-whnen-workign-with-k8s/2278 ## Headings Structure: H1: Number of connections and CPUs/VCPUs whnen workign with K8s H3: Related topics ## Main Content: H1: Number of connections and CPUs/VCPUs whnen workign with K8s H3: Related topics Originally from the User Slack @J: How should I think about connections and cpus/ vcpus when working in K8s. 3 per Scylla cpu. In a 3 node cluster each with 16 vcpus (K8s). Should that be 3x16 or 3x8 for the logical core? @avi: I don’t understand the question. 3x16 what? @J: Number of connections. Your best practices say number of connections is 1-3 per Scylla cpu per node. But I wanted to check this is per vcpu or per logical core. >> As a rule of thumb, for Scylla’s best performance, each client needs at least 1-3 connections per Scylla core. For example, in a cluster with three nodes, each node with 16 cores, each client application should open 32 (2x16) connections to each Scylla node. @avi: In general the drivers we provide take care of this and generate one connection per shard vcpu == logical core (physical core == 2vcpu == 2lcore on x86) --- ### Page: https://forum.scylladb.com/t/use-case-for-replacing-cassandra-with-scylladb-read-and-write-performance/2279 Title: Use case for replacing Cassandra with ScyllaDB, read and write performance - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/use-case-for-replacing-cassandra-with-scylladb-read-and-write-performance/2279 ## Headings Structure: H1: Use case for replacing Cassandra with ScyllaDB, read and write performance H3: Related topics ## Main Content: H1: Use case for replacing Cassandra with ScyllaDB, read and write performance H3: Related topics Originally from the User Slack @Narayan_Ramamurthi: HI! I had a nice chat with @Lawrence_Wan today. Basically here is our use case. We are a telecom billing product development company with over 200+ installations across the globe. Our billing product is primarily written in C/C++ on unix like platforms including RHEL linux, sun (, which is now oracle) solaris, aix, and hp-ux. We also have parallel teams who write the surrounding software in various technologies including Java, Python and Kotilin. We have traditionally been an Oracle-Database company But over time in the recent years we started using Apache-/Datastacks-Cassandra both for ourselves and also our clients. When Cassandra was being proposed, I also proposed using ScyllaDB as an alternative but our management did not consider my request at that time. After happily using Cassandra for years now, we are facing troubles with Datastacks support and Apache open-source support to the cassandra-cpp-driver. So, I brought up this topic of Scylla with the management. They seem to be more willing to listen this time. I should also give some more background - our parallel Java/Go team has done a recent POC with the ScyllaDB and their results indicate that Cassandra performs 3x better than Scylla. Our management is quoting this as an example to me now. I am looking to prove them wrong, I am here seeking support to learn what can be done to prove them wrong. Basically what settings can be done on the cluster and on the client side to make sure that Scylla beats Cassandra in fair competition. I seek this vibrant community’s support to kindly suggest the needful. @avi: Please specify the hardware on which you deployed. To understand your performance problem, it’s best to set up scylla-monitoring and work with us to identify the problem. @Narayan_Ramamurthi: intel xeon server family Both Cassandra and ScyllaDB clusters were deployed using similar hardware @avi: Please provide more information, number of cores, amount of memory, disk types And set up scylla-monitoring @Narayan_Ramamurthi: We have been using apache cassandra for quite a while now, and our experience says that cassandra is possibly good at doing writes but it is not really good at doing reads. On the other hand, Oracle (the legacy SQL datastore) does it way better for either, but on a single node. Can we expect something better with ScyllaDb? If so, can you provide me information where we can do both reads and writes much better, possibly unlike cassandra? Regarding number of cores, hardware, etc: for the sake of comparison we shall setup identical hardware for cassandra and scylla: Haswell-based processors, something like Xeon E5-2683 v4 or the likes. Note that this is being pursued in the POC (proof of concept) mode at the moment, so foresight is important. @avi: Thank you in advance for your ever prompt response. I will be eagerly watching this space for help. @avi: ScyllaDB should be faster than Cassandra for both reads and writes --- ### Page: https://forum.scylladb.com/t/scylla-manager-issue-connecting-to-scylla-agents-from-sctool/2280 Title: Scylla Manager, issue connecting to Scylla agents from sctool - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/scylla-manager-issue-connecting-to-scylla-agents-from-sctool/2280 ## Headings Structure: H1: Scylla Manager, issue connecting to Scylla agents from sctool H3: Related topics ## Main Content: H1: Scylla Manager, issue connecting to Scylla agents from sctool H3: Related topics Originally from the User Slack @John_McDonald: I’m having an issue connecting to scylla agents from sctool to setup a cluster for scylla manager. The agents are running and listening on 10001, as evidence by the curl command, yet sctool claims that it can’t connect. Any assistance would be appreciated as I’m stuck. Seems like a bug. Im long past tired of running repairs manually. Any help would be appreciated @avi: It’s complaining about reachability to other nodes, not the one you curled, no? @John_McDonald: Avi, you’re absolutely correct. Thank you for your help! Avi, I verified that the agent is running on all 5 nodes, yet the curl command fails. What am I missing? @avi: I know almost nothing about manager-agent, so can’t help here. I don’t even know what’s supposed to listen on port 10001. @John_McDonald: Avi, we found the problem. For the 4 nodes to which it couldn’t connect, for some reason we had the api_address set to the internal IP of the node, but it’s supposed to be 127.0.0.1. :flag_il: --- ### Page: https://forum.scylladb.com/t/different-running-compactions-between-multi-dc/2282 Title: Different running compactions between multi dc - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, we’re running two clusters scyllaDB with 3 nodes. DC1 (3 nodes) DC2 (3 nodes) CREATE KEYSPACE test_keyspace WITH replication = {‘class’: ‘NetworkTopologyStrategy’, ‘dc1’: ‘3’, ‘dc2’: ‘3’}; Table is using STCS As… Language: en Canonical URL: https://forum.scylladb.com/t/different-running-compactions-between-multi-dc/2282 ## Headings Structure: H1: Different running compactions between multi dc H3: Related topics ## Main Content: H1: Different running compactions between multi dc H3: Related topics Hi, we’re running two clusters scyllaDB with 3 nodes. CREATE KEYSPACE test_keyspace WITH replication = {‘class’: ‘NetworkTopologyStrategy’, ‘dc1’: ‘3’, ‘dc2’: ‘3’}; As daily batch job. we bulk write to DC1 using spark, and data are replicated to DC2. While batch write is on going, running compactions occurs different in DC1 and DC2 While compaction occurs evenly in DC1, in DC2, compaction happens in large batches all at once. I don’t know why compaction pattern is different. And due to heavy compaction on DC2, durning batch operations, read p95, p99 latency spikes occur. DC1 read latency has no problem. I think DC1’s coordinator will relay to remote coordinator in DC2. So write pattern will be same. Therefore I think that compaction pattern should be same in DC1 and DC2. But It shows different. Is there anybody who knows why and how to reduce running compaction on DC2? So if upgrade to Enterprise is an option, consider it and switch to ICS, it’s superior to STCS. DC2 will recieve just one copy from DC1 coordinator and has to replicate it (to save bandwidth between DCs), this I think can create delays in how and when compactions happen in DC2. Also IO scheduler will impact this, so being on Scylla 5.4 or newer will also help. Older IO schedulers might not be that effective. (or even better Enterprise 2024 will be just faster out of box, it has lots of improvements and optimizations and should be 30% faster than OSS) Also if you are on latest Scylla and if disks/instance types are the same in both DCs you can try to decrease compaction_static_shares in scylla.yaml on all nodes in DC2 (which will make compactions run longer and be less aggressive to take resources away from queries) e.g. to 500? 300? or 200? this number is hard to tell without understanding the workload. If it’s write heavy and you want data compacted asap, then it should be higher. --- ### Page: https://forum.scylladb.com/t/error-during-truncate/2284 Title: Error during truncate - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: We are getting below error while truncating table. And truncate failed. Jul 10 12:16:20 Ubuntu-2204-jammy-amd64-base scylla: [shard 29:stat] database - Truncating databasename.tablename with auto-snapshot Jul 10 12:27… Language: en Canonical URL: https://forum.scylladb.com/t/error-during-truncate/2284 ## Headings Structure: H1: Error during truncate H3: Related topics ## Main Content: H1: Error during truncate H3: Related topics We are getting below error while truncating table. And truncate failed. Jul 10 12:16:20 Ubuntu-2204-jammy-amd64-base scylla: [shard 29:stat] database - Truncating databasename.tablename with auto-snapshot Jul 10 12:27:36 Ubuntu-2204-jammy-amd64-base scylla: [shard 0:stat] mutation_partition - Memory usage of unpaged query exceeds soft limit of 1048576 (configured via max_memory_for_unlimited_query_soft_limit) Please suggest solution around it. @sachinkumar So tips when truncating - make sure you have BIG --request-timeout if you use cqlsh to run it (or in client set request timeout to be > 60s ) If it fails(times out), run it AGAIN until it succeeds. Once it’s done, on all nodes clear the snapshot it created (or all of them, or if you don’t need the safety check disable auto snapshot feature) OH and " mutation_partition - Memory usage of unpaged query exceeds soft limit of 1048576 (configured via max_memory_for_unlimited_query_soft_limit) " is unrelated and basically about your queries and has nothing to do with truncate, as you see it’s a warning that one of your queries was unpaged and its result was bigger than 1MB, which certainly impacts performance - you should page in your queries, unpaged queries are antipattern --- ### Page: https://forum.scylladb.com/t/range-queries-paging-and-tombstones-in-scylladb-and-cassandra/2287 Title: Range queries, paging, and tombstones in ScyllaDB and Cassandra - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/range-queries-paging-and-tombstones-in-scylladb-and-cassandra/2287 ## Headings Structure: H1: Range queries, paging, and tombstones in ScyllaDB and Cassandra H3: Related topics ## Main Content: H1: Range queries, paging, and tombstones in ScyllaDB and Cassandra H3: Related topics Originally from the User Slack @sandr8: Are range queries over clustering columns potentially scanning a lot of rows because of having to leap over tombstones or this issue only affects Cassandra and ScyllaDB has a way to avoid it? @avi: It affects ScyllaDB too but we have some mitigation. Most important is the ability to return empty pages, so if you have a really long run of tombstones, the query doesn’t time out. @sandr8: Thank you! I would love to understand better what you mean by the ability to return empty pages. @avi: Some more info here, but I don’t think we have a user level explanation: https://github.com/scylladb/scylladb/commit/e9cbc9ee85c25deb5ba9ee67ffa6e3ca9f904660 GitHub: Merge ‘Add support for empty replica pages’ from Botond Dénes · scylladb/scylladb@e9cbc9e --- ### Page: https://forum.scylladb.com/t/scylladb-labs-july-17th-2024/2288 Title: ScyllaDB Labs - July 17th 2024 - Announcements - ScyllaDB Community NoSQL Forum Meta Description: Next week, we’re hosting ScyllaDB Labs, a hands-on, online, training event. The focus is on learning by doing, so you’ll have a chance to learn some theory and put it into practice with hands-on labs. You can save your … Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-labs-july-17th-2024/2288 ## Headings Structure: H1: ScyllaDB Labs - July 17th 2024 H3: Related topics ## Main Content: H1: ScyllaDB Labs - July 17th 2024 H3: Related topics Next week, we’re hosting ScyllaDB Labs, a hands-on, online, training event. The focus is on learning by doing, so you’ll have a chance to learn some theory and put it into practice with hands-on labs. You can save your free spot here. Hope to see you there! --- ### Page: https://forum.scylladb.com/t/release-introducing-layout-changes-in-scylladb-cloud-app-11-jul-2024/2289 Title: [RELEASE] Introducing Layout Changes in ScyllaDB Cloud App - 11 Jul 2024 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: We are happy to announce the latest version of ScyllaDB Cloud, which includes significant layout redesign changes aimed at enhancing your user experience. Below are the key updates: Removal of the Left Sidebar We have… Language: en Canonical URL: https://forum.scylladb.com/t/release-introducing-layout-changes-in-scylladb-cloud-app-11-jul-2024/2289 ## Headings Structure: H1: [RELEASE] Introducing Layout Changes in ScyllaDB Cloud App - 11 Jul 2024 H3: Related topics ## Main Content: H1: [RELEASE] Introducing Layout Changes in ScyllaDB Cloud App - 11 Jul 2024 H3: Related topics We are happy to announce the latest version of ScyllaDB Cloud, which includes significant layout redesign changes aimed at enhancing your user experience. Below are the key updates: Removal of the Left Sidebar Relocation of External Resource Links Introduction of Breadcrumbs Navigation We hope these changes will improve your overall experience with our app. As always, we welcome your feedback and suggestions. --- ### Page: https://forum.scylladb.com/t/rebalance-the-cluster-token-range/2291 Title: Rebalance the cluster token range - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I’ve a test cluster and it’s load is imbalance. We have keyspace with RF=1. -- Address Load Tokens UN 127.0.0.2 146.7 GB 256 UN 127.0.0.3 2.85 GB 256 UN 127.0.0.1 36.45 MB 256 How can i rebala… Language: en Canonical URL: https://forum.scylladb.com/t/rebalance-the-cluster-token-range/2291 ## Headings Structure: H1: Rebalance the cluster token range H3: Related topics ## Main Content: H1: Rebalance the cluster token range H3: Related topics I’ve a test cluster and it’s load is imbalance. We have keyspace with RF=1. How can i rebalance this cluster token range? The Load column indicates the size on disk that ScyllaDB uses. It depends on your data model. Please share more details about it (tables, primary key, etc.). Basic data modeling is covered in this ScyllaDB University lesson. Also make sure you understand tokens and data distribution covered in the ScyllaDB Essentials course. --- ### Page: https://forum.scylladb.com/t/release-scylladb-5-4-9/2292 Title: [RELEASE] ScyllaDB 5.4.9 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.4.9, a bugfix release of the ScyllaDB 5.4 stable branch. ScyllaDB Open Source 5.4.9, like all past and future 5.x.y releases, is backward compatible and supports rolling… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-4-9/2292 ## Headings Structure: H1: [RELEASE] ScyllaDB 5.4.9 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 5.4.9 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 5.4.9, a bugfix release of the ScyllaDB 5.4 stable branch. ScyllaDB Open Source 5.4.9, like all past and future 5.x.y releases, is backward compatible and supports rolling upgrades. Note that the latest ScyllaDB Open Source stable release is 6.0 and you are encouraged to upgrade to it. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/facing-errors-in-repair-jobs/2293 Title: Facing errors in repair jobs - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Getting below error in logs while repair job. This is causing repair job to run for long hours. Error : Jul 9 05:55:58 nodename scylla: [shard 25] repair - repair[f06d168e-664f-4b3f-9717-5abc0e4cbf01]: shard=25, keys… Language: en Canonical URL: https://forum.scylladb.com/t/facing-errors-in-repair-jobs/2293 ## Headings Structure: H1: Facing errors in repair jobs H3: Related topics ## Main Content: H1: Facing errors in repair jobs H3: Related topics Getting below error in logs while repair job. This is causing repair job to run for long hours. Error : Jul 9 05:55:58 nodename scylla: [shard 25] repair - repair[f06d168e-664f-4b3f-9717-5abc0e4cbf01]: shard=25, keyspace=time_series_interaction, cf=click_logged_in, range=(-7693897 20434297988, -551877587955844925], got error in row level repair: std::runtime_error (put_row_diff: Repair follower=10.x.x.x failed in put_row_diff hanlder, status=0) Jul 9 05:55:58 nodename scylla: [shard 28] repair - repair[f06d168e-664f-4b3f-9717-5abc0e4cbf01]: shard=28, keyspace=time_series_interaction, cf=click_logged_in, range=(-7693897 20434297988, -551877587955844925], got error in row level repair: std::runtime_error (put_row_diff: Repair follower=10.x.x.x failed in put_row_diff hanlder, status=0) Jul 9 05:55:58 nodename scylla: [shard 11] repair - repair[f06d168e-664f-4b3f-9717-5abc0e4cbf01]: shard=11, keyspace=time_series_interaction, cf=click_logged_in, range=(-7693897 Jul 9 06:04:41 nodename scylla: [shard 0] repair - repair[f06d168e-664f-4b3f-9717-5abc0e4cbf01]: repair_tracker run failed: std::runtime_error ({shard 0: seastar::rpc::closed_e rror (connection is closed), shard 1: seastar::rpc::closed_error (connection is closed), shard 2: seastar::rpc::closed_error (connection is closed), shard 3: seastar::rpc::closed_error (connection i s closed), shard 4: seastar::rpc::closed_error (connection is closed), shard 5: seastar::rpc::remote_verb_error (seastar::nested_exception), shard 6: seastar::rpc::closed_error (connection is closed ), shard 7: seastar::rpc::closed_error (connection is closed), shard 8: seastar::rpc::closed_error (connection is closed), shard 9: seastar::rpc::closed_error (connection is closed), shard 10: seast ar::rpc::closed_error (connection is closed), shard 11: seastar::rpc::remote_verb_error (seastar::nested_exception), shard 12: seastar::rpc::closed_error (connection is closed), shard 13: seastar::r pc::closed_error (connection is closed), shard 14: seastar::rpc::closed_error (connection is closed), shard 15: seastar::rpc::closed_error (connection is closed), shard 16: seastar::rpc::closed_erro Please share as many details as possible: ScyllaDB Version, cluster size, OS, hardware details, cluster size. Did you check the Monitoring dashboards? ScyllaDB version : 5.2.14 Cluster size : 1.5 TB with 2 nodes and RF as 2. OS : ubuntu 20.04 focal Hardware : n2-highmem-32 GCP with 16 NVME disks Monitoring dashboards seems showing spikes in latencies in read and write. It looks like the error happened on one of the peer nodes. Please check the logs of the other nodes participating in the repair, to see the real error. error in row level repair Errors : Jul 12 01:53:54 nodename scylla: [shard 2] repair - repair[b3fbc952-171d-4420-9a82-fb7981a32d9f]: Started to repair 1 out of 1 tables in keyspace=, table=, table_id=c195a300-baaa-11ee-b6 9d-54c54cd21d49, repair_reason=repair Jul 12 01:53:54 nodename scylla: [shard 17] repair - repair[b3fbc952-171d-4420-9a82-fb7981a32d9f]: Started to repair 1 out of 1 tables in keyspace=, table=, table_id=c195a300-baaa-11ee-b6 9d-54c54cd21d49, repair_reason=repair Jul 12 01:53:54 nodename scylla: [shard 28] repair - repair[b3fbc952-171d-4420-9a82-fb7981a32d9f]: Started to repair 1 out of 1 tables in keyspace=, table=, table_id=c195a300-baaa-11ee-b6 9d-54c54cd21d49, repair_reason=repair Jul 12 01:53:54 nodename scylla: [shard 5] repair - repair[b3fbc952-171d-4420-9a82-fb7981a32d9f]: Started to repair 1 out of 1 tables in keyspace=, table=, table_id=c195a300-baaa-11ee-b6 9d-54c54cd21d49, repair_reason=repair Jul 12 01:53:54 nodename scylla: [shard 20] repair - repair[b3fbc952-171d-4420-9a82-fb7981a32d9f]: Started to repair 1 out of 1 tables in keyspace=, table=, table_id=c195a300-baaa-11ee-b6 9d-54c54cd21d49, repair_reason=repair Jul 12 01:53:54 nodename scylla: [shard 9] repair - repair[b3fbc952-171d-4420-9a82-fb7981a32d9f]: Started to repair 1 out of 1 tables in keyspace=, table=, table_id=c195a300-baaa-11ee-b6 9d-54c54cd21d49, repair_reason=repair Jul 12 01:53:58 nodename scylla: [shard 24] storage_proxy - Exception when communicating with 10.x.x.x, to read from .click_anon: std::bad_alloc Jul 12 01:53:59 nodename scylla: [shard 17] compaction - [Compact system.compaction_history 9b0d8f00-3ff1-11ef-ae9c-8bda48236624] Compacting [/var/lib/scylla/data/system/compaction_history-b4dbb7b4dc493fb5b3bfce6e434832ca/me- 183887-big-Data.db:level=0:origin=memtable,/var/lib/scylla/data/system/compaction_history-b4dbb7b4dc493fb5b3bfce6e434832ca/me-183857-big-Data.db:level=0:origin=compaction] Jul 12 01:53:59 nodename scylla: [shard 17] compaction - [Compact system.compaction_history 9b0d8f00-3ff1-11ef-ae9c-8bda48236624] Compacted 2 sstables to [/var/lib/scylla/data/system/compaction_history-b4dbb7b4dc493fb5b3bfce6 e434832ca/me-183917-big-Data.db:level=0]. 299kB to 192kB (~64% of original) in 107ms = 2MB/s. ~1024 total partitions merged to 844. Jul 12 01:53:59 nodename scylla: [shard 7] repair - repair[b3fbc952-171d-4420-9a82-fb7981a32d9f]: shard=7, keyspace=, cf=, range=(-2136115943339218610, -2031250356567146432], got error i n row level repair: seastar::rpc::remote_verb_error (std::bad_alloc) Jul 12 01:54:00 nodename scylla: [shard 20] repair - repair[b3fbc952-171d-4420-9a82-fb7981a32d9f]: shard=20, keyspace=, cf=, range=(-2136115943339218610, -2031250356567146432], got error in row level repair: seastar::rpc::remote_verb_error (std::bad_alloc) Jul 12 01:54:00 nodename scylla: [shard 3] storage_proxy - Exception when communicating with 10.x.x.x, to read from .: std::bad_alloc Jul 12 01:54:00 nodename scylla: [shard 3] storage_proxy - Exception when communicating with 10.x.x.x, to read from .: std::bad_alloc Jul 12 01:54:00 nodename scylla: [shard 18] storage_proxy - Exception when communicating with 10.x.x.x, to read from .: std::bad_alloc Jul 12 01:54:00 nodename scylla: [shard 27] storage_proxy - Exception when communicating with 10.x.x.x, to read from .: std::bad_alloc Jul 12 01:54:00 nodename scylla: [shard 18] storage_proxy - Exception when communicating with 10.x.x.x, to read from .: std::bad_alloc Jul 12 01:54:00 nodename scylla: [shard 20] repair - repair[b3fbc952-171d-4420-9a82-fb7981a32d9f]: shard=20, keyspace=, cf=, range=(-2141427493649988622, -2136115943339218610], got error in row level repair: seastar::rpc::remote_verb_error (std::bad_alloc) Jul 12 01:54:00 nodename scylla: [shard 18] storage_proxy - Exception when communicating with 10.x.x.x, to read from .: std::bad_alloc Jul 12 01:54:00 nodename scylla: [shard 3] storage_proxy - Exception when communicating with 10.x.x.x, to read from .: std::bad_alloc Jul 12 01:54:00 nodename scylla: [shard 18] storage_proxy - Exception when communicating with 10.x.x.x, to read from .: std::bad_alloc Jul 12 01:54:00 nodename scylla: [shard 26] storage_proxy - Exception when communicating with 10.x.x.x, to read from .: std::bad_alloc Jul 12 01:54:01 nodename scylla: [shard 18] storage_proxy - Exception when communicating with 10.x.x.x, to read from .: std::bad_alloc Jul 12 01:54:01 nodename scylla: [shard 27] storage_proxy - Exception when communicating with 10.x.x.x, to read from .: std::bad_alloc Jul 12 01:54:01 nodename scylla: [shard 26] storage_proxy - Exception when communicating with 10.x.x.x, to read from .: std::bad_alloc Jul 12 01:54:01 nodename scylla: [shard 26] storage_proxy - Exception when communicating with 10.x.x.x, to read from .: std::bad_alloc Jul 12 01:54:01 nodename scylla: [shard 3] repair - repair[b3fbc952-171d-4420-9a82-fb7981a32d9f]: shard=3, keyspace=, cf=, range=(-2141427493649988622, -2136115943339218610], got error in row level repair: seastar::rpc::remote_verb_error (std::bad_alloc)` Jul 12 01:54:01 nodename scylla: [shard 3] repair - repair[b3fbc952-171d-4420-9a82-fb7981a32d9f]: shard=3, keyspace=, cf=, range=(-2136115943339218610, -2031250356567146432], got error in row level repair: seastar::rpc::remote_verb_error (std::bad_alloc) Jul 12 01:54:02 nodename scylla: [shard 10] repair - repair[b3fbc952-171d-4420-9a82-fb7981a32d9f]: shard=10, keyspace=, cf=, range=(-2141427493649988622, -2136115943339218610], got error in row level repair: seastar::rpc::remote_verb_error (std::bad_alloc) Jul 12 01:54:02 nodename scylla: [shard 4] storage_proxy - Exception when communicating with 10.x.x.x, to read from .: std::bad_alloc Jul 12 01:54:02 nodename scylla: [shard 26] storage_proxy - Exception when communicating with 10.x.x.x, to read from .: std::bad_alloc Jul 12 01:54:02 nodename scylla: [shard 26] storage_proxy - Exception when communicating with 10.x.x.x, to read from .: std::bad_alloc Jul 12 01:54:02 nodename scylla: [shard 27] storage_proxy - Exception when communicating with 10.x.x.x, to read from .: std::bad_alloc Jul 12 01:54:02 nodename scylla: [shard 23] repair - repair[b3fbc952-171d-4420-9a82-fb7981a32d9f]: shard=23, keyspace=, cf=, range=(-2141427493649988622, -2136115943339218610], got error in row level repair: seastar::rpc::remote_verb_error (std::bad_alloc) Jul 12 01:54:03 nodename scylla: [shard 0] repair - repair[b3fbc952-171d-4420-9a82-fb7981a32d9f]: shard=0, keyspace=, cf=, range=(-2136115943339218610, -2031250356567146432], got error in row level repair: std::runtime_error (put_row_diff: Repair follower=10.x.x.x failed in put_row_diff hanlder, status=0) Jul 12 01:54:03 nodename scylla: [shard 18] storage_proxy - Exception when communicating with 10.x.x.x, to read from .: std::bad_alloc Jul 12 01:54:03 nodename scylla: [shard 27] storage_proxy - Exception when communicating with 10.x.x.x, to read from .: std::bad_alloc Jul 12 01:54:03 nodename scylla: message repeated 2 times: [ [shard 27] storage_proxy - Exception when communicating with 10.x.x.x, to read from .: std::bad_alloc] Jul 12 01:54:03 nodename scylla: [shard 24] storage_proxy - Exception when communicating with 10.x.x.x, to read from .click_anon: std::bad_alloc Jul 12 01:54:03 nodename scylla: [shard 17] compaction - [Compact .click_anon 9db30c80-3ff1-11ef-ae9c-8bda48236624] Compacting [/var/lib/scylla/data//click_anon-c576fbe0baaa11eeb69d54c54cd21d49/me-588917-big-Data.db:level=0:origin=memtable,/var/lib/scylla/data//click_anon-c576fbe0baaa11eeb69d54c54cd21d49/me-588887-big-Data.db:level=0:origin=compaction] Jul 12 01:54:03 nodename scylla: [shard 17] compaction - [Compact . 9dba1160-3ff1-11ef-ae9c-8bda48236624] Compacting [/var/lib/scylla/data//-bdb00460baaa11eeb69d54c54cd21d49/me-577157-big-Data.db:level=0:origin=memtable,/var/lib/scylla/data//-bdb00460baaa11eeb69d54c54cd21d49/me-577127-big-Data.db:level=0:origin=compaction] Also we tried changing repair job intensity from 1 to 2 and then 4, resulted in restart of scylla-server service. I see std::bad_alloc:s in the logs. This is the reason repair failed. Increasing repair intensity will make this problem worse, that is why you had a restart (crash). Check the Bloom Filter memory usage (percentage) graph, on the Detailed dashboard in monitoring. What does it show? There is a fix for this in 5.2.18. Upgrade to either the latest 5.2, or better, to 5.4 or 6.0 (5.2 is not supported anymore). So you are saying its because of a bug which has hit on 5.2.14 ? The bug was not introduced in 5.2.14 (as far as I remember), but it was fixed in 5.2.18. Thanks Botond_Denes. This was helpful. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-55-2024-07-12/2294 Title: Last week in scylla-cluster-tests.git master (issue #55; 2024-07-12) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 909a185e…75d03380 range are covered. There were 19 non-merge commits from 5 authors in tha… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-55-2024-07-12/2294 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #55; 2024-07-12) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #55; 2024-07-12) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 909a185e…75d03380 range are covered. There were 19 non-merge commits from 5 authors in that period. Some notable commits: Debian10 is deprecated as of June 30, 2024, and we have removed all jobs based on it. We’ve introduced the perf_extra_jobs_to_compare test parameter, which enables comparison of staging performance jobs with others (e.g., weekly jobs). This feature is also available in the hydra perf-regression-report command for manual report creation when needed. We now cache pull request details for the SkipPerIssues utility since tests are sometimes skipped based on PRs rather than issues. KMS is enabled by default for enterprise tests. To disable it, set the enterprise_disable_kms test parameter to true. We encountered issues with dropped SSH connections for long-running commands. The run_nodetool method now has a long_running feature to address this by running nodetool in the background on the remote host and periodically checking its status. This feature has been enabled in several nemeses. For certain cases, like parallel nemeses, we don’t want to fail a test based on an error event alone. Instead, we look for coredumps or critical-level events. A new teardown-validator was introduced to determine test failure based on predefined error events only. This has been enabled in the longevity-2TB-48h-authorization-and-tls-ssl-1dis-2nondis-nemesis test. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/spark-connector-ttl-on-a-per-element-basis-within-the-map-collection-type-which-compaction-strategy/2297 Title: Spark Connector, TTL on a per-element basis within the map collection type, which compaction strategy - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/spark-connector-ttl-on-a-per-element-basis-within-the-map-collection-type-which-compaction-strategy/2297 ## Headings Structure: H1: Spark Connector, TTL on a per-element basis within the map collection type, which compaction strategy H3: Related topics ## Main Content: H1: Spark Connector, TTL on a per-element basis within the map collection type, which compaction strategy H3: Related topics Originally from the User Slack @Ritesh: Hello ScyllaDB users! I am currently working with Spark and the Cassandra connector to store and modify map collections in a ScyllaDB table. So far, I have successfully implemented and run the following code: // Write the result DataFrame to a new table in ScyllaDB final_df.rdd.saveToCassandra(“mykeyspace”, “mytable”, SomeColumns(“id”, “map1” append, “map2” append)) However, I would like to specify a TTL (time-to-live) on a per-element basis within the map collection type. I am aware of the example code in the spark-cassandra-connector repository, which applies TTL at the row level: import com.datastax.spark.connector.writer._ … rdd.saveToCassandra(“test”, “tab”, writeConf = WriteConf(ttl = TTLOption.constant(100))) rdd.saveToCassandra(“test”, “tab”, writeConf = WriteConf(timestamp = TimestampOption.constant(ts))) Is there a way to specify TTL for each element in a map collection rather than at the row level? Any help would be greatly appreciated! @avi: You could write each element in a separate saveToCassandra call perhaps @Ritesh: @avi Thanks for the suggestion! I used an older version of Spark and was able to set a TTL on the map column in ScyllaDB. However, despite setting the TTL to 200 seconds, my data does not expire as expected. I am still able to retrieve the data in query results using cqlsh command even after the specified TTL has passed. @avi: You can try to debug it by using the ttl() function to see what the database thinks your ttl is @Ritesh: I tried using TTL() function on map type column(non-frozen) but I got below error- “TTL expects an atomic column, but it is a non-frozen collection” @avi: What about TTL(my_map[‘key’])? @Ritesh: No such luck! @avi: I see. Please file an enhancement request. @Ritesh: Okay Hi @avi, I have a question regarding our previous discussion. Since I am applying a TTL of 2 months to each element in the map, which compaction strategy would be most suitable for this scenario? Additionally, there might be monthly updates to the existing map if needed @Ritesh: Can’t we use TWCS for this? @avi: TWCS is suitable for insert-only, no updates, and with the clustering key corresponding to time @Ritesh: Okay, got it Thanks! Hi @avi! Could you please help me with the following question: Is there a way to temporarily pause TTL in ScyllaDB without re-inserting the same data with a new TTL? @avi: No (you can try moving your clock backwards, but likely many things will break) --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-238-2024-07-14/2298 Title: Last week in scylladb.git master (issue #238; 2024-07-14) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 407274e828…53a6ec05ed range are covered. There were 55 non-merge commits from 19 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-238-2024-07-14/2298 ## Headings Structure: H1: Last week in scylladb.git master (issue #238; 2024-07-14) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #238; 2024-07-14) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 407274e828…53a6ec05ed range are covered. There were 55 non-merge commits from 19 authors in that period. Some notable commits: There are now metrics keeping track of incoming hints, in addition to the existing metrics for outgoing hints. A regression the the Lightweight Transaction (LWT) contention metric has been fixed. The regression would shop contentions increasing even when none were happening. While it’s just a metrics, it’s one of the more important ones for LWT users. There is now a REST API for triggering a Raft group 0 read barrier. This is useful for making sure all nodes have caught up with the Raft leader and see the latest schema and topology. In ScyllaDB writes have a server memory footprint even after an acknowledgement is returned to the client, in order to track writes to replica past the consistency level requirement. In one case, CL=ANY and all the target replicas DOWN, writes were kept even after all replicas acknowledged, increasing server memory load, and possibly preventing topology changes from making progress. This scenario can be generated internally when writing to a materialized view. We now clean up the internal structures immediately. In rare cases, ScyllaDB might crash while inserting a mutation into memtable or cache. This is now fixed. Keyspaces with tablets enabled will now reject tables with counter columns, as counters aren’t yet supported with tablets. A bug in DESCRIBE SCHEMA when describing indexes on collection columns was fixed. Changes to service levels are now reflected immediately after the change, rather than via a polling loop with a cycle time of 10 seconds. A lock over the node tablet replica map was removed, as it was causing topology changes to be delayed. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/nodes-not-joining-a-cluster-incrementally-adding-nodes-to-a-cluster/2299 Title: Nodes not joining a cluster, incrementally adding nodes to a cluster - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/nodes-not-joining-a-cluster-incrementally-adding-nodes-to-a-cluster/2299 ## Headings Structure: H1: Nodes not joining a cluster, incrementally adding nodes to a cluster H3: Related topics ## Main Content: H1: Nodes not joining a cluster, incrementally adding nodes to a cluster H3: Related topics Originally from the User Slack @Pham_William: Hey everyone, setting up a test cluster in a virtual environment, and I have a few questions. Based on the configuration guide for a single data center: https://opensource.docs.scylladb.com/stable/operating-scylla/procedures/cluster-management/create-cluster.html I just have 2 nodes: | scylladb | RUNNING | 10.70.188.92 (eth0) | fd42:ab78:d54:2140:216:3eff:fe59:33a7 (eth0) | CONTAINER | 0 | ±----------±--------±---------------------±---------------------------------------------±----------------±----------+ | scylladb2 | RUNNING | 10.70.188.39 (eth0) | fd42:ab78:d54:2140:216:3eff:feeb:2be7 (eth0) | CONTAINER | 0 | scylladb (10.70.188.92 - seed): scylladb2 (10.70.188.39): I tried to confirm the client can connect to it, so using cqlsh , I got Connection error: ('Unable to connect to any servers', {'127.0.0.1:9042': ConnectionRefusedError(111, "Tried connecting to [('127.0.0.1', 9042)]. Last error: Connection refused")}). cqlsh --help then suggested I to change the host with $CQLSH_HOST. So I did: scylladb2 (10.70.188.39): I also tried to connect to node 1 with: scylladb2 (10.70.188.39): And the documentation says I should run nodetool status to check the status of the cluster. But it failed with error running operation: std::system_error (error system:111, Connection refused) . So I checked the config which you can specify -h for host, so I did: scylladb2 (10.70.188.39): This is a little bit odd for me because I set multiple nodes, but there’s only one node there. I can connect to node 1 as well, but I can only see one node there. scylladb2 (10.70.188.39): Create a ScyllaDB Cluster - Single Datacenter (DC) | ScyllaDB Docs @Felipe_Cardeneti_Mendes: 1. It seems like node2 bootstrapped as a separate cluster and is unaware of node1. Wipe it clean and retry. 2. Set rpc_address to 0.0.0.0, and rpc_broadcast_address to the node relevant IP 3. You should specify the IP you set as RPC address as a contact point @Pham_William: @Felipe_Cardeneti_Mendes That works! Thank you very much. I would assume I should be using the RPC address of the seed? Or it can be whatever node? @Felipe_Cardeneti_Mendes: Any node is fine if you refer to the contact point for the app, but you more often than not want to specify more than a single node as a contact point @Pham_William: Hi, @Felipe_Cardeneti_Mendes, there was some issue, so I started from scratch. With 3 nodes as follows: ±----------±--------±---------------------±---------------------------------------------±----------------±----------+ | scylladb1 | RUNNING | 10.70.188.63 (eth0) | fd42:ab78:d54:2140:216:3eff:fef1:2d6a (eth0) | CONTAINER | 0 | ±----------±--------±---------------------±---------------------------------------------±----------------±----------+ | scylladb2 | RUNNING | 10.70.188.242 (eth0) | fd42:ab78:d54:2140:216:3eff:fe67:72ef (eth0) | CONTAINER | 0 | ±----------±--------±---------------------±---------------------------------------------±----------------±----------+ | scylladb3 | RUNNING | 10.70.188.145 (eth0) | fd42:ab78:d54:2140:216:3eff:fea9:4149 (eth0) | CONTAINER | 0 | ±----------±--------±---------------------±---------------------------------------------±----------------±----------+ They all have the same /etc/scylla/cassandra-rackdc.properties (only difference in rack number 1,2,3) Now, if I run nodetool status on whatever node, will all return node 1 status only. node 2 and 3 will show node 1 only but without information on its load I have also tried to remove the data from nodes 2 and 3, but no luck. Is there a reason why it’s only showing node 1 with status DN “Down Normal”, instead of up?: Note that, systemctl status scylla-server all returns active (running) @Felipe_Cardeneti_Mendes: Is this a “local” docker deployment? How are you making changes to the config file? In general, it’s best to start with a single node, then add nodes incrementally until you get the config and steps right. For example, if you add the 2nd node and it doesn’t work, then probably it bootstrapped early and you have to clear it up. Check its logs, see what happened, then correct whatever insights the logs might be telling you. Once you’ve got 2 nodes up, just repeat the same for other particular nodes set rpc_address to 0.0.0.0 and broadcast_rpc_address to private_ip. restart the cluster. this is how i solved it. --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-7/2300 Title: [RELEASE] ScyllaDB Enterprise 2024.1.7 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2024.1.7, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2024.1 LTS Release. Related Links Get ScyllaDB Enterprise 2024.1 (customer… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-7/2300 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2024.1.7 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2024.1.7 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2024.1.7, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2024.1 LTS Release. The following issues are fixed in this release (with an open-source reference, if available): --- ### Page: https://forum.scylladb.com/t/what-is-the-intended-way-to-handle-errors-from-insert-update-using-the-rust-driver/2301 Title: What is the intended way to handle errors from insert/update using the rust driver? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I notice many helpful examples for the rust driver here, it for the most part has examples for almost everything I want to do: Unless I missed it, none of the examples describe how to handle the result of using sessio… Language: en Canonical URL: https://forum.scylladb.com/t/what-is-the-intended-way-to-handle-errors-from-insert-update-using-the-rust-driver/2301 ## Headings Structure: H1: What is the intended way to handle errors from insert/update using the rust driver? H3: scylla-rust-driver/examples at main · scylladb/scylla-rust-driver H3: Related topics ## Main Content: H1: What is the intended way to handle errors from insert/update using the rust driver? H3: scylla-rust-driver/examples at main · scylladb/scylla-rust-driver H3: Related topics I notice many helpful examples for the rust driver here, it for the most part has examples for almost everything I want to do: Async CQL driver for Rust, optimized for ScyllaDB! - scylladb/scylla-rust-driver Unless I missed it, none of the examples describe how to handle the result of using session.query to insert or update except for just throwing it upwards. I am trying to work out the best way to read the result of an insert/update and do different things depending on the result. In the documentation, I notice that the query() function does contain some helper functions like rows() or rows_or_none() and things like that. But nothing obvious regarding explicitly checking the results of an insert/update. I need to check that the update reached the server and the server indicated that the update succeeded. But the sample code seems to ignore error handling for the most part, i.e. What is the best way to capture if the query definitely failed? I can see this in the source of the driver: But this seems to be designed to do something assuming the query was successful, but the response was not what was expected? UPDATE Never mind, I worked it out in the end. It’d be neat to see some examples of handling errors at some point. --- ### Page: https://forum.scylladb.com/t/how-to-create-a-table-with-2-primary-keys-and-0-clustering-keys/2303 Title: How to create a table with 2 primary keys and 0 clustering keys - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello all, In the tutorials provided in the scylla university, it says that it is possible to have one or more primary keys and 0 or more clustering keys. In the example provided for compound keys in scylladb you can d… Language: en Canonical URL: https://forum.scylladb.com/t/how-to-create-a-table-with-2-primary-keys-and-0-clustering-keys/2303 ## Headings Structure: H1: How to create a table with 2 primary keys and 0 clustering keys H3: Related topics ## Main Content: H1: How to create a table with 2 primary keys and 0 clustering keys H3: Related topics Hello all, In the tutorials provided in the scylla university, it says that it is possible to have one or more primary keys and 0 or more clustering keys. In the example provided for compound keys in scylladb you can define a table like this: In this example, it creates a table with a primary key with pet_chip_id and with a clustering key on time column. I have two questions: Add one additional set of parenthesis, like this: In general, an extra set of parenthesis can be used to delineate the partition key from the clustering key: PRIMARY KEY ((partition_key_col1, partition_key_col2), clustering_key_col1, clustering_key_col2). The part in the inner parenthesis is the partition key. Clustering keys are used for more than just sorting, they identify rows, just like partition keys identify partitions. Clustering keys allow you to have many rows within the same partition. If you want a single row in a partition, don’t define a clustering key. --- ### Page: https://forum.scylladb.com/t/is-scylladb-a-suitable-starting-point-for-learning-nosql-databases/2306 Title: Is ScyllaDB a Suitable Starting Point for Learning NoSQL Databases? - University and Training - ScyllaDB Community NoSQL Forum Meta Description: Hello ScyllaDB Community, I’m relatively new to the world of NoSQL databases and am considering ScyllaDB as my entry point. Having already gained a foundational understanding of SQL databases, I’m interested in expandin… Language: en Canonical URL: https://forum.scylladb.com/t/is-scylladb-a-suitable-starting-point-for-learning-nosql-databases/2306 ## Headings Structure: H1: Is ScyllaDB a Suitable Starting Point for Learning NoSQL Databases? H3: Related topics ## Main Content: H1: Is ScyllaDB a Suitable Starting Point for Learning NoSQL Databases? H3: Related topics Hello ScyllaDB Community, I’m relatively new to the world of NoSQL databases and am considering ScyllaDB as my entry point. Having already gained a foundational understanding of SQL databases, I’m interested in expanding my skills into NoSQL technologies. Given ScyllaDB’s unique characteristics and performance advantages, I’m curious to know if it would be a good choice for someone who is just beginning to explore NoSQL databases. Here are a few specific questions I have: I appreciate any guidance or experiences you can share that would help make this learning transition as smooth as possible. I can tell you that if you come from a SQL background, ScyllaDB in a first moment will trick your mind. But it’s way easier to understand than relational databases IMHO. Now, about your questions: I usually do livecoding session daily at twitch (@danielhe4rt) about ScyllaDB environment. We can chat there after you do some classes at our university. --- ### Page: https://forum.scylladb.com/t/upcoming-scylladb-cloud-portal-maintenance-on-july-29-2024/2307 Title: Upcoming ScyllaDB Cloud Portal Maintenance on July 29, 2024 - Announcements - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB Cloud Portal will undergo maintenance on July 29, 2024 between 09:00 to 12:00 UTC. This transition is aimed at enhancing our scalability and reliability as well as improving ScyllaDB Cloud’s control plane in… Language: en Canonical URL: https://forum.scylladb.com/t/upcoming-scylladb-cloud-portal-maintenance-on-july-29-2024/2307 ## Headings Structure: H1: Upcoming ScyllaDB Cloud Portal Maintenance on July 29, 2024 H3: Related topics ## Main Content: H1: Upcoming ScyllaDB Cloud Portal Maintenance on July 29, 2024 H3: Related topics The ScyllaDB Cloud Portal will undergo maintenance on July 29, 2024 between 09:00 to 12:00 UTC. This transition is aimed at enhancing our scalability and reliability as well as improving ScyllaDB Cloud’s control plane infrastructure. There is no impact to database availability. Planned Cut Over Time: July 29, 2024 09:00 to 12:00 UTC Planned Downtime: up to 3 hours During this period, the following services will be affected: This migration is an important step towards enabling our infrastructure to support future feature enhancements . We apologize for any inconvenience during this scheduled maintenance. Please reach out to ScyllaDB Support with any questions or concerns at cloud-support@scylladb.com. Thank you for your attention to this matter. – The ScyllaDB Cloud team --- ### Page: https://forum.scylladb.com/t/debezium-query-cdc-streams/2308 Title: Debezium query CDC streams - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi I am using debezium Kafka connect to read CDC log table of ScyllaDB . I am willing to understand , how does debezium query CDC streams for the FIRST TIME ,somewhere i have read it queries now() - ttl of cdc table ,… Language: en Canonical URL: https://forum.scylladb.com/t/debezium-query-cdc-streams/2308 ## Headings Structure: H1: Debezium query CDC streams H3: Related topics ## Main Content: H1: Debezium query CDC streams H3: Related topics Hi I am using debezium Kafka connect to read CDC log table of ScyllaDB . I am willing to understand , how does debezium query CDC streams for the FIRST TIME ,somewhere i have read it queries now() - ttl of cdc table , how does it maintains and use timestamp in connect-offset topic ? how it uses offset to polls queries after scylla.query.time.window.size ? If there is any link or tutorial on this , it would be helpful to understand Hi, Currently available resources I know of right now are the following: The readme of CDC connector, readme of scylla-cdc-java (which is used underneath), scylla-cdc-java printer readme (recommended read; points to replicator example as a follow up). Additionally there is Scylla CDC documentation and “ScyllaDB university” resources about CDC In regards to specifics of connector your best bet is probably looking at the source code directly. Many of the configuration options have wordy descriptions, but if the Kafka platform you’re using does not provide GUI, you may haven’t had the opportunity to see them. how does debezium query CDC streams for the FIRST TIME ,somewhere i have read it queries now() - ttl of cdc table That would be correct. There is no use to query earlier data anyway - there shouldn’t be any. I believe this is relevant section from scylla-cdc-java - Worker.java#createTasksWithState() For more insight on how does the cdc connector maintains offsets see the TaskStateOffsetContext class. I believe this is what holds the relevant information about the current progress. If you want to trace how a singular row is processed by connector see consume method of ScyllaChangesConsumer. how it uses offset to polls queries after scylla.query.time.window.size ? After processing current window connector should proceed to the next window of the same size. --- ### Page: https://forum.scylladb.com/t/nodes-stuck-in-start-up-phase-seed-ip-and-listen-address/2309 Title: Nodes stuck in start up phase, seed ip and listen address - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/nodes-stuck-in-start-up-phase-seed-ip-and-listen-address/2309 ## Headings Structure: H1: Nodes stuck in start up phase, seed ip and listen address H3: Related topics ## Main Content: H1: Nodes stuck in start up phase, seed ip and listen address H3: Related topics Originally from the User Slack @Joakim_Lindqvist: Hey I am attempting to test scylla 6.0.1 on a new cluster and running into an issue were my nodes seems stuck starting up in phase starting the view builder. This node that is starting up has ip 192.168.2.239 and is set to be the seed node. This is on AWS using the scylla amis in us-east-1, using i4i.large instances. This is meant as a small dev deployment thus the small scale. The scylla monitoring is reporting the status as Starting. I have given this node about 2 hours and it still hasn’t moved. Scylla.yaml cassandra-rackdc.properties @dor: I am not sure but maybe it’s related to Raft not finding a majority, as it’s a single node cluster. It should work but maybe it’s an issue. If you’ll add 2 more nodes, you’ll see whether I was right or not @Joakim_Lindqvist: I attempted to add one more node and that didn’t seem to help. I can add another one to see if it fixes itself. The 3rd node is up, doesn’t seem to be making progress. @dor: hmm, sorry, I didn’t see anything special in the logs. I assume the config is completely fresh (minus edits) @Joakim_Lindqvist: Yeah the config is all generated by the machine image except for my edits to the seed ip @avi: @Kamil_Braun please help @Kamil_Braun: > I am not sure but maybe it’s related to Raft not finding a majority, as it’s a single node cluster. “majority” is always calculated from total number of nodes. If your total is 1 then your majority is 1. If total is 3 then majority is 2 etc. So single node clusters work as usual. I suspect that you tried to first boot the node with a different seed by mistake, then stopped it and changed the seed. But the old seed was persisted in the cluster discovery algorithm which now requires to contact that seed in order to proceed OR you concurrently tried to start a second node which contacted this node and inserted itself as one of participants of cluster discovery algorithm, and then shut down. But it’s also one of the required contact points now. In any case the node looks to be stuck inside cluster discovery algorithm. stop it, delete everything (workdir i.e. data, commitlog, etc.), make sure you have correct seed setup (see documentation), try again @Joakim_Lindqvist: I did initially start it with 127.0.0.1 as the seed ip, yes. I did start the second node after a while (as I suspected a similar issue to Dor were we might need more nodes for it to start) but it should have started up in time for that. I will tear it all down and try again and see if I can repro it. Thanks for the suggestion! I have tested tearing this down and setting up a new cluster. This time with a single node, this is set to have 127.0.0.1 as its seed node (seems to be what happens by default when I do not specify any seed ips). There are no other nodes attempting to connect. I am still seeing the same issue with it stopping at the same point. @Kamil_Braun: You cannot use 127.0.0.1 if your listen_address / broadcast_address etc. is different you must use the same IP everywhere https://opensource.docs.scylladb.com/stable/operating-scylla/procedures/cluster-management/create-cluster.html Create a ScyllaDB Cluster - Single Datacenter (DC) | ScyllaDB Docs so if you used listen_address: 192.168.2.239 then you must use but now you have to teardown again. If you use seeds: "127.0.0.1" and then change later to seeds: "192.168.2.239" it won’t work because the cluster discovery algorithm already persisted 127.0.0.1 and will continue trying to contact it (and never finish) @Joakim_Lindqvist: I didn’t explicitly set 127.0.0.1 as the seed or listen address, I set the seed to empty string which seems to then default to that ip. The machine image generated yaml sets the listen address to the ip of the machine. But I guess this may have been a result of my changes to how our bootstrapping works were I do not allocate a ec2 ip ahead of time for our machines and instead try to update it after the machine has started and the ip has been allocated. But I am fairly sure we have been able to setup new clusters like this in 5.x (we never tried this with 4.x as we had predefined ips for those). But this gives me some data to go on, I will try and see if I can rework our bootstrapping to work better without having to rely on predefined ips (as that results in special case handling of this initial seed node which doesn’t really make sense with how seeds are not really a thing anymore once the cluster starts up). I resolved this now. Avoiding specifying a seed ip for the first node, e.g. a machine image user data like this: Resolves this issue as the machine image defaults to the private ip of the ec2 instance. So this was really a problem in our ec2 setup, I am unclear on why this was not a issue for the new clusters we have setup with 5.x but fundamentally we were not following the setup recommendations and doing so resolves the issue. --- ### Page: https://forum.scylladb.com/t/cant-start-on-vps-that-does-enough-memory/2310 Title: Cant start on VPS that does enough memory? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I am starting scylla for on a VPS for a very small project, when I configure with 1 cpu and 500M or even 750M scylla refuses to start, even with --cpuset 1 setup and with/without --overprovisioned scylla[5761]: command … Language: en Canonical URL: https://forum.scylladb.com/t/cant-start-on-vps-that-does-enough-memory/2310 ## Headings Structure: H1: Cant start on VPS that does enough memory? H3: Related topics ## Main Content: H1: Cant start on VPS that does enough memory? H3: Related topics I am starting scylla for on a VPS for a very small project, when I configure with 1 cpu and 500M or even 750M scylla refuses to start, even with --cpuset 1 setup and with/without --overprovisioned As far as I know the VPS should have enough memory: Im still trying to poke around and work out why, but I thought I would post here in the mean time to see if anyone has some useful pointers in the right direction for me. I’ve never had this problem with Cassandra before. After some googling I have discovered that this will increase the reported “free” memory, but scylla still will not start: There still appears to be some strange hard limit somewhere: Seastar will not consider swap when calculating available memory, it will only consider RAM. Furthermore, when considering the vailable RAM, it will substract the kernel reserve, which it reads from /proc/sys/vm/min_free_kbytes. Looks like on your system this amount is 500 MB, which is less them the minimal amount seastar was told to allocate for itself (--memory parameter). See seastar/src/core/resource.cc at bd27a45289ec13fddddcf383436cd6141a32c1b0 · scylladb/seastar · GitHub for the full calculation. --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2023-1-10/2318 Title: [RELEASE] ScyllaDB Enterprise 2023.1.10 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2023.1.10, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2023.1 LTS Release. 2023.1.10 patch release includes multiple minor bug fix… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2023-1-10/2318 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2023.1.10 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2023.1.10 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2023.1.10, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2023.1 LTS Release. 2023.1.10 patch release includes multiple minor bug fixes. ScyllaDB Enterprise’s latest Long Term Support (LTS) is 2024.1. While we continue to support 2023.1, you are encouraged to upgrade to it in coordination with the ScyllaDB Support team. The following issues are fixed in this release (with an open-source reference, if available): --- ### Page: https://forum.scylladb.com/t/release-scylladb-cloud-16-july-2024/2319 Title: [RELEASE] ScyllaDB Cloud - 16 July 2024 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: We are happy to announce a new update to the ScyllaDB Cloud service. New features and enhancements: Added support for ScyllaDB Enterprise version 2024.1.7 Added me-central-1 (Middle East (UAE)) region support for the … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-cloud-16-july-2024/2319 ## Headings Structure: H1: [RELEASE] ScyllaDB Cloud - 16 July 2024 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Cloud - 16 July 2024 H3: Related topics We are happy to announce a new update to the ScyllaDB Cloud service. New features and enhancements: --- ### Page: https://forum.scylladb.com/t/scylla-manager-getting-error-when-trying-to-restore-backup-from-s3-no-valid-connect-address-for-host/2323 Title: Scylla Manager getting error when trying to restore backup from S3: "no valid connect address for host" - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/scylla-manager-getting-error-when-trying-to-restore-backup-from-s3-no-valid-connect-address-for-host/2323 ## Headings Structure: H1: Scylla Manager getting error when trying to restore backup from S3: "no valid connect address for host" H3: Related topics ## Main Content: H1: Scylla Manager getting error when trying to restore backup from S3: "no valid connect address for host" H3: Related topics Originally from the User Slack @Sabin: Hi, whenever I try to add a cluster on Scylla Manager with --username and --password, I am greeted with this error message: But if I remove --username and --password, it is added without any errors. I am trying to restore backup from S3 repository and restore fails if I do not provide the credentials: I have configured the default credentials on Scylla Manager’s configuration file. @Felipe_Cardeneti_Mendes: what’s the full command line you are trying to run? 0.0.0.0 isn’t a valid address you know @Sabin: This is the command that fails: and this one works fine except I see CQL TIMEOUT For the last command, the output looks like this: @Felipe_Cardeneti_Mendes: does any of your nodes have rpc_address: 0.0.0.0 ? show all rpc_* entries on your scylla.yaml definitions @Sabin: Hmm. All of them have that: listen_on_broadcast_address is set to true however and it looks like this on one of the node @Felipe_Cardeneti_Mendes: that’s the problem then. You need to set broadcast_rpc_address to the actual node IP, not 0.0.0.0 @Sabin: okay let me try Thanks @Felipe_Cardeneti_Mendes it worked. I do not understand why RPC played such a role there? Is there any document that I can read to improve my knowledge upon this? @Felipe_Cardeneti_Mendes: this is the IP broadcasted to drivers when they connect to the node. if its 0.0.0.0 the drivers dont know how to route traffic to it --- ### Page: https://forum.scylladb.com/t/release-scylladb-java-driver-4-18-0-1/2325 Title: [RELEASE] ScyllaDB Java Driver 4.18.0.1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Java Driver 4.18.0.1, a feature-rich and highly tunable Java client library for Scylla and [Apache Cassandra®] (2.1+), using exclusively Cassandra’s binary protocol and Cassandra Quer… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-java-driver-4-18-0-1/2325 ## Headings Structure: H1: [RELEASE] ScyllaDB Java Driver 4.18.0.1 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Java Driver 4.18.0.1 H3: Related topics The ScyllaDB team announces ScyllaDB Java Driver 4.18.0.1, a feature-rich and highly tunable Java client library for Scylla and [Apache Cassandra®] (2.1+), using exclusively Cassandra’s binary protocol and Cassandra Query Language (CQL) v3 This version brings just an enhancement to schema queries on Scylla clusters. Before the schema queries had configurable timeouts that were purely driver-side. Now when querying Scylla the driver will also make use of ScyllaDB CQL “Using Timeout” extension by applying additional clause to schema queries. This is a first release post here for the java driver, so it only makes sense to also mention that version 4.18.0.0 introduced the support for tablets - a new data distribution mechanism introduced in ScyllaDB 6.0. If you are using this mechanism with your tables it is highly encouraged to switch to newer Java Driver version. Itemized changelog is available on releases page https://github.com/scylladb/java-driver/releases Format of these release posts is subject to change. Let me know if you have any suggestions and please look forward to future releases. You can check out the source code on GitHub https://github.com/scylladb/java-driver/tree/scylla-4.x For how to use the driver please refer to “Getting the driver” section. --- ### Page: https://forum.scylladb.com/t/running-a-cluster-with-docker-compose/2328 Title: Running a cluster with Docker Compose - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello everyone! I’m trying to deploy a ScyllaDB cluster using docker-compose. I have the following YAML file: name: scylla-cluster services: scylla-node-1: container_name: scylla-node-1 image: scylladb/scyl… Language: en Canonical URL: https://forum.scylladb.com/t/running-a-cluster-with-docker-compose/2328 ## Headings Structure: H1: Running a cluster with Docker Compose H3: Related topics ## Main Content: H1: Running a cluster with Docker Compose H3: Related topics I’m trying to deploy a ScyllaDB cluster using docker-compose. I have the following YAML file: When I run docker-compose up, I get the following error from each node: Now, if I run cat /proc/sys/fs/aio-max-nr on the host machine, I get 1048576. If I run the same command inside one of the containers, I get 65536.What should I do to increase the aio-max-nr inside the container? Or what should I do to make it work in general? This is what running lscpu gives me inside a container: Apparently, the issue was that I was using Docker Desktop for Linux, that runs inside a VM instead of the host OS. This meant that my containerized /proc tree was different than my host OS /proc. I tried to search of a way to edit the /proc of the VM but did not find anything, so I just installed regular docker engine on the host and tried it that way, and it finally worked: --- ### Page: https://forum.scylladb.com/t/upgrading-scylladb-version-downtime-and-raft-configuration/2331 Title: Upgrading ScyllaDB version, downtime, and Raft configuration - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/upgrading-scylladb-version-downtime-and-raft-configuration/2331 ## Headings Structure: H1: Upgrading ScyllaDB version, downtime, and Raft configuration H3: Related topics ## Main Content: H1: Upgrading ScyllaDB version, downtime, and Raft configuration H3: Related topics Originally from the User Slack @Carter: unless i’m misunderstanding, you cannot upgrade from open source 5.2 to 5.4 without downtime for the raft upgrade? @Felipe_Cardeneti_Mendes: You’re misunderstanding. Unless you strictly opted out of Raft in the configuration, Raft will be enabled as you upgrade. All you need to check is whether it completed successfully. @Carter: I tried the upgrade on staging – just changing our image version from 5.2 to 5.4, no config changes. It required all nodes to be upgraded and online before the upgrade could commence, and only then, it would open the port for CQL. consistent_cluster_management was not set to true nor false when I did the upgrade – is that why? I should mention that our clusters are tiny at 3 nodes each @Felipe_Cardeneti_Mendes: Well, yes all nodes must be up and running before you start the upgrade. Not sure WDYM on the last statement for CQL. Are you saying you upgraded and the nodes wouldn’t get start CQL until you finalized the entire upgrade? That shouldn’t happen. @Carter: Hmm, I stand corrected! I double checked my logs and the way I was streaming them, it was coming from all pods and it caused some confusion. Thanks for your help, @Felipe_Leme_Gasparini! Error with upgrading … I have followed the instructions from here i.e. applied changes to scylla.yaml before restart and then restarted the nodes. --- ### Page: https://forum.scylladb.com/t/online-migration-from-amazon-mysql-to-scylladb/2333 Title: Online Migration from Amazon MySQL to ScyllaDB - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi team, We have a requirement to migrate the data from Amazon MySQL to ScyllaDB. This should be online migration. Do we have any tools to perform this activity? Can we leverage the ScyallDB migrator for this task? Coul… Language: en Canonical URL: https://forum.scylladb.com/t/online-migration-from-amazon-mysql-to-scylladb/2333 ## Headings Structure: H1: Online Migration from Amazon MySQL to ScyllaDB H3: Related topics ## Main Content: H1: Online Migration from Amazon MySQL to ScyllaDB H3: Related topics We have a requirement to migrate the data from Amazon MySQL to ScyllaDB. This should be online migration. Do we have any tools to perform this activity? Can we leverage the ScyallDB migrator for this task? Could you please provide the suggestions on this? There are different options for performing a migration. Which version of ScyllaDB are you migrating to? A good starting point is the Migrating to ScyllaDB lesson on ScyllaDB University. You can also see an example of real world MySQL to ScyllaDB migration here. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-56-2024-07-19/2334 Title: Last week in scylla-cluster-tests.git master (issue #56; 2024-07-19) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 2cf61970…baf5542e range are covered. There were 36 non-merge commits from 7 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-56-2024-07-19/2334 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #56; 2024-07-19) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #56; 2024-07-19) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 2cf61970…baf5542e range are covered. There were 36 non-merge commits from 7 authors in that period. Some notable commits: Added support for TEST_ERROR Argus status, indicating a run failed due to a test error, such as problems with cloud resources (e.g., insufficient capacity or spot termination). Performance tests can now use BYO ScyllaDB. The RPC transport error that occurs when one node boot is interrupted and Scylla does not start is ignored in BootstrapStreamingError nemesis. New jobs were created for tablet-enabled runs. We created a new branch for performance tests from master and refactored the directory structure. All the perf regression tests were moved into their own folder in each release (master/enterprise), based on the branch they originated from. We fixed BYO ScyllaDB to work with multidc setups, which required copying AMIs to multiple regions. A new feature in ScyllaDB enables advanced compression of internode RPC messaging, now supported by SCT, and was enabled in two tests: multidc rolling-upgrade and one longevity. An elasticity test measuring the time to scale out/in a cluster with parallel nodes bootstrap/decommission and doubling the load after growing was finalized. It supports both cql-stress and cassandra-stress tools as loaders. Doubling the load in GrowShrinkNemesis can also be set for longevity tests with the nemesis_double_load_during_grow_shrink_duration parameter. The default version of Scylla Manager was updated to 3.3.0. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/maximum-number-of-tables-in-a-cluster-std-bad-alloc-issue-dropping-tables/2335 Title: Maximum number of tables in a cluster, std::bad_alloc issue, dropping tables - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/maximum-number-of-tables-in-a-cluster-std-bad-alloc-issue-dropping-tables/2335 ## Headings Structure: H1: Maximum number of tables in a cluster, std::bad_alloc issue, dropping tables H3: Related topics ## Main Content: H1: Maximum number of tables in a cluster, std::bad_alloc issue, dropping tables H3: Related topics Originally from the User Slack @AmiSri: I am using Scylla version 5.4.3 I have 5882 tables till now And recently I faced std::bad_alloc and writes failed Is there a limit of number of tables in a keyspace Or limit of total number of tables in a cluster How to overcome the issue?? @Piotr_Smaroń: It may be that you created too many tables within one KS, per the doc: https://opensource.docs.scylladb.com/master/reference/limits.html, we “support” up to 5k tablets per KS Limits | ScyllaDB Docs @AmiSri: Yes its in a single keyspace with cdc enabled Is this relevant for all versions of ScyllaDB?? @Piotr_Smaroń @Piotr_Smaroń: we’re not testing limits in each version, but you can assume they apply to the recent versions @AmiSri: On dropping the tables, i found load of only one node is seems to be reduced while load on all other nodes in the cluster is same as earlier. I have cleared snapshots too Is a repair or rebuild required?? What is the correct way to drop tables in a cluster?? @Piotr_Smaroń @Piotr_Smaroń: just DROP TABLE should be enough if the disk is still full, you may take a look here https://opensource.docs.scylladb.com/stable/troubleshooting/drop-table-space-up.html Dropping tables doesn’t mean your load will decrease, for load to decrease you must lower the number of queries to the db. Perhaps the tables you dropped had partitions located only on a single node, and since you stopped querying these, your load for this single node has dropped Dropped (or truncated) Table (or keyspace) and Disk Space is not Reclaimed | ScyllaDB Docs @AmiSri: This is a restore server and no quering is going on here, do you mean partition key Please help me understanding “number of queries” mean @Piotr_Smaroń: by queries I meant both reading and writing to a database --- ### Page: https://forum.scylladb.com/t/unable-to-config-kafka-scylla-sink-connector-kafka-topic-to-scylla-table/2336 Title: Unable to config kafka-scylla-sink connector(Kafka topic to Scylla Table) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I have made the basic setup for syncing data from Kafka-topic to my scylla Topic. While setting up a connector for Kafka topic to a particular topic, I got stuck somewhere and was not able to debug it. If someone else h… Language: en Canonical URL: https://forum.scylladb.com/t/unable-to-config-kafka-scylla-sink-connector-kafka-topic-to-scylla-table/2336 ## Headings Structure: H1: Unable to config kafka-scylla-sink connector(Kafka topic to Scylla Table) H3: Related topics ## Main Content: H1: Unable to config kafka-scylla-sink connector(Kafka topic to Scylla Table) H3: Related topics I have made the basic setup for syncing data from Kafka-topic to my scylla Topic. While setting up a connector for Kafka topic to a particular topic, I got stuck somewhere and was not able to debug it. If someone else has encountered the same earlier, please help here with your expertise. Using following command: ./kafka-console-consumer.sh --bootstrap-server b-11.prod-eva-kafka.w5mjg98.c3.kafka.ap-south-1.amazonaws.com:9092 --topic ap-south-1-sentinel-WhiteListingForGoodHvpUsers --from-beginning I am getting list of json message. One of them is like this: Now I am required to sync all these message to my scyllaTable. Following is description of table: And I am using the following the kafka-scylla-sink.json config for controller: While hitting the POST request to start this connector, I am getting following error: Could you please help me in resolving the error. I tried a lot but didn’t find any promising solution for the same. P.S.: Thanks for your valuable time . I believe the root cause of the problem here is that the topic name does not exactly match the table name. From the connectors point of view there is no table named ap-south-1-sentinel-WhiteListingForGoodHvpUsers inside kafka_test keyspace in your database. Usually the connector would just create the table, however it does not do that when running in schemaless mode as is indicated by the warning you can see in the logs: The json data on your topic does not contain the schema information from which the connector would pull necessary information to build the table. Recently small fix that adds clearer log in such circumstances was added, so hopefully that will help troubleshooting those cases. The solution here is to either create the table with the exact name as the name of the topic or add an SMT that will change the topic name (of course, only inside messages) to the name of the table you’ve created. The table has to reasonably match the types of the columns with the types of the JSON fields. I’m not sure what type would match the nested “data” field that is inside your message. It seems your table has some columns named similarly to the fields of this nested structure, so one way to solve this is to transform the data on this topic to flatten the nested “data” field and leave only the subfields you are interested in at the same level as other fields in your json. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-239-2024-07-21/2337 Title: Last week in scylladb.git master (issue #239; 2024-07-21) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 53a6ec05ed…ad68a7f799 range are covered. There were 45 non-merge commits from 14 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-239-2024-07-21/2337 ## Headings Structure: H1: Last week in scylladb.git master (issue #239; 2024-07-21) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #239; 2024-07-21) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 53a6ec05ed…ad68a7f799 range are covered. There were 45 non-merge commits from 14 authors in that period. Some notable commits: A crash in the REST API call to get the Raft group 0 leader was fixed. A crash in certain cases of schema changes while an sstable was being written was fixed. The container image is now based on Ubuntu 24.04. The compiler toolchain used to build ScyllaDB is now itself optimized using profile-guided optimization, resulting in faster build speeds. Lightweight transactions (LWT) on the same partition are serialized on the coordinator since this generates less wasted work when Paxos transactions contend. A bug in this serialization, if we timed out while waiting to acquire the lock, was fixed. The master branch version was updated to 6.2, signifying the start of the 6.1 stabilization cycle. The sstable primary index reader will now respond to service shutdown requests. This can happen if we’re rebuilding the bloom filter for a large sstable when the service is shut down. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/what-is-the-size-limit-of-the-text-data-type-in-scylladb/2339 Title: What is the size limit of the text data type in ScyllaDB - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/what-is-the-size-limit-of-the-text-data-type-in-scylladb/2339 ## Headings Structure: H1: What is the size limit of the text data type in ScyllaDB H3: Related topics ## Main Content: H1: What is the size limit of the text data type in ScyllaDB H3: Related topics Originally from the User Slack @Kishore: Hi, What is the size limit on text data type in Scylla? @Felipe_Cardeneti_Mendes: Very big. By default you shouldn’t be able to write entries (entire mutation) larger than 16MB tho, which is half a commitlog segment size. You can bump it but then you would probably have not so great latency. @Kishore: got it. thank you. --- ### Page: https://forum.scylladb.com/t/using-in-in-a-query-for-a-specific-partition-is-the-entire-partition-fetched/2340 Title: Using IN in a query for a specific partition, is the entire partition fetched? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/using-in-in-a-query-for-a-specific-partition-is-the-entire-partition-fetched/2340 ## Headings Structure: H1: Using IN in a query for a specific partition, is the entire partition fetched? H3: Related topics ## Main Content: H1: Using IN in a query for a specific partition, is the entire partition fetched? H3: Related topics Originally from the User Slack @raghav_tandon: Hi, have a small question… If i have a partition with key/value pair and partition size will never go beyond 200kB. So, if i do IN query (for specific partition) to fetch only some keys would that be beneficial or scylla will always fetch entire partition and we will only save on network. @avi: It will likely fetch the entire partition, but you can check with tracing @raghav_tandon: so we have a design decision to make like a blob with a everything in one column or as a key/value store… if entire partition is fetched then both would give same performance? @avi: It’s generally better to have separate rows/columns rather than blobs (esp if they’re modified separately) @raghav_tandon: sure, Thanks for the clarification --- ### Page: https://forum.scylladb.com/t/why-do-my-clients-need-to-register-for-change-events/2342 Title: Why do my clients need to register for CHANGE events? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I’ve noticed when using the gocql driver that, by default, my client will register for TOPOLOGY_CHANGE, STATUS_CHANGE, and SCHEMA_CHANGE events. I can disable them using the driver, but I fear that might be a bad idea. H… Language: en Canonical URL: https://forum.scylladb.com/t/why-do-my-clients-need-to-register-for-change-events/2342 ## Headings Structure: H1: Why do my clients need to register for CHANGE events? H3: Related topics ## Main Content: H1: Why do my clients need to register for CHANGE events? H3: Related topics I’ve noticed when using the gocql driver that, by default, my client will register for TOPOLOGY_CHANGE, STATUS_CHANGE, and SCHEMA_CHANGE events. I can disable them using the driver, but I fear that might be a bad idea. How exactly do clients use these events? Doesn’t statement caching solve most of the problems with topology/schema changes? These events are part of CQL standard, see section 4.2.6. “Event” in Apache Cassandra | Apache Cassandra Documentation. This way the driver stays in sync with the db. If by statement caching you mean that some CQL queries are cached, then it’s not enough, because these don’t reflect the cluster state. Hey Piotr, thanks for taking the time to reply! I do have one follow up question: is the driver keeping track of schema changes? It seems weird to me that the driver is notified of all schema changes, even if they may be related to a table/view that the driver is not using. Is there a reason why it needs to know everything that’s going on across the cluster? The driver is notified about all schema changes, but what it is doing this this information is different for every type of event. For most of them it just clears the information about schema for given keyspace (see: gocql/events.go at master · scylladb/gocql · GitHub). Only for schemaChangeKeyspace it does some additional work if tokenAwarePolicy is used (for other policies this is ignored). --- ### Page: https://forum.scylladb.com/t/release-scylla-doctor-v1-1/2344 Title: [RELEASE]: Scylla Doctor v1.1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: Scylla Doctor v1.1 is released. Fixes: CPUSetCollector was getting stuck if perftune.py execution was failing. New features: New ClientConnectionCollector and DriverVersionAnalyzer were added. Artifacts can be dow… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-doctor-v1-1/2344 ## Headings Structure: H1: [RELEASE]: Scylla Doctor v1.1 H3: Related topics ## Main Content: H1: [RELEASE]: Scylla Doctor v1.1 H3: Related topics Scylla Doctor v1.1 is released. Artifacts can be downloaded from https://downloads.scylladb.com/downloads/scylla-doctor/ --- ### Page: https://forum.scylladb.com/t/coming-from-an-sql-background-new-to-nosql-scylladb-for-beginners-getting-started/2345 Title: Coming from an SQL background, new to NoSQL, ScyllaDB for beginners, getting started - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/coming-from-an-sql-background-new-to-nosql-scylladb-for-beginners-getting-started/2345 ## Headings Structure: H1: Coming from an SQL background, new to NoSQL, ScyllaDB for beginners, getting started H3: Related topics ## Main Content: H1: Coming from an SQL background, new to NoSQL, ScyllaDB for beginners, getting started H3: Related topics Originally from the User Slack @Sultan_Al-otaibi: Hello ScyllaDB Community, I’m relatively new to the world of NoSQL databases and am considering ScyllaDB as my entry point. Having already gained a foundational understanding of SQL databases, I’m interested in expanding my skills into NoSQL technologies. Given ScyllaDB’s unique characteristics and performance advantages, I’m curious to know if it would be a good choice for someone who is just beginning to explore NoSQL databases. Here are a few specific questions I have: @Felipe_Cardeneti_Mendes: You may want to join our ScyllaDB Labs event tomorrow where we’ll basically get all these answered. https://lp.scylladb.com/virtual-labs-2024-07-registration https://www.scylladb.com: ScyllaDB Labs: Building High-Performance Apps @Sultan_Al-otaibi: Thank you for directing me to the ScyllaDB Labs event. I’ve already registered and I am looking forward to it. I’m eager to learn more about transitioning from SQL to NoSQL with ScyllaDB, and I hope to gain deeper insights into my questions @Guy: Another helpful resource is ScyllaDB University, specifically the ScyllaDB Essentials course. It covers basic concepts of NoSQL databases and includes videos, quizzes, and labs covering ScyllaDB features, architecture, high availability, getting started, installation, and more. ScyllaDB University: S101: ScyllaDB Essentials – Overview of ScyllaDB and NoSQL Basics --- ### Page: https://forum.scylladb.com/t/cqlsh-keeps-reporting-errors-and-timeout-when-querying-empty-tables/2347 Title: Cqlsh keeps reporting errors and timeout when querying empty tables - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I found a very strange phenomenon during the test. Even when I query an empty table, it actually takes a long time. I have already performed a compact on this table. ]# du -sh /var/lib/scylla/data/alternator_test_5 0 … Language: en Canonical URL: https://forum.scylladb.com/t/cqlsh-keeps-reporting-errors-and-timeout-when-querying-empty-tables/2347 ## Headings Structure: H1: Cqlsh keeps reporting errors and timeout when querying empty tables H3: Related topics ## Main Content: H1: Cqlsh keeps reporting errors and timeout when querying empty tables H3: Related topics I found a very strange phenomenon during the test. Even when I query an empty table, it actually takes a long time. I have already performed a compact on this table. There are other tables with a large number of records in the cluster, but it should not affect the query of empty tables. Please provide as many details as possible so people can help you. This includes things like the ScyllaDB version, hardware, OS, and the data model you’re using. This is a known issue in ScyllaDB. Scans on sparse (nearly empty) tables often time out. The reason for this is in the vnode architecture. Every database cluster has number_of_nodes * 256 vnodes. When scanning a table, each one of these has to be scanned individually, in ring order, to read the entire content of the table. With a dense table, that has a lot of data, the query will quickly fill up the page and return to the client, maybe the page will be filled after reading 1-2 vnodes only. In a sparse table, it may take hundreds or thousands of vnodes to fill a page. In the meanwhile, ScyllaDB will not return to the client, because it wants to fill the page first, so the query will end up timing out. This can be worked around by increasing the range_request_timeout_in_ms in ScyllaDB configuration, and increasing the timeout for these queries on the client side as well. --- ### Page: https://forum.scylladb.com/t/workaround-for-compatibility-issue-backup-and-restore-using-an-older-scylladb-version/2354 Title: Workaround for compatibility issue, backup and restore using an older ScyllaDB version - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/workaround-for-compatibility-issue-backup-and-restore-using-an-older-scylladb-version/2354 ## Headings Structure: H1: Workaround for compatibility issue, backup and restore using an older ScyllaDB version H3: Related topics ## Main Content: H1: Workaround for compatibility issue, backup and restore using an older ScyllaDB version H3: Related topics Originally from the User Slack @Sabin: Hi Team, After restoring the backup on entire different cluster using the script (https://github.com/GoogleCloudPlatform/cassandra-cloud-backup), we noticed that the cluster status is DN as it picked up IP address of the previous cluster which is not accessible from this newly created cluster where restore operation is being done. Is there way to change the IP address somehow? The script that is used backs up entire data directory along with commit logs and saved caches directory for each node of the cluster. Old cluster and it’s IP looks like this: GitHub: GitHub - GoogleCloudPlatform/cassandra-cloud-backup: Cassandra backups to Google Cloud Storage @Chaitanya_Tondlekar: Better to use scylla manager 3.2 for better and easy data restoratiom @Sabin: We already have Scylla Manager v2.2.0-0.20201103.a3fdb862. I could configure it to push backup data to S3 but I am not sure if same data can be restored on newer cluster. There’s already an issue with backup/restore for they take way too much time (400 GB in 6 hours) to complete @Chaitanya_Tondlekar: 1. upgrade the scylla manager to 3.2 atleast. 2. If not, you can use https://github.com/scylladb/scylla-manager/tree/master/ansible/restore GitHub: scylla-manager/ansible/restore at master · scylladb/scylla-manager @Sabin: how is the compatibility? Scylla version itself is 4.2.1, will Scylla Manager function with that version? @Chaitanya_Tondlekar: you can check the compalibility but ansible should work if you have taken backup from scylla manager @Sabin: Sadly it is not from Scylla Manager. It’s backed up using the script I mentioned above, which is why the issue @avi: ScyllaDB 4.2.1 has reached end-of-life, upgrade to a supported version @Sabin: It partly for that purpose @avi Backup is weird and we inherited the system. ScyllaDB’s knowledge isn’t that much in our case. So I am just trying to figure out what we can do. @avi: I guess you can take snapshots and copy the files away manually, if you can’t find your arms and legs @Sabin: Can’t I just copy the entire data directory and restore it on new cluster? @avi: You can restore it with nodetool refresh --load-and-stream @Sabin: okay. I will give it a try once more and report back Apart from the approach listed above, I gave it a go with Scylla Manager as well. On old cluster running version 4.2.1 of ScyllaDB, after configuring Wasabi bucket to store backups and verifying that the setup works, I took a backup of a keyspace (productdb) via Scylla Manager (v2.2) using following command: I, then, used DESCRIBE KEYSPACE productdb query via cqlsh to get the schema and stored it for later use. On new cluster running ScyllaDB v6.0.1, and Scylla Manager v3.3, I set the Wasabi bucket to the same bucket as with old cluster. I SOURCEed the schema dumped previously to prepare for restore. Afterwards, I ran the following command to restore the keyspace It went well for 4% then it failed. On one of the node, I saw following error: I don’t know what to try next or where to look. Entire idea was to at least have a backup which can be restored so that we could migrate to newer cluster avoiding potential disaster of data loss. I don’t know what to try next or where to look. It looks like the error points to a different schema than the one existing in the backup. In this case, column cvr is missing. You should restore to the same schema as the source SSTables expect. IIRC, back in SM 2.2 days we used to dump the schema cql alongside the SSTables. --- ### Page: https://forum.scylladb.com/t/scylla-dashboards-custom-labels-not-available/2356 Title: Scylla dashboards: custom labels not available - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello there, I have a question related to monitoring. I ported your dashboards to my external monitoring stack (thanks for this!)(victoriametrics,alertmaganer,grafana) but dashboards do not provide filtering by Cluster… Language: en Canonical URL: https://forum.scylladb.com/t/scylla-dashboards-custom-labels-not-available/2356 ## Headings Structure: H1: Scylla dashboards: custom labels not available H3: Related topics ## Main Content: H1: Scylla dashboards: custom labels not available H3: Related topics Hello there, I have a question related to monitoring. I ported your dashboards to my external monitoring stack (thanks for this!)(victoriametrics,alertmaganer,grafana) but dashboards do not provide filtering by ClusterName or DC. I have two different environments , stage and prod which monitors by one vicroria metrics so it leads for me combining both. How you additionally prepare labeling metrics to differentiate by cluster/dc? Example: on your monitoring stack I see scylla_reactor_utilization{cluster=“prod_cluster”, dc=“my-prod-dc”, dd=“2”, instance=“IP_ADDR_OF_NODE”, job=“scylla”, shard=“0”} on my monitoring stack I do not have cluster or other stuff: scylla_reactor_utilization{host_name=“prod-scylladb-01”,instance=“IP:9180”,job=“scylladb_exporter”,shard=“0”} median:0.5203, min:0.3810, max:0.6867, last:0.5679 scylladb_exporter is a job for victoriametrics which simply checks 9180 port cluster and dc are labels set in prometheus/scylla_servers.yml, which are then copied to the Prometheus container when you use the default docker installation. --- ### Page: https://forum.scylladb.com/t/is-there-a-per-table-per-database-way-to-store-the-database-by-per-directory/2358 Title: Is there a per table / per database way to store the database by per directory? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: e.g. 1 db or 1 table stores in 1 database folder or 1 table folder. if i want to delete a database / table, i can just delete the folder Language: en Canonical URL: https://forum.scylladb.com/t/is-there-a-per-table-per-database-way-to-store-the-database-by-per-directory/2358 ## Headings Structure: H1: Is there a per table / per database way to store the database by per directory? H3: Related topics ## Main Content: H1: Is there a per table / per database way to store the database by per directory? H3: Related topics e.g. 1 db or 1 table stores in 1 database folder or 1 table folder. if i want to delete a database / table, i can just delete the folder if i want to delete a database / table, i can just delete the folder You should definitely not do that, you should DROP it before-hand. e.g. 1 db or 1 table stores in 1 database folder or 1 table folder. Kinda. workdir specifies which folder the database will use to store its files (commitlog, hints, etc). Although unusual, there’s nothing preventing you from creating different filesystems on a directory basis. I didn’t test it, but you should be able to assign a different filesystem to different keyspaces, tables are a bit different as they require assigning a random UUID at creation time. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-57-2024-07-26/2361 Title: Last week in scylla-cluster-tests.git master (issue #57; 2024-07-26) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the d6fc1d7a…a56b93f8 range are covered. There were 27 non-merge commits from 7 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-57-2024-07-26/2361 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #57; 2024-07-26) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #57; 2024-07-26) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the d6fc1d7a…a56b93f8 range are covered. There were 27 non-merge commits from 7 authors in that period. Some notable commits: We updated cql-stress to the latest version containing a fix for creating proper tables based on workflow. When we only have the owner ID and the AMI is not published in the marketplace (e.g., Oracle DB), we cannot use SSM to find it. We introduced a way to find them by owner and can use it by providing, e.g., resolve:owner:131827586825/x86_64/OL8.* for the ami_id_db_scylla parameter. Scylla-bench version was bumped to v0.1.22 containing a fix for counter tables creation based on workflow. A new job for specific test resources cleanup was added and can be easily used from Argus ‘Resources’ tab. A new nemesis_grow_shrink_instance_type test parameter was introduced to control the instance type the grow-shrink nemesis uses. When this option is used, the shrink phase will decommission exactly the node that was added, leaving the cluster the same size as before nemesis. Perf-simple-query results gained the perf_extra_jobs_to_compare Jenkins option so results can be compared across different jobs (releases). We replaced pip and pip-compile with uv (a package installer written in Rust) which is much faster. This reduced the time to build out requirements.txt from 16s to less than a second (with a warm cache). See docs for instructions on how to use it. See updated documentation (Notion) on how to configure AWS using OKTA in the local SCT git repo. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/missing-nodes-in-the-monitoring-dashboards-lack-of-resources/2362 Title: Missing nodes in the Monitoring dashboards, lack of resources? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/missing-nodes-in-the-monitoring-dashboards-lack-of-resources/2362 ## Headings Structure: H1: Missing nodes in the Monitoring dashboards, lack of resources? H3: Related topics ## Main Content: H1: Missing nodes in the Monitoring dashboards, lack of resources? H3: Related topics Originally from the User Slack @serg: Hi to everyone! I have the monitoring stack that works in Docker. But adding nodes up to 5 for 2 clusters I see some slowness and even misses of objects in Grafana - for example, only 3 nodes of a cluster I see in the dashbord and after refresh I see 3 another ones, or 4 or 2. And such behavour for many objects in Grafana, seems as lack of resources. On the server with containers I see about 15% of available memory (from 32 Gb) and about 3-5% of CPU load. So the server itself has plenty of resources. The container with Prometheus consumes 2+ Gb, so I doubt there are memore limits for containers. Are they and where? For ScyllaDB instance for monitoring I see such startup options (from ps aux): /usr/bin/scylla --memory 250M --log-to-syslog 0 --log-to-stdout 1 --default-log-level info --network-stack posix --developer-mode=1 --smp 1 But cannot find where these parameters are wrintten, I’d like to change them - maybe it helps to encrease performance and rendering in Grafana . Could you help me how to increase monitoring performance? @avi: The scylla instance serves scylla-manager, not monitoring @serg: Yes, Avi, you’re right. I have Manager on the same host. Sorry, for miseading. Would you recommend some settings for monitoring of many servers/clusters? May be some default docker parameters are not enough? @avi: I’m not an expert in monitoring, you can try giving it more memory and cpu. @serg: Thank you. I’ll try. By the way, It’s interesting where scylla parameters are written - in scylla* service files on /lib/systemd/system/ I’ve not found memory limit to 250M. Thanks for help. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-240-2024-07-28/2363 Title: Last week in scylladb.git master (issue #240; 2024-07-28) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the ad68a7f799…27b305b9d1 range are covered. There were 69 non-merge commits from 17 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-240-2024-07-28/2363 ## Headings Structure: H1: Last week in scylladb.git master (issue #240; 2024-07-28) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #240; 2024-07-28) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the ad68a7f799…27b305b9d1 range are covered. There were 69 non-merge commits from 17 authors in that period. Some notable commits: There is now a REST API to detect the sstable format version supported across the cluster. In order to make sstables durable, the directory where they are placed must be flushed after they are sealed. This is now done without re-opening the directory each time, saving some cycles. The system may sometimes drop the bloom filter of some sstables to save memory, and then reload it when memory is available. We no no longer reload bloom filters for sstables that are queued for deletion. Commitlog segments older than 24 hours will now flush corresponding memtables regardless of memory pressure. This allows more timely garbage collection of tombstones. ScyllaDB tracks internal maintenance work, as well as work requested by the user (for example, repair), as tasks. These tasks can now be virtual to reduce their memory footprint. Alternator, ScyllaDB’s implementation of the DynamoDB API, will now reject authentication from roles that do not have the LOGIN attribute. Materialized view updates destined to a node that has left the cluster are now dropped. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/what-is-the-recommended-way-to-manage-schema-evolution-migrations/2365 Title: What is the recommended way to manage schema evolution/migrations? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: When using an SQL database, it is quite common to use a tool, framework or library that handles database migrations, i.e. schema changes over time. Often these tools keep track of the schema version by using an additiona… Language: en Canonical URL: https://forum.scylladb.com/t/what-is-the-recommended-way-to-manage-schema-evolution-migrations/2365 ## Headings Structure: H1: What is the recommended way to manage schema evolution/migrations? H3: Related topics ## Main Content: H1: What is the recommended way to manage schema evolution/migrations? H3: Related topics When using an SQL database, it is quite common to use a tool, framework or library that handles database migrations, i.e. schema changes over time. Often these tools keep track of the schema version by using an additional table. This ensures for instance that the schema changes (migrations) are only run once and aren’t applied multiple times (which could lead to problems). What is the equivalent of this when using ScyllaDB, if any? Are there any tools or libraries to manage migrations and their versioning? Are there any issues with schema versioning due to the eventual consistency of ScyllaDB? I.e. how would one ensure that schema evolutions are not applied multiple times, or is this perhaps not a problem to begin with? It’s not a common problem when we’re talking about NoSQL and ScyllaDB. These Databases database are Query Driven, so you’re creating a leaderboard query like: SELECT * FROM leaderboard WHERE tier_id = 123 AND map_id = 1234 your modeling will be resulted on partition/clusterization plus other things related. Usually you only do partition/clusterization key once and then after that, if you miss something related to the keys you should probably create a new table to solve this new problem. But… I like the idea of tracking the changes inside a table for a tool but it still depends on which language are you using tbh. --- ### Page: https://forum.scylladb.com/t/lightweight-transactions-lwt-consistency-partitions-and-lock-granularity/2366 Title: Lightweight Transactions (LWT) consistency, partitions and lock granularity - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/lightweight-transactions-lwt-consistency-partitions-and-lock-granularity/2366 ## Headings Structure: H1: Lightweight Transactions (LWT) consistency, partitions and lock granularity H3: Related topics ## Main Content: H1: Lightweight Transactions (LWT) consistency, partitions and lock granularity H3: Related topics Originally from the User Slack @Jeffery_Utter: Do LWT’s lock across a larger area than the partition key? Like does it end up hashing the partition key into a bucket and then locking the bucket? (oversimplified, I’m sure) @dor: LWT only guarantee a single partitions consistency. One can batch but multiple coordinators can exists. So only single partition consistency really exists with LWT @Jeffery_Utter: Hmm, I think I follow that, but If I have two partition keys that map to the same token, can only one of those be executed concurrently? @avi: The lock has token granularity, so if two partition keys hash to the same token, they will indeed be serialized @Jeffery_Utter: Got it. And that defaults to 256 tokens per node? Is that correct? So if we’re seeing contention on lwts adding nodes would help? @avi: No, the token-per-node are just range boundaries. A partition key’s token is a 64-bit hash of the key, and always has the full 64-bit range. Accidental collisions are very rare. If you’re seeing contention, it’s likely on the same key --- ### Page: https://forum.scylladb.com/t/release-scylladb-6-0-2/2368 Title: [RELEASE] ScyllaDB 6.0.2 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 6.0.2, a bugfix patch release of the ScyllaDB 6.0 stable branch. ScyllaDB Open Source 6.0.2, like all past and future 6.x.y releases, is backward compatible and supports r… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-6-0-2/2368 ## Headings Structure: H1: [RELEASE] ScyllaDB 6.0.2 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 6.0.2 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 6.0.2, a bugfix patch release of the ScyllaDB 6.0 stable branch. ScyllaDB Open Source 6.0.2, like all past and future 6.x.y releases, is backward compatible and supports rolling upgrades. Users are encouraged to upgrade to 6.0.2. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/data-model-for-frequent-deletes-with-partition-key/2370 Title: Data model for frequent deletes with partition key - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello! I want to ask for a suggestion on how to rethink my data model for a use case in which I’m using ScyllaDB, if anyone can help me out with a suggestion. I’m using ScyllaDB to store the data generated by running m… Language: en Canonical URL: https://forum.scylladb.com/t/data-model-for-frequent-deletes-with-partition-key/2370 ## Headings Structure: H1: Data model for frequent deletes with partition key H3: Related topics ## Main Content: H1: Data model for frequent deletes with partition key H3: Related topics I want to ask for a suggestion on how to rethink my data model for a use case in which I’m using ScyllaDB, if anyone can help me out with a suggestion. I’m using ScyllaDB to store the data generated by running multiple scripts. All of them run in parallel, and they generate a different amount of data. Up until now I’ve been using a bad partition key as the run_id for each script run. However, since they generate different amount of data, that caused some issues because some partitions are bigger than others and some nodes in my cluster have run out of disk space because of that. What I want to know is how to deal with this, keeping in mind that: My question is, what is better to do about DELETES like this that need to target data from multiple partitions, especially when the partition key has high cardinality? Thanks I can’t imagine how is your modeling fully but here it goes: maybe you can use a partition key with more than one value. As I could understand from your modeling is that you’re adding everything inside the partition with run_id. If you can share a snippet of your data modeling would be good either (don’t need to be a real one, but something that shows the problem you’re facing). hello! thanks and ofc, here is how the table is created: what I need to do throughout the script run is : and at the end I need to be able to do the values of the request_id column are hashed values that would be really good for spreading the data as a partition key, however with it I won’t be able to correctly target the delete for all the data saved throughout the run. Your developer environment doesn’t have something that can get merged at the partition key to make it traceable? But you can also add a field like timeuuid as a clustering and uses run_id and request_id as your partition ((run_id, request_id), ts_uuid) so you will never have duplicates during the run and still traceable. Hello, as I mentioned previously, what could be merged to the partition key to really increase the cardinality and split the data better is the values from the request_id column. But that won’t really help me with targeting all the data of a scrip in the DELETE at the end. The name of the scripts is what’s being used as a partition key now, as I’ve mentioned, but since scripts generate different amount of data, that’s not a good partition key since the partitions are really different in size. I’m not really sure how CDC might help me in this situation. Also, I don’t have problems with duplicated entries so adding a timeuuid will not help me either. Adding that won’t help with either the SELECTS, nor the DELETES. One more idea: If you do not have many scripts, you can create a table per script and use the run_id as the partition of this table. Once you are done with the script, drop the table. thank you! this sounds good, I think I might try this but I have one more question, will multiple tables be a in any way problematic? cus I’ll eventually need to run a couple of hundreds of scripts and I’m not really sure if this will badly affect the db You are right; the DB was designed for only a few tables, but a couple hundred should work. --- ### Page: https://forum.scylladb.com/t/release-scylladb-6-1-rc1/2372 Title: [RELEASE] ScyllaDB 6.1 RC1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce ScyllaDB Open Source 6.1 RC1, the first Release Candidate for the ScyllaDB Open Source 6.1 minor release. ScyllaDB 6.1 includes many improvements in functionality, stability, UX,… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-6-1-rc1/2372 ## Headings Structure: H1: [RELEASE] ScyllaDB 6.1 RC1 H2: Related Links H3: Alternator H3: Stability H3: Bloom Filter H3: Tools H3: Tracing H3: Monitoring H3: Performance H3: Tablets H3: Configuration H3: Build H3: Deprecated and removed features H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 6.1 RC1 H2: Related Links H3: Alternator H3: Stability H3: Bloom Filter H3: Tools H3: Tracing H3: Monitoring H3: Performance H3: Tablets H3: Configuration H3: Build H3: Deprecated and removed features H3: Related topics The ScyllaDB team is pleased to announce ScyllaDB Open Source 6.1 RC1, the first Release Candidate for the ScyllaDB Open Source 6.1 minor release. ScyllaDB 6.1 includes many improvements in functionality, stability, UX, and performance, particularly for the new Tablet features introduced in 6.0. Only the latest two minor releases of the ScyllaDB Open Source project are supported. Once 6.1 is released, only ScyllaDB Open Source 6.0 and 6.1 are supported. Get ScyllaDB Open Source 6.1 as binary packages (RPM/DEB), AWS AMI, GCP Image and Docker image --- ### Page: https://forum.scylladb.com/t/error-when-deploying-scylladb-to-aws-eks-scylla-operator-webhook-scylla-operator-svc-context-deadline-exceeded/2375 Title: Error when deploying scylladb to AWS EKS scylla-operator-webhook.scylla-operator.svc context deadline exceeded - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, I’ve tried to deploy Scylladb into AWS EKS Kubernetes with yaml manifests and helm charts, but I’m encountering an error Internal error occurred: failed calling webhook "webhook.scylla.scylladb.com": failed to call … Language: en Canonical URL: https://forum.scylladb.com/t/error-when-deploying-scylladb-to-aws-eks-scylla-operator-webhook-scylla-operator-svc-context-deadline-exceeded/2375 ## Headings Structure: H1: Error when deploying scylladb to AWS EKS scylla-operator-webhook.scylla-operator.svc context deadline exceeded H3: Related topics ## Main Content: H1: Error when deploying scylladb to AWS EKS scylla-operator-webhook.scylla-operator.svc context deadline exceeded H3: Related topics Hi, I’ve tried to deploy Scylladb into AWS EKS Kubernetes with yaml manifests and helm charts, but I’m encountering an error Internal error occurred: failed calling webhook "webhook.scylla.scylladb.com": failed to call webhook: Post "https://scylla-operator-webhook.scylla-operator.svc:443/validate?timeout=10s": context deadline exceeded Deployment process with manifests: I’ve seen in couple threads that firewall port 9443 should be open from cluster to the nodes. I’ve checked the rules with this process: Cluster security group sg-06a93cc5cd09cf20a has access to the nodes via port 9443, so I think the firewall rule is fine. Any ideas what I should check? Webhook Service listens on port 443. You have to make sure traffic on 443 port between Kubernetes master nodes and nodes where Operator Webhook Pods are running is allowed Shouldn’t this rule already allow port 443 from the cluster API (= master nodes?) to the nodes? Traffic is from random port to 443 port, maybe rule has to specify that instead of FromPort:443. Thanks for advice so far! The output of AWS CLI is bit confusing. The output says FromPort:443 and it means that the security group will let port 443 through. This screenshot shows the same information as the AWS CLI returns. Port 443 is open on node security group and allowed source is cluster security group. Would there be some commands that I could run to test connections between pods or something? Hey @iisti, did you find a solution? I have the same issue as you. Port 443 is open, and I actually have other webhooks configured on this cluster, like the one for the Prometheus Operator, and it’s working fine. I’m out of ideas. Thanks! Sorry I never got it working, and I stopped working on the project. I also could configure other webhooks, but didn’t get scylladb working. Alright, thanks. There must be something I’m missing, but I can’t figure out what. Maybe someone who got it working will comment on this post! Check if webhook is up and running by checking logs from two webhook pods. One of them should complain about not being able to become a leader, and second one should mostly be silent (leader one). If they are up and running, then you have to make sure kube-apiserver has access to those pods. They are usually running on worker clusters, so you have to make sure that master nodes where kube-apiserver is running has access to worker clusters on port 443. You may also temporarily allow entire traffic to check if firewall rules are indeed an issue. Okay, allowing all ports on the nodes with the Kubernetes API as the source makes it work. However, all ports need to be open on the nodes — not just port 443 — and I don’t understand how this webhook is different from the one for ingress-nginx or the Prometheus operator, which both work and also listen on port 443. I think I’ve figured it out: port 5000 also needs to be open. I suppose the service is only used to provide the IP and port of a pod, and the API eventually connects directly to port 5000 of the webhook. Here’s the Terraform block to add if you’re using the Terraform EKS module — and the operator will work! --- ### Page: https://forum.scylladb.com/t/facing-issue-in-exploding-json-for-kafka-scylla-sink-connector/2376 Title: Facing issue in Exploding JSON for kafka-scylla sink connector - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hey everyone, Can anyone help me with some documentation for kafka-scylla-sink connector for some issues which I am facing. How can I configure these things in connector config json file. { "name": "user-deatils-conn… Language: en Canonical URL: https://forum.scylladb.com/t/facing-issue-in-exploding-json-for-kafka-scylla-sink-connector/2376 ## Headings Structure: H1: Facing issue in Exploding JSON for kafka-scylla sink connector H3: Related topics ## Main Content: H1: Facing issue in Exploding JSON for kafka-scylla sink connector H3: Related topics Hey everyone, Can anyone help me with some documentation for kafka-scylla-sink connector for some issues which I am facing. How can I configure these things in connector config json file. Issues: In my kafka Json message I am getting some field as eventName while the table has column name as event_name . In my kafka message, I am getting nested JSON, how can I map these to column names. Example: Now I have to insert data into the table such that for each element in data there should be a new unique row in table. The primary key for row is dealId, userId. For example, new rows will be like: Can anyone help with there expertise/documentation. Thanks in Advance for your valuable time . Sorry for not replying for too long. Can I ask you to create an issue on GitHub - scylladb/kafka-connect-scylladb: Kafka Connect Scylladb Sink and attache table scheme, including cdc configuration --- ### Page: https://forum.scylladb.com/t/how-can-i-determine-if-my-data-will-be-available-on-node-failure-depending-on-the-replication-factor-consistency-level-etc/2378 Title: How can I determine if my data will be available on node failure depending on the Replication Factor, Consistency Level, etc.? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Is there a way to know in advance if my data will be available in case of node failure, with different values of Replication Factor (RF), Cluster Size, Consistency Level (CL), and so on? Language: en Canonical URL: https://forum.scylladb.com/t/how-can-i-determine-if-my-data-will-be-available-on-node-failure-depending-on-the-replication-factor-consistency-level-etc/2378 ## Headings Structure: H1: How can I determine if my data will be available on node failure depending on the Replication Factor, Consistency Level, etc.? H3: Related topics ## Main Content: H1: How can I determine if my data will be available on node failure depending on the Replication Factor, Consistency Level, etc.? H3: Related topics Is there a way to know in advance if my data will be available in case of node failure, with different values of Replication Factor (RF), Cluster Size, Consistency Level (CL), and so on? The short answer is that you can use the Consistency Calculator. To elaborate, high-availability (HA) databases, such as ScyllaDB, aim for continuous operation, allowing applications to remain “always-on” despite failures by leveraging architectural designs to ensure data availability. ScyllaDB achieves high availability by eliminating single points of failure, implementing failover mechanisms, and being topology-aware. Being topology-aware means that even if an entire rack or datacenter fails, there is no downtime. Data is automatically replicated across multiple nodes. For example, a Replication Factor (RF) of three (RF=3) means that each piece of data is replicated to three different nodes. Different tradeoffs apply between performance, consistency, availability, and partition tolerance. Also, see the related PACELC theorem. The above calculator will help you determine the appropriate values for your specific use case. NoSQL High Availability explanation High Availability lesson on ScyllaDB University and blog post --- ### Page: https://forum.scylladb.com/t/what-is-eventual-consistency-and-how-is-it-different-from-strong-consistency/2379 Title: What is Eventual Consistency, and how is it different from Strong Consistency? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Can you please provide information on Eventual Consistency and how exactly it differs from Strong Consistency? Language: en Canonical URL: https://forum.scylladb.com/t/what-is-eventual-consistency-and-how-is-it-different-from-strong-consistency/2379 ## Headings Structure: H1: What is Eventual Consistency, and how is it different from Strong Consistency? H3: Related topics ## Main Content: H1: What is Eventual Consistency, and how is it different from Strong Consistency? H3: Related topics Can you please provide information on Eventual Consistency and how exactly it differs from Strong Consistency? It’s a big topic, and I’ll try to answer it and give an overview. Check the links below to learn more. To understand the difference, I’ll start with defining consistency. In Database Management Systems, consistency (sometimes also called correctness) means that after a successful write (or update or delete) request of a value, any read request receives the latest value. Another way to understand it is that any given database transaction can only change the affected data in allowed ways. Any written data has to be valid according to the defined rules, constraints, and triggers. Consistency is one of the guarantees defined in relational, transactional databases. These provide ‘ACID guarantees’ and are sometimes called strongly consistent databases. There are ambiguities in the definition of these guarantees. One definition for ACID transactions is: ACID compliance is a complex and often contested topic. Delivering the consistency guarantee is incredibly difficult in a globally distributed database topology involving multiple clusters, each containing many nodes. For this reason, ACID-compliant databases are usually slower, more rigid, and difficult to scale. Since SQL databases are all ACID compliant to varying degrees, they also share these downsides. Some relational database systems enable ACID guarantees to be relaxed to offset these downsides. In contrast to SQL’s ACID guarantees, NoSQL databases provide BASE guarantees: The above is also related to the CAP theorem, which states that in a distributed data system, only two out of the following three guarantees can be satisfied: An extension to the CAP theorem, originally coined by Dr. Daniel Abadi, is the PACELC theorem. This theorem states that in case of partitioning (P) in a distributed system, a choice has to be made between Availability (A) and Consistency (C). ScyllaDB, Apache Cassandra, Amazon DynamoDB, and other NoSQL databases sacrifice a degree of consistency to increase availability. Rather than providing strong consistency, they provide eventual consistency. This means that in some cases, a read request will fail to return the result of the latest WRITE. In ScyllaDB (and Apache Cassandra), consistency is tunable. For a given query, the client can specify a Consistency Level, which refers to the number of replicas required to respond to a request for it to be considered successful. This is also known as Tunable Consistency. A related feature is Lightweight Transactions (LWT). LWT in ScyllaDB allow the client to modify data based on its current state: that is, to perform an update that is executed only if a row does not exist or contains a certain value. LWT are limited to a single conditional statement, which allows an “atomic compare and set” operation. That is, it checks if a condition is true, and if so, it conducts the transaction. If the condition is not met, the transaction does not go through. --- ### Page: https://forum.scylladb.com/t/release-added-support-for-aws-ap-south-2-region-1-august-2024/2380 Title: [RELEASE] Added support for AWS ap-south-2 region - 1 August 2024 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: We are happy to announce a new update to the ScyllaDB Cloud service. New features and enhancements: We’ve added support for the AWS ap-south-2 (Hyderabad) region. Language: en Canonical URL: https://forum.scylladb.com/t/release-added-support-for-aws-ap-south-2-region-1-august-2024/2380 ## Headings Structure: H1: [RELEASE] Added support for AWS ap-south-2 region - 1 August 2024 H3: Related topics ## Main Content: H1: [RELEASE] Added support for AWS ap-south-2 region - 1 August 2024 H3: Related topics We are happy to announce a new update to the ScyllaDB Cloud service. New features and enhancements: --- ### Page: https://forum.scylladb.com/t/unable-to-explode-json-array-for-kafka-scylla-sink-connector/2381 Title: Unable to explode JSON Array for kafka-scylla sink connector - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hey everyone, Can anyone help me with some documentation for kafka-scylla-sink connector for some issues which I am facing. How can I configure these things in connector config json file. { "name": "user-deatils-con… Language: en Canonical URL: https://forum.scylladb.com/t/unable-to-explode-json-array-for-kafka-scylla-sink-connector/2381 ## Headings Structure: H1: Unable to explode JSON Array for kafka-scylla sink connector H3: Related topics ## Main Content: H1: Unable to explode JSON Array for kafka-scylla sink connector H3: Related topics Can anyone help me with some documentation for kafka-scylla-sink connector for some issues which I am facing. How can I configure these things in connector config json file. Issues: In some kafka topic I am getting Array of JSON for a particular key. Like this: Example: Now I have to insert data into the table such that for each element in data there should be a new unique row in table. The primary key for row is dealId, userId. For example, new rows will be like: Can anyone help with there expertise/documentation. Thanks in Advance for your valuable time . The Kafka-Scylla sink connector does not natively support exploding JSON arrays in a Kafka message into multiple rows in a Scylla table. To write each element of a JSON array as a separate row (with unique keys like dealId and userId), you typically need a preprocessing step before the data reaches the sink connector. Common approaches include: Use Kafka Streams or ksqlDB to consume the original topic, explode the array (using functions like EXPLODE), and produce a new Kafka topic with one message per array element. The sink connector then writes these flattened messages directly to Scylla. If not using ksqlDB or Kafka Streams, implement a custom Kafka Connect Single Message Transform (SMT) or a processor that flattens/nests the JSON arrays before the sink connector inserts into Scylla. Keep your sink connector configuration focused on single-level JSON objects that match directly to your Scylla table schema. The key point is the Kafka-Scylla sink connector itself expects one message per row, so flattening must happen upstream in your Kafka topic pipeline to insert JSON array elements as individual rows with composite keys (dealId, userId). --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-58-2024-08-02/2382 Title: Last week in scylla-cluster-tests.git master (issue #58; 2024-08-02) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the d1e2ac5b…3eefb3c6 range are covered. There were 21 non-merge commits from 5 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-58-2024-08-02/2382 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #58; 2024-08-02) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #58; 2024-08-02) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the d1e2ac5b…3eefb3c6 range are covered. There were 21 non-merge commits from 5 authors in that period. Some notable commits: The investigate show-jepsen-results issue was fixed by creating a new Docker image due to the original image format’s deprecation. We’ve resolved the issue with running Python code outside of the Hydra image, so we no longer make assumptions about the user’s environment. New pipelines have been introduced for LWT/CDC with tablets disabled. The test_method automatic parameter was added to the test config, making it easier to understand which test method was used. This is especially useful for jobs utilizing various test methods (e.g., in performance tests) within the same Jenkins job. It is also visible in Argus’s ‘Run Details’ section. The Argus Client has been updated to 0.12.5, including support for the generic-results feature. This allows defining custom results data in SCT (and other test tools) and sending them to Argus, where they can be displayed in tabular form in the test run’s ‘Results’ tab. As a first step, we have started sending latencies, durations, and stalls events counts as part of performance with nemesis tests. More results will be sent to Argus in the coming weeks, and additional features for presenting them will be added. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/restore-backup-with-expired-ttl/2385 Title: Restore backup with expired TTL - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I have a backup, but TTL for records in this backup already expired. If I’ll try to restore backup using scylla manager, it will download sstables and run load&stream, but that means during compaction all rows will be de… Language: en Canonical URL: https://forum.scylladb.com/t/restore-backup-with-expired-ttl/2385 ## Headings Structure: H1: Restore backup with expired TTL H3: Related topics ## Main Content: H1: Restore backup with expired TTL H3: Related topics I have a backup, but TTL for records in this backup already expired. If I’ll try to restore backup using scylla manager, it will download sstables and run load&stream, but that means during compaction all rows will be deleted because TTL for this records is expired. setting TTL = 0 on the table level will not help, because if the table have TTL set the got written to it with an expiration date set, and nothing can remove it. I have 2 ideas in mind are any other options available? sstableloader is not an option because if tables have uuid_sstable_identifiers_enabled: true Hi, speaking from ScyllaDB Manager POV, I believe that the trick with the clock is the only way. There are ways to ensure that the data with expired TTL remains on the disk (by disabling table’s tombstone_gc and compaction options), but you still need to somehow force ScyllaDB to return this data despite the expired TTL, and I don’t think that’s possible (you can’t use USING TIMESTAMP on select queries). --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-241-2024-08-04/2389 Title: Last week in scylladb.git master (issue #241; 2024-08-04) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 27b305b9d1…39b49a41cc range are covered. There were 61 non-merge commits from 10 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-241-2024-08-04/2389 ## Headings Structure: H1: Last week in scylladb.git master (issue #241; 2024-08-04) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #241; 2024-08-04) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 27b305b9d1…39b49a41cc range are covered. There were 61 non-merge commits from 10 authors in that period. Some notable commits: The tablet load balancer now tries to ensure that not only are tablet distributed evenly among nodes and starts, but that tablets for any particular table are evenly distributed. This prevents a hot table that is unevenly distributed from causing hot nodes or hot shards. When a materialized view’s primary key has the same columns as the base table primary key, we now optimize deletions by deleting an an entire partition when possible. A race between table drop and a counter column update was fixed. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2023-1-11/2396 Title: [RELEASE] ScyllaDB Enterprise 2023.1.11 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2023.1.11, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2023.1 LTS Release. 2023.1.11 patch release includes multiple minor bug fix… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2023-1-11/2396 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2023.1.11 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2023.1.11 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2023.1.11, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2023.1 LTS Release. 2023.1.11 patch release includes multiple minor bug fixes. ScyllaDB Enterprise’s latest Long Term Support (LTS) is 2024.1. While we continue to support 2023.1, you are encouraged to upgrade to it in coordination with the ScyllaDB Support team. The following issues are fixed in this release (with an open-source reference, if available): --- ### Page: https://forum.scylladb.com/t/scylladb-community-image-on-azure/2400 Title: ScyllaDB community image on Azure - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello. I have attended the ScyllaDB getting started webinar on 07/17 (or 07/18?). I attempted to create a scyllaDB cluster on Azure. I have observed trouble with the ScyllaDB community images on Azure. As I have rememb… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-community-image-on-azure/2400 ## Headings Structure: H1: ScyllaDB community image on Azure H3: Related topics ## Main Content: H1: ScyllaDB community image on Azure H3: Related topics Hello. I have attended the ScyllaDB getting started webinar on 07/17 (or 07/18?). I attempted to create a scyllaDB cluster on Azure. I have observed trouble with the ScyllaDB community images on Azure. As I have remembered from Felipe’s presentation, scylla db offers the Azure community image on select regions in Azure cloud. Currently, the azure team is limiting the LSV2 disk sizes on the us east region, so they have given me allocations on the us east 2 region. When I attempted to run the script from documentation (Launch ScyllaDB on Azure | ScyllaDB Docs) , azure renders an error. Azure cli says there is no scyllaDB images on us east 2 region. I have included a screenshot of the error. Felipe mentioned that the Azure community image was available in select Azure regions. I forgot the region names. I perused the documentation, and the regions were not mentioned in the documentation. I emailed Guy Schtub, and he told me to post the question on the forum. I have launched the scyllaDB installation using the scyllaDB web installer. The correct region should be eastus - as mention in ScyllaDB | Get Started with ScyllaDB Please let me know if you have any other questions You are correct that the azure community image lives on the east-us region. However, if Azure’s east us region becomes saturated, does the scyllaDB team anticipate launching the azure community image on other Azure regions? I hope the community image could be launched on additional Azure regions. --- ### Page: https://forum.scylladb.com/t/probabilistic-tracing-with-scylladb/2401 Title: Probabilistic Tracing with ScyllaDB - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I’ve been exploring different debugging techniques to improve the performance and reliability of my applications, specifically those using ScyllaDB. Recently, I came across the concept of Probabilistic Tracing, but I’m … Language: en Canonical URL: https://forum.scylladb.com/t/probabilistic-tracing-with-scylladb/2401 ## Headings Structure: H1: Probabilistic Tracing with ScyllaDB H3: Related topics ## Main Content: H1: Probabilistic Tracing with ScyllaDB H3: Related topics I’ve been exploring different debugging techniques to improve the performance and reliability of my applications, specifically those using ScyllaDB. Recently, I came across the concept of Probabilistic Tracing, but I’m not entirely clear on how to effectively implement and utilize it. Could someone explain how Probabilistic Tracing can be used for debugging purposes? Additionally, I would appreciate some concrete examples demonstrating its application in real-world scenarios. Tracing is a tool to debug and analyze internal flows in the cluster. useful for observing behaviors of specific queries and troubleshooting network issues, data transfer, and data replication problems. There are different types of tracing that you can use: The tracing documentation section includes instructions and examples for debugging common issues (e.g., Hot Partition, large partition, slow queries). Tracing rows are inserted in the system_traces.sessions and system_traces.events tables when probabilistic tracing is enabled. In order to minimize the impact of this procedure on cluster performance, the following must be considered when enabling it: It is important to remember that this is a node-level analysis, so if you want to get info for the entire cluster, the command must run for all cluster nodes. The following steps are an example of collecting information for 5 minutes and then disabling the tracing (value 0 stops collecting information): Connect to the cluster via ssh For each node, open tmux (Getting Started · tmux/tmux Wiki · GitHub) to run the command listed below. nodetool settraceprobability 0.001; sleep 5m; nodetool settraceprobability 0 SELECT*FROM system_traces.sessions SELECT*FROM system_traces.events The commands below expect a CSV file with a “;” delimiter and the parameters in the 6th column. That can be verified with: session_id;client;command;coordinator;duration;parameters;request;request_size;response_size;started_at;username awk 'BEGIN{FS=";"} {print $6}' sessions.csv | sed "s/.*'query': '//" | sed 's/\\n//g' | grep -i 'INSERT INTO' | sed 's/.*INSERT INTO//i' | sed 's/(.*//i' | sort | uniq -c | sort -nr awk 'BEGIN{FS=";"} {print $6}' sessions.csv | sed "s/.*'query': '//" | sed 's/\\n//g' | grep -i 'select' | sed 's/.*FROM//i' | sed 's/.WHERE.*//i' | sort | uniq -c | sort -nr awk 'BEGIN{FS=";"} {print $6}' sessions.csv | grep -i 'SELECT' | sed "s/.*'query': '//" | sed 's/\\n//g | sort | uniq > queries.txt awk 'BEGIN{FS=";"} {print $6}' sessions.csv | grep -i 'SELECT' | sed "s/.*'query': '//" | sed 's/\\n//g' | sort | uniq -c | sort -nr > queries.txt --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-59-2024-08-09/2409 Title: Last week in scylla-cluster-tests.git master (issue #59; 2024-08-09) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 8bd7d9a0…23ccb24e range are covered. There were 19 non-merge commits from 7 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-59-2024-08-09/2409 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #59; 2024-08-09) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #59; 2024-08-09) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 8bd7d9a0…23ccb24e range are covered. There were 19 non-merge commits from 7 authors in that period. Some notable commits: We added configuration and pipeline files for artifacts tests for Debian 12. We started to collect cloud-init logs on DB nodes to speed up investigations of cloud-init issues. Note that the docker_backend_local.yaml config file was renamed to docker_backend.yaml, as it’s now used not only for local tests but also in the CI. Keep this in mind when running tests locally. Developers can now easily update Scylla DB packages on a running cluster without needing to recreate it. New packages can be uploaded from S3, GS, or a local workstation, and tests can be rerun using the reuse-cluster capability. Run hydra update-scylla-packages --help for more details on this feature. Upgrade tests now execute the Raft topology upgrade procedure if the new Scylla version supports it. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/upgrade-scylla-version-from-2023-1-to-2024-1/2415 Title: Upgrade Scylla Version from 2023.1 to 2024.1 - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: So i am trying to update from scylla enterprise 2023.1 to 2024.1. I followed the below steps as mentioned in official scylla document nodetool describecluster (Check that the cluster’s schema is synchronized) nodetool … Language: en Canonical URL: https://forum.scylladb.com/t/upgrade-scylla-version-from-2023-1-to-2024-1/2415 ## Headings Structure: H1: Upgrade Scylla Version from 2023.1 to 2024.1 H3: Related topics ## Main Content: H1: Upgrade Scylla Version from 2023.1 to 2024.1 H3: Related topics So i am trying to update from scylla enterprise 2023.1 to 2024.1. I followed the below steps as mentioned in official scylla document Now when i run the command sudo add-apt-repository -y ppa:scylladb/ppa i am getting the below issue in the screenshotNote: i am using ubuntu 20.04 OS Also there are some additional question Hey Pushpak, are you an Enterprise customer? --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-242-2024-08-11/2448 Title: Last week in scylladb.git master (issue #242; 2024-08-11) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 39b49a41cc…e18a855abe range are covered. There were 101 non-merge commits from 19 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-242-2024-08-11/2448 ## Headings Structure: H1: Last week in scylladb.git master (issue #242; 2024-08-11) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #242; 2024-08-11) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 39b49a41cc…e18a855abe range are covered. There were 101 non-merge commits from 19 authors in that period. Some notable commits: Alternator, ScyllaDB’s implementation of the DynamoDB API, uses JSON to communicate with the client. Previously, sending very large JSON values was adjusted to avoid stalls. The destruction of these large JSON values now also avoids stalls. The tablet allocator will now refrain from allocating tablets on a table created concurrently with decommission. We now collect cell statistics in addition to row and tombstone statistics for result pages. When service level parameters are modified, connections are adjusted in real time. Hinted handoff writes to local storage now use the commitlog scheduling group. The driver for S3 access is now optimized for throughput. The internal ‘cluster feature’ mechanism now supports suppressing features, enabling simulation of upgrades. This should catch version upgrade problems earlier. A race between tablet split and migration has been fixed. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/hi-team-i-wanted-to-upgrading-the-scylla-5-0-to-5-1in-aws-by-using-ami/2458 Title: Hi Team, I wanted to Upgrading the scylla 5.0 to 5.1in AWS by using AMI - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi All, My Scylla DB runs with a single DC OS version of Ubuntu “22.04.3 LTS” Scylla version “5.0.1-0.20220719.b177dacd3”, We went to upgrade the 5.1 AMI in AWS. If it is possible or not by using AMI. Language: en Canonical URL: https://forum.scylladb.com/t/hi-team-i-wanted-to-upgrading-the-scylla-5-0-to-5-1in-aws-by-using-ami/2458 ## Headings Structure: H1: Hi Team, I wanted to Upgrading the scylla 5.0 to 5.1in AWS by using AMI H3: Upgrade ScyllaDB Image: EC2 AMI, GCP, and Azure Images | ScyllaDB Docs H3: Related topics ## Main Content: H1: Hi Team, I wanted to Upgrading the scylla 5.0 to 5.1in AWS by using AMI H3: Upgrade ScyllaDB Image: EC2 AMI, GCP, and Azure Images | ScyllaDB Docs H3: Related topics Hi All, My Scylla DB runs with a single DC OS version of Ubuntu “22.04.3 LTS” Scylla version “5.0.1-0.20220719.b177dacd3”, We went to upgrade the 5.1 AMI in AWS. If it is possible or not by using AMI. Yes, you can upgrade Scylla AMI from 5.0 to 5.1. See ScyllaDB is an Apache Cassandra-compatible NoSQL data store that can handle 1 million transactions per second on a single server. Note you are using very old, unsupported versions of Scylla. Best will be to upgrade from 5.0, to 5.1 and then 5.1->5.2->5.3->5.4-> 6.0 → 6.1 --- ### Page: https://forum.scylladb.com/t/raft-server-id-cannot-be-translated-to-an-ip-address/2462 Title: Raft server id ... cannot be translated to an IP address - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, yesterday I updated a cluster to Scylla 6.0.2. I am seeing quite a lot of these errors: Aug 12 07:12:47 osdev-6 scylla[3776943]: [shard 2:main] raft_group_registry - (rate limiting dropped 2997 similar messages)… Language: en Canonical URL: https://forum.scylladb.com/t/raft-server-id-cannot-be-translated-to-an-ip-address/2462 ## Headings Structure: H1: Raft server id ... cannot be translated to an IP address H3: Related topics ## Main Content: H1: Raft server id ... cannot be translated to an IP address H3: Related topics yesterday I updated a cluster to Scylla 6.0.2. I am seeing quite a lot of these errors: They appear in bursts. It even seems like queries are being affected by this as I had my application hang while this was happening, but I need more investigation to be sure about this. Is this normal or might there be a bug? Is there really no way to return a cluster to gossip? For my understanding: If Raft is down, are only schema/topology changes down, or also read/writes? As per https://www.scylladb.com/2023/05/04/scylladbs-path-to-strong-consistency-a-new-milestone/:slight_smile : " If a node is partitioned away from the cluster, it can’t perform schema changes. That’s the main difference, or limitation, from the pre-Raft clusters that you should keep in mind. You can still perform other operations with such nodes (such as reads and writes) so data availability is unaffected." I am relieved to read this. There are still some things to test, e.g.: What if a DC is nuked and never comes back. Will I be able to forcefully remove the nodes from the nuked DC? Otherwise there is no way to ever recover from this. Would be nice if there was some kind of force option. Hi, if you still have quorum I believe you can remove nodes normally, for lost quorum you can do recovery procedure: Handling Node Failures | ScyllaDB Docs --- ### Page: https://forum.scylladb.com/t/release-scylladb-6-1/2486 Title: [RELEASE] ScyllaDB 6.1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce ScyllaDB Open Source 6.1.0, a production-ready minor release. ScyllaDB 6.1 includes many improvements in functionality, stability, UX and performance, in particular for the new T… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-6-1/2486 ## Headings Structure: H1: [RELEASE] ScyllaDB 6.1 H2: Related Links H3: Alternator H3: Stability H3: Bloom Filter H3: Tools H3: Tracing H3: Monitoring H3: Performance H3: Tablets H3: Configuration H3: Build H3: Deprecated and removed features H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 6.1 H2: Related Links H3: Alternator H3: Stability H3: Bloom Filter H3: Tools H3: Tracing H3: Monitoring H3: Performance H3: Tablets H3: Configuration H3: Build H3: Deprecated and removed features H3: Related topics The ScyllaDB team is pleased to announce ScyllaDB Open Source 6.1.0, a production-ready minor release. ScyllaDB 6.1 includes many improvements in functionality, stability, UX and performance, in particular for the new Tablets features, introduced in ScyllaDB 6.0. Only the latest two minor releases of the ScyllaDB Open Source project are supported. With this release, only ScyllaDB Open Source 6.0 and 6.1 are supported. Users running earlier releases are encouraged to upgrade to one of these two releases. Get ScyllaDB Open Source 6.1 as binary packages (RPM/DEB), AWS AMI, GCP Image and Docker image Upgrade from ScyllaDB 6.0 to ScyllaDB 6.1 Scylla Monitoring Stack released 4.8 and later supports ScyllaDB 6.1. See metrics update between 6.0 and 6.1 here, as well as the new, beta, metrics reference here. More monitoring related updates: There are now metrics keeping track of incoming hints, in addition to the existing metrics for outgoing hints. #10987 A regression in Lightweight Transaction (LWT) contention metric has been fixed. The regression would shop contentions increasing even when none were happening. While it’s just a metric, it’s one of the more important ones for LWT users. #19625 --- ### Page: https://forum.scylladb.com/t/nodes-become-unusable-when-split-compactions-are-going-on/2489 Title: Nodes become unusable when SPLIT compactions are going on - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello, I recently started migrating my production cluster from Cassandra 4 to Scylla 6.0.2. After having had a great success with our dev cluster, we decided to go to prod. I switched the stack around 10 days ago, and … Language: en Canonical URL: https://forum.scylladb.com/t/nodes-become-unusable-when-split-compactions-are-going-on/2489 ## Headings Structure: H1: Nodes become unusable when SPLIT compactions are going on H3: Related topics ## Main Content: H1: Nodes become unusable when SPLIT compactions are going on H3: Related topics I recently started migrating my production cluster from Cassandra 4 to Scylla 6.0.2. After having had a great success with our dev cluster, we decided to go to prod. I switched the stack around 10 days ago, and started copying data over. I’m currently facing two problems: My keyspace has tablets enabled. My production workload is currently writing around 60k writes/sec at peak hours, and around 10k at off-peak hours. My copy program is able to copy at around 200k rows/s for the current table (speed varies from table to table). I have around 6TB of data to be copied. Machines are aws m6a.2xlarge (8c, 32GB ram) with 4 TB of EBS storage, 12k provisioned IOPS, 1000 MiB/s provisioned throughput. Cluster currently have 8 nodes. I have tried a bunch of things to fix this problem, but none of them seem to fix it. I’ve tried: During data copy, I will observe that the metrics “Writes currently blocked on commitlog” and “Writes blocked on commitlog” will go high, and I’ll have to reduce the copy throuput until I don’t see anything on these graphs anymore. When the split compactions fire and the cluster gets unusable, I’ll only see “Writes timed out” and “Writes failed”. EBS graphs don’t show usage above 600 IOPS at any time, despite being provisioned for 12000. Bandwidth also won’t go above 40MB/s, despite being provisioned for 1000. CPU does not appear to be maxed. What can I do to diagnose the problem with the compactions? the underwhelming throughput on the copy is bad, but tolerable since it’s a one off thing, but these SPLIT compactions keep bringing my production pipeline down for many hours at a time and I don’t know what to do. One node just restarted while doing a SPLIT compaction. this was in the logs: @matheus2740 do you have monitoring data at the time of the incident? logs can also help. You can upload the requested data using these instructions: How to Report a ScyllaDB Problem | ScyllaDB Docs Split happens in the same I/O class as maintenance ops like repair / streaming. So you can try tuning stream_io_throughput_mb_per_sec. @raphaelsc I got the same recommendation about throttling streaming in slack some weeks ago, and it seems to have mitigated the issue. it isn’t a fix, but at least the cluster stopped crashing. Monitoring data can be found at the github ticket: Cluster becomes unusable when SPLIT compactions are happening · Issue #20211 · scylladb/scylladb · GitHub Unfortunately I don’t have more than that since our log/metric retention is not that long and it’s been a while. --- ### Page: https://forum.scylladb.com/t/auth-doesnt-work-after-removing-nodes/2492 Title: Auth doesn't work after removing nodes - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I recently removed some nodes from my staging cluster (3>1) and may have missed a step somewhere. Now auth won’t work. If i set the authenticator/authorizer to the all versions, I can get in but if I set it to Password… Language: en Canonical URL: https://forum.scylladb.com/t/auth-doesnt-work-after-removing-nodes/2492 ## Headings Structure: H1: Auth doesn't work after removing nodes H3: Related topics ## Main Content: H1: Auth doesn't work after removing nodes H3: Related topics I recently removed some nodes from my staging cluster (3>1) and may have missed a step somewhere. Now auth won’t work. If i set the authenticator/authorizer to the all versions, I can get in but if I set it to PasswordAuthentication, I can’t log in. I’ve tried following this procedure but it doesn’t work. on the opensource docs site /branch-5.2/troubleshooting/password-reset.html “I’ve tried following this procedure but it doesn’t work.” Please share as many details/logs/errors as possible so that others can help you. --- ### Page: https://forum.scylladb.com/t/scylladb-query-issue/2495 Title: Scylladb query issue - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I encountered difficulties while fetching data, resulting in inconsistent information. Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-query-issue/2495 ## Headings Structure: H1: Scylladb query issue H3: Related topics ## Main Content: H1: Scylladb query issue H3: Related topics I encountered difficulties while fetching data, resulting in inconsistent information. --- ### Page: https://forum.scylladb.com/t/testing-scylladb-version-upgrade-upgrade-path-and-old-versions-download/2501 Title: Testing ScyllaDB version upgrade, upgrade path, and old versions download - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: where can I find older versions of scylla? specifically 5.2.2. I’m trying to test the upgrade process to 6.1 and i’m currently running 5.2.2 Alternatively, maybe you could offer a better suggestion for testing the upgr… Language: en Canonical URL: https://forum.scylladb.com/t/testing-scylladb-version-upgrade-upgrade-path-and-old-versions-download/2501 ## Headings Structure: H1: Testing ScyllaDB version upgrade, upgrade path, and old versions download H3: Related topics ## Main Content: H1: Testing ScyllaDB version upgrade, upgrade path, and old versions download H3: Related topics where can I find older versions of scylla? specifically 5.2.2. I’m trying to test the upgrade process to 6.1 and i’m currently running 5.2.2 Alternatively, maybe you could offer a better suggestion for testing the upgrade. It’s recommended that you upgrade sequentially, going through each major release. So, for example: 5.2.2 → 5.4 → 6.0 → 6.1 This path is the one that’s tested. To find and download older versions, you can look at the relevant version of the documentation: for example here. When you download ScyllaDB you can specify the version (depends on which OS you use), for example with Docker you can see the versions here. --- ### Page: https://forum.scylladb.com/t/bootstrap-failed/2502 Title: Bootstrap failed - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi everyone scylaa version 5.2.13-0.20240103.c57a0a7a46c6 I trying to add host to cluster scylla after host decommission. I start scylla-server and after 2.5 hours it is fails with error. Aug 16 15:28:59 scylla115 scy… Language: en Canonical URL: https://forum.scylladb.com/t/bootstrap-failed/2502 ## Headings Structure: H1: Bootstrap failed H3: Related topics ## Main Content: H1: Bootstrap failed H3: Related topics Hi everyone scylaa version 5.2.13-0.20240103.c57a0a7a46c6 I trying to add host to cluster scylla after host decommission. I start scylla-server and after 2.5 hours it is fails with error. Aug 16 15:28:59 scylla115 scylla[61440]: [shard 0] range_streamer - Finished 925 out of 3572 ranges for bootstrap, finished percentage=0.25895858 Aug 16 15:28:59 scylla115 scylla[61440]: [shard 0] range_streamer - Bootstrap with 10.3.0.111 for keyspace=*** succeeded, took 9275.1 seconds Aug 16 15:28:59 scylla115 scylla[61440]: [shard 0] range_streamer - Bootstrap failed, took 9278 seconds, nr_ranges_remaining=2647 Aug 16 15:28:59 scylla115 scylla[61440]: [shard 0] boot_strapper - Error during bootstrap: streaming::stream_exception (Stream failed) Have everyones any idea? Version 5.2 has reached end-of-life, a first step would be to upgrade to the latest version. --- ### Page: https://forum.scylladb.com/t/query-error-clustering-key-cartesian-product-is-greater-than-maximum-max-clustering-key-restrictions-per-query-and-max-partition-key-restrictions-per-query/2503 Title: Query error, clustering key cartesian product is greater than maximum, max_clustering_key_restrictions_per_query and max_partition_key_restrictions_per_query - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/query-error-clustering-key-cartesian-product-is-greater-than-maximum-max-clustering-key-restrictions-per-query-and-max-partition-key-restrictions-per-query/2503 ## Headings Structure: H1: Query error, clustering key cartesian product is greater than maximum, max_clustering_key_restrictions_per_query and max_partition_key_restrictions_per_query H3: Related topics ## Main Content: H1: Query error, clustering key cartesian product is greater than maximum, max_clustering_key_restrictions_per_query and max_partition_key_restrictions_per_query H3: Related topics Originally from the User Slack @Bhardwaj_Thummar: Hey guys, I am getting “Query 1 ERROR at Line 1: : clustering-key cartesian product size 236 is greater than maximum 100” from scylladb when using below query, my scylladb server version is 5.4.3-0.20240211.cf42ca0c2a65 I have read scylla - Scylladb : clustering key cartesian product size 600 is greater than maximum 100 - Stack Overflow and applied max_clustering_key_restrictions_per_query in scylla.yaml although I could not find the same flag in the scylla.yaml.example. What are your thoughts ? Thanks in Advance !!! select zpid from forsaleproperty where property_type IN ('single-family home','apartment','condo','coops','townhouse','multi-family','mobile/manufactured','lot/land','other') AND state = 'New York' AND city = 'New York' AND neighborhood IN ('','Washington Heights','Inwood','Hamilton Heights','Highbridge','Harlem','Upper East Side','Upper West Side','East Harlem','Sutton Place','Roosevelt Island','Astoria','Battery Park','Financial District','Chelsea','Hell''s Kitchen','Hudson Yards','Flatiron District','Turtle Bay','Kips Bay','Murray Hill','Midtown','Midtown South','Greenwich Village','West Village','SoHo','Gramercy','East Village','Greenpoint','Long Island City','Sunnyside','Lower East Side','Nolita','Tribeca','Civic Center','Little Italy','Chinatown','Brooklyn Heights','DUMBO','Fort Greene','Downtown','Williamsburg','Bedford-Stuyvesant','Bushwick','Clinton Hill','Edenwald','Woodlawn','Wakefield','Riverdale','Van Cortlandt Park','Kingsbridge','University Heights','Fordham','Bedford Park','Norwood','Marble Hill','Morningside Heights','West Harlem','Concourse','Williamsbridge','Laconia','Bronxwood','Pelham Gardens','Baychester','East Tremont','Crotona Park East','Morrisania','Tremont','Belmont','Morris Heights','Soundview','Van Nest','Parkchester','Pelham Parkway','Morris Park','Castle Hill','Westchester Village','Schuyerville','Throggs Neck','City Island','Country Club','Pelham Bay','Eastchester','Pelham Bay Park','Jackson Heights','College Point','Woodstock','Melrose','East Elmhurst','Mott Haven','Hunts Point','Longwood','Whitestone','Flushing','Bayside','Auburndale','Douglaston','Woodside','Elmhurst','Maspeth','Corona','North Corona','Forest Hills','Rego Park','Middle Village','Ridgewood','Glendale','Richmond Hill','Woodhaven','Forest Park','Kew Gardens','Kew Gardens Hills','Fresh Meadows','Oakland Gardens','Queens Village','Jamaica Estates','Briarwood','Jamaica','Jamaica Hills','Hollis','St. Albans','Bellerose','Glen Oaks','Little Neck','Floral park','Cambria Heights','New Dorp Beach','Port Richmond','Westerleigh','New Dorp','Grant City','Mariner''s Harbor','Castleton Corners','West Brighton','Lighthouse Hill','Graniteville','New Springville','Meiers Corners','Oakwood','Arlington','Todt Hill','Greenridge','Silver Lake','Richmond Town','Arden Heights','Bulls Head','Willowbrook','New Brighton','Travis','Dongan Hills','Great Kills','Egbertsville','Elm Park','Bay Terrace','Manor Heights','St. George','Bay Ridge','Sunset Park','Tompkinsville','Dyker Heights','Red Hook','Stapleton','Grymes Hill','Boerum Hill','Park Slope','Greenwood','Gowanus','Columbia Street Waterfront District','Prospect Heights','Cobble Hill','Carroll Gardens','Crown Heights','Wingate','Brownsville','Prospect Lefferts Gardens','Kensington','Windsor Terrace','Borough Park','Midwood','Prospect Park South','East Flatbush','Flatbush','Ditmas Park','South Beach','Bath Beach','Rosebank','Midland Beach','Clifton','Park Hill','Arrochar','Shore Acres','Grasmere','Emerson Hill','Gravesend','Bensonhurst','Marine Park','Flatlands','Sheepshead Bay','Gerritsen Beach','Coney Island','Seagate','Brighton Beach','Manhattan Beach','Tottenville','Eltingville','Prince''s Bay','Huguenot','Rossville','Annadale','Charleston','Pleasant Plains','Woodrow','Richmond Valley','East New York','Ozone Park','Howard Beach','South Ozone Park','South Richmond Hill','Highland Park','Canarsie','Bergen Beach','Brookville','Springfield Gardens','South Jamaica','Laurelton','Old Mill Basin','Mill Basin','Belle Harbor','Neponsit','Rockaway Park','Far Rockaway','Rockaway Beach','Broad Channel','Arverne','Rosedale','Central Park South') limit 1000000; @Bhardwaj_Thummar: The possible solutions i have looked up is • use multiple smaller queries • make better database schema max_clustering_key_restrictions_per_query: 5000 max_partition_key_restrictions_per_query: 5000 set above flags in scylla.yaml, its working though not sure its recommended. @avi: It’s not recommended, but can be used while you work on a better solution @Bhardwaj_Thummar: Thanks for the reply @avi. so restricting the search by 3 partition keys (table has 4 partition keys) and not restrict this 4th key (which has lots of params for IN option) with allow filtering enabled, is better than using below flags maxed beyond recommendation? max_clustering_key_restrictions_per_query: 5000 max_partition_key_restrictions_per_query: 5000 here is the table schema for reference @avi: Better to use individual queries for each partition key combination --- ### Page: https://forum.scylladb.com/t/deleting-materialized-view-and-affect-on-tombstone-data-disk-usage/2504 Title: Deleting Materialized View and affect on tombstone data, disk usage - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/deleting-materialized-view-and-affect-on-tombstone-data-disk-usage/2504 ## Headings Structure: H1: Deleting Materialized View and affect on tombstone data, disk usage H3: Related topics ## Main Content: H1: Deleting Materialized View and affect on tombstone data, disk usage H3: Related topics Originally from the User Slack @Peter_Flockhart: When deleting a Materialized View, what are the variables that affect the timeline of all that tombstoned data? I.e. when can I expect to see a reduction in disk use after deleting a materialized view? and additionally, will nodetool cleanup result in the cleanup of this data - or does that only affect data that is no longer part of a node’s token range? @avi: With DROP MATERIALIZED VIEW the effect should be immediate. You may want to check if any snapshots were created and delete them too. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-243-2024-08-18/2505 Title: Last week in scylladb.git master (issue #243; 2024-08-18) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the e18a855abe…afee3924b3 range are covered. There were 69 non-merge commits from 15 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-243-2024-08-18/2505 ## Headings Structure: H1: Last week in scylladb.git master (issue #243; 2024-08-18) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #243; 2024-08-18) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the e18a855abe…afee3924b3 range are covered. There were 69 non-merge commits from 15 authors in that period. Some notable commits: When tablet metadata (system.tablets table) changes, we now reload only the changed rows. A bug when ALTERing a keyspace that doesn’t exist, with tablets, was fixed. Reversed queries (WITH CLUSTERING ORDER BY) are already quite efficient in ScyllaDB, yet the wire protocol between nodes was kept unaware of reversed queries in order to maintain compatibility; result sets were un-reversed before sending over the wire, then re-reversed. We now support an alternate protocol where these wasteful transformations are avoided. A regression in processing limits for the GROUP BY clause was fixed. Service levels are used to group and classify sessions. Service level names beginning with $ are now reserved. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/alternator-mode-doesnt-start-in-docker/2509 Title: Alternator mode doesn't start in Docker - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I’m trying to run the alternator in Docker with: docker run --name some-scylla --hostname some-scylla -p 8000:8000 -d scylladb/scylla --smp 1 --memory=750M --overprovisioned 1 --alternator-port=8000 and when I watch th… Language: en Canonical URL: https://forum.scylladb.com/t/alternator-mode-doesnt-start-in-docker/2509 ## Headings Structure: H1: Alternator mode doesn't start in Docker H3: Related topics ## Main Content: H1: Alternator mode doesn't start in Docker H3: Related topics I’m trying to run the alternator in Docker with: docker run --name some-scylla --hostname some-scylla -p 8000:8000 -d scylladb/scylla --smp 1 --memory=750M --overprovisioned 1 --alternator-port=8000 and when I watch the log, scylla starts, then shuts down, then restarts, wash, rinse, repeat. What am I missing? I’m running on a M3 MacBook with Sonoma 14.6.1 Can you please add the relevant logs (first error?) docker run --name scylla -p 8000:8000 -d scylladb/scylla:latest --overprovisioned 1 --smp 1 --alternator-port=8000 --alternator-write-isolation=always I guess I needed the --alternator-write-isolation=always --- ### Page: https://forum.scylladb.com/t/release-scylla-manager-3-3-1/2516 Title: [RELEASE] Scylla Manager 3.3.1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces the release of ScyllaDB Manager 3.3.1, a production-ready patch release of the stable 3.3 branch. This release brings performance and stability improvements, and a new way to tag clusters and… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-manager-3-3-1/2516 ## Headings Structure: H1: [RELEASE] Scylla Manager 3.3.1 H3: Related topics ## Main Content: H1: [RELEASE] Scylla Manager 3.3.1 H3: Related topics The ScyllaDB team announces the release of ScyllaDB Manager 3.3.1, a production-ready patch release of the stable 3.3 branch. This release brings performance and stability improvements, and a new way to tag clusters and tasks. ScyllaDB Manager is a centralized cluster administration and recurrent tasks automation tool. ScyllaDB Manager is available for ScyllaDB Enterprise customers and ScyllaDB Open Source users. With ScyllaDB Open Source, ScyllaDB Manager is limited to 5 nodes. See the ScyllaDB Manager Proprietary Software License Agreement for details. In version 3.3.1, we added a dedicated “deduplication” stage to the backup task (#3827). Previously, SSTables in snapshot directory and backup location were compared by their calculated hashes. Now, in order to improve performance, they are compared based on their UUID generation ID or the contents of the crc32 file. Backup task no longer uploads Secondary Indexes and Materialized Views SSTables (#3771). When ScyllaDB Manager has credentials to a managed cluster (see sctool cluster add), it is able to differentiate Secondary Indexes and Materialized Views tables from regular tables and skips their backup. Their SSTables are not needed for the restore purposes, as views and secondary indexes are restored by recreating them on a restored base table. It is now possible to add set of labels to both clusters and tasks (#3219). It can be used to document changes to the task specification, or to track specific groups of tasks. You can specify labels when adding/updating cluster or tasks with the ‘–label’ flag (e.g. ‘sctool repair -c cluster_id --label k1=v1,k2=v2’). Currently set labels can be seen in the output of ‘sctool tasks’, ‘sctool info’, ‘sctool cluster list’ commands. Healthcheck metrics have been enriched with the datacenter and rack info labels (#3908) ScyllaDB Enterprise customers are encouraged to upgrade to ScyllaDB Manager 3.3.1 in coordination with the ScyllaDB support team. The new release includes upgrades of both ScyllaDB Manager Server and Agent. Scylla Manager 3.3.1 supports the following Scylla Enterprise releases: And the following Open Source release (limited to 5 nodes see license 1): You can install and run Scylla Manager on Kubernetes using Scylla Operator. More here. We updated the GPG key of the deb repository together with the release 3.3.1. The following steps are required on client side to make it working correctly: --- ### Page: https://forum.scylladb.com/t/read-performance-and-selecting-the-consistency-level-time-series/2517 Title: Read performance and selecting the consistency level, time series - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/read-performance-and-selecting-the-consistency-level-time-series/2517 ## Headings Structure: H1: Read performance and selecting the consistency level, time series H3: Related topics ## Main Content: H1: Read performance and selecting the consistency level, time series H3: Related topics Originally from the User Slack @Sven: I want to maximize read performance. I have an existing Scylla database with time series. I am streaming the data from it. I am using 1 query SELECT * from table_name and I am letting it run, having set client-side request timeout to 10 years. I am processing page by page. In this scenario, should I use consistency level TWO or THREE in order to distribute reads across replicas? Or maybe I am mistaken, perhaps TWO would mean that two replicas need to answer and agree on the result before the read is completed, which would probably lead to slower reading? @avi: CL=LOCAL_QUORUM is recommended. CL=LOCAL_ONE is faster, but can return stale data. If you can tolerate stale data that’s okay. --- ### Page: https://forum.scylladb.com/t/migrating-from-cassandra-and-dynamodb-using-the-migrator/2519 Title: Migrating from Cassandra and DynamoDB using the Migrator - Blog Posts - ScyllaDB Community NoSQL Forum Meta Description: The new blog post covers the latest changes in the open source Migrator project. The Migrator simplifies the migration process and gives you more control over how it’s performed. The post covers the architecture, sample… Language: en Canonical URL: https://forum.scylladb.com/t/migrating-from-cassandra-and-dynamodb-using-the-migrator/2519 ## Headings Structure: H1: Migrating from Cassandra and DynamoDB using the Migrator H3: Related topics ## Main Content: H1: Migrating from Cassandra and DynamoDB using the Migrator H3: Related topics The new blog post covers the latest changes in the open source Migrator project. The Migrator simplifies the migration process and gives you more control over how it’s performed. The post covers the architecture, sample use cases, and a hands-on tutorial you can run yourself. Do you have any experience with the Migrator you want to share with the community? Do you have any feature requests or questions? --- ### Page: https://forum.scylladb.com/t/error-while-adding-scylladb-as-datasource-on-grafana/2520 Title: Error while adding scylladb as datasource on Grafana - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi Team, I am trying to add scylladb as datasource on Grafana(Scylladb Monitoring Stack). But Getting error “gocql: unable to create session: unable to discover protocol version: authentication required (using “org.apac… Language: en Canonical URL: https://forum.scylladb.com/t/error-while-adding-scylladb-as-datasource-on-grafana/2520 ## Headings Structure: H1: Error while adding scylladb as datasource on Grafana H3: Related topics ## Main Content: H1: Error while adding scylladb as datasource on Grafana H3: Related topics I am trying to add scylladb as datasource on Grafana(Scylladb Monitoring Stack). But Getting error “gocql: unable to create session: unable to discover protocol version: authentication required (using “org.apache.cassandra.auth.PasswordAuthenticator”)”. Could you please help me to resolve this error. I am trying to trace the query response on Grafana dashboard. See Deploying Scylla Monitoring Stack Without Docker | ScyllaDB Docs - there’s a datasource.yml file you may use to specify credentials. I am trying to add scylladb as datasource on Grafana(Scylladb Monitoring Stack). But Getting error “gocql: unable to create session: unable to discover protocol version: authentication required (using “org.apache.cassandra.auth.PasswordAuthenticator”)”. Could you please help me to resolve this error. The error suggests an authentication problem between Grafana and ScyllaDB. Ensure you have the correct username and password set up in Grafana under the ScyllaDB datasource configuration. --- ### Page: https://forum.scylladb.com/t/cqlsh-connection-error-unable-to-connect-to-any-servers/2521 Title: Cqlsh Connection error: 'Unable to connect to any servers' - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: When I trying to connect to CQL Shell, I get an error below. Anyone knows the resolution? Or how can I find out what’s wrong? $ cqlsh Connection error: ('Unable to connect to any servers', {'127.0.0.1:9042': Connection… Language: en Canonical URL: https://forum.scylladb.com/t/cqlsh-connection-error-unable-to-connect-to-any-servers/2521 ## Headings Structure: H1: Cqlsh Connection error: 'Unable to connect to any servers' H3: Related topics ## Main Content: H1: Cqlsh Connection error: 'Unable to connect to any servers' H3: Related topics When I trying to connect to CQL Shell, I get an error below. Anyone knows the resolution? Or how can I find out what’s wrong? Environment and background Started scylla-server and checked the status. Checked the Node status. Tried to connect CQL shell. Then, connection error occurred. I tried to show the TCP port condition. Indeed, port 9042 is being listened, but it’s different IP than the one the process is trying to connect to with cqlsh. I could connect to CQL shell when I used the ScyllaDB v5.4.6, though I configured no special settings to connect to CQL shell. What’s different or wrong now? what version of CQLsh are you using? Maybe try to upgrade to the latest GitHub - scylladb/scylla-cqlsh: A fork of the cqlsh code Scylla will use the address specified in rpc_address, as the address to listen on, for CQL connections. You mention above that you have changed this configuration, this is probably why your ScyllaDB node now doesn’t listen on the (default) 127.0.0.1 address anymore. Also note that cqlsh, when launched without arguments, will attempt to connect to localhost (127.0.0.1). If your ScyllaDB node listens on another address, you can provide this as a positional command-line argument: Thanks for the suggestion. I got the following log in installing. I thought it was enough updated, but it’s not enough? Anyway, I will keep in mind that I can upgrade it individually from there. Thanks for the valuable information! Thanks for the info! Yeah, I also thought maybe that changing the settings in scylla.yaml would affect it. I got that the setting of rpc_address is also reflected in the listen address of cqlsh. And I didn’t know you could specify arguments to cqlsh command. I’ll give it a try! Thanks alot. $ cqlsh 172.12.34.567 I tried to this way, but doesn’t work. I got the same error as before. I’m so glad to let me know anything If you guys have another idea. Thank you! I copied from your log output the address which seemed to be the one on which CQL listens. But I may have copied the wrong address. You can find out the correct address yourself: when a ScyllaDB node starts up, it will print the address it listens on: This is the address you have to pass to cqlsh and other clients. Note that you have to make sure that this address is one that is reachable from where your client is running. if you are running ScyllaDB in a docker container, you might have to open the appropriate ports. --- ### Page: https://forum.scylladb.com/t/data-size-gets-larger-after-a-compaction-is-this-normal/2523 Title: Data size gets larger after a compaction - Is this normal? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I found some compaction log records like this: compaction … 1GB to 1GB (~102% of original) in 140469ms = 9MB/s. ~10368 total partitions merged to 2459. I thought it was some rounding issue because compacted data shou… Language: en Canonical URL: https://forum.scylladb.com/t/data-size-gets-larger-after-a-compaction-is-this-normal/2523 ## Headings Structure: H1: Data size gets larger after a compaction - Is this normal? H3: Related topics ## Main Content: H1: Data size gets larger after a compaction - Is this normal? H3: Related topics I found some compaction log records like this: compaction … 1GB to 1GB (~102% of original) in 140469ms = 9MB/s. ~10368 total partitions merged to 2459. I thought it was some rounding issue because compacted data should be smaller than the original. So I checked the nodetool compactionhistory. And it shows the data size indeed gets larger after compaction. In the worst case, it can get 20% larger. The table in question is actually a materialized view. The workload is very overwrite heavy. It’s using the default STCS with zstd compression with 128k chunk size. I tried changing the compaction class to LCS and did a major compaction. But the results are the same. The only way I can see this happening is if the output data of the compaction compresses worst than the input data did. There is no way compaction produces more data then its input is. --- ### Page: https://forum.scylladb.com/t/replacing-a-dead-dn-node-with-the-same-ip-syncing-data/2524 Title: Replacing a dead (DN) node with the same IP, syncing data - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/replacing-a-dead-dn-node-with-the-same-ip-syncing-data/2524 ## Headings Structure: H1: Replacing a dead (DN) node with the same IP, syncing data H3: Related topics ## Main Content: H1: Replacing a dead (DN) node with the same IP, syncing data H3: Related topics Originally from the User Slack @Dominik_Mankowski: Is it normal, that during replace operation (https://opensource.docs.scylladb.com/branch-5.4/operating-scylla/procedures/cluster-management/replace-dead-node.html), a node that is replacing the dead node (the new node has the same IP as the old one) has status (nodetool status) UN, even though it hasn’t synced all the data? i.e. nodetool status on the new node looks like this (all the other nodes in the cluster also report this node as UN, while I had expected UJ status): @avi: @Kamil_Braun do you know? @Felipe_Cardeneti_Mendes: It is normal because the node isn’t joining. It was dead (DN), and now you are replacing it with another. So think about it this way: The node has already joined, you are bringing it back up @Dominik_Mankowski: > It is normal because the node isn’t joining. This is in contradiction to what the metrics/dashboard did show (that node was reported as Joining) @Felipe_Cardeneti_Mendes: Well — then that’s a monitoring issue — probably replacing would be more accurate @Kamil_Braun: in short: don’t use replace-with-same-IP, because it’s dangerous and has a bunch of these stupid quirks one recent issue we found with replace-with-same-IP: https://github.com/scylladb/scylladb/issues/19975 GitHub: Failure during replace-with-same-IP leaves the node without STATUS application_state (permanently), and token_metadata inconsistent (until restart) (applies to gossiper / “node-ops” based topology changes) · Issue #19975 · scylladb/scylladb well, if it completes, then it will be fine but generally, try avoiding it, use replace-with-different-IP instead @Dominik_Mankowski: @Kamil_Braun thanks for the hint. Would it be ok if we first removed a dead node (nodetool removenode) from the cluster (scale in), and then just simply add it to the cluster (scale out), with the same IP? @Kamil_Braun: yes, but it will take 2x much time as replace, data streaming phase will have to be done twice even more since at the end you should run cleanup (and IIRC cleanup is not really necessary if you use replace. But it is if you use remove + add) @Dominik_Mankowski: got it, thanks --- ### Page: https://forum.scylladb.com/t/why-is-a-minimum-of-three-nodes-recommended/2561 Title: Why is a minimum of three nodes recommended - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi Why is a minimum of three nodes recommended I’m using Docker to distribute two nodes to two servers and a single cluster I wanted to adjust the number of nodes to suit the server environment, but you can’t find eac… Language: en Canonical URL: https://forum.scylladb.com/t/why-is-a-minimum-of-three-nodes-recommended/2561 ## Headings Structure: H1: Why is a minimum of three nodes recommended H3: ScyllaDB Architecture - Fault Tolerance | ScyllaDB Docs H3: Related topics ## Main Content: H1: Why is a minimum of three nodes recommended H3: ScyllaDB Architecture - Fault Tolerance | ScyllaDB Docs H3: Related topics Why is a minimum of three nodes recommended I’m using Docker to distribute two nodes to two servers and a single cluster I wanted to adjust the number of nodes to suit the server environment, but you can’t find each node when a specific node died. When I looked up the document, I found that it recommended 3 nodes, but I didn’t find why. (I understand that scylladb uses the Gossip protocol instead of the raft algorithm. Is there a reason why nodes should be odd in Gossip protocol? As far as I understand the document, I know that each node is equal to the raft algorithm.) Is the issue of nodes not being able to connect to each other after they die because they are made up of two nodes? TL;DR A Replication Factor (RF) of 3 is recommended, so the Quorum’s Consistency Level(CL) will work even if one node is down. To have a RF of 3 you need at least 3 nodes. ScyllaDB is an Apache Cassandra-compatible NoSQL data store that can handle 1 million transactions per second on a single server. --- ### Page: https://forum.scylladb.com/t/unable-to-do-complex-searches-in-scylladb/2575 Title: Unable to do complex searches in ScyllaDB - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Good morning, How can ScyllaDB help with our use case in a most optimal way, what is the best approach? I was hoping we could do everything in ScyllaDB and eliminate having to use PostgreSQL(as a separate DB for the fi… Language: en Canonical URL: https://forum.scylladb.com/t/unable-to-do-complex-searches-in-scylladb/2575 ## Headings Structure: H1: Unable to do complex searches in ScyllaDB H3: CDC Overview | ScyllaDB Docs H3: Related topics ## Main Content: H1: Unable to do complex searches in ScyllaDB H3: CDC Overview | ScyllaDB Docs H3: Related topics Good morning, How can ScyllaDB help with our use case in a most optimal way, what is the best approach? I was hoping we could do everything in ScyllaDB and eliminate having to use PostgreSQL(as a separate DB for the final stage of reporting) but it seems that’s not feasible. Therefore our approach is still to stream the large volumes of data in real-time from the NoSQL database(ScyllaDB) out to Postgres and do our queries from there. Note that: 1.Reporting is not in real-time. 2.source systems that do searches is in real-time(complex queries in the sense that different systems come in to the scylladb cluster to do queries their own way, some do joins and also temporary tables) 3. A microservice does the transformation. Take a look at ScyllaDB CDC Feature ScyllaDB is an Apache Cassandra-compatible NoSQL data store that can handle 1 million transactions per second on a single server. You can use it to read updates from ScyllaDB and feed them to external systems (like PostgreSQL). Keep in mind: --- ### Page: https://forum.scylladb.com/t/consistency-level-issues/2581 Title: Consistency level Issues - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, My cluster is a multi-dc setup with 4 nodes and a replication factor (RF) of 2. I’m using the default consistency level (ONE) for both read and write operations. However, for a very small portion of reads (less than… Language: en Canonical URL: https://forum.scylladb.com/t/consistency-level-issues/2581 ## Headings Structure: H1: Consistency level Issues H3: ScyllaDB Architecture - Fault Tolerance | ScyllaDB Docs H3: ScyllaDB Repair | ScyllaDB Docs H3: Related topics ## Main Content: H1: Consistency level Issues H3: ScyllaDB Architecture - Fault Tolerance | ScyllaDB Docs H3: ScyllaDB Repair | ScyllaDB Docs H3: Related topics Hi, My cluster is a multi-dc setup with 4 nodes and a replication factor (RF) of 2. I’m using the default consistency level (ONE) for both read and write operations. However, for a very small portion of reads (less than 0.1%), I’m experiencing inconsistencies where the same request sometimes returns a null value and other times a non-null value. Although I am using consistency level ONE, I expect scylla to be eventually consistent. But I am experiencing this inconsistent behaviour in reads even after 5-6 hours of writing the data. Is there a max time after which read consistency can be guaranteed? And what is the recommended consistency level for production systems? My use-case requires the data to be eventually consistent after 1-2 hours of writes. Are you using RF of 2 per DC? With CL=1, if a write request (mutation) got lost somewhere between the coordinator and a replica, the coordinator will still return a success ack based on the one replica that did succeed. This means a replica still needs to be updated; there is no way to know that. In this case, a Read with CL=1 might get a stalled value. Only a repair operation will bring all the replicas back into sync. You can mitigate this issue using RF=3 and CL=QOURUM for reads and writes. ScyllaDB is an Apache Cassandra-compatible NoSQL data store that can handle 1 million transactions per second on a single server. ScyllaDB is an Apache Cassandra-compatible NoSQL data store that can handle 1 million transactions per second on a single server. --- ### Page: https://forum.scylladb.com/t/using-udf-and-wasm-for-a-specific-use-case-internal-lwt/2658 Title: Using UDF and WASM for a specific use case, internal LWT - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/using-udf-and-wasm-for-a-specific-use-case-internal-lwt/2658 ## Headings Structure: H1: Using UDF and WASM for a specific use case, internal LWT H3: Related topics ## Main Content: H1: Using UDF and WASM for a specific use case, internal LWT H3: Related topics Originally from the User Slack @Mikael_Hedberg: Hi, I have some questions regarding UDF and WASM. I think it’s a really exciting feature and I had the following use case in mind. Could someone maybe shed some light if this is a good idea or not and if UDFs are really intended for such a use case. • I store a complex (not big) graph as a serialized object. We have a bunch of operations on that object that can affect many nodes or none (reading the whole thing is a must). • At the moment we are reading it from scylla, performing the patch on the object and then set it again using LWT (IF condition on etag) What would be awesome is that if we could simply compile some C or Rust code that performs the patch operations and use it as UDFs inside the DB. This would for example require that we deserialize a protobuf (or a flatbuffer), perform the algorithms and then update the object. Would you consider this a good approach? I expect it to greatly reduce the latency since we can get rid of the LWT operation all together and only perform a simple update (possible a read when we need to return the data). Both operations are preferable to LWT. Any thoughts? Are there stable WASM bindings available (i.e. encapsulating the ABI), if so where? @Piotr_Smaroń: cc @Wojciech_Mitros @avi: UDFs can’t be used for update-in-place yet; even if they could, they’d have to perform a LWT internally. @Wojciech_Mitros: Right, currently you can only use them to read the calculated result or aggregate results from many rows using UDAs. The deserialization and other algorithms should be possible to implement in them - the Rust bindings can be found at https://github.com/scylladb/scylla-rust-udf Note that they’re still experimental - we haven’t tested them well yet @Mikael_Hedberg: Great, thanks a lot for your answers. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-244-2024-08-25/2703 Title: Last week in scylladb.git master (issue #244; 2024-08-25) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the afee3924b3…4823a1e203 range are covered. There were 95 non-merge commits from 14 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-244-2024-08-25/2703 ## Headings Structure: H1: Last week in scylladb.git master (issue #244; 2024-08-25) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #244; 2024-08-25) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the afee3924b3…4823a1e203 range are covered. There were 95 non-merge commits from 14 authors in that period. Some notable commits: Alternator, ScyllaDB’s implementation of the DynamoDB API, now supports Role-Based Access Control (RBAC). Control is done via CQL. Alternator, ScyllaDB’s implementation of the DynamoDB API, now has metrics for batch latency and size. A race condition between tablet repair and tablet split (the latter happens when a table grows) has been fixed. Raft uses log truncation to limit memory consumption. A mismatch between in-memory log truncation and on-disk log truncation was fixed. When communicating with older versions of ScyllaDB, the server uses a schema digest to see whether there is a schema mismatch or not. This is now less likely to stall when processing large schemas. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/failed-to-type-check-query-arguments-which-does-not-exist/2760 Title: Failed to type check query arguments which does not exist - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello all. I’m using rust lang to connect to scylla DB, however I always run into the warning when I INSERT: lib_core::model::activity::ActivityModel: value for column user_date_time was provided, but there is no bind … Language: en Canonical URL: https://forum.scylladb.com/t/failed-to-type-check-query-arguments-which-does-not-exist/2760 ## Headings Structure: H1: Failed to type check query arguments which does not exist H3: Related topics ## Main Content: H1: Failed to type check query arguments which does not exist H3: Related topics Hello all. I’m using rust lang to connect to scylla DB, however I always run into the warning when I INSERT: I’m pretty sure I’m not using user_date_time anywhere in the ActivityModel data model or queries. I have no idea about how scylladb does type checks using rust lang. The user_date_time is a field in MoodModel and the query with binding mark for that field is used there. Why would it affect ActivityModel? I have used common trait with generics for those data models, though. It should be the prepared statements which cause this problem. I should not another way rather than global static variables to hold prepared statements. --- ### Page: https://forum.scylladb.com/t/release-scylladb-cloud-26-august-2024/2776 Title: [RELEASE] ScyllaDB Cloud - 26 August 2024 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: We are happy to announce a new update to the ScyllaDB Cloud service. New features and enhancements: The Notification Preferences page can now be accessed via the top-right user dropdown menu. The Connect page Java tab… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-cloud-26-august-2024/2776 ## Headings Structure: H1: [RELEASE] ScyllaDB Cloud - 26 August 2024 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Cloud - 26 August 2024 H3: Related topics We are happy to announce a new update to the ScyllaDB Cloud service. New features and enhancements: --- ### Page: https://forum.scylladb.com/t/adding-multiple-new-nodes-into-an-existing-cluster-tablets-and-consistent-topology/2785 Title: Adding multiple new nodes into an existing cluster, tablets and consistent topology - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/adding-multiple-new-nodes-into-an-existing-cluster-tablets-and-consistent-topology/2785 ## Headings Structure: H1: Adding multiple new nodes into an existing cluster, tablets and consistent topology H3: Related topics ## Main Content: H1: Adding multiple new nodes into an existing cluster, tablets and consistent topology H3: Related topics Originally from the User Slack @Johannes_Gilger: Can we add multiple new nodes into an existing scylla cluster at once or do we have to do it one-by-one? @avi: When consistent topology is enabled, you can add nodes simultaneously. It’s only an improvement when your tables use tablets, however. @Johannes_Gilger: @avi Thanks, that makes sense. We’ll go one-by-one for now. Still having some issues with the fs.aio-max-nr setting anyway that we need to iron out @avi: Just increase it 10x --- ### Page: https://forum.scylladb.com/t/high-read-latency/2792 Title: High read latency - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, I’ve scylla cluster deployed on self hosted kubernetes cluster with k8s operator. I have 3 nodes of 16 CPUs and 32GB of mem. Each node is using node local storage with local-path-provisioneron ssd. When client ser… Language: en Canonical URL: https://forum.scylladb.com/t/high-read-latency/2792 ## Headings Structure: H1: High read latency H3: Related topics ## Main Content: H1: High read latency H3: Related topics I’ve scylla cluster deployed on self hosted kubernetes cluster with k8s operator. I have 3 nodes of 16 CPUs and 32GB of mem. Each node is using node local storage with local-path-provisioneron ssd. When client service is started read latency reaches 4s at peak and hangs at 1-2s on average. That is with reads at 6kosp/s and that does not seem like a lot. I’ve found this in the logs: INFO 2024-08-26 14:10:20,132 [shard 8:stat] reader_concurrency_semaphore - (rate limiting dropped 1 similar messages) Semaphore _read_concurrency_sem with 100/100 count and 2280960/38755368 memory resources: timed out, dumping permit diagnostics: permits count memory table/operation/state 61 61 1299K user_activity.user_internal_ids/data-query/active/await 39 39 928K user_activity.device_internal_ids/data-query/active/await 1 0 0B user_activity.device_internal_ids/mutation-query/waiting_for_admission 132 0 0B user_activity.device_internal_ids/data-query/waiting_for_admission 151 0 0B user_activity.user_internal_ids/data-query/waiting_for_admission 384 100 2228K total Stats: permit_based_evictions: 15 time_based_evictions: 0 inactive_reads: 0 total_successful_reads: 298208 total_failed_reads: 2729 total_reads_shed_due_to_overload: 0 total_reads_killed_due_to_kill_limit: 0 reads_admitted: 299593 reads_enqueued_for_admission: 67237 reads_enqueued_for_memory: 0 reads_admitted_immediately: 234191 reads_queued_because_ready_list: 29875 reads_queued_because_need_cpu_permits: 7514 reads_queued_because_memory_resources: 29848 reads_queued_because_count_resources: 0 reads_queued_with_eviction: 4 total_permits: 301438 current_permits: 384 need_cpu_permits: 100 awaits_permits: 100 disk_reads: 100 sstables_read: 115 I’m not sure how to interpret that, especially “100/100 count”, cpu usage is low, memory is not saturated, I want to blame IO but can’t find any metrics to support it. 32GB of RAM for 16 CPUs is very little. ScyllaDB will use most of that, leaving very little for cache, which will make most of your reads go to disk. ScyllaDB has a hard-limit on the maximum amount of concurrent disk reads: 100. Once you have this many disk reads, new read request will be queued. This is what we see above. Some read requests sit in this queue for so long, waiting for their turn, that they time out while waiting. We usually provide around 8GB of RAM per CPU. That is a much more comfortable amount, which leaves room for a healthy amount of cache, which improves latencies a lot. Thank you, I’m starting to understand some aspects better like the reads limit and queue. That was my first thought that there is not enough resources for scylla to run stable, but when I check metrics in grafana scylla uses only 8 CPUs and 25 GB of memory. As a new scylla user it’s hard for me to know what performance should I be expecting from scylla in terms of reads/writes per second depending on what resources I assign to scylla pods. I know it depends on the actual queries but is there like a rule of thumb like each 2 CPUs and 16 GB or ram should result in at least 1000 reads per second so I know at least it I’m getting performance in order of values I should be. --- ### Page: https://forum.scylladb.com/t/scylladb-monitoring-data-vanish-after-restarting-start-all-sh/2795 Title: Scylladb monitoring data vanish after restarting start-all.sh - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, Scylladb monitoring data on grafana vanish after restarting start-all.sh Could you please help me to keep it even after restart. Thanks Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-monitoring-data-vanish-after-restarting-start-all-sh/2795 ## Headings Structure: H1: Scylladb monitoring data vanish after restarting start-all.sh H3: The start-all.sh Command | ScyllaDB Docs H3: Related topics ## Main Content: H1: Scylladb monitoring data vanish after restarting start-all.sh H3: The start-all.sh Command | ScyllaDB Docs H3: Related topics Scylladb monitoring data on grafana vanish after restarting start-all.sh Could you please help me to keep it even after restart. “Note Specifying an external directory is important for systems in production. Without it, every restart of the monitoring stack will result in metrics lost.” Scylla Monitoring Stack is a full stack for Scylla monitoring and alerting. The stack contains open source tools including Prometheus and Grafana, as well as custom Scylla dashboards and tooling. Hi, Thanks for quick response. I am using command “./start-all.sh -d //prometheus_data” Anything else I have to take care ? --- ### Page: https://forum.scylladb.com/t/consistency-level-cl-rollback-on-failure-retries-and-repair/2813 Title: Consistency LeveL (CL), rollback on failure, retries and repair - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/consistency-level-cl-rollback-on-failure-retries-and-repair/2813 ## Headings Structure: H1: Consistency LeveL (CL), rollback on failure, retries and repair H3: Related topics ## Main Content: H1: Consistency LeveL (CL), rollback on failure, retries and repair H3: Related topics Originally from the User Slack @Daria_Fedorova: Hi I want to ask some questions about Consistency Level (CL) Does it affect anything except what scylla returns to client ? If CL for write was not reached , will coordinator try to rollback the operation or mark it as failed to be repaired in accordance with what we returned to client ? Will repair be affected by CL ? For example if write requests CL=ALL and it failed but most of the nodes succeeded, Will repair know that operation overall failed and roll it back or will it try to add this write on rest of the nodes ? The question arose when I found DowngradingConsistencyRetryPolicy If CL only affects the answer to client then it seems to me that there is no use in retrying with lower CL , if lower CL is acceptable for a client it can set lower CL from the beginning. @Felipe_Cardeneti_Mendes: There’s no rollback. The idea is that your client must retry as it receives an exception. @Daria_Fedorova: so repair will repair to quorum value ? @Felipe_Cardeneti_Mendes: repair will sync all replicas. For example, if you write n=1 with Quorum and this write gets accepted to one out of three replicas, you will receive an exception. Considering you did nothing and “let it go”, later - as you repair, n=1 will be propagated to the remaining 2 replicas which missed the write. also see hinted handoff and read repair, other anti-entropy mechanisms. @Daria_Fedorova: Thanks I will read it too This leaves the question - what is DowngradingConsistencyRetryPolicy used for ? Is it legacy ? why would a client need to rerty with lower CL if first request failed • if first request reached lower CL write will be propagated eventually Is it a way to effectually use lower CL but still monitor errors with high CL. ? @Felipe_Cardeneti_Mendes: In general this is down to application semantics. Some apps are fine with a lower consistency level momentarily, others aren’t. Imagine you lose quorum for whatever reason. In this situation, you’d have 2 nodes down (considering RF=3). If this is fine, then Downgrading the consistency would allow you to continue serving traffic with a lower consistency, given that only a single replica is available. Some apps can’t tolerate even that, because there’s no way to guarantee that just because 1 replica is available that it has all the data the other 2 had. So at this point, even though you are still operating fine, you are ok with serving stale data. @Karol_Baryła: I may be misremembering, but I think DowngradingConsistencyRetryPolicy is actually deprecated in Java Driver 3.x or 4.x (don’t remember which one) @Felipe_Cardeneti_Mendes: V4 --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-8/2823 Title: [RELEASE] ScyllaDB Enterprise 2024.1.8 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2024.1.8, a bug-fix production-ready patch release for ScyllaDB Enterprise 2024.1 LTS Release. This release introduces the zstd compression algorithm for node-to-node comm… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-8/2823 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2024.1.8 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2024.1.8 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2024.1.8, a bug-fix production-ready patch release for ScyllaDB Enterprise 2024.1 LTS Release. This release introduces the zstd compression algorithm for node-to-node communication, in addition to the existing lz4 compression. The efficiency of each compression algorithm is workload-dependent. Zstd compression is an Enterprise-only feature. To enable zstd compression, set: The following issues are fixed in this release (with an open-source reference, if available): --- ### Page: https://forum.scylladb.com/t/release-scylladb-6-1-1/2824 Title: [RELEASE] ScyllaDB 6.1.1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 6.1.1, a bugfix patch release of the ScyllaDB 6.1 stable branch. ScyllaDB Open Source 6.1.1, like all past and future 6.x.y releases, is backward compatible and supports r… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-6-1-1/2824 ## Headings Structure: H1: [RELEASE] ScyllaDB 6.1.1 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 6.1.1 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 6.1.1, a bugfix patch release of the ScyllaDB 6.1 stable branch. ScyllaDB Open Source 6.1.1, like all past and future 6.x.y releases, is backward compatible and supports rolling upgrades. Users are encouraged to upgrade to 6.1.1. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/last-3-weeks-in-scylla-cluster-tests-git-master-issue-60-2024-08-30/2825 Title: Last 3 weeks in scylla-cluster-tests.git master (issue #60; 2024-08-30) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last 3 weeks. Commits in the bc58a9a1…5e46b04f range are covered. There were 57 non-merge commits from 13 authors in… Language: en Canonical URL: https://forum.scylladb.com/t/last-3-weeks-in-scylla-cluster-tests-git-master-issue-60-2024-08-30/2825 ## Headings Structure: H1: Last 3 weeks in scylla-cluster-tests.git master (issue #60; 2024-08-30) H3: Related topics ## Main Content: H1: Last 3 weeks in scylla-cluster-tests.git master (issue #60; 2024-08-30) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last 3 weeks. Commits in the bc58a9a1…5e46b04f range are covered. There were 57 non-merge commits from 13 authors in that period. Some notable commits: Monitoring was updated to version 4.8.0. The SCT runner image (and SCT Jenkins builders) received an update to version 1.8, which includes an OS upgrade to Ubuntu 24.04, Java updated to version 21, and stronger SSH keys. Nosqlbench was updated to version 5.21.2 and configured to use the Java 4.x Scylla driver to facilitate testing with the Java v4.x driver, as no other tool currently supports it. Several fixes and documentation updates were made regarding resource cleanup and SCT local setup. With the migration of cassandra-stress from scylla-tools-java to a new repository, a new Docker image and DockerHub release have been created. The new image is now used instead of Scylla’s older Docker images, currently at version 3.13.0. An issue with reuse-cluster was fixed, addressing leftover autossh hook containers that were blocking syslogng initialization. These are now properly cleaned up. Additionally, reusing clusters from the k8s-local-kind backend has been fixed. A new AWS ASG was introduced for testing Docker artifacts on FIPS machines, using a specific SCT runner and Jenkins labels, separate from the regular ones. Latte was upgraded to version ‘0.26.3-scylladb’, incorporating the newer Scylla-Rust driver ‘v0.13.2’ and compiled with Rust ‘1.79’. This update includes a necessary bug fix for customer-oriented tests. In addition to linking GitHub issues to specific reactor_stall errors in SCT events, the severity of these errors is now adjusted based on the linked issue status. If an issue is resolved but not labeled with sct--skip (indicating the fix was not backported), the severity is increased to Error. A follow-up script was added to migrate latency decorator results to Argus to maintain historical data for graphing purposes. We introduced a method to inject specific SCT_* environment variables into Jenkins jobs using the extra_environment_variables parameter, allowing configuration via environment variables without needing to create a pipeline parameter for each one. Perf-simple-query performance results are now sent to Argus. We reworked and introduced several Scylla Manager tests, including no-delta backup, restoring from a pre-created backup, and added new test parameters and job improvements. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/reader-concurrency-semaphore/2833 Title: Reader_concurrency_semaphore - Database Community - ScyllaDB Community NoSQL Forum Meta Description: Hi all, We are running 8 node cluster with scylladb 6.0 version. We could see below message from logs across all the nodes . But overall cpu(8 cores) usage was not much. And also we observed most of the times all 8 shar… Language: en Canonical URL: https://forum.scylladb.com/t/reader-concurrency-semaphore/2833 ## Headings Structure: H1: Reader_concurrency_semaphore H3: Related topics ## Main Content: H1: Reader_concurrency_semaphore H3: Related topics Hi all, We are running 8 node cluster with scylladb 6.0 version. We could see below message from logs across all the nodes . But overall cpu(8 cores) usage was not much. And also we observed most of the times all 8 shards were 100% from the Load metrics(Scylla dashboard). Can someone help share your thoughts and how to fix it. Log message [shard 5:stmt] reader_concurrency_semaphore - (rate limiting dropped 584 similar messages) Semaphore user with 61/100 count and 11331744/166807470 memory resources: timed out, dumping permit diagnostics: permits count memory table/operation/state 2 2 581K in_user_history.recent_recommended_items/data-query/active/need_cpu 1 0 0B in_user_history.recent_page_views/mutation-query/waiting_for_admission 27 0 0B in_user_history.recent_recommended_items/mutation-query/waiting_for_admission 132 0 0B in_user_history.recent_page_views/data-query/waiting_for_admission 228 0 0B in_user_history.recent_recommended_items/data-query/waiting_for_admission Looks like you overloaded ScyllaDB, at least when it comes to CPU. Try adding more nodes to the cluster, or reducing the load. Also, the report you posted seems incomplete. I says there are 61 count resources, but I only see 2 in the table below. Either parts are missing, or the table is from the a different report? --- ### Page: https://forum.scylladb.com/t/how-to-find-all-of-the-databases-partitions-efficiently-full-table-scan/2834 Title: How to find all of the database's partitions efficiently, full table scan - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/how-to-find-all-of-the-databases-partitions-efficiently-full-table-scan/2834 ## Headings Structure: H1: How to find all of the database's partitions efficiently, full table scan H3: Related topics ## Main Content: H1: How to find all of the database's partitions efficiently, full table scan H3: Related topics Originally from the User Slack @Dylan_Piette: Hello to the scylladb team, I’m currently facing the issue of needing to know all the current partitions in the db in a performant way Is there any recommandations on how to do it ? I need to know the values of the pk, not just the count or something like that https://github.com/scylladb/scylladb/issues/2066 GitHub: Really slow select distinct · Issue #2066 · scylladb/scylladb @Felipe_Cardeneti_Mendes: https://www.scylladb.com/2017/03/28/parallel-efficient-full-table-scan-scylla/ there’s some boilerplate Go code in the relevant github repo, it originally just counts IIRC - but should be straightforward to iterate and save values. @Sven: @Dylan_Piette For years we have just used SELECT DISTINCT our_partition_key_column_name FROM our_table_name and it has worked without problems. However, note that we do not have a big database. In our case, the result set to this query within any keyspace contains at most 100000 partition key values, which in our case are UUIDs. @Piette_Dylan: Well in our use case we have millions of partitions and the data is volatile, a partition can be gone in a few seconds (we deleted all the rows for that partition) So we need a fast way to get the most accurate representation of what is actually inside the table The solution @Felipe_Cardeneti_Mendes linked is a great one right now, after adapting the code I’m now able to query all the partitions 6x faster which is already great ! But I can’t help wishing for a near instant way to get this data haha --- ### Page: https://forum.scylladb.com/t/release-new-usage-dashboard-beta-in-scylladb-cloud-1-september-2024/2835 Title: [RELEASE] New Usage Dashboard (Beta) in ScyllaDB Cloud! - 1 September 2024 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: We’re excited to introduce the latest update to ScyllaDB Cloud, featuring a new Usage Dashboard—now available (from the the top bar menu) for exploration ! This feature, currently in Beta, is designed to give you enhanc… Language: en Canonical URL: https://forum.scylladb.com/t/release-new-usage-dashboard-beta-in-scylladb-cloud-1-september-2024/2835 ## Headings Structure: H1: [RELEASE] New Usage Dashboard (Beta) in ScyllaDB Cloud! - 1 September 2024 H3: Related topics ## Main Content: H1: [RELEASE] New Usage Dashboard (Beta) in ScyllaDB Cloud! - 1 September 2024 H3: Related topics We’re excited to introduce the latest update to ScyllaDB Cloud, featuring a new Usage Dashboard—now available (from the the top bar menu) for exploration ! This feature, currently in Beta, is designed to give you enhanced visibility into your cluster’s usage. With the Usage Dashboard, you can: While ScyllaDB takes care of optimizing your cluster, the Usage Dashboard provides you with the insights you need to understand your cluster’s resource consumption. This transparency helps you stay informed and make data-driven decisions as your needs evolve. For more details on how to leverage the Usage Dashboard, visit our documentation. We welcome your feedback and suggestions as we continue to enhance ScyllaDB Cloud! --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-245-2024-09-01/2836 Title: Last week in scylladb.git master (issue #245; 2024-09-01) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 4823a1e203…e01cef01a6 range are covered. There were 142 non-merge commits from 20 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-245-2024-09-01/2836 ## Headings Structure: H1: Last week in scylladb.git master (issue #245; 2024-09-01) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #245; 2024-09-01) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 4823a1e203…e01cef01a6 range are covered. There were 142 non-merge commits from 20 authors in that period. Some notable commits: Integrated backup has been merged. A new nodetool backup command (and corresponding REST API endpoint) will copy a snapshot to an S3 compatible endpoint. Integrated restore has been merged. A new nodetool restore command (and corresponding REST API endpoint) will restore a backup taken by nodetool backup into an empty table. A new nodetool tasks command can be used to view and manage maintenance tasks running on the node. There is now support for zero-token nodes. Such nodes do not replicate any data, but can participate in query coordination, and in Raft quorum voting. This can help in maintaining Raft quorum in two-datacenter clusters when one datacenter has been lost. A write to a base table will now be rejected by the coordinator when one or more of the replicas has a full view update backlog. This reduces inconsistencies in materialized views. Alternator, ScyllaDB’s implementation of the DynamoDB API, has more efficient reverse queries now, reducing the gap from CQL. A bug in replacing a node with inter-datacenter network encryption enabled has been fixed. A node will now ignore dns name resolution errors of seeds when restarting, as those seed names could be referring to nodes that were removed. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/nosql-column-oriented-database-tutorial/2842 Title: Nosql Column-oriented database tutorial - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Does anybody have any good links to code snippets on how to create a nosql Column-oriented database as well as crud code? I’m new and this would help me get started. Language: en Canonical URL: https://forum.scylladb.com/t/nosql-column-oriented-database-tutorial/2842 ## Headings Structure: H1: Nosql Column-oriented database tutorial H3: Related topics ## Main Content: H1: Nosql Column-oriented database tutorial H3: Related topics Does anybody have any good links to code snippets on how to create a nosql Column-oriented database as well as crud code? I’m new and this would help me get started. I found a link for " Compress data with columnar tables in Azure Cosmos DB for PostgreSQL" but I’m getting an error when I try to post it here. CREATE TABLE contestant ( handle TEXT, birthdate DATE, rating INT, percentile FLOAT, country CHAR(3), achievements TEXT ) USING columnar; Incorrect syntax near ‘USING’. This is on Windows 10 SQL Server I need to use Azure Cosmos DB Try the Essentials course on ScyllaDB University. It includes some hands-on labs where you can create a Table and do some basic queries. --- ### Page: https://forum.scylladb.com/t/queries-taking-longer-after-migrating-to-a-new-cluster-how-can-i-debug-and-analyze-this/2845 Title: Queries taking longer after migrating to a new cluster, how can I debug and analyze this? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/queries-taking-longer-after-migrating-to-a-new-cluster-how-can-i-debug-and-analyze-this/2845 ## Headings Structure: H1: Queries taking longer after migrating to a new cluster, how can I debug and analyze this? H3: Related topics ## Main Content: H1: Queries taking longer after migrating to a new cluster, how can I debug and analyze this? H3: Related topics Originally from the User Slack @FRED_FAT: Hi Gurus, recently I have migrated data from an old cluster to a new one, however I find the avg query (by primary key) performance is degrade around 30%, from 22ms to 26ms. Compare with other cluster(old cluster is destroyed) with exactly same data, there is no difference on hardware and network configuration. The only difference we see is the old cluster has more sstables (around 300 sstables each node) than other cluster (less than 100 sstables each node) . I am guessing whether query read more record version between multiple sstables which could cause more disk reads than normal cluster and I want to check further on it. But I have no idea how to check the avg disk reads for average or specific CQL from any internal virtual table or from tracing output, anyone has suggestion ? Or anyone know whether more sstables (without enough compaction) could cause additional disk reads ? @avi: Try nodetool tablehistograms and tracing to see how many sstables are read per query @FRED_FAT: thanks, I tested tracing with ‘bypass cache’ and can see some file io operation. will try nodetool tablehistograms --- ### Page: https://forum.scylladb.com/t/unable-to-implement-function-in-group-by-clause/2846 Title: Unable to implement function in group by clause - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi scylla team, Can anyone help me with following issue: I am required to execute the following query, but was not able to execute. Seems like we can’t use function in group by clause. Can anyone help me to confirm we… Language: en Canonical URL: https://forum.scylladb.com/t/unable-to-implement-function-in-group-by-clause/2846 ## Headings Structure: H1: Unable to implement function in group by clause H3: Related topics ## Main Content: H1: Unable to implement function in group by clause H3: Related topics Hi scylla team, Can anyone help me with following issue: I am required to execute the following query, but was not able to execute. Seems like we can’t use function in group by clause. Can anyone help me to confirm weather we can use function in group clause or not. Can you post the error? How the table is defined? Why are you using ALLOW FILTERING? Maybe you want to group by a column that doesn’t belong to the partition key, which is not possible, per doc: " Using the GROUP BY option, it is only possible to group rows at the partition key level or at a clustering column level" Also, you’re repeating the toDate function call in the group by clause. --- ### Page: https://forum.scylladb.com/t/release-scylladb-6-0-3/2847 Title: [RELEASE] ScyllaDB 6.0.3 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 6.0.3, a bugfix patch release of the ScyllaDB 6.0 stable branch. ScyllaDB Open Source 6.0.3, like all past and future 6.x.y releases, is backward compatible and supports r… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-6-0-3/2847 ## Headings Structure: H1: [RELEASE] ScyllaDB 6.0.3 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 6.0.3 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 6.0.3, a bugfix patch release of the ScyllaDB 6.0 stable branch. ScyllaDB Open Source 6.0.3, like all past and future 6.x.y releases, is backward compatible and supports rolling upgrades. Note that the latest open-source stable branch is 6.1, and you are encouraged to upgrade to it. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/release-scylladb-rust-driver-0-14-0/2848 Title: [RELEASE] ScyllaDB Rust Driver 0.14.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce ScyllaDB Rust Driver 0.14.0, an asynchronous CQL driver for Rust, optimized for Scylla, but also compatible with Apache Cassandra! Some interesting statistics: over 2.103k dow… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-rust-driver-0-14-0/2848 ## Headings Structure: H1: [RELEASE] ScyllaDB Rust Driver 0.14.0 H2: Changes H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Rust Driver 0.14.0 H2: Changes H3: Related topics The ScyllaDB team is pleased to announce ScyllaDB Rust Driver 0.14.0, an asynchronous CQL driver for Rust, optimized for Scylla, but also compatible with Apache Cassandra! Some interesting statistics: API cleanups / breaking changes: New features / enhancements: CI / developer tool improvements: Congrats to all contributors and thanks everyone for using our driver! The source code of the driver can be found here: The official crates.io registry entry is here: Thank you for your attention, please do not hesitate to contact us if you have any questions, issues, feature requests, or are simply interested in our driver! Contributors since the last release: --- ### Page: https://forum.scylladb.com/t/insert-query-execution-time-suddenly-slows-down/2849 Title: INSERT query execution time suddenly slows down - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: INSERT query execution time suddenly slows down when the number of columns is approximately 235 or more. I would like to know whether this is a problem that can be solved by data modeling or other techniques, or whether… Language: en Canonical URL: https://forum.scylladb.com/t/insert-query-execution-time-suddenly-slows-down/2849 ## Headings Structure: H1: INSERT query execution time suddenly slows down H3: Related topics ## Main Content: H1: INSERT query execution time suddenly slows down H3: Related topics INSERT query execution time suddenly slows down when the number of columns is approximately 235 or more. I would like to know whether this is a problem that can be solved by data modeling or other techniques, or whether it is a performance limitation of ScyllaDB. If the problem can be solved, I would appreciate it if you could let me know how to do it too. The time of INSERT query execution [*1] Most times it takes only 0.0007 sec, but sometimes it takes 0.04 sec. ScyllaDB is not desgien for 3000 columns per table (See Limits | ScyllaDB Docs) I suggest revisiting the data modeling. See here for Wide Table Design Vs. Narrow table design Data model | ScyllaDB Docs Thanks alot for your info. I see that there are CQL limits. Column count has an effect on how expensive it is to merge writes – which translates to time it takes to commit the write to the memtable, as well as to commitlog. You can use UDT:s to improve your latencies, see this blog post for more details: If You Care About Performance Use UDT's sorry for late reply and thank you for letting me know about UDT. We found out that inserting one record by 200 columns at several times doesn’t take long time. So, we’ll take it method for now. But I know it’s not a really correct resolution, I will keep researching how to use UDT on our system. Thank you so much!! --- ### Page: https://forum.scylladb.com/t/alternator-returns-a-cursor-with-empty-item-set/2851 Title: Alternator returns a cursor with empty Item set - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I’m in the process of transitioning away from Dynamodb and I’m using Alternator to handle our queries / entries to Scylladb. I’m running into an issue where sometimes Alternator will return and empty array but also ha… Language: en Canonical URL: https://forum.scylladb.com/t/alternator-returns-a-cursor-with-empty-item-set/2851 ## Headings Structure: H1: Alternator returns a cursor with empty Item set H3: Related topics ## Main Content: H1: Alternator returns a cursor with empty Item set H3: Related topics I’m in the process of transitioning away from Dynamodb and I’m using Alternator to handle our queries / entries to Scylladb. I’m running into an issue where sometimes Alternator will return and empty array but also have a cursor (LastEvaluatedKey). There are items returned after using the cursor to paginate through the db, but the expected behavior is that Alternator returns all the items it can first (like it does in Dynamodb). Anyone have an idea what’s going on here? Hi, I believe current behavior is valid and compatible with dynamodb, see note: --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-61-2024-09-06/2853 Title: Last week in scylla-cluster-tests.git master (issue #61; 2024-09-06) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the a6c03ca2…64fa8b39 range are covered. There were 17 non-merge commits from 6 authors in tha… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-61-2024-09-06/2853 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #61; 2024-09-06) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #61; 2024-09-06) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the a6c03ca2…64fa8b39 range are covered. There were 17 non-merge commits from 6 authors in that period. Some notable commits: SCT PR pipeline can now verify the reuse_cluster functionality. To trigger it, simply add a *-reuse label (e.g. test-provision-aws-reuse) to the PR. Reuse cluster functionality is now supported on the GCE backend. Scylla Manager tests now support Ubuntu 24. Configs and sanity/installation/upgrade tests added. Azure DB/loader instances no longer provide public IPs by default, reducing exposure to the internet and saving costs. Several adjustments were made for tablets testing, including setting a reliable replication factor, skipping decommissioning if RF >= number of DB nodes, or decreasing RF before decommission to allow it. The collect-logs command will now resend SCT events to Argus when they are missing. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/no-code-scylla-api/2856 Title: No code scylla api - Database Community - ScyllaDB Community NoSQL Forum Meta Description: Hi everyone, I’m currently working on a no-code Web-API generator backed by ScyllaDB. Features implemented: CRUD OPS (insert, bulk_insert, all selects possible with primary key, possible updates based on specified fie… Language: en Canonical URL: https://forum.scylladb.com/t/no-code-scylla-api/2856 ## Headings Structure: H1: No code scylla api H3: Related topics ## Main Content: H1: No code scylla api H3: Related topics Hi everyone, I’m currently working on a no-code Web-API generator backed by ScyllaDB. Features implemented: Interesting, thanks! Can you share a link to the project? Can you show it? I am interested in such thing… --- ### Page: https://forum.scylladb.com/t/connecting-with-ssl-and-non-ssl-at-the-same-time/2860 Title: Connecting with SSL and Non-SSL at the same time - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/connecting-with-ssl-and-non-ssl-at-the-same-time/2860 ## Headings Structure: H1: Connecting with SSL and Non-SSL at the same time H3: Related topics ## Main Content: H1: Connecting with SSL and Non-SSL at the same time H3: Related topics Originally from the User Slack @bryam_castillo: Hello team, for client side, is it possible to support connection with SSL and not-SSL simultaneously? @avi: native_transport_port_ssl @bryam_castillo: got it now. have to use different ports for encrypted and unencrypted. thank you --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-246-2024-09-08/2861 Title: Last week in scylladb.git master (issue #246; 2024-09-08) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the e01cef01a6…ab32ce6b45 range are covered. There were 93 non-merge commits from 16 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-246-2024-09-08/2861 ## Headings Structure: H1: Last week in scylladb.git master (issue #246; 2024-09-08) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #246; 2024-09-08) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the e01cef01a6…ab32ce6b45 range are covered. There were 93 non-merge commits from 16 authors in that period. Some notable commits: CQL has two styles of bind variables: ? and :var. The first is indented to be used with positional parameter passing style, the second with keyword argument style. In ScyllaDB 5.1 we started unifying multiple references to the same :name, requiring them to have the same type. However, it turns out that some users us named bind variables with positional parameter passing style, so the unification can now be disabled. Metadata queries on GCP and Azure (to determine the region and availability zone) will now be retried. Heat weighted load balancing is used to direct replica reads towards nodes with better cache hit rates. In some cases it can miscompute nodes, returning the same node twice. We now detect those cases and disable the optimization. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/memory-leak-cleaning-up/2863 Title: Memory leak cleaning up - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: hello I am running scylla 5.1.0 on linux centos7 on a 3 node cluster 128 GB RAM , 32 cores My application is writing around 500 million writes and reads per day in 1 table per hour Peaks at 20k /s with key being jus… Language: en Canonical URL: https://forum.scylladb.com/t/memory-leak-cleaning-up/2863 ## Headings Structure: H1: Memory leak cleaning up H3: Related topics ## Main Content: H1: Memory leak cleaning up H3: Related topics hello I am running scylla 5.1.0 on linux centos7 on a 3 node cluster 128 GB RAM , 32 cores My application is writing around 500 million writes and reads per day in 1 table per hour Peaks at 20k /s with key being just 32 bytes and value ranges from 1kb to 200kb in blobs All reads and writes are thru multi routine golang programs My question is Is there a limit on number of golang connections ScyllaDB FAQ | ScyllaDB Docs states max 32 connections 2. How to clean up the memory leaks ? After every 5-6 days the memory of the machine is exhausted I need to restart scylladb node by node. Is there a better way ? What do you mean by “memory of the machine is exhausted”? What are the symptoms you observe? Also, 5.1.0 reached its end of life. Please upgrade to the latest version and see if you are still having issues. --- ### Page: https://forum.scylladb.com/t/release-scylla-manager-3-3-2/2864 Title: [RELEASE] Scylla Manager 3.3.2 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces the release of ScyllaDB Manager 3.3.2, a production-ready patch release of the stable 3.3 branch. ScyllaDB Manager is a centralized cluster administration and recurrent tasks automation tool. … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-manager-3-3-2/2864 ## Headings Structure: H1: [RELEASE] Scylla Manager 3.3.2 H3: Related topics ## Main Content: H1: [RELEASE] Scylla Manager 3.3.2 H3: Related topics The ScyllaDB team announces the release of ScyllaDB Manager 3.3.2, a production-ready patch release of the stable 3.3 branch. ScyllaDB Manager is a centralized cluster administration and recurrent tasks automation tool. ScyllaDB Manager is available for ScyllaDB Enterprise customers and ScyllaDB Open Source users. With ScyllaDB Open Source, ScyllaDB Manager is limited to 5 nodes. See the ScyllaDB Manager Proprietary Software License Agreement for details. In version 3.3.2, we fixed a regression issue that could occur on single-node clusters. More details can be found in #3989. Additionally, there is an improvement in how the DNS address, when passed as a host, is handled by Scylla Manager #4020. Last but not least, a new flag, --skip-schema, has been introduced for the backup task. This flag can be used to skip the backup of the schema, as the schema backup will fail if CQL credentials are not provided #4008. ScyllaDB Enterprise customers are encouraged to upgrade to ScyllaDB Manager 3.3.2 in coordination with the ScyllaDB support team. The new release includes upgrades of both ScyllaDB Manager Server and Agent. Scylla Manager 3.3.2 supports the following Scylla Enterprise releases: And the following Open Source release (limited to 5 nodes see license 1): You can install and run Scylla Manager on Kubernetes using Scylla Operator. More here. --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-9/2869 Title: [RELEASE] ScyllaDB Enterprise 2024.1.9 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2024.1.9, a bug-fix production-ready patch release for ScyllaDB Enterprise 2024.1 LTS Release. Related Links Get ScyllaDB Enterprise 2024.1 (customers only, or 30-day ev… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-9/2869 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2024.1.9 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2024.1.9 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2024.1.9, a bug-fix production-ready patch release for ScyllaDB Enterprise 2024.1 LTS Release. The following issues are fixed in this release (with an open-source reference, if available): --- ### Page: https://forum.scylladb.com/t/sstable-resharding-how-much-free-disk-space-required/2872 Title: SSTable Resharding, how much free disk space required? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/sstable-resharding-how-much-free-disk-space-required/2872 ## Headings Structure: H1: SSTable Resharding, how much free disk space required? H3: Related topics ## Main Content: H1: SSTable Resharding, how much free disk space required? H3: Related topics Originally from the User Slack @Mahdi_Kamali: How much free disk space is required for resharding? @avi: Good question. We have this commit: To be safe, pick the largest sstable, multiple by the number of shards, and again by a factor of 2. @raphaelsc may know better @raphaelsc: A small correction on free space required: size of largest table * 2. That’s because in the worst case, all the sstables may be resharded in parallel, across all the shards. To reduce the space requirement, schema’s max_threshold can be tweaked temporarily to something like 2, so each shard will work on, at most 2 sstables at a time. @Mahdi_Kamali: Thanks @raphaelsc You mean the max_threshold property of STCS? Does it impact major compaction? @raphaelsc: all compaction strategies inherit that property. no, major ignores max_threshold. @Mahdi_Kamali: Thank you very much! I think that will be good if major compaction considers max_threshold Is there anyway to decrease required storage for major compaction? @raphaelsc: the idea of major is to compact everything. major space requirement is fixed in enterprise with ICS. in OSS, the only strategy that has low space req during major is LCS. but it has high write amplification, so might not fit your use case. --- ### Page: https://forum.scylladb.com/t/wrong-response-message-for-local-quorum-cl-with-datastax-java-driver-3-11-x/2877 Title: Wrong response message for Local_Quorum CL with datastax java driver 3.11.x - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I have 33 nodes Syclla cluster. I have keyspace which has 3 replication factor. I use datastax-java-driver 3.11.0 version and Scylla 4.6.x version. Due to some issue, 2 nodes were down from rack C. I get error message be… Language: en Canonical URL: https://forum.scylladb.com/t/wrong-response-message-for-local-quorum-cl-with-datastax-java-driver-3-11-x/2877 ## Headings Structure: H1: Wrong response message for Local_Quorum CL with datastax java driver 3.11.x H3: Related topics ## Main Content: H1: Wrong response message for Local_Quorum CL with datastax java driver 3.11.x H3: Related topics I have 33 nodes Syclla cluster. I have keyspace which has 3 replication factor. I use datastax-java-driver 3.11.0 version and Scylla 4.6.x version. Due to some issue, 2 nodes were down from rack C. I get error message below which is impossible in my opinion. As far as If I use LOCAL_QUORUM 2 nodes is enough for 3 replication factor. “Operation timed out for - received only 2 responses from 3 CL=LOCAL_QUORUM.” info={‘consistency’: ‘LOCAL_QUORUM’, ‘required_responses’: 3, ‘received_responses’: 2, ‘write_type’: ‘SIMPLE’ Is it a bug or something wrong with driver configuration? There are cases where the effective replication increases (during streaming to new or replacement nodes), it could be the case. I agree that the error message is confusing --- ### Page: https://forum.scylladb.com/t/error-when-replacing-a-node-init-bad-configuration-consistent-cluster-management-requires-schema-commit-log-to-be-enabled/2879 Title: Error when replacing a node - init - Bad configuration: consistent_cluster_management requires schema commit log to be enabled - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/error-when-replacing-a-node-init-bad-configuration-consistent-cluster-management-requires-schema-commit-log-to-be-enabled/2879 ## Headings Structure: H1: Error when replacing a node - init - Bad configuration: consistent_cluster_management requires schema commit log to be enabled H3: Related topics ## Main Content: H1: Error when replacing a node - init - Bad configuration: consistent_cluster_management requires schema commit log to be enabled H3: Related topics Originally from the User Slack @Mahdi_Kamali: I got this error when I was replacing a node. init - Bad configuration: consistent_cluster_management requires schema commit log to be enabled What is the problem? What can I do? Should every node in cluster be up and normal? @Robert: put inside scylla.yaml (the node which is under replacement) @Mahdi_Kamali: Thanks @Robert. Should I remove this config after replacement? @Robert: I guess not, I even thinking about adding that to all nodes inside cluster, but better to ask some more experience, for example @avi, to know why it’s required and what is a general purpose of that. From my PoV it sth new in 5.4.x and I didn’t have a time to dig into as long as I stuck in C* . It’s also a reason why I postpone upgrade to 6.x, need to have a time to read about all changes. Jumping, reading and testing all new features from Scylla and C* is really time consuming… btw. @Mahdi_Kamali did it solves Your problem? @Mahdi_Kamali: Yes it does. I wait for @avi to give more information @avi: Don’t remove after replacement. In general you should make sure that when adding a node that the configuration of the new node is consistent with the other nodes. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-62-2024-09-13/2891 Title: Last week in scylla-cluster-tests.git master (issue #62; 2024-09-13) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 8d33b6a4…34b86ab9 range are covered. There were 30 non-merge commits from 10 authors in t… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-62-2024-09-13/2891 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #62; 2024-09-13) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #62; 2024-09-13) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 8d33b6a4…34b86ab9 range are covered. There were 30 non-merge commits from 10 authors in that period. Some notable commits: A new test for restoring pre-created backups without using Scylla Manager is now available via the nodetool refresh command. The reuse cluster flow is now supported on the Azure backend. Since the tablets feature is disabled by default in Alternator, the create table method now supports enabling tablets explicitly via a parameter. A new method decorator allows users to apply specific contexts, such as decreasing the severity of events, based on a specified Scylla issue using the SkipPerIssues mechanism. When cleaning instances with cloud cleanup scripts (periodical ones outside the test scope), Argus resources information is updated. The latency_calculator_decorator, which tracks latency and duration, can now be used dynamically, even outside the Nemesis scope. Basic test details (timestamp, job name, build URL/number, status) are sent to Elasticsearch for future statistical analysis. The latency decorator now sends screenshots for each cycle to Argus, which can be accessed in results tables. Scylla Manager installation now supports patch version selection, allowing users to specify any available patch version (e.g., 3.2.1). This feature extends coverage of Scylla Manager upgrade tests. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/release-scylla-manager-3-3-3/2896 Title: [RELEASE] Scylla Manager 3.3.3 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces the release of ScyllaDB Manager 3.3.3, a production-ready patch release of the stable 3.3 branch. ScyllaDB Manager is a centralized cluster administration and recurrent tasks automation tool. … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-manager-3-3-3/2896 ## Headings Structure: H1: [RELEASE] Scylla Manager 3.3.3 H3: Related topics ## Main Content: H1: [RELEASE] Scylla Manager 3.3.3 H3: Related topics The ScyllaDB team announces the release of ScyllaDB Manager 3.3.3, a production-ready patch release of the stable 3.3 branch. ScyllaDB Manager is a centralized cluster administration and recurrent tasks automation tool. ScyllaDB Manager is available for ScyllaDB Enterprise customers and ScyllaDB Open Source users. With ScyllaDB Open Source, ScyllaDB Manager is limited to 5 nodes. See the ScyllaDB Manager Proprietary Software License Agreement for details. In version 3.3.3, we fixed a regression issue that was introduced in the 3.3.2 release, which affected clusters added to the Scylla Manager with a manager server version below 3.2.6. More details can be found in issue #4028. ScyllaDB Enterprise customers are encouraged to upgrade to ScyllaDB Manager 3.3.3 in coordination with the ScyllaDB support team. The new release includes upgrades of both ScyllaDB Manager Server and Agent. Scylla Manager 3.3.3 supports the following Scylla Enterprise releases: And the following Open Source release (limited to 5 nodes see license 1): You can install and run Scylla Manager on Kubernetes using Scylla Operator. More here. --- ### Page: https://forum.scylladb.com/t/changing-table-compression-from-lz4-to-zstd-to-decrease-disk-usage-chunk-size/2899 Title: Changing table compression from LZ4 to Zstd to decrease disk usage, chunk size - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/changing-table-compression-from-lz4-to-zstd-to-decrease-disk-usage-chunk-size/2899 ## Headings Structure: H1: Changing table compression from LZ4 to Zstd to decrease disk usage, chunk size H3: Related topics ## Main Content: H1: Changing table compression from LZ4 to Zstd to decrease disk usage, chunk size H3: Related topics Originally from the User Slack @Mahdi_Kamali: Do we have Zstd compression in ScyllaDB now? The only page I found in doc that tells about compression options is this. But Zstd is not mentioned Data Definition | ScyllaDB Docs @avi: It’s supported, I’ll fix the documentation @Mahdi_Kamali: Thanks @avi. Do you recommend that I change a table (with a lot of blob contents, byte arrays) compression from LZ4 to Zstd ? My goal is to decrease disk space usage. Is there anything I should consider? After changing this property, Should I run compaction, upgrade SSTables, etc.? @avi: It’s worth to try it. Also consider increasing chunk_size_in_kb (will make reads slower) @Yequ_Sun: I did some test on a sample dataset. With default options the disk space usage is around 104MB. Switching to lz4 it reduces to 75MB. With lz4 and 128kB chunk size it is 16.61MB. @avi: Large chunk size will severely impact reads that miss the cache --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-247-2024-09-15/2900 Title: Last week in scylladb.git master (issue #247; 2024-09-15) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the ab32ce6b45…6d8e9645ce range are covered. There were 185 non-merge commits from 23 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-247-2024-09-15/2900 ## Headings Structure: H1: Last week in scylladb.git master (issue #247; 2024-09-15) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #247; 2024-09-15) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the ab32ce6b45…6d8e9645ce range are covered. There were 185 non-merge commits from 23 authors in that period. Some notable commits: ScyllaDB will now tune the number of allowed open files descriptors for very large nodes. When replaying hints for a node, we now consider whether the node is leaving the cluster. If it is, we send the hints to all replicas to avoid losing the data when the node leaves. Major compaction now supports a new option to only check existing sstables during tombstone garbage collection; this can increase the effectiveness of garbage collection for partitions that are updated frequently. Scrub, split, and upgrade compactions increase the compaction group shares in order to guarantee progress; however this is unnecessary as they now run in the maintenance group. The share increase is therefore removed. Commitlog is now able to store entries larger than half a commitlog segment. This limitation caused problems with large clusters, as cluster metadata could exceed this limit. Large entries are now fragmented and split over multiple segments. A bug in computing whether to flush all memtables was fixed. The system_distributed.view_build_status was moved to the system keyspace and is now managed by Raft. Scrub/validate compactions will now verify checksums for uncompressed sstables. The jmx submodule was removed from the source tree. With nodetool now talking directly to the REST API, it is no longer necessary. JMX is still available as a separate package. The heuristics for purging tombstones during compaction were improved, leading to less tombstone accumulation. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/node-bootstrap-progress/2904 Title: Node Bootstrap Progress - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, I’ve looked around the API and online but can’t find it, is there a way to determine the progress of a node bootstrap? I know I can see the status is joining, but a rough percentage or way to figure out N keyspaces l… Language: en Canonical URL: https://forum.scylladb.com/t/node-bootstrap-progress/2904 ## Headings Structure: H1: Node Bootstrap Progress H3: Related topics ## Main Content: H1: Node Bootstrap Progress H3: Related topics Hi, I’ve looked around the API and online but can’t find it, is there a way to determine the progress of a node bootstrap? I know I can see the status is joining, but a rough percentage or way to figure out N keyspaces left would be a nice estimation. If you are using RBNO (it is the default for some time), you can look in the logs to see which tables are already repaired, which are in progress and which are not started yet. For in-progress repairs, you will see logs like “repaired X out Y range”, which can give you a rough idea of progress. This is not really convenient to do, but it can give you a rough idea. We are working on better ways to track tasks such as bootstrap, have a look at Task manager tasks | ScyllaDB Docs. This is still work in progress, so watch this space. Thanks Botond, I do see the repair messages and always took that as a “bootstrapping is progressing” status, but was hoping for some more details than that. I have ~190 keyspaces on this cluster with a wild range of tables from a few to dozens, so tracking the progress from the logs would require a bit of work. The task manager is good to know if we’re watching something specific though. --- ### Page: https://forum.scylladb.com/t/scylladb-university-live-september-24/2907 Title: ScyllaDB University LIVE - September 24 - University and Training - ScyllaDB Community NoSQL Forum Meta Description: Next week, we’re hosting our next ScyllaDB University LIVE online training event. It’s an instructor-led NoSQL event with live sessions conducted by our lead engineers and architects. You can save your free spot here a… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-university-live-september-24/2907 ## Headings Structure: H1: ScyllaDB University LIVE - September 24 H3: Related topics ## Main Content: H1: ScyllaDB University LIVE - September 24 H3: Related topics Next week, we’re hosting our next ScyllaDB University LIVE online training event. It’s an instructor-led NoSQL event with live sessions conducted by our lead engineers and architects. You can save your free spot here and read more about it in this blog post. Hope to see you there! --- ### Page: https://forum.scylladb.com/t/will-inconsistencies-between-the-base-table-and-the-gsi-table-be-detected/2910 Title: Will inconsistencies between the base table and the GSI table be detected? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I tried deleting some SSTables files of the base table, then restarting scylla and performing repair on only the base table. At this time, I found that new data will be generated in the memtable, and the data is written … Language: en Canonical URL: https://forum.scylladb.com/t/will-inconsistencies-between-the-base-table-and-the-gsi-table-be-detected/2910 ## Headings Structure: H1: Will inconsistencies between the base table and the GSI table be detected? H3: Related topics ## Main Content: H1: Will inconsistencies between the base table and the GSI table be detected? H3: Related topics I tried deleting some SSTables files of the base table, then restarting scylla and performing repair on only the base table. At this time, I found that new data will be generated in the memtable, and the data is written to the GSI table. When the base table data changes, will a copy of the data be sent to the GSI synchronously? This behavior looks identical to the call chain for put_item. database::do_apply(schema_ptr s, → push_view_replica_updates - > do_push_view_replica_updates(schema_ptr s, mutation m, → generate_and_propagate_view_updates(const schema_ptr& base How does nodetool repair relate to database::do_apply? Repair does not use the regular write path, so data written by repair will not go through database::apply(). Repair writes data to sstables directly. If the repaired table has views or indexes attached to it, the sstable will be registered with the view builder. See make_streaming_consumer() in streaming/consumer.cc, in particular the call to register_staging_sstable(). Sstables which are registered with the view builder, will be processed and the view builder will generate view updates for all the writes contained in them. get_missing_rows_from_follower_nodes → flush_rows_in_working_row_buf → flush_rows → repair_writer_impl::create_writer → streaming::make_streaming_consumer → view_update_generator::register_staging_sstable → ??? → database::do_apply(schema_ptr s, @Botond_Denes How is register_staging_sstable propagated to database::do_apply? It is known that when I perform repair on the base table, it will inevitably use the database::do_apply function, as shown in the figure above. View updates generated by the view_builder, will be sent to the relevant view replicas, where they will go through database::apply(). But the repaired data of the repaired table (base table) will not go through database::apply(). There is a problem of non-key attributes being lost when the repair base table propagates data to GSI, see Major bug: Repair will cause index table data loss!. Can you help me take a look? The entire call chain logic is quite complicated, and I haven’t understood this code yet. The call chain is as follows: As you can see, the code-path joins-in with that of the normal write path. They all end up calling mutate_MV() in the end. Hi @denesb , I found that view_updates::update_entry(const partition_key& base_key, const clustering_or_static_row& update, const clustering_or_static_row& existing, gc_clock::time_point now) was called in function view_updates::generate_update. There is a line of statement auto diff = update.cells().difference(*_base, kind, existing.cells()); in generate_update. I found through reproducing problem Major bug: Repair will cause index table data loss! that the diff (actually a row type) here is empty, which results in the cells of the generated GSI being empty. If the data modified from the base table(update) is the same as the data existing in the node(existing), then diff is empty, resulting in the record only having a key and no cells in GSI. Why use update_entry, can’t we just use create_entry directly? I’m out of my depth here. We need help from @nyh or @wmitros. --- ### Page: https://forum.scylladb.com/t/raft-majority-loss-issue-no-raft-quorum-after-node-failure/2911 Title: Raft majority loss issue - no raft quorum after node failure - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/raft-majority-loss-issue-no-raft-quorum-after-node-failure/2911 ## Headings Structure: H1: Raft majority loss issue - no raft quorum after node failure H1: @saranya_R_B: yes all nodes are down in one dc. we are running multi-dc architecture . Datacenter: SSD_datacenter H1: Status=Up/Down |/ State=Normal/Leaving/Joining/Moving – Address Load Tokens Owns Host ID Rack UN 10.92.3.5 94.17 GB 256 ? a86af8b2-7b56-44a9-80b6-c3fb1852ad01 rack_ssd UN 10.92.3.6 93.28 GB 256 ? b4362683-532d-4185-a5b9-e2b8acb0658e rack_ssd UN 10.92.3.7 93.92 GB 256 ? e1284d67-ca46-4e4b-a396-75fdd8142e61 rack_ssd Datacenter: local_ssd_dc H3: Related topics ## Main Content: H1: Raft majority loss issue - no raft quorum after node failure H1: @saranya_R_B: yes all nodes are down in one dc. we are running multi-dc architecture . Datacenter: SSD_datacenter H1: Status=Up/Down |/ State=Normal/Leaving/Joining/Moving – Address Load Tokens Owns Host ID Rack UN 10.92.3.5 94.17 GB 256 ? a86af8b2-7b56-44a9-80b6-c3fb1852ad01 rack_ssd UN 10.92.3.6 93.28 GB 256 ? b4362683-532d-4185-a5b9-e2b8acb0658e rack_ssd UN 10.92.3.7 93.92 GB 256 ? e1284d67-ca46-4e4b-a396-75fdd8142e61 rack_ssd Datacenter: local_ssd_dc H3: Related topics Originally from the User Slack @saranya_R_B: Hello scylla Team, Scylla_version: 6.0 keyspace enabled with tablets We have configured two data centers with three nodes each, but one data center has completely failed and cannot be recovered. We’ve tried using removenode with the --ignore-dead-nodes option and the replace_node_at_first_boot procedure, but neither has been successful in removing or adding new nodes. Instead, we’re encountering the following error. Can anyone help with this? "there is no raft quorum, total voters count 6, alive voters count 3, dead voters f-390e10ddc662] raft operation [read_barrier] timed out, there is no raft quorum, total voters count 6, alive voters count 3, dead voters" @Piotr_Smaroń: cc @Kamil_Braun @Botond_Dénes: What do you mean exactly by “one data center has completely failed”? Did you loose all nodes in that DC? Status=Up/Down |/ State=Normal/Leaving/Joining/Moving – Address Load Tokens Owns Host ID Rack DN 10.92.3.8 81.08 GB 256 ? 7deb42d7-dd1d-4848-97d0-88462cf442ae local_ssd_rack DN 10.92.3.9 82.17 GB 256 ? 26c1a8d3-e30f-496e-8509-d658ad758eda local_ssd_rack DN 10.92.3.10 82.79 GB 256 ? fcde1ff7-05f4-44bd-b3d6-496f7a409ece local_ssd_rack Using local ssd for reads and peristant disk for writes .ALll nodes in local_ssd_dc went down. Removal of node We’ve tried using removenode with the --ignore-dead-nodes option and the replace_node_at_first_boot procedure, but neither has been successful in removing or adding new nodes. Instead, we’re encountering error. @Botond_Dénes: If you have two DCs, then you ran into the problem of raft majority loss. This is why it is advisable to always have odd number DCs, so you always have a majority left. @avi: You need to run the raft recovery procedure to reassemble the cluster @Kamil_Braun: https://opensource.docs.scylladb.com/branch-6.0/troubleshooting/handling-node-failures.html#manual-recovery-procedure Handling Node Failures | ScyllaDB Docs ah actually it won’t work with tablets > The manual recovery procedure is not supported if tablets are enabled on any of your keyspaces. In such a case, you need to restore from backup. Data Distribution with Tablets | ScyllaDB Docs Restore from a Backup and Incremental Backup | ScyllaDB Docs in this case the only way is to setup a new cluster and move the data over, unfortunately --- ### Page: https://forum.scylladb.com/t/release-scylla-doctor-v1-2/2913 Title: [RELEASE]: Scylla Doctor v1.2 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: Scylla Doctor v1.2 is released. Fixes: SD fails to collect info using CQL on hosts that have different values for broadcast_rpc_address and rpc_address. Added new Collectors: GossipInfoCollector TokenMetadataHos… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-doctor-v1-2/2913 ## Headings Structure: H1: [RELEASE]: Scylla Doctor v1.2 H3: Related topics ## Main Content: H1: [RELEASE]: Scylla Doctor v1.2 H3: Related topics Scylla Doctor v1.2 is released. SD fails to collect info using CQL on hosts that have different values for broadcast_rpc_address and rpc_address. Added new Collectors: Added new TopologyConsistencyAnalyzer. Artifacts can be downloaded from https://downloads.scylladb.com/downloads/scylla-doctor/ or installed from Scylla OSS or Scylla Enterprise repositories. --- ### Page: https://forum.scylladb.com/t/querying-an-index-via-alternator-results-in-a-different-behavior-than-dynamodb/2917 Title: Querying an Index via Alternator results in a different behavior than DynamoDB - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: @Joseph_Stroman: The expected behavior is Alternator would return all the items it can and if there are leftovers then provide the last evaluated key (cursor) to continue “paginating” through the table. It looks like t… Language: en Canonical URL: https://forum.scylladb.com/t/querying-an-index-via-alternator-results-in-a-different-behavior-than-dynamodb/2917 ## Headings Structure: H1: Querying an Index via Alternator results in a different behavior than DynamoDB H3: Related topics ## Main Content: H1: Querying an Index via Alternator results in a different behavior than DynamoDB H3: Related topics Originally from the User Slack @Joseph_Stroman: Hello all, I’m running into an issue with Alternator. I recently switch from dynamodb, but am still using the dynamodb sdk to perform all the operations in scylladb. Everything works as expected except sometimes I am getting a LastEvaluatedKey, with an empty array when I’m querying an index. This isn’t the same behavior I was getting with Dynamodb so I’m curious if anyone knows what could be happening. @avi: Please file an issue with a reproducer @Felipe_Cardeneti_Mendes: https://github.com/scylladb/scylladb/issues/20474 GitHub: Alternator - Query/Scan Operations return an empty LastEvaluatedKey field when scanning through many tombstones · Issue #20474 · scylladb/scylladb actually the issue description is misleading, let me fix it. The problem isnt with LastEvalKey, but with an Empty Item array @Joseph_Stroman: The expected behavior is Alternator would return all the items it can and if there are leftovers then provide the last evaluated key (cursor) to continue “paginating” through the table. It looks like the cursor still works, because when I use it, it finally returns the orders but they should be returned initially (without using the cursor) @Felipe_Cardeneti_Mendes: Interesting. I wonder if this has anything to do with tombstones — I would guess yes because the paging infrastructure is the same Plus it’s an index so the likeliness of scanning many tombstones is higher than the base table So … can you check it? Look at a setting called query_tombstone_page_limit — increase it to see if you stop receiving an empty response @Joseph_Stroman: It worked! So I assume I should decrease the gc_grace_seconds for this table? Anything I should be aware of with that? (Im guessing its more resource intensive) @Felipe_Cardeneti_Mendes: Ok, so you just found an implementation detail. Dynamo doesn’t have a tombstone concept as they don’t use LSM — but we do. The idea is to return a quick response back to the client so it doesn’t timeout the request and allow for concurrency of other queries to proceed So I think the default setting is sane, we need just to document this finding. I will open an issue (and maybe send a PR) Tombstones slow down your read path. Many tombstones can even cause timeouts. So the balance is how many tombstones Scylla needs to scan through before it decides to save an empty page back to your client — as a resort to prevent it from timing out (like: I am still processing your query, which is expensive because of these tombstones) Ideally you want to compact them away faster, so either tune compaction settings, or gc_grace or both — tombstone_gc=repair is also an option @Joseph_Stroman: Thanks for the explanation, I’m reading over these docs links to get more familiar with it: https://opensource.docs.scylladb.com/stable/kb/gc-grace-seconds.html https://opensource.docs.scylladb.com/stable/architecture/compaction/compaction-strategies.html#which-strategy-is-best https://opensource.docs.scylladb.com/stable/kb/compaction.html Also this one but not sure if it’s relevant if im using the open source version: https://enterprise.docs.scylladb.com/stable/kb/garbage-collection-ics.html If you know any other useful docs let me know, and thanks again for the response @Felipe_Cardeneti_Mendes: I guess before even reading docs you may want to first see “how bad” it is. When you read many tombstones typically you will see warnings in logs telling you about it. This is controlled by the tombstone_warn_threshold setting and is logged on a per page basis. For example, if most of your reads are fine, but only a few ones are showing this, maybe you dont need to bother If most of the reads are empty paging, then that’s when you would perhaps optimize it a bit in general - these docs work - see also https://opensource.docs.scylladb.com/stable/cql/ddl.html#ddl-tombstones-gc at the end of the day, it is about wisely choosing your battles. eg: Dynamo doesn’t have tombstones, but writes are super expensive. @Joseph_Stroman: Ok so just for reference, we have are storing an “orderbook”, where we expect many entries and in our case each entry will be definitely be updated once, and in rare cases updated twice. I’m seeing two different tombstone related logs for the same “index”: I know there is a difference in how indexing is handled compared to dynamodb, but not sure if that’s affecting this I’m guessing there are two logs because one partition reached the tombstone limit (10000) I set it back to the default so I can figure out the correct strategy, and in my case once it reaches that limit I think it should just be compacted away because we don’t care at all about data that was there before an update I also wouldn’t mind immediate tombstone deletion, just unsure of the consequences of that @Felipe_Cardeneti_Mendes: oh, well … this is probably going to take long for me to advise - even worse via slack Note it isn’t a support channel, but a place for you to openly discuss your use case, ask questions, strategies, and such. https://github.com/scylladb/scylladb/issues/20474 @Joseph_Stroman: Ok I think I found a fix that works for us for now, but if not I will schedule a strategy session. Since we are only using one node and need to iterate fast I think this will suffice: I moved grace period to 1 hr for all tables and indexes (and set default in the yaml file to 1 hr) And set the tombstone page limit to 1000000, which may be overkill but it prevents the client from needing to paginate, and if the tombstones ever accumulate that much we have bigger problems I assume. And I’ll be following the issue to see what you guys recommend, and I can add any details if you need @Felipe_Cardeneti_Mendes: question - are you receiving all items in the initial pages and then nothing in the subsequent ones? Or are you getting an empty page at the beginning until you reach “live data” ? in my short-lived test, for whatever reason I dont yet understand I get all live items, and subsequent empty pages after so I am trying to understand if what I am seeing aligns to what you are reporting. Nvm, I figured it out. @Joseph_Stroman: cool cool, but just for clarity yes it was the first thing, initial pages were empty (i assume because there were tombstones at the top) also sorting by creation time and descending made it return what we wanted which makes sense we just have indexes that sort on other keys so we couldn’t just do that as a fix --- ### Page: https://forum.scylladb.com/t/release-scylla-operator-v1-14-0/2919 Title: [RELEASE] Scylla Operator v1.14.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of Scylla Operator 1.14.0. Scylla Operator is an open-source project that helps ScyllaDB Open Source and ScyllaDB Enterprise users run ScyllaDB on Kubernetes. The Scy… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-operator-v1-14-0/2919 ## Headings Structure: H1: [RELEASE] Scylla Operator v1.14.0 H1: Notable changes H1: Supported versions H1: Upgrade instructions H2: Getting started with Scylla Operator H2: Related Links H3: Related topics ## Main Content: H1: [RELEASE] Scylla Operator v1.14.0 H1: Notable changes H1: Supported versions H1: Upgrade instructions H2: Getting started with Scylla Operator H2: Related Links H3: Related topics The ScyllaDB team is pleased to announce the release of Scylla Operator 1.14.0. Scylla Operator is an open-source project that helps ScyllaDB Open Source and ScyllaDB Enterprise users run ScyllaDB on Kubernetes. The Scylla Operator manages ScyllaDB clusters deployed to Kubernetes and automates tasks related to operating a ScyllaDB cluster, like installation, vertical and horizontal scaling, as well as rolling upgrades. Scylla Operator 1.14.0 improves stability and brings new features. As with all of our releases, all API changes are backward compatible. For more changes and details check out the GitHub release notes. Upgrading from v1.13.x with kubectl apply doesn’t require any extra action, just take the manifest from v1.14.0 tag and substitute the released image. Using helm requires a mandatory manual step for every release because helm can’t handle CRDs updates. For details, see our upgrade documentation. We’ll welcome your feedback! Feel free to open an issue or reach out on the #scylla-operator channel in Scylla User Slack. --- ### Page: https://forum.scylladb.com/t/update-network-usage-data-issue-in-scylladb-cloud-usage-dashboard-beta/2920 Title: Update: Network Usage Data Issue in ScyllaDB Cloud Usage Dashboard (Beta) - Announcements - ScyllaDB Community NoSQL Forum Meta Description: Dear ScyllaDB Cloud Users, Following the release of the ScyllaDB Cloud Usage Dashboard (Beta), we would like to inform you of a known issue where Network Usage data is currently not being displayed for periods starting … Language: en Canonical URL: https://forum.scylladb.com/t/update-network-usage-data-issue-in-scylladb-cloud-usage-dashboard-beta/2920 ## Headings Structure: H1: Update: Network Usage Data Issue in ScyllaDB Cloud Usage Dashboard (Beta) H3: Related topics ## Main Content: H1: Update: Network Usage Data Issue in ScyllaDB Cloud Usage Dashboard (Beta) H3: Related topics Dear ScyllaDB Cloud Users, Following the release of the ScyllaDB Cloud Usage Dashboard (Beta), we would like to inform you of a known issue where Network Usage data is currently not being displayed for periods starting from September 1, 2024. Our team is actively working on a fix, and we expect this issue to be resolved by September 26, 2024. We apologize for any inconvenience this may cause and appreciate your understanding. Thank you for your continued support! Best regards, The ScyllaDB Team --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-63-2024-09-20/2923 Title: Last week in scylla-cluster-tests.git master (issue #63; 2024-09-20) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 92febe89…447dff1e range are covered. There were 27 non-merge commits from 8 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-63-2024-09-20/2923 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #63; 2024-09-20) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #63; 2024-09-20) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 92febe89…447dff1e range are covered. There were 27 non-merge commits from 8 authors in that period. Some notable commits: Adjusting the replication factor for the system_auth keyspace is skipped when auth-v2 is enabled. The latency decorator used in performance tests now also reports throughput stats, based on HDR files returned by cassandra-stress. These metrics are reported in Argus. Before terminating a node during tests (e.g., decommissioning), SCT collects all its logs, similar to the collect-log CI stage. These logs are listed separately in Argus/show-logs, aiding in debugging terminated nodes. The reuse cluster feature, which enables faster retry cycles, now has its own documentation, including configuration steps and usage examples. Upgrade tests for different operating systems were sped up by limiting stress and focusing on ScyllaDB artifacts. Fully featured upgrade tests with high stress are reserved for crafted cloud images. The result of latency_decorator_calculator can now be named dynamically, instead of using the function name. This enables storing results for different use cases of the same function (nemesis). A new ‘gradual’ performance test was introduced, replacing both the simple latency and max throughput tests with less flakiness. It runs steps with varying fixed throughput, collecting latency based on HDR files, with the final step untrottled to determine max throughput. Results are reported in Argus and via email. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/seastar-assertion-thread-failed/2926 Title: Seastar - assertion `thread' failed - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello everyone, The question is more about seastar programming but since it connected to ScyllaDB development would ask here. Could please somebody explain me the situation: when I use construction similar to this sea… Language: en Canonical URL: https://forum.scylladb.com/t/seastar-assertion-thread-failed/2926 ## Headings Structure: H1: Seastar - assertion `thread' failed H3: Related topics ## Main Content: H1: Seastar - assertion `thread' failed H3: Related topics The question is more about seastar programming but since it connected to ScyllaDB development would ask here. Could please somebody explain me the situation: when I use construction similar to this seastar::future<> foo3() { return seastar::make_ready_future<>(); } seastar::future<> foo1() { return seastar::make_ready_future<>(); } seastar::future<> foo2() { return foo3().then( { }); } seastar::future<> work() { return foo1().then( { foo2().get(); //… at runtime I get an assertion and the process ends “future.cc:248: void seastar::internal::future_base::do_wait(): Assertion `thread’ failed” I suspect the source of problem but can’t guess how to resolve it. Tried to wrap functions to seastar::async but it doesn’t work. Would appreciate any help. Thank you. You can also try on the seastar group https://groups.google.com/g/seastar-dev seastar::future<> work() { return foo1().then( { foo2().get(); //… You cannot use get() outside of a seastar thread context. A seastar thread can be created by: --- ### Page: https://forum.scylladb.com/t/rebuilt-node-missing-from-raft-state-not-starting/2927 Title: Rebuilt Node Missing from Raft State, Not Starting - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Ok, so it’s been a week, I’m going to give the full timeline of events, in case something that I don’t think is relevant actually is. A node (let’s call it Node A) of our cluster filled disk space 100% (we missed alarm… Language: en Canonical URL: https://forum.scylladb.com/t/rebuilt-node-missing-from-raft-state-not-starting/2927 ## Headings Structure: H1: Rebuilt Node Missing from Raft State, Not Starting H3: Related topics ## Main Content: H1: Rebuilt Node Missing from Raft State, Not Starting H3: Related topics Ok, so it’s been a week, I’m going to give the full timeline of events, in case something that I don’t think is relevant actually is. That brought me to this afternoon and reading through this documentation on failed membership changes. I ran both of the queries on a good node to get the cluster state, both returned 10 rows (which matches our node count currently). I saw the below state: 49deb87f-adbf-4524-9728-9d49aa58e36b was only showing up in the raft_state and was not in nodetool status or nodetool gossipinfo, and only referenced in some log warnings about not being able to hit it. So, it matched the definition of a ghost member, so I ran a nodetool removenode on it, which quickly completed and removed it from the above table. I gave Scylla on Node A a restart a while afterwards, as nothing on it changed and there were no logs coming out of it. It’s back to where it was before, just logging a ton about compactions, but still stuck at starting system distributed keyspace. So, the current state: Is there some way to get Node A back into Raft? I rebuilt Node A again, supplying 36593f45-cba6-4368-a287-b5e692b8bb7e as the id, and had no hiccups during the bootstraps. I had identified that Node B had a regression of our configuration and had enable_node_aggregated_table_metrics enabled still, where it will cause massive memory contention for us, since these nodes have tons of keyspaces/tables. I’d still be curious if there is a quicker way to resolve this than waiting another day for a rebuild, but it’s not urgent anymore. What version of Scylla are you using? This cluster is using version 5.4.9 What is the host ID of node A? It prints it during startup. By “rebuilt” do you mean that you bootstrapped the node from scratch, performing a replace operation? Was 49deb87f-adbf-4524-9728-9d49aa58e36b perhaps the host ID of node A at some point? In general, it appears you were using replace-with-same-IP operation, as I understand from your description you continued to use the same AWS instance where previous incarnation of node A was, so it was the same IP. Unfortunately replace-with-same-IP is full of problems, and I recommend against using it in the future. If you lose disk from an instance, either drop that instance (and start a new one) or if it’s possible, restart it with different IP (I’m not familiar with AWS enough to know if changing instance IP is possible like that). You can learn more details from this issue: Failure during replace-with-same-IP leaves the node without `STATUS` application_state (permanently), and `token_metadata` inconsistent (until restart) (applies to gossiper / "node-ops" based topology changes) · Issue #19975 · scylladb/scylladb · GitHub @kbr, The core steps in the op are essentially the Setup RAID Following a Restart instructions from the Scylla docs. After losing the ephemeral storage of an i3 instance, it says to recreate the RAID volume and use replace_node_first_boot with the old Host ID. I was aware that using IP addresses to replace nodes was discouraged, and host IDs were much more reliable. But are you saying that the “Setup RAID Following a Restart” process is still replace-with-same-ip because the IP is unchanged, in spite of using Host IDs? Yes @prtolo is correct, we used the replace_node_first_boot option and have for a while. I think the replace by IP is actually disabled and yells at you to use the ID if attempted (we found that out when it stopped working). I do believe 49deb87f-adbf-4524-9728-9d49aa58e36b was an old ID of Node A, and so that was why I tried removing it, hoping the new ID would join the raft in it’s stead. I was aware that using IP addresses to replace nodes was discouraged, and host IDs were much more reliable. But are you saying that the “Setup RAID Following a Restart” process is still replace-with-same-ip because the IP is unchanged, in spite of using Host IDs? Whether you use host ID or IP to define the node you’re replacing in configuration matters little here. Internally Scylla will translate between them. The issue I linked to happens when the IP address of the node that is being replaced, and the node that is replacing it, is the same. And in 5.4 and earlier Scylla releases, and corresponding Enterprise releases, unfortunately, this operation has some quirks and handles failures non-gracefully (cluster is put into a weird state that may require deep knowledge and manual steps to recover from). The “Setup RAID Following a Restart” guide is precisely a guide to perform a replace operation of a node with the same IP address. I think we should update the documentation to recommend against it, or include a step to change the IP address of the instance which lost its disk (if it is possible; if not, use a new instance). same. And in 5.4 and earlier Scylla releases, and corresponding Enterprise releases, unfortunately, this CC @Anna with regards to documentation filed docs: recommend against replacing node with the same IP addresses · Issue #23884 · scylladb/scylladb · GitHub --- ### Page: https://forum.scylladb.com/t/operator-memory-management-issue-when-ingesting-data-writes-blocked-on-dirty/2939 Title: Operator - memory management issue when ingesting data, Writes blocked on dirty - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/operator-memory-management-issue-when-ingesting-data-writes-blocked-on-dirty/2939 ## Headings Structure: H1: Operator - memory management issue when ingesting data, Writes blocked on dirty H3: Related topics ## Main Content: H1: Operator - memory management issue when ingesting data, Writes blocked on dirty H3: Related topics Originally from the User Slack @Łukasz_Sanokowski: Hello! I have question about memory management, namely during data ingestion we see a large number of Writes blocked on dirty , despite that there is plenty of free ram. Stack consist of latest Scylla operator running on recent GKE. The average load is around 50%: (please ignore the fact that we do lack node exporter metrics here, screenshot with CPU / RAM is pasted below) Config of running scylla process: And the memory utilisation: We suppose that because of this issue(?), the performance of data ingestion is affected. @Maciej_Zimnoch: make sure to follow performance tuning documentation, you’re not getting most out of your cluster. https://operator.docs.scylladb.com/stable/performance.html Performance tuning | ScyllaDB Docs @Łukasz_Sanokowski: Thanks @Maciej_Zimnoch, you are right: our pods are not QoS Guaranteed class. Quick question: after editing the cluster.yaml resource requests and limits (namely: adding agentResources section), and applying the changes, we can see that scylla operator itself is displaying the correct value: namely here 4CPU / 16 GB of ram && 1CPU and 1GB of ram, but the changes are not being propagated to the underlying statefulset: where scylla container limits are still: @Maciej_Zimnoch: please collect must-gather dump and attach it here so i can have a look at entire picture Gathering data with must-gather | ScyllaDB Docs @Łukasz_Sanokowski: Sure thing: @Maciej_Zimnoch: it’s because we apply statefulset changes only once they are fully rolled out. Your first pod in -a rack is missing @Łukasz_Sanokowski: The pod is missing since the attempt of setting: which eventually has resulted with lack of ram available on the node, so the second attempt was to reduce the: which is failing because (likely) we apply statefulset changes only once they are fully rolled out Shall I edit the statefulset directly, to make it healthy and operator to accept and propagate changes? Update: it helped Alright Maciej, i think that for now I know where and how to proceed, thanks for your help! @Maciej_Zimnoch: i’m not sure your initial issue with memory will be solved, but lets see if it helped. @Łukasz_Sanokowski: Sure, let me tweak around for a while, will come back with the results --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-248-2024-09-22/2940 Title: Last week in scylladb.git master (issue #248; 2024-09-22) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 6d8e9645ce…3d781c4fc8 range are covered. There were 119 non-merge commits from 18 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-248-2024-09-22/2940 ## Headings Structure: H1: Last week in scylladb.git master (issue #248; 2024-09-22) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #248; 2024-09-22) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 6d8e9645ce…3d781c4fc8 range are covered. There were 119 non-merge commits from 18 authors in that period. Some notable commits: A CQL filtering bug when a regular column was filtered but no regular columns were selected was fixed. A recently-introduced memory leak in the Paxos implementation was fixed. Compaction CLEANUP jobs now run under the maintenance/streaming scheduling/group. The version in the master branch was updated to 6.3, indicating the beginning of the 6.2 stabilization cycle. The test suite now collects resource consumption metrics and stores them in a local database. This will be used in the future to improve resource utilization. The system will now negotiate the ability to run cross-node queries under the maintenance scheduling group. The backup API now supports a prefix parameter, allowing more fine-grained of where the backup is stored. Some bugs were fixed in Alternator role-based access control. The diagnostics dump from reader_concurrency_semaphore was improved. Lightweight transaction query processing will move to the shard that owns the data, if it is not already there. It now recognizes that the owning shard could have moved again, due to tablet migration, and retries. It is now possible to reroute an individual statement to a different service level; previously a different login session was required. This is useful for drivers to reduce the strain from the login queries. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/scylladb-image-does-not-exist-in-aws-china-regions/2941 Title: Scylladb image does not exist in aws China regions - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, I want to deploy scylladb cluster in AWS China regions (which are isolated with aws global regions), but I don’t see scylladb images in EC2 image catalog in BJS and NingXia regions. Do we have plan to provide pre-con… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-image-does-not-exist-in-aws-china-regions/2941 ## Headings Structure: H1: Scylladb image does not exist in aws China regions H3: Related topics ## Main Content: H1: Scylladb image does not exist in aws China regions H3: Related topics Hi, I want to deploy scylladb cluster in AWS China regions (which are isolated with aws global regions), but I don’t see scylladb images in EC2 image catalog in BJS and NingXia regions. Do we have plan to provide pre-configured image in aws Ningxia and BJS regions? Unfortunately, there is no ETA for ScyllaDB to create images in AWS China. Reasone is the regulatory differences between global AWS and AWS China. We will get to it. --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-10/2946 Title: [RELEASE] ScyllaDB Enterprise 2024.1.10 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2024.1.10, a bug-fix production-ready patch release for ScyllaDB Enterprise 2024.1 LTS Release. Related Links Get ScyllaDB Enterprise 2024.1 (customers only, or 30-day e… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-10/2946 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2024.1.10 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2024.1.10 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2024.1.10, a bug-fix production-ready patch release for ScyllaDB Enterprise 2024.1 LTS Release. The following issues are fixed in this release (with an open-source reference, if available): --- ### Page: https://forum.scylladb.com/t/how-do-many-small-partitions-influence-memory-usage-in-scylladb/2948 Title: How Do Many Small Partitions Influence Memory Usage in ScyllaDB? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello everyone, In ScyllaDB, does having a large number of small partitions significantly increase Bloom filter memory consumption? I found a post discussing this issue in Cassandra, as shown in the following link: H… Language: en Canonical URL: https://forum.scylladb.com/t/how-do-many-small-partitions-influence-memory-usage-in-scylladb/2948 ## Headings Structure: H1: How Do Many Small Partitions Influence Memory Usage in ScyllaDB? H3: Advanced Primary Key Selection H3: Large Partitions Support in ScyllaDB 2.3 and Beyond H3: ScyllaDB Large Partitions Table | ScyllaDB Docs H3: Related topics ## Main Content: H1: How Do Many Small Partitions Influence Memory Usage in ScyllaDB? H4: Is it a bad practice to have a Cassandra table with partitions of a single row? H3: Advanced Primary Key Selection H3: Large Partitions Support in ScyllaDB 2.3 and Beyond H3: ScyllaDB Large Partitions Table | ScyllaDB Docs H4: Wasteful storage of table with very short partitions H4: [Epic] Bloom Filter H3: Related topics In ScyllaDB, does having a large number of small partitions significantly increase Bloom filter memory consumption? I found a post discussing this issue in Cassandra, as shown in the following link: However, I couldn’t find similar information specific to ScyllaDB. On the other hand, ScyllaDB strongly recommends creating small partitions to avoid hotspots. Would it be acceptable to reduce the number of partitions and make them larger, as long as hotspots can still be avoided? Thank you in advance! It seems I was confusing the primary key and partition key. While the partition key should ideally have high cardinality, the primary key itself does not need to have high cardinality. ScyllaDB University recommends designing partitions to be neither too small nor too large. 8 min to complete In the previous lesson, we learned about partition and clustering keys and that each one of them can be composed of multiple columns. Now that we understand the concepts, how do we choose a suitable partition key? To recap, in the... The maximum number of rows per partition is not universally determined, but varies depending on the actual workload and queries. https://groups.google.com/g/scylladb-users/c/I_qHdQV5u1Q/m/BIjBveceCQAJ The official blog below suggests keeping the partition size below 100MB. Currently, it seems that warning logs are generated for sizes above 1000MB, as the limit has been raised. ScyllaDB release 2.3 allows the discovery and investigation of large partitions present in your cluster -- system.large_partitions table. ScyllaDB is an Apache Cassandra-compatible NoSQL data store that can handle 1 million transactions per second on a single server. It appears that having a very large number of small partitions significantly increases the memory consumption of Bloom filters, which is true not only for Cassandra but also for ScyllaDB. https://groups.google.com/g/scylladb-users/c/3d-GhCl6x3U/m/V_YUpv3vAwAJ A user complained that a table with a huge number of very short partitions was s…urprisingly big - perhaps as much as 4 times larger than if the same data is stored in a modest number of large partitions. Let's look at a simple example. Consider two tables, each of them has three integer columns, `p, c, x`: * In table1, `p` is the partition key, `c` is the clustering key, `x` is a regular column. * In table2, `(p,c)` is a compound partition key, `x` is a regular column. We write a million rows to each table, with `(p, c, x) = (1, i, 1)` for one million i's. So both tables have exactly the same data - table1 is one partition with one million rows, and table2 in one million partitions, but both have exactly the same rows. It turns out that after compaction, the size of table1's sstables is 8.2 MB, and the size of table2 is 34 MB. table2 is more than 4 times larger than table1! To understand why, let's look at the size of the individual sstable components: 1. For table1, almost the entire 8.2 MB size is the "Data" component. The "Index" component is almost empty (just one partition). 2. For table2, the "Data" component is 12.6 MB, the "Index" component is 20MB. The "Filter" is 1.2 MB. It is not surprising that table2's Data component is slightly larger than table1's (12.6 MB vs 8.2 MB) - after all the individual partitions do have some overhead (e.g., a tombstone), and this overhead is noticable when the partitions are so tiny (just a single integer). It's also not surprising that the Bloom filter (the "Filter" file) takes more space when we have many partitions. But what is really surprising, and really frustrating is the size of the Index file, which is almost twice bigger than the Data file, which we didn't expect. The idea I want to propose in this issue is that when partitions are very short, it would be better not to have an Index file at all. The Summary file could point directly into the Data file instead of the Index file. I don't know what should be the threshold for dropping the Index file - for larger partitions, it may still be good. Maybe we can write the Data and Index file as we do today, and after-the-fact, if we notice that Index is larger than Data (or even if it is larger than half of Data), we delete the Index component and rewrite the Summary. Instead of dropping the Index file, and even more efficient thing to do can be to create in Index a level between the Data and Summary - in other words, Index will be a sample of Data's keys - not every paritition as in Data but not a sparse a sample as Summary, but something in the middle. But this will require more work to implement. This use case, of very short partitions, may seem artificial, but a real user encountered it with a materialized view - the user had a base table with reasonably-long partitions, but then had a view with a compound partitions, where each row of data was in its own partition - and each row was also very short. The user didn't even realize that this very-short-partitions case was happening, but was surprised that the view was 4 times larger than the base table. Code for the tests described above: ```python @pytest.fixture(scope="function") def table1(cql, test_keyspace): t = f'{test_keyspace}.{unique_name()}' cql.execute(f'CREATE TABLE {t}(p int, c int, x int, PRIMARY KEY (p, c))') yield t cql.execute(f'DROP TABLE {t}') @pytest.fixture(scope="function") def table2(cql, test_keyspace): t = f'{test_keyspace}.{unique_name()}' cql.execute(f'CREATE TABLE {t}(p int, c int, x int, PRIMARY KEY ((p, c)))') yield t cql.execute(f'DROP TABLE {t}') # Table with a single 1-million-row partition, each row has, in addition # to the key, a single int. def test_single_partition(cql, table1): table = table1 stmt = cql.prepare(f"INSERT INTO {table} (p, c, x) VALUES (1, ?, 1)") for i in range(1000000): cql.execute(stmt, [i]) if (i%100000)==0: print(i) nodetool.flush(cql, table) nodetool.compact(cql, table) nodetool.flush(cql, table) nodetool.compact(cql, table) print('going to sleep\n') time.sleep(10000) # Table with a single 1-million partitions, each with a single row # (no clustering key), and each row has additionally an int value def test_million_partitions(cql, table2): table = table2 stmt = cql.prepare(f"INSERT INTO {table} (p, c, x) VALUES (1, ?, 1)") for i in range(1000000): cql.execute(stmt, [i]) if (i%100000)==0: print(i) nodetool.flush(cql, table) nodetool.compact(cql, table) nodetool.flush(cql, table) nodetool.compact(cql, table) print('going to sleep\n') time.sleep(10000) ``` Regarding the issue of Bloom filter size, various efforts are still ongoing. Bloom filters allow for quick and cheap checks for whether a partition exists in… an sstable or not, without doing any I/O. Bloom filters are allowed to provide false-positive answers (check says that the partition is in the sstable, but its not), but it can never provide false-negative answers. They are created based on two parameters: partition count (often an estimate) and require false-positive chance (probability of false-positive results). To satisfy these two requirements, a filter of a certain size is required. The filter size is proportional to the partition count and inversely proportional to the false-positive chance. Too large or too small filters are problematic. Too small filters will result in too many false-positives, increasing read latency and read memory consumption (extra I/O). Too large filters consume a lot of extra memory (filters are always kept in memory). Mis-sized filters are a result of bad partition estimates: our estimate of how many partitions are in an sstable is either an over or under estimate. In the extreme case, either severe under or over estimates can cause OOM. Under estimates because a lot of sstables are opened for reads (this is rare) and over estimates because the filters end up using up all memory of the shard. Design Document: https://docs.google.com/document/d/1LHuDCF2YbTsXBLG3l6KH5RZVJmq_Y9GrVj3zvPHUiKc/edit?usp=sharing This issue is an epic, covering the effort of resolving all our current issues with bloom filters. List of issues currently part of this epic: ```[tasklist] ### Tasks - [ ] https://github.com/scylladb/scylladb/issues/17747 - [ ] https://github.com/scylladb/scylladb/pull/18141 - [ ] https://github.com/scylladb/scylladb/issues/18398 - [ ] https://github.com/scylladb/scylladb/issues/18283 - [ ] https://github.com/scylladb/scylladb/issues/18607 - [ ] https://github.com/scylladb/scylladb/issues/19049 - [ ] https://github.com/scylladb/scylladb/issues/2024 ``` Based on this information, my conclusions are as follows: Since my original question has been resolved, I consider this issue closed, but if there are any inaccuracies in the information, please let me know. --- ### Page: https://forum.scylladb.com/t/running-with-operator-on-gke-cluster-wont-start-with-the-guaranteed-qos-class/2949 Title: Running with Operator on GKE, cluster won't start with the Guaranteed QoS class - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/running-with-operator-on-gke-cluster-wont-start-with-the-guaranteed-qos-class/2949 ## Headings Structure: H1: Running with Operator on GKE, cluster won't start with the Guaranteed QoS class H3: Related topics ## Main Content: H1: Running with Operator on GKE, cluster won't start with the Guaranteed QoS class H3: Related topics Originally from the User Slack @Łukasz_Sanokowski: I have a problem with Scylla spinned up via Operator in latest version 1.13.0 on GKE. It won’t start with Guaranteed QoS class, while it works on exactly the same stack, but on Burstable class. The log of scylla pod running in Guranteed class is following (thats whole, it is getting stuck there): After the simple change of disabling guaranteed class by commenting out: exactly the same setup works, scylla pod is loading successfully as the other containers are. @Maciej_Zimnoch: Is your kubelet cpuManagerPolicy set to static? @Łukasz_Sanokowski: You are right @Maciej_Zimnoch, setting it to static did the job, thank you --- ### Page: https://forum.scylladb.com/t/release-scylladb-6-1-2/2951 Title: [RELEASE] ScyllaDB 6.1.2 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 6.1.2, a bugfix patch release of the ScyllaDB 6.1 stable branch. ScyllaDB Open Source 6.1.2, like all past and future 6.x.y releases, is backward compatible and supports r… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-6-1-2/2951 ## Headings Structure: H1: [RELEASE] ScyllaDB 6.1.2 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 6.1.2 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 6.1.2, a bugfix patch release of the ScyllaDB 6.1 stable branch. ScyllaDB Open Source 6.1.2, like all past and future 6.x.y releases, is backward compatible and supports rolling upgrades. Users are encouraged to upgrade to 6.1.2. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/where-can-i-find-a-docker-compose-for-scylladb-with-spark-analytical-workloads/2957 Title: Where can I find a docker-compose for ScyllaDB with Spark - Analytical workloads? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/where-can-i-find-a-docker-compose-for-scylladb-with-spark-analytical-workloads/2957 ## Headings Structure: H1: Where can I find a docker-compose for ScyllaDB with Spark - Analytical workloads? H3: Related topics ## Main Content: H1: Where can I find a docker-compose for ScyllaDB with Spark - Analytical workloads? H3: Related topics Originally from the User Slack @IO: anyone has a good docker-compose file for running opensource scylla (latest stable) with spark for an analitical POC exercise before I start compiling one myself? @Felipe_Cardeneti_Mendes: good starting point: https://migrator.docs.scylladb.com/stable/tutorials/dynamodb-to-scylladb-alternator/index.html#set-up-the-services-and-populate-the-source-database Originally for the ScyllaDB Migrator. Dockerfiles & other stuff under https://github.com/scylladb/scylla-migrator --- ### Page: https://forum.scylladb.com/t/compaction-storm-slows-down-scylla/2958 Title: Compaction Storm slows down Scylla - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Today we observed a weird behavior of ScyllaDB: All of a sudden, the cluster became noticeable slow. Logs were not very interesting, but showed alot of compactions going on on some nodes of the cluster. Looking into the… Language: en Canonical URL: https://forum.scylladb.com/t/compaction-storm-slows-down-scylla/2958 ## Headings Structure: H1: Compaction Storm slows down Scylla H3: Related topics ## Main Content: H1: Compaction Storm slows down Scylla H4: scylla_commitlog_memory_buffer_bytes ever growing since 6.0 H3: Related topics Today we observed a weird behavior of ScyllaDB: All of a sudden, the cluster became noticeable slow. Logs were not very interesting, but showed alot of compactions going on on some nodes of the cluster. Looking into the metrics in our prometheus/grafana, it seems to show that more and more nodes suffered from these intense compactions (going at IOPS-Limit of the underlying disks). A little excerpt from the Logs (going on like this for hours): The scylla-version running is 6.0.2-0.20240703.c9cd171f426e-1. For some reasons the memtable-I/O started to freak out on a few nodes and the memtables-shares went up. Our application behavior has not changed and we have the same load as usual. Restarting the whole cluster (every node, one by one) seems to have temporarily solved the issue, but we’re simply unsure about the root-cause and how to get there. Maybe someone also seen this pattern appearing or has an idea what we’re seeing here? Adding some screenshots of the Scylla-Advanced Dashboard. Might this issue be caused by #16514 ? Stability: Off-strategy compaction is used to make sstables conform to the compaction strategy after an operation such as repair. Off-strategy compaction for TWCS will now have less storage space overhead. [#16514] Could this trigger anything, that might result in this weird memtable-compaction above? Sep 26 16:22:30 o-p-L3-3 scylla[1381236]: [shard 5:stmt] lsa - Standard allocator failure, increasing head-room in section 0x604006144670 to 2048 [B]; trace: 0x645df8e 0x645e5a0 0x645e888 0x215033f 0x20f504e 0x20b6c0c 0x20b7000 0x1caeb23 0x1cadc83 0x1cab21d 0x1c70767 0x1c6b81d 0x1b8c108 0x47dde0a 0x144ccda 0x5f56c3f 0x5f57f27 0x5f7be90 0x5f1734a /opt/scylladb/libreloc/libc.so.6+0x8c946 /opt/scylladb/libreloc/libc.so.6+0x11296f -------- seastar::internal::coroutine_traits_base::promise_type Sep 26 16:22:30 o-p-L3-3 scylla[1381236]: [shard 5:stmt] lsa - Standard allocator failure, increasing head-room in section 0x604006144670 to 4096 [B]; trace: 0x645df8e 0x645e5a0 0x645e888 0x215033f 0x20f504e 0x20b6c0c 0x20b7000 0x1caeb23 0x1cadc83 0x1cab21d 0x1c70767 0x1c6b81d 0x1b8c108 0x47dde0a 0x144ccda 0x5f56c3f 0x5f57f27 0x5f7be90 0x5f1734a /opt/scylladb/libreloc/libc.so.6+0x8c946 /opt/scylladb/libreloc/libc.so.6+0x11296f -------- seastar::internal::coroutine_traits_base::promise_type Sep 26 16:22:30 o-p-L3-3 scylla[1381236]: [shard 5:stmt] lsa - Standard allocator failure, increasing head-room in section 0x604006144670 to 8192 [B]; trace: 0x645df8e 0x645e5a0 0x645e888 0x215033f 0x20f504e 0x20b6c0c 0x20b7000 0x1caeb23 0x1cadc83 0x1cab21d 0x1c70767 0x1c6b81d 0x1b8c108 0x47dde0a 0x144ccda 0x5f56c3f 0x5f57f27 0x5f7be90 0x5f1734a /opt/scylladb/libreloc/libc.so.6+0x8c946 /opt/scylladb/libreloc/libc.so.6+0x11296f -------- seastar::internal::coroutine_traits_base::promise_type Sep 26 16:22:30 o-p-L3-3 scylla[1381236]: [shard 5:stmt] lsa - Standard allocator failure, increasing head-room in section 0x604006144670 to 16384 [B]; trace: 0x645df8e 0x645e5a0 0x645e888 0x215033f 0x20f504e 0x20b6c0c 0x20b7000 0x1caeb23 0x1cadc83 0x1cab21d 0x1c70767 0x1c6b81d 0x1b8c108 0x47dde0a 0x144ccda 0x5f56c3f 0x5f57f27 0x5f7be90 0x5f1734a /opt/scylladb/libreloc/libc.so.6+0x8c946 /opt/scylladb/libreloc/libc.so.6+0x11296f Interestingly, during the “Compaction Storm” we see alot of very small memtable flushes and compactions. An even better screenshot “scylla_memtables_pending_flushes_bytes > 0”: During the times of massive compactions occurring, it showed many small memtable flushes. Question is: What can suddenly cause many small flushes to occur? I think we had a problem with repair flushing at the beginning. Does time where the problem start match the start or a repair? I was thinking about repairs triggering the flushes too, but it does not look like it (see graphs below). But we have found something weird, that I think must be either a bug in the metrics, or perhaps some kinf of memory leak: This is Scylla's bug tracker, to be used for reporting bugs only. If you have a… question about Scylla, and not a bug, please ask it in our mailing-list at scylladb-dev@googlegroups.com or in our slack channel. - [x] I have read the disclaimer above, and I am reporting a suspected malfunction in Scylla. *Installation details* Scylla version (or git commit hash): 6.0.2 Cluster size: Same issue with 6 node or 20 node cluster OS (RHEL/CentOS/Ubuntu/AWS AMI): Ubuntu 22.04 The metric "scylla_commitlog_memory_buffer_bytes" seems to be ever growing since we have updated to 6.0.2: ![image](https://github.com/user-attachments/assets/b4121b84-6b36-40c9-9453-261a3b07a365) 9/22 is the day the cluster was upgraded from 5.4 to 6.0. The times when the lines drop to 0 is when the cluster got restarted. This is the case on various different landscapes. The metric is defined as `sm::make_gauge("memory_buffer_bytes", totals.buffer_list_bytes, sm::description("Holds the total number of bytes in internal memory buffers.")),` Regarding the github issue: It seems the scylla_commitlog_memory_buffer_bytes is constantly going up since we updated to 6.0. Perhaps if that fillls up , it keeps flushing? It’s probably related. elcallio says that the commitlog-buffer-bytes cannot be responsible for any flushes. I wonder what else can cause lots and lots of tiny flushes: Please share commitlog metrics. Also: does nodetool flush help? Also we are quite a lot of memory related warnings: My previous comment is still waiting for approval. One metrics I find particular interesting is: Another idea: Can the schema-commitlog cause flushes of our data tables? For some reason the scylla_schema_commitlog_active_allocations was pretty high after the 6.0 upgrade. On 25th when we did a rolling restart it went down a lot and stayed like that ever since: Is it possible that some action was still pending from the upgrade, that caused many schema commitlog activity? I do not think that nodetool flush helps, as it seems it was flushing like crazy anyway. Also I am 99% sure we did some flushes in between and it did not help. Only the rolling restart did. I have no cluster where I can reproduce it, so I cannot say 100% for sure. Lots of small flushes: As for commitlog metrics. Here are the metrics around the time when the compactions started going up (~ 24th 11:00 UTC): “Merged from Memtable to Cache” is going up for the hosts having the issue: Memtable switches goes up for the hot tables: Schema commitlog latency goes up on some of the hosts, which I think is a symptom (IO being bottlenecked): What I find interesing is that the “reserved disk space” is similar in size Commitlog buffer size (this is what should be fixed in the ticket) is similar in size. My initial theory was, that this puts pressure on scylla to flush to free up memory. But elcallio does not think so: Let me know if you are looking for something in particular… You can try running with the commitlog logger at debug mode. I’m looking at this log line: If we correlate it with the memtable flush activity, we can focus on commitlog, and if not, it’s somewhere else. Maybe that’s the answer. We keep increasing the memory reserve. But it’s not quite right, since the memtable size isn’t affected by the reserves. In any case, nothing can work like this. Suggest looking at the system.large* tables to look for large rows/cells/collections (large partitions are fine). Also, please decode the backtraces. I would say its definetaly related to flushes: I guess the question is: Were these flushes triggered by commitlog or memory? Backtrace decode may help here --- ### Page: https://forum.scylladb.com/t/getting-error-in-the-lab-for-essentials-training/2962 Title: Getting Error in the Lab for Essentials Training - University and Training - ScyllaDB Community NoSQL Forum Meta Description: Hi, I’m trying to complete the training for the Essentials track and it seems that executable ‘migrate’ is not available on the VM. Language: en Canonical URL: https://forum.scylladb.com/t/getting-error-in-the-lab-for-essentials-training/2962 ## Headings Structure: H1: Getting Error in the Lab for Essentials Training H3: Related topics ## Main Content: H1: Getting Error in the Lab for Essentials Training H3: Related topics Hi, I’m trying to complete the training for the Essentials track and it seems that executable ‘migrate’ is not available on the VM. I also have an issue in the first lab “Building a Real-World Application”. In the very first step of starting clusters. I am able to create the 3 nodes and the status command says they are “Up and Normal” but the check button says “Your challenge failed because the Docker container ‘carepet-node1’ is not running or does not exist.” What am I doing wrong? @danielhe4rt can you have a look? There were 3 Essentials sessions, can you indicate which one you got the error in? Seems to be fixed now! Somehow the pipeline didn’t uploaded the latest releases… @thedeepman and @Nikhilanand_Vithalka let me know if something goes wrong! My issue was fixed and I was able to finish the entire Essentials course. Thank you! --- ### Page: https://forum.scylladb.com/t/can-you-change-the-topology-with-down-nodes/2964 Title: Can you change the topology with down nodes - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: The recommendation is to not make schema changes. We are looking at having to make schema changes, but have 3 nodes with an inability to come back into the cluster. We do have quorum. What are the real risks to making to… Language: en Canonical URL: https://forum.scylladb.com/t/can-you-change-the-topology-with-down-nodes/2964 ## Headings Structure: H1: Can you change the topology with down nodes H3: Related topics ## Main Content: H1: Can you change the topology with down nodes H3: Related topics The recommendation is to not make schema changes. We are looking at having to make schema changes, but have 3 nodes with an inability to come back into the cluster. We do have quorum. What are the real risks to making topology changes as we in a rock and hard place now between those 2 things. Which version are you using? --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-64-2024-09-27/2965 Title: Last week in scylla-cluster-tests.git master (issue #64; 2024-09-27) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the bc2c3ec0…09ffaf8b range are covered. There were 12 non-merge commits from 5 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-64-2024-09-27/2965 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #64; 2024-09-27) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #64; 2024-09-27) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the bc2c3ec0…09ffaf8b range are covered. There were 12 non-merge commits from 5 authors in that period. Some notable commits: A new upgrade test verifying latency with the ‘latte’ load tool was added, allowing us to test latency changes post-upgrade using custom workloads. A performance test measuring latency before, during, and after materialized view creation was added. It tests a challenging case where a regular column in the base table becomes a primary key in the materialized view. When users configure tests with incorrect config files, SCT now raises clear error messages, improving clarity on what went wrong. We encourage users to report misleading errors so we can address them. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-249-2024-09-29/2969 Title: Last week in scylladb.git master (issue #249; 2024-09-29) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 3d781c4fc8…c17d353718 range are covered. There were 110 non-merge commits from 27 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-249-2024-09-29/2969 ## Headings Structure: H1: Last week in scylladb.git master (issue #249; 2024-09-29) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #249; 2024-09-29) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 3d781c4fc8…c17d353718 range are covered. There were 110 non-merge commits from 27 authors in that period. Some notable commits: The DESCRIBE SCHEMA statement is now extended with statements to re-create roles and grants. This can be used to re-create not only the schema, but also the user and permission structure when restoring from backup. A race between tablet split and repair could cause an sstable not to be split among two new tablets, which in turn could prevent the sstable from being loaded on restart. This is now fixed. The backup API can now operate on a single table. A node that is being replaced is now marked so earlier so it does not get unexpected traffic. The bundled node_exporter Prometheus metrics exporter was updated to version 1.8.2. Alternator, ScyllaDB’s implementation of the DynamoDB API, can now stored information about provisioned throughput for the table. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-sphinx-theme-1-8/2970 Title: [RELEASE] ScyllaDB Sphinx Theme 1.8 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: To all ScyllaDB maintainers, we are pleased to announce that ScyllaDB Sphinx Theme 1.8 is now available. You are welcome to start upgrading your documentation projects, and if you need any assistance, don’t hesitate … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-sphinx-theme-1-8/2970 ## Headings Structure: H1: [RELEASE] ScyllaDB Sphinx Theme 1.8 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Sphinx Theme 1.8 H3: Related topics To all ScyllaDB maintainers, we are pleased to announce that ScyllaDB Sphinx Theme 1.8 is now available. You are welcome to start upgrading your documentation projects, and if you need any assistance, don’t hesitate to contact me or @Anna. You can read more about all the notable changes here For a migration guide, check out Upgrading from 1.7 to 1.8. --- ### Page: https://forum.scylladb.com/t/release-scylladb-6-0-4/2971 Title: [RELEASE] ScyllaDB 6.0.4 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 6.0.4, a bugfix patch release of the ScyllaDB 6.0 stable branch. ScyllaDB Open Source 6.0.4, like all past and future 6.x.y releases, is backward compatible and supports r… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-6-0-4/2971 ## Headings Structure: H1: [RELEASE] ScyllaDB 6.0.4 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 6.0.4 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 6.0.4, a bugfix patch release of the ScyllaDB 6.0 stable branch. ScyllaDB Open Source 6.0.4, like all past and future 6.x.y releases, is backward compatible and supports rolling upgrades. Note that the latest open-source stable branch is 6.1, and you are encouraged to upgrade to it. Issue fixed in this release: See more in the CQL extension docs. #15559 --- ### Page: https://forum.scylladb.com/t/upgrade-process-stuck-on-build-coordinator-state-replication-settings-for-system-tables-issue/2972 Title: Upgrade process stuck on build_coordinator_state - replication settings for system tables issue - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/upgrade-process-stuck-on-build-coordinator-state-replication-settings-for-system-tables-issue/2972 ## Headings Structure: H1: Upgrade process stuck on build_coordinator_state - replication settings for system tables issue H3: Related topics ## Main Content: H1: Upgrade process stuck on build_coordinator_state - replication settings for system tables issue H3: Related topics Originally from the User Slack @rfurmanski: Hi, I was trying to enable consistent topology updates in my cluster (scylla open source 6.0.3 everywhere, 3 dcs, 26 nodes) and it is stuck on build_coordinator_state. What might be a problem? @Piotr_Smaroń: CC @Kamil_Braun @rfurmanski: I verified all prerequisites and started upgrade via curl (http://127.0.0.1:10000/storage_service/raft_topology/upgrade) after migrating from 5.4.7 to 6.0.3. I followed the procedure described in here: https://opensource.docs.scylladb.com/branch-6.0/upgrade/upgrade-opensource/upgrade-guide-from-5.4-to-6.0/enable-consistent-topology.html Enable Consistent Topology Updates | ScyllaDB Docs @Kamil_Braun: > What might be a problem? Anything could be a problem. Without logs, it’s impossible to say. cc @Piotr_Dulikowski (author of the upgrade procedure) @rfurmanski: sure. Plz let me know what to provide. I don’t see anything wrong in logs. Just @Kamil_Braun: First try the recommendations from https://opensource.docs.scylladb.com/branch-6.0/upgrade/upgrade-opensource/upgrade-guide-from-5.4-to-6.0/enable-consistent-topology.html#what[…]tuck If nothing works (including rolling restart) – I recommend opening an issue and attaching logs from all of your nodes from the moment you started the upgrade procedure, ±1h Also nodetool status output from one of the nodes Enable Consistent Topology Updates | ScyllaDB Docs @Piotr_Dulikowski: > upgrade to topology on raft is scheduled After this, the topology coordinator should have started. Can you see a “start topology coordinator fiber” log message on any of the nodes? @rfurmanski: yes I see this on some of the nodes, but definitely not on all of them @Piotr_Dulikowski: It’s expected - the topology coordinator fiber runs on the current raft leader, so this won’t be printed on all nodes. So, most likely the upgrade process has started but got stuck. I think the best way forward would be to do what @Kamil_Braun suggested, so that we can analyze logs in more detail ourselves and look for more clues on what made it stuck. @rfurmanski: looks like raft still sees 2 recently removed nodes: nodetool status gives 26. these 2 nodes were removed before upgrade to 6.0.3 how to convince raft that these 2 nodes are removed? ahhh! After changing replication settings for system tables upgrade was successful. Thank you guys! @Piotr_Smaroń: @rfurmanski what exactly have you changed, can you please share the details? previously in ams3 I had 8 nodes after executing these commands raft upgrade went through @Kamil_Braun: ah, so it was probably trying to migrate from system_auth to the new auth v2 tables and it was trying to read system_auth using CL=ALL @Kamil_Braun: we should probably improve that error message so it’s more clear which table it is trying to query --- ### Page: https://forum.scylladb.com/t/release-scylladb-monitoring-4-8-1/2973 Title: [RELEASE] ScyllaDB Monitoring 4.8.1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce a patch release of ScyllaDB Monitoring Stack 4.8.1 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB Enterprise and ScyllaDB Open Source, based on Prometh… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-monitoring-4-8-1/2973 ## Headings Structure: H1: [RELEASE] ScyllaDB Monitoring 4.8.1 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Monitoring 4.8.1 H3: Related topics The ScyllaDB team is pleased to announce a patch release of ScyllaDB Monitoring Stack 4.8.1 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB Enterprise and ScyllaDB Open Source, based on Prometheus and Grafana. ScyllaDB Monitoring Stack 4.8.1 supports: --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-10-is-now-available-in-scylladb-cloud/2978 Title: [RELEASE] ScyllaDB Enterprise 2024.1.10 is now available in ScyllaDB Cloud! - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: For more information on this version, see here. Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-10-is-now-available-in-scylladb-cloud/2978 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2024.1.10 is now available in ScyllaDB Cloud! H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2024.1.10 is now available in ScyllaDB Cloud! H3: Related topics For more information on this version, see here. --- ### Page: https://forum.scylladb.com/t/setting-up-scylladb-with-two-different-ips-and-two-different-nics-for-cql-and-intra-node-calls/2990 Title: Setting up Scylladb with two different IPs and two different NICs for CQL and intra-node calls - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I have a server with two nics and two IPs and have configured it as per the documentation. After editing the scylla.yaml file and reviving scylla-server, doing a netstat -antp I have the following: tcp IP INTERNO:7… Language: en Canonical URL: https://forum.scylladb.com/t/setting-up-scylladb-with-two-different-ips-and-two-different-nics-for-cql-and-intra-node-calls/2990 ## Headings Structure: H1: Setting up Scylladb with two different IPs and two different NICs for CQL and intra-node calls H3: Related topics ## Main Content: H1: Setting up Scylladb with two different IPs and two different NICs for CQL and intra-node calls H3: Related topics I have a server with two nics and two IPs and have configured it as per the documentation. After editing the scylla.yaml file and reviving scylla-server, doing a netstat -antp I have the following: tcp IP INTERNO:7000 tcp IP INTERNO:9180 tcp 127.0.0.1:10000 Now I ask myself these two things, why don’t I see port 9042 for making calls from outside? What am I doing wrong in the configuration? How can I configure it correctly? The other question is how can I close the ports I’m not using? I tried commenting them out in the YAML file but it doesn’t seem to be working. I am creating scylladb with two nodes and Open-Source. I thank those who can help me understand where the problem lies Now I ask myself these two things, why don’t I see port 9042 for making calls from outside? What am I doing wrong in the configuration? How can I configure it correctly? Hard to say without seeing the actual YAML file and a trimmed down netstat output. Look into rpc options. If using IPv6, also state it accordingly. The other question is how can I close the ports I’m not using? I tried commenting them out in the YAML file but it doesn’t seem to be working. What do you mean by ports you ain’t using? Considering you read the Administration Guide | ScyllaDB Docs and you are sure you don’t need something, you may try setting it to 0. Commenting out will just leave their defaults. --- ### Page: https://forum.scylladb.com/t/secondary-indexes-vs-denormalized-tables/2997 Title: Secondary indexes vs. denormalized tables - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hey. I’m writing a database that contains guilds and channels. Guilds can have multiple channels, but channel IDs are unique and each channel can only belong to a single guild. I want my channel table to be indexable by … Language: en Canonical URL: https://forum.scylladb.com/t/secondary-indexes-vs-denormalized-tables/2997 ## Headings Structure: H1: Secondary indexes vs. denormalized tables H3: Related topics ## Main Content: H1: Secondary indexes vs. denormalized tables H3: Related topics Hey. I’m writing a database that contains guilds and channels. Guilds can have multiple channels, but channel IDs are unique and each channel can only belong to a single guild. I want my channel table to be indexable by either channel ID or guild ID, because I will often want to find the guild that a given channel is part of. If I establish guild ID as the primary key for my table and channel ID as a secondary index, Scylla will handle the creation of a reverse lookup table that maps channel ID → guild ID and then index the ‘channels’ table using its guild ID. However, I could also create a custom denormalized table that does the same thing (maps channel ID → guild ID). Then, whenever I want to retrieve a channel by channel ID, I’ll query that table, retrieve the guild ID, and then index the original table by guild ID. This is prone to more errors because I will have to manually update the denormalized table whenever a channel is added/removed. Are there any performance benefits to using a denormalized table vs. the built-in secondary index? I have read that denormalized tables are “more efficient for predefined, specific access patterns” but I don’t see how so. I assume that Scylla’s reverse lookup table and a custom denormalized table will have the same storage overhead. Views (thus indexes) updates are done asynchronously by the query coordinator. This - in turn - means that under rare circumstances the view update in question may fail and get inconsistent. You may overcome this with synchronous_updates = true, at the expense of more cycles to acknowledge a request back to your clients. Denormalized tables allow for more efficient routing, at the expense of increased client-side complexity. As you realized, both methods will work. --- ### Page: https://forum.scylladb.com/t/seastar-addr2line-appears-broken-in-official-images/2999 Title: Seastar-addr2line appears broken in official images - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Trying to run seastar-addr2line to decode a traceback in logs, It fails to run with the error below. Is this something that broke with the recent python3 changes, or am I holding it wrong? Tried on 6.0.3 docker image, 6… Language: en Canonical URL: https://forum.scylladb.com/t/seastar-addr2line-appears-broken-in-official-images/2999 ## Headings Structure: H1: Seastar-addr2line appears broken in official images H3: Related topics ## Main Content: H1: Seastar-addr2line appears broken in official images H3: Related topics Trying to run seastar-addr2line to decode a traceback in logs, It fails to run with the error below. Is this something that broke with the recent python3 changes, or am I holding it wrong? Tried on 6.0.3 docker image, 6.0.3 AMI, and 6.0.4 AMI (all x86_64). Also, the decoding kb article’s method for finding the debug symbols seems to be out of date: the location no longer looks like /usr/lib/debug/opt/scylladb/libexec/scylla-[version].x86_64.debug - is there a better heuristic for finding the right file? addr2line.py is in seastar/scripts/addr2line.py at master · scylladb/seastar · GitHub I don’t remember if we ship it in debug packages, but the link should get you going. Please open an issue. We also have https://backtrace.scylladb.com (which interestingly is down for me I as write this, but it works) We also have https://backtrace.scylladb.com (which interestingly is down for me I as write this, but it works) Fellipe, try http://backtrace.scylladb.com rather than https. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-65-2024-10-04/3000 Title: Last week in scylla-cluster-tests.git master (issue #65; 2024-10-04) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the dd6d58ab…2e9e77c0 range are covered. There were 19 non-merge commits from 7 authors in tha… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-65-2024-10-04/3000 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #65; 2024-10-04) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #65; 2024-10-04) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the dd6d58ab…2e9e77c0 range are covered. There were 19 non-merge commits from 7 authors in that period. Some notable commits: Improved code and logic behind the NodeBootstrapAbortManager, which orchestrates abort bootstrap test scenarios. It was reorganized, addressing issues with rebootstrap by cleaning data beforehand. Scylla Manager’s restore benchmark results are now sent to Argus, including three duration metrics: restore, repair, and total time. A new test was added to CI to monitor restore speed changes across releases. A new force_run_iotune config option has been introduced. If enabled, iotune runs before starting Scylla for the first time, overriding any pre-defined results in the image. This option is now enabled in our PR-provision test. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-cpp-over-rust-driver-0-2-0/3004 Title: [RELEASE] ScyllaDB CPP-over-Rust Driver 0.2.0 - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce ScyllaDB CPP Rust Driver 0.2.0, an API-compatible rewrite of GitHub - scylladb/cpp-driver: Scylla C/C++ Driver as a wrapper for the Rust driver. It will fully replace the CPP driv… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-cpp-over-rust-driver-0-2-0/3004 ## Headings Structure: H1: [RELEASE] ScyllaDB CPP-over-Rust Driver 0.2.0 H2: Changes H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB CPP-over-Rust Driver 0.2.0 H2: Changes H3: Related topics The ScyllaDB team is pleased to announce ScyllaDB CPP Rust Driver 0.2.0, an API-compatible rewrite of GitHub - scylladb/cpp-driver: Scylla C/C++ Driver as a wrapper for the Rust driver. It will fully replace the CPP driver, which is already reaching its End of Life. The current driver version should be considered Alpha. Some minor features still need to be included. See Limitations section in README.md. The underlying Rust driver used version: 0.14.0. Implemented API functions: CassCollection: (#143) Token awareness: (#170) Client’s identity: (#167) Duration CQL type support: (#135) Decimal CQL type support: (#146) New features / enhancements CI / developer tool improvements: Congrats to all contributors and thanks everyone for using our driver! The source code of the driver can be found here: Thank you for your attention, please do not hesitate to contact us if you have any questions, issues, feature requests, or are simply interested in our driver! Contributors since the last release: --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-250-2024-10-06/3010 Title: Last week in scylladb.git master (issue #250; 2024-10-06) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the c17d353718…882a3c60e4 range are covered. There were 77 non-merge commits from 23 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-250-2024-10-06/3010 ## Headings Structure: H1: Last week in scylladb.git master (issue #250; 2024-10-06) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #250; 2024-10-06) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the c17d353718…882a3c60e4 range are covered. There were 77 non-merge commits from 23 authors in that period. Some notable commits: After a tablet is migrated away from a node, a cleanup process takes place to remove the tablet’s footprint from the replica. During this process a race can happen the tablet’s compaction strategy is changed concurrently. This is now fixed. The metric for the amount of memory used by commitog was updated incorrectly, indicating a leak that wasn’t there. This is now fixed. A rare case where an sstable promoted index lookup would return in correct information was fixed. The system.snapshots virtual table now lists snapshots of dropped tables. The CREATE MATERIALIZED VIEW statement now supports the undocumented WITH ID clause, improving compatibility with Cassandra. The nodetool replace and remove operations now deprecate using IP addresses to specify nodes. Use host IDs to specify nodes. Repair has a read timeout in order to protect against deadlocks, but in practice the timeout caused repairs to fail. The timeout is now removed. When a tablet is split into two, tombstone garbage collection becomes more complicated, since sstables for the tablet exist in both pre-split and post-split state. We are now more careful with tombstone garbage collection during this operation. The restore REST API can now accept a list of sstables to restore. Restore will keep sstables on object storage in place after restore. The index page cache will now generate fewer disk IOPS if an index read is partially cached. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/relase-scylla-6-2-rc1/3011 Title: [RELASE] Scylla 6.2 RC1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce ScyllaDB Open Source 6.2 RC1, the first Release Candidate for the ScyllaDB Open Source 6.2 minor release. ScyllaDB 6.2 introduces many Tablets improvements, new zero-token nodes … Language: en Canonical URL: https://forum.scylladb.com/t/relase-scylla-6-2-rc1/3011 ## Headings Structure: H1: [RELASE] Scylla 6.2 RC1 H3: High Availability - Arbiter H3: Alternator RBAC H3: More updated H3: Related topics ## Main Content: H1: [RELASE] Scylla 6.2 RC1 H3: High Availability - Arbiter H3: Alternator RBAC H3: More updated H4: Tablets H4: Tracing H4: Stability H4: Admin H4: Alternator H4: Performance H4: CQL H4: Materialized view H4: Packaging H4: Config H4: Monitoring H3: Related topics The ScyllaDB team is pleased to announce ScyllaDB Open Source 6.2 RC1, the first Release Candidate for the ScyllaDB Open Source 6.2 minor release. ScyllaDB 6.2 introduces many Tablets improvements, new zero-token nodes (Arbiter), Alternator RBAC support and many other bug fixes and stabilizations. We encourage you to run ScyllaDB 6.2 release candidates on your test environments; this will help ensure that an upgrade to ScyllaDB 6.2 General Availability will proceed smoothly with your workload. Only the last two minor releases of the ScyllaDB Open Source project are supported. Once ScyllaDB Open Source 6.2 is officially released, only ScyllaDB Open Source 6.2 and ScyllaDB 6.1 will be supported, and ScyllaDB 6.0 will be retired. Get ScyllaDB Open Source 6.2 as binary packages (RPM/DEB), AWS AMI, GCP Image and Docker image Upgrade from ScyllaDB 6.1 to ScyllaDB 6.2 There is now support for zero-token nodes. Such nodes do not replicate any data, but can participate in query coordination, and in Raft quorum voting. One can use this to create an Arbiter: a tiebreaker node, with no data, that can help maintain quorum in the case of two symmetrical two-datacenter clusters. If one of the data centers fails, the Arbiter, deployed on a 3rd datacenter, keeps quorum on the node alive. Since the Arbiter has zero token, it does not replicate user data, and does not come with network and storage costs. #15360 Authorization: Alternator supports Role-Based Access Control (RBAC). Control is done via CQL. #5047 Scylla Monitoring stack 4.8.1 and later support ScyllaDB 6.2 release. See upgrade docs for Metrics update in ScyllaDB 6.2. --- ### Page: https://forum.scylladb.com/t/django-cant-make-first-get-post-function/3013 Title: Django can't make first get post function - Database Community - ScyllaDB Community NoSQL Forum Meta Description: i can’t make first get post function Language: en Canonical URL: https://forum.scylladb.com/t/django-cant-make-first-get-post-function/3013 ## Headings Structure: H1: Django can't make first get post function H1: Create your views here. H3: Related topics ## Main Content: H1: Django can't make first get post function H1: Create your views here. H3: Related topics i can’t make first get post function Hard to assist with so few details. Please state exactly what you are trying to do, which steps you’ve taken, what is going wrong, and any other relevant info as needed. If you are using the django-scylla Pypi, you may also want to check out their own docs and ask in their discord server. See django-scylla · PyPI thanks for replaying to me from rest_framework import serializers from .models import * class BlogPostSerializer(serializers.ModelSerializer): class Meta: model = BlogPost fields = ‘all’ def create(self): return BlogPost.title from django.shortcuts import render from django.http import JsonResponse from rest_framework import status, generics,request from rest_framework.decorators import api_view from rest_framework.response import Response from .models import BlogPost from .serializers import * class BlogPostListCreate(generics.ListCreateAPIView): serializer_class = BlogPostSerializer import uuid from datetime import datetime from django.db import models from cassandra.cqlengine import columns from django_cassandra_engine.models import DjangoCassandraModel class BlogPost(DjangoCassandraModel): id = columns.UUID(primary_key=True, default=uuid.uuid4) title = columns.Text(required=True) content = columns.Text() created_at = columns.DateTime(default=datetime.now) class Meta: app_label = ‘galo’ DATABASES = { ‘default’: { ‘ENGINE’: ‘django_cassandra_engine’, ‘NAME’: ‘my_keyspace’, # Keyspace in ScyllaDB ‘HOST’: ‘127.0.0.1’, # ScyllaDB instance ‘PORT’: ‘9042’, ‘USER’: ‘’, # Leave empty if no authentication ‘PASSWORD’: ‘’, ‘OPTIONS’: { ‘replication’: { ‘strategy_class’: ‘SimpleStrategy’, ‘replication_factor’: 1 } } } } There’s no error, so we still can’t crystal ball this gives me keysapce problem i forget to tell i solved it for a while --- ### Page: https://forum.scylladb.com/t/release-scylladb-cloud-7-october-2024/3014 Title: [RELEASE] ScyllaDB Cloud - 7 October 2024 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: We are excited to announce a new update to the ScyllaDB Cloud service! New features and enhancements: Extended Cluster Name Limit: You can now name clusters with up to 63 characters, an increase from the previous limi… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-cloud-7-october-2024/3014 ## Headings Structure: H1: [RELEASE] ScyllaDB Cloud - 7 October 2024 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Cloud - 7 October 2024 H3: Related topics We are excited to announce a new update to the ScyllaDB Cloud service! New features and enhancements: --- ### Page: https://forum.scylladb.com/t/error-read-concurrency-sem-wait-queue-overload/3022 Title: Error _read_concurrency_sem: wait queue overload - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi Team, We are seeing lots of “_read_concurrency_sem: wait queue overload” exception in scylladb logs. Any specific reason for that. And how to resolve this. Oct 9 10:35:45 Cass-1 scylla: message repeated 22 times: [… Language: en Canonical URL: https://forum.scylladb.com/t/error-read-concurrency-sem-wait-queue-overload/3022 ## Headings Structure: H1: Error _read_concurrency_sem: wait queue overload H3: Related topics ## Main Content: H1: Error _read_concurrency_sem: wait queue overload H3: Related topics We are seeing lots of “_read_concurrency_sem: wait queue overload” exception in scylladb logs. Any specific reason for that. And how to resolve this. Oct 9 10:35:45 Cass-1 scylla: message repeated 22 times: [ [shard 28:stat] storage_proxy - Exception when communicating with 10.0.1.4, to read from . In general @denesb keeps a very good detailed doc on the reader concurrency semaphore internals at scylladb/docs/dev/reader-concurrency-semaphore.md at master · scylladb/scylladb · GitHub You should have received a diagnostics dump with detailed information on what happened the moment the overload manifested. Most of the time (though not always) this happens due to clients overloading the database, so Monitoring should also be your ally there. Look at the by-shard data on foreground and background reads, latencies, the advanced dashboard and take it from there. I can not find coredump in /var/lib/scylla/coredump/ Any idea how to setup coredump for scylladb. I tried with scylla_coredump_setup and following output. root@(Cass-6)[/var/log/scylla]->> scylla_coredump_setup kernel.core_pattern = |/lib/systemd/systemd-coredump %P %u %g %s %t 9223372036854775808 %h kernel.core_pipe_limit = 16 fs.suid_dumpable = 2 Generating coredump to test systemd-coredump… PID: 1041734 (bash) UID: 0 (root) GID: 0 (root) Signal: 11 (SEGV) Timestamp: Thu 2024-10-10 06:31:38 UTC (3s ago) Command Line: /bin/bash /tmp/tmps43gavyw Executable: /usr/bin/bash Control Group: /user.slice/user-0.slice/session-34040.scope Unit: session-34040.scope Slice: user-0.slice Session: 34040 Owner UID: 0 (root) Boot ID: 026dc33a62f84052af04918dffe1170d Machine ID: 299f966af56e46e5b31a8e89157e832e Hostname: Cas05 Storage: /var/lib/systemd/coredump/core.bash.0.026dc33a62f84052af04918dffe1170d.1041734.1728541898000000 (present) Disk Size: 544.0K Message: Process 1041734 (bash) of user 0 dumped core. systemd-coredump is working finely. root@(Cass-6)[/var/log/scylla]->> ll /var/lib/scylla/coredump/ total 8 drwxr-xr-x 2 scylla scylla 4096 Aug 24 2015 ./ drwxr-xr-x 7 scylla scylla 4096 Sep 1 02:10 …/ root@(Cass-6)[/var/log/scylla]->> This looks fine. We delete the generated core to save space. You may trigger a core yourself and check for its presence. --- ### Page: https://forum.scylladb.com/t/relase-scylla-6-2-rc2/3023 Title: [RELASE] Scylla 6.2 RC2 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The Scylla team is pleased to announce ScyllaDB Open Source 6.2 RC2, the second Release Candidate for the Scylla Open Source 6.2 minor release. ScyllaDB 6.2 introduces many Tablets improvements, new zero-token nodes (Ar… Language: en Canonical URL: https://forum.scylladb.com/t/relase-scylla-6-2-rc2/3023 ## Headings Structure: H1: [RELASE] Scylla 6.2 RC2 H3: Related topics ## Main Content: H1: [RELASE] Scylla 6.2 RC2 H3: Related topics The Scylla team is pleased to announce ScyllaDB Open Source 6.2 RC2, the second Release Candidate for the Scylla Open Source 6.2 minor release. ScyllaDB 6.2 introduces many Tablets improvements, new zero-token nodes (Arbiter), Alternator RBAC support and many other bug fixes and stabilizations. Only the last two minor releases of the ScyllaDB Open Source project are supported. Once Scylla Open Source 6.2 is officially released, ScyllaDB Open Source 6.2 and 6.1 will be supported, and ScyllaDB 6.0 will be retired. For a complete description of ScyllaDB 6.2 see ScyllaDB 6.2 RC1. Full description of the release in 6.2 RC1 notes Get ScyllaDB Open Source 6.2 as binary packages (RPM/DEB), AWS AMI, GCP Image and Docker image Upgrade from ScyllaDB 6.1 to ScyllaDB 6.2 Updates and bug fixes since 6.2 RC1 (not including tests and docs updates) --- ### Page: https://forum.scylladb.com/t/regarding-the-create-table-query/3025 Title: Regarding the create table query - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: When I tried creating a table, I wasn’t able to keep certain variable names like token, from etc… I am adding the CQL for create table. Could you please look into it and suggest some alternatives that we can use if we wa… Language: en Canonical URL: https://forum.scylladb.com/t/regarding-the-create-table-query/3025 ## Headings Structure: H1: Regarding the create table query H3: Related topics ## Main Content: H1: Regarding the create table query H3: Related topics When I tried creating a table, I wasn’t able to keep certain variable names like token, from etc… I am adding the CQL for create table. Could you please look into it and suggest some alternatives that we can use if we want to keep the field name as token, from and similar to these? CREATE TABLE email_template ( name text PRIMARY KEY, service text, text_body text, html_body text, subject text, priority text, from text, country text, sender_name text, email_type text, created_by text, created_at bigint, updated_at bigint ) WITH compaction = { ‘class’: ‘SizeTieredCompactionStrategy’ }; On executing the query getting an error: SyntaxException: line 8:4 : Missing ‘)’ I found some alternative of keeping the variable name as ‘_from’ but this also fails and gives SyntaxException as mentioned above. From is a reserved keyword Thanks Felipe for helping out here. --- ### Page: https://forum.scylladb.com/t/time-based-bucketing-issue-compromise-between-storage-ram/3026 Title: Time based bucketing issue- compromise between storage & RAM - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I need to bucket my data so as to fit it in partitions. Im using time based bucketing. & the current modelling is as follows- Im counting the number of buckets that are unused. Example- bucket 10, counter 100 means buck… Language: en Canonical URL: https://forum.scylladb.com/t/time-based-bucketing-issue-compromise-between-storage-ram/3026 ## Headings Structure: H1: Time based bucketing issue- compromise between storage & RAM H3: Related topics ## Main Content: H1: Time based bucketing issue- compromise between storage & RAM H3: Related topics I need to bucket my data so as to fit it in partitions. Im using time based bucketing. & the current modelling is as follows- Im counting the number of buckets that are unused. Example- bucket 10, counter 100 means bucket 10-110 are unused. bucket 3 counter 0, means only bucket 3 is unused. So in the db its stored as- bucket 10 counter 10, bucket 22 counter 1, bucket 50 counter 40, bucket 92 counter 8. Thus, used buckets will be- 0-9, 21, 24-49, 91 but since im paginating I need only 5 buckets at a time. using this modal, saaves storage but I cant get 5 buckets from current bucket offset from db itself, I need to write business logic for that & I suspect that will impact RAM. Is there any better method to track unused buckets such that I can get buckets that are used, so as to query Is there any better method to track unused buckets such that I can get buckets that are used, so as to query I am probably missing some details, but here goes a long shot. IIUC you are using an auxiliary table for that. Wouldn’t simply clustering by used buckets work instead? That way, you simply read from whatever you retrieve as part of a scan. there are pros & cons to that. Im using time based bucketing so simplest would be to store bucket once used, every week. But then for high traffic clusters, which are always seeing inserts, itll be a waste of storage. Unused buckets are being used since we expect the volume of inserts per partition key to be greater than 0 for majority clients. Hence only a few ppl will leave buckets unused by creating a stream & forgetting about it hence not inserting anything. So tracking unused buckets, only saves some storage. Hope you r understanding? Yeah, I think I get it now. You’ll definitely need to make some choices there. Ideally, you could perhaps never bucket these outliers if you had the option to. ie: if its users running a “free-trial” chances are they will forget about it, when regular users will benefit from bucketing. Feel free to share schema and describe the use case in question if you feel like you need more assistance for example, take discords time based partitioning. They track unused buckets for storing available buckets to query against for channel messages so as not to iterate over every bucket. More specifically this- We noticed Cassandra was running 10 second “stop-the-world” GC constantly but we had no idea why. We started digging and found a Discord channel that was taking 20 seconds to load. The Puzzles & Dragons Subreddit public Discord server was the culprit. Since it was public we joined it to take a look. To our surprise, the channel had only 1 message in it. It was at that moment that it became obvious they deleted millions of messages using our API, leaving only 1 message in the channel. If you have been paying attention you might remember how Cassandra handles deletes using tombstones (mentioned in Eventual Consistency). When a user loaded this channel, even though there was only 1 message, Cassandra had to effectively scan millions of message tombstones (generating garbage faster than the JVM could collect it). We solved this by doing the following: I decided to follow in their foorsteps & do the same. I tried amount based bucketing but that requires me to setup kafka & run background jobs per week to check every channel (in discords terminology) in existence to see their current count. So yeah, not as low mentainance as I needed it to be. So I was looking for such a method, where I could query efficiently without a lot of storage wastage & have a decent bucketing strat thats low maintenance --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-11/3028 Title: [RELEASE] ScyllaDB Enterprise 2024.1.11 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2024.1.11, a bug-fix production-ready patch release for ScyllaDB Enterprise 2024.1 LTS Release. Related Links Get ScyllaDB Enterprise 2024.1 (customers only, or 30-day e… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-11/3028 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2024.1.11 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2024.1.11 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2024.1.11, a bug-fix production-ready patch release for ScyllaDB Enterprise 2024.1 LTS Release. The following issues are fixed in this release (with an open-source reference, if available): --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-66-2024-10-10/3029 Title: Last week in scylla-cluster-tests.git master (issue #66; 2024-10-10) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 99ee7b77…2bf819cb range are covered. There were 17 non-merge commits from 8 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-66-2024-10-10/3029 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #66; 2024-10-10) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #66; 2024-10-10) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 99ee7b77…2bf819cb range are covered. There were 17 non-merge commits from 8 authors in that period. Some notable commits: The Argus Client has been updated to support best results tracking and data validation by Argus. A new gradual throughput test will be executed from the new branch-perf-v16 to match the results from SCT’s master branch. We’ve added a rocky9-selinux artifact test. Test result emails and Elasticsearch data now contain links to the Argus run. A new restore benchmark test running on a 9-node cluster, testing TWCS with a dataset of two tables (700GB, 300GB). When restoring monitoring for new runs, dashboard time ranges now default to the test duration period. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/node-crashing-after-adding-new-nodes-in-scylla-cluster/3036 Title: Node crashing after adding new nodes in scylla cluster - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello Team, We are facing issue of old nodes getting crashed after adding new nodes. Details : Scylla version : 6.1.2-0.20240915.b60f9ef4c223 Backup restoration performed on 4 node scylla cluster running on 6.1. So… Language: en Canonical URL: https://forum.scylladb.com/t/node-crashing-after-adding-new-nodes-in-scylla-cluster/3036 ## Headings Structure: H1: Node crashing after adding new nodes in scylla cluster H3: Related topics ## Main Content: H1: Node crashing after adding new nodes in scylla cluster H4: Node crashing after adding new nodes in scylla cluster 6.1 with tablets enabled. H3: Related topics We are facing issue of old nodes getting crashed after adding new nodes. Scylla version : 6.1.2-0.20240915.b60f9ef4c223 Backup restoration performed on 4 node scylla cluster running on 6.1. Source : Scylla version : 5.2.19 Scylla manager : 3.2 Nodes : 4 Size : 1.5 TB each node Destination : Scylla version : 6.1.2 Scylla manager : 3.3 Nodes : 4 Tablets enabled cluster. Restoration and repair post restoration was successfully completed. Once 4 node cluster of 6.1 was running with Size 1.5 TB each node, we tried doing elastic scaling by adding 4 more nodes by starting scylla-server service simultaneously. All 4 new nodes got added in 2 mins by coming UN state but with minimal 4-5 GB of data. Tablets redistribution started. While redistribution was going on, observed that we old two nodes got abruptly crashed and not able to come back up. Attaching logs here of one of the nodes which got crashed. Tried taking service reboot but didnt helped. Any resolution for such cases where [shard 5: gms] table - Found that storage of group 21 for table 51d04920-86fc-11ef-8be6-7886065e5abf wasn’t split correctly, therefore groups cannot be remapped with the new tablet count. Let us know if anything is needed for debugging ? Thanks for reporting. Look like a bug. Please open an issue. Hello Team, We are facing issue of old nodes getting crashed after adding new n…odes. Details : Scylla version : 6.1.2-0.20240915.b60f9ef4c223 Backup restoration performed on 4 node scylla cluster running on 6.1. Source : Scylla version : 5.2.19 Scylla manager : 3.2 Nodes : 4 Size : 1.5 TB each node Destination : Scylla version : 6.1.2 Scylla manager : 3.3 Nodes : 4 Tablets enabled cluster. Restoration and repair post restoration was successfully completed. Once 4 node cluster of 6.1 was running with Size 1.5 TB each node, we tried doing elastic scaling by adding 4 more nodes by starting scylla-server service simultaneously. All 4 new nodes got added in 2 mins by coming UN state but with minimal 4-5 GB of data. Tablets redistribution started. While redistribution was going on, observed that we old two nodes got abruptly crashed and not able to come back up. Attaching logs here of one of the nodes which got crashed. ` Oct 11 12:37:27 NODENAME scylla[60269]: [shard 0:main] table - Unable to load SSTable /var/lib/scylla/data/vss/embeddings-51d0492086fc11ef8be67886065e5abf/me-3gk9_1nef_1r88g2c1uc12qoe97m-big-Data.db that belongs to tablets 2 and 3, at: 0x5e9e> -------- seastar::internal::coroutine_traits_base::promise_type -------- seastar::internal::coroutine_traits_base::promise_type -------- seastar::internal::coroutine_traits_base::promise_type -------- seastar::coroutine::parallel_for_each&, seastar::sharded&, replica::keyspace&, seastar::basic_sst> -------- seastar::internal::coroutine_traits_base::promise_type -------- seastar::coroutine::parallel_for_each&, seastar::sharded&, replica::keyspace&, seastar::basic_sst> -------- seastar::internal::coroutine_traits_base::promise_type -------- seastar::continuation, seastar::internal::complete_when_all >, seastar::futur> -------- seastar::continuation, seastar::future::discard_result()::{lambda((auto:1&&)...)#1}, seastar::future::then_impl_nrvo::d> -------- seastar::(anonymous namespace)::thread_wake_task -------- seastar::continuation, seastar::async&, seastar::sharded -------- seastar::continuation, seastar::future::finally_body -------- seastar::(anonymous namespace)::thread_wake_task -------- seastar::continuation, seastar::async(seastar::thread_attributes, scylla_main(int, char**> -------- seastar::continuation, seastar::future::finally_body(seastar::thread_> -------- seastar::continuation, seastar::future::finally_body ()>&&)::$_0::operator> Oct 11 12:37:27 NODENAME scylla[60269]: terminate called after throwing an instance of 'seastar::internal::backtraced' Oct 11 12:37:27 NODENAME scylla[60269]: what(): Unable to load SSTable /var/lib/scylla/data/vss/embeddings-51d0492086fc11ef8be67886065e5abf/me-3gk9_1nef_1r88g2c1uc12qoe97m-big-Data.db that belongs to tablets 2 and 3 Backtrace: 0x5e9e22e 0x5e> Oct 11 12:37:27 NODENAME scylla[60269]: -------- Oct 11 12:37:27 NODENAME scylla[60269]: seastar::internal::coroutine_traits_base::promise_type Oct 11 12:37:27 NODENAME scylla[60269]: -------- Oct 11 12:37:27 NODENAME scylla[60269]: seastar::internal::coroutine_traits_base::promise_type Oct 11 12:37:27 NODENAME scylla[60269]: -------- Oct 11 12:37:27 NODENAME scylla[60269]: seastar::internal::coroutine_traits_base::promise_type Oct 11 12:37:27 NODENAME scylla[60269]: -------- Oct 11 12:37:27 NODENAME scylla[60269]: seastar::coroutine::parallel_for_each&, seastar::sharded&, replica::keyspace&, seastar::basic_sst> Oct 11 12:37:27 NODENAME scylla[60269]: -------- Oct 11 12:37:27 NODENAME scylla[60269]: seastar::internal::coroutine_traits_base::promise_type Oct 11 12:37:27 NODENAME scylla[60269]: -------- Oct 11 12:37:27 NODENAME scylla[60269]: seastar::coroutine::parallel_for_each&, seastar::sharded&, replica::keyspace&, seastar::basic_sst> Oct 11 12:33:35 NODENAME scylla[18799]: 0x13842a5 Oct 11 12:33:35 NODENAME scylla[18799]: 0x1385c60 Oct 11 11:40:01 NODENAME scylla[18799]: [shard 0:comp] compaction - [Compact system.peers 8d28e7b0-87c5-11ef-9987-878599d64622] Compacted 2 sstables to [/var/lib/scylla/data/system/peers-37f71aca7dc2383ba70672528af04d4f/me-3gka_0wep_4yywh2c1u> Oct 11 11:40:01 NODENAME scylla[18799]: [shard 0:comp] compaction - [Compact system.group0_history 8d2a9560-87c5-11ef-9987-878599d64622] Compacting [/var/lib/scylla/data/system/group0_history-027e42f5683a3ed7b404a0100762063c/me-3gka_055x_02sb> Oct 11 11:40:01 NODENAME scylla[18799]: [shard 0:comp] sstable - Rebuilding bloom filter /var/lib/scylla/data/system/group0_history-027e42f5683a3ed7b404a0100762063c/me-3gka_0wep_51jhs2c1uc12qoe97m-big-Filter.db: resizing bitset from 328 bytes> Oct 11 11:40:01 NODENAME scylla[18799]: [shard 0:comp] compaction - [Compact system.group0_history 8d2a9560-87c5-11ef-9987-878599d64622] Compacted 2 sstables to [/var/lib/scylla/data/system/group0_history-027e42f5683a3ed7b404a0100762063c/me-3> Oct 11 11:40:02 NODENAME scylla[18799]: [shard 0:strm] stream_session - [Stream #779698c4-87c5-11ef-bc3c-ec5f9328e09e] Streaming plan for Tablet migration-vss-index-0 succeeded, peers={10.138.64.129}, tx=1205197 KiB, 33087.57 KiB/s, rx=0 KiB,> Oct 11 11:40:02 NODENAME scylla[18799]: [shard 30:strm] table - Cleaned up tablet 0 of table vss.catalog_realtime_accumulator successfully. Oct 11 11:40:02 NODENAME scylla[18799]: [shard 12:strm] table - Cleaned up tablet 23 of table vss.catalog_realtime_accumulator successfully. Oct 11 12:31:43 NODENAME scylla[18799]: [shard 0:comp] compaction - [Compact system.peers c59f6310-87cc-11ef-9987-878599d64622] Compacted 2 sstables to [/var/lib/scylla/data/system/peers-37f71aca7dc2383ba70672528af04d4f/me-3gka_0ysv_08scx2c1u> Oct 11 12:33:35 NODENAME scylla[18799]: [shard 0: gms] table - Detected tablet split for table vss.embeddings, increasing from 64 to 128 tablets Oct 11 12:33:35 NODENAME scylla[18799]: [shard 0: gms] table - Found that storage of group 1 for table 51d04920-86fc-11ef-8be6-7886065e5abf wasn't split correctly, therefore groups cannot be remapped with the new tablet count., at: 0x5e9e22e > -------- seastar::internal::coroutine_traits_base::promise_type -------- seastar::internal::coroutine_traits_base::promise_type -------- seastar::internal::coroutine_traits_base::promise_type -------- seastar::internal::coroutine_traits_base::promise_type -------- seastar::internal::coroutine_traits_base::promise_type -------- seastar::internal::coroutine_traits_base::promise_type -------- seastar::internal::coroutine_traits_base::promise_type -------- seastar::internal::coroutine_traits_base::promise_type Oct 11 12:33:35 NODENAME scylla[18799]: [shard 1: gms] table - Found that storage of group 34 for table 51d04920-86fc-11ef-8be6-7886065e5abf wasn't split correctly, therefore groups cannot be remapped with the new tablet count., at: 0x5e9e22e> -------- seastar::smp_message_queue::async_work_item::invoke_on_all(seastar::smp_submit_to_options, std::function (service::storage_service&)>)::> Oct 11 12:33:35 NODENAME scylla[18799]: [shard 3: gms] table - Found that storage of group 50 for table 51d04920-86fc-11ef-8be6-7886065e5abf wasn't split correctly, therefore groups cannot be remapped with the new tablet count., at: 0x5e9e22e> -------- seastar::smp_message_queue::async_work_item::invoke_on_all(seastar::smp_submit_to_options, std::function (service::storage_service&)>)::> Oct 11 12:33:35 NODENAME scylla[18799]: [shard 29: gms] table - Found that storage of group 13 for table 51d04920-86fc-11ef-8be6-7886065e5abf wasn't split correctly, therefore groups cannot be remapped with the new tablet count., at: 0x5e9e22e> -------- seastar::smp_message_queue::async_work_item::invoke_on_all(seastar::smp_submit_to_options, std::function (service::storage_service&)>)::> Oct 11 12:33:35 NODENAME scylla[18799]: [shard 20: gms] table - Found that storage of group 46 for table 51d04920-86fc-11ef-8be6-7886065e5abf wasn't split correctly, therefore groups cannot be remapped with the new tablet count., at: 0x5e9e22e> -------- seastar::smp_message_queue::async_work_item::invoke_on_all(seastar::smp_submit_to_options, std::function (service::storage_service&)>)::> Oct 11 12:33:35 NODENAME scylla[18799]: [shard 6: gms] table - Found that storage of group 4 for table 51d04920-86fc-11ef-8be6-7886065e5abf wasn't split correctly, therefore groups cannot be remapped with the new tablet count., at: 0x5e9e22e > -------- seastar::smp_message_queue::async_work_item::invoke_on_all(seastar::smp_submit_to_options, std::function (service::storage_service&)>)::> Oct 11 12:33:35 NODENAME scylla[18799]: [shard 26: gms] table - Found that storage of group 24 for table 51d04920-86fc-11ef-8be6-7886065e5abf wasn't split correctly, therefore groups cannot be remapped with the new tablet count., at: 0x5e9e22e> -------- seastar::smp_message_queue::async_work_item::invoke_on_all(seastar::smp_submit_to_options, std::function (service::storage_service&)>)::> Oct 11 12:33:35 NODENAME scylla[18799]: [shard 7: gms] table - Found that storage of group 59 for table 51d04920-86fc-11ef-8be6-7886065e5abf wasn't split correctly, therefore groups cannot be remapped with the new tablet count., at: 0x5e9e22e> -------- seastar::smp_message_queue::async_work_item::invoke_on_all(seastar::smp_submit_to_options, std::function (service::storage_service&)>)::> Oct 11 12:33:35 NODENAME scylla[18799]: [shard 27: gms] table - Found that storage of group 19 for table 51d04920-86fc-11ef-8be6-7886065e5abf wasn't split correctly, therefore groups cannot be remapped with the new tablet count., at: 0x5e9e22e> -------- seastar::smp_message_queue::async_work_item::invoke_on_all(seastar::smp_submit_to_options, std::function (service::storage_service&)>)::> Oct 11 12:33:35 NODENAME scylla[18799]: [shard 24: gms] table - Found that storage of group 31 for table 51d04920-86fc-11ef-8be6-7886065e5abf wasn't split correctly, therefore groups cannot be remapped with the new tablet count., at: 0x5e9e22e> -------- seastar::smp_message_queue::async_work_item::invoke_on_all(seastar::smp_submit_to_options, std::function (service::storage_service&)>)::> Oct 11 12:33:35 NODENAME scylla[18799]: [shard 10: gms] table - Found that storage of group 38 for table 51d04920-86fc-11ef-8be6-7886065e5abf wasn't split correctly, therefore groups cannot be remapped with the new tablet count., at: 0x5e9e22e> -------- seastar::smp_message_queue::async_work_item::invoke_on_all(seastar::smp_submit_to_options, std::function (service::storage_service&)>)::> Oct 11 12:33:35 NODENAME scylla[18799]: [shard 11: gms] table - Found that storage of group 29 for table 51d04920-86fc-11ef-8be6-7886065e5abf wasn't split correctly, therefore groups cannot be remapped with the new tablet count., at: 0x5e9e22e> -------- seastar::smp_message_queue::async_work_item::invoke_on_all(seastar::smp_submit_to_options, std::function (service::storage_service&)>)::> Oct 11 12:33:35 NODENAME scylla[18799]: [shard 25: gms] table - Found that storage of group 27 for table 51d04920-86fc-11ef-8be6-7886065e5abf wasn't split correctly, therefore groups cannot be remapped with the new tablet count., at: 0x5e9e22e> -------- seastar::smp_message_queue::async_work_item::invoke_on_all(seastar::smp_submit_to_options, std::function (service::storage_service&)>)::> Oct 11 12:33:35 NODENAME scylla[18799]: [shard 22: gms] table - Found that storage of group 41 for table 51d04920-86fc-11ef-8be6-7886065e5abf wasn't split correctly, therefore groups cannot be remapped with the new tablet count., at: 0x5e9e22e> -------- seastar::smp_message_queue::async_work_item::invoke_on_all(seastar::smp_submit_to_options, std::function (service::storage_service&)>)::> Oct 11 12:33:35 NODENAME scylla[18799]: [shard 23: gms] table - Found that storage of group 32 for table 51d04920-86fc-11ef-8be6-7886065e5abf wasn't split correctly, therefore groups cannot be remapped with the new tablet count., at: 0x5e9e22e> -------- seastar::smp_message_queue::async_work_item::invoke_on_all(seastar::smp_submit_to_options, std::function (service::storage_service&)>)::> Oct 11 12:33:35 NODENAME scylla[18799]: [shard 13: gms] table - Found that storage of group 15 for table 51d04920-86fc-11ef-8be6-7886065e5abf wasn't split correctly, therefore groups cannot be remapped with the new tablet count., at: 0x5e9e22e> -------- seastar::smp_message_queue::async_work_item::invoke_on_all(seastar::smp_submit_to_options, std::function (service::storage_service&)>)::> Oct 11 12:33:35 NODENAME scylla[18799]: [shard 12: gms] table - Found that storage of group 23 for table 51d04920-86fc-11ef-8be6-7886065e5abf wasn't split correctly, therefore groups cannot be remapped with the new tablet count., at: 0x5e9e22e> -------- seastar::smp_message_queue::async_work_item::invoke_on_all(seastar::smp_submit_to_options, std::function (service::storage_service&)>)::> Oct 11 12:33:35 NODENAME scylla[18799]: Aborting on shard 0, in scheduling group gossip. Oct 11 12:33:35 NODENAME scylla[18799]: Backtrace: Oct 11 12:33:35 NODENAME scylla[18799]: 0x59d7144 Oct 11 12:33:35 NODENAME scylla[18799]: 0x59965bb Oct 11 12:33:35 NODENAME scylla[18799]: 0x59cbf16 Oct 11 12:33:35 NODENAME scylla[18799]: /opt/scylladb/libreloc/libc.so.6+0x40cff Oct 11 12:33:35 NODENAME scylla[18799]: /opt/scylladb/libreloc/libc.so.6+0x994a3 Oct 11 12:33:35 NODENAME scylla[18799]: /opt/scylladb/libreloc/libc.so.6+0x40c4d Oct 11 12:33:35 NODENAME scylla[18799]: /opt/scylladb/libreloc/libc.so.6+0x28901 Oct 11 12:33:35 NODENAME scylla[18799]: 0x3e22ba7 Oct 11 12:33:35 NODENAME scylla[18799]: 0x13d93ea Oct 11 12:33:35 NODENAME scylla[18799]: 0x59a691f Oct 11 12:33:35 NODENAME scylla[18799]: 0x59a7e8a Oct 11 12:33:35 NODENAME scylla[18799]: 0x59a9077 Oct 11 12:33:35 NODENAME scylla[18799]: 0x59a8428 Oct 11 12:33:35 NODENAME scylla[18799]: 0x5938773 Oct 11 12:33:35 NODENAME scylla[18799]: 0x5937ad3 Oct 11 12:33:35 NODENAME scylla[18799]: 0x13842a5 Oct 11 12:33:35 NODENAME scylla[18799]: 0x1385c60 Oct 11 12:33:35 NODENAME scylla[18799]: 0x13826c3 Oct 11 12:33:35 NODENAME scylla[18799]: /opt/scylladb/libreloc/libc.so.6+0x2a087 Oct 11 12:33:35 NODENAME scylla[18799]: /opt/scylladb/libreloc/libc.so.6+0x2a14a Oct 11 12:33:35 NODENAME scylla[18799]: 0x137fd44 Oct 11 12:33:35 NODENAME scylla[18799]: [shard 5: gms] table - Found that storage of group 21 for table 51d04920-86fc-11ef-8be6-7886065e5abf wasn't split correctly, therefore groups cannot be remapped with the new tablet count., at: 0x5e9e22e> -------- seastar::smp_message_queue::async_work_item::invoke_on_all(seastar::smp_submit_to_options, std::function (service::storage_service&)>)::> Oct 11 12:33:35 NODENAME scylla[18799]: [shard 15: gms] table - Found that storage of group 62 for table 51d04920-86fc-11ef-8be6-7886065e5abf wasn't split correctly, therefore groups cannot be remapped with the new tablet count., at: 0x5e9e22e> -------- seastar::smp_message_queue::async_work_item::invoke_on_all(seastar::smp_submit_to_options, std::function (service::storage_service&)>)::> Oct 11 12:33:35 NODENAME scylla[18799]: [shard 14: gms] table - Found that storage of group 6 for table 51d04920-86fc-11ef-8be6-7886065e5abf wasn't split correctly, therefore groups cannot be remapped with the new tablet count., at: 0x5e9e22e > -------- seastar::smp_message_queue::async_work_item::invoke_on_all(seastar::smp_submit_to_options, std::function (service::storage_service&)>)::> Oct 11 12:33:35 NODENAME scylla[18799]: [shard 28: gms] table - Found that storage of group 17 for table 51d04920-86fc-11ef-8be6-7886065e5abf wasn't split correctly, therefore groups cannot be remapped with the new tablet count., at: 0x5e9e22e> -------- seastar::smp_message_queue::async_work_item::invoke_on_all(seastar::smp_submit_to_options, std::function (service::storage_service&)>)::> Oct 11 12:33:35 NODENAME scylla[18799]: [shard 9: gms] table - Found that storage of group 45 for table 51d04920-86fc-11ef-8be6-7886065e5abf wasn't split correctly, therefore groups cannot be remapped with the new tablet count., at: 0x5e9e22e> -------- seastar::smp_message_queue::async_work_item::invoke_on_all(seastar::smp_submit_to_options, std::function (service::storage_service&)>)::> Oct 11 12:33:35 NODENAME scylla[18799]: [shard 30: gms] table - Found that storage of group 8 for table 51d04920-86fc-11ef-8be6-7886065e5abf wasn't split correctly, therefore groups cannot be remapped with the new tablet count., at: 0x5e9e22e > -------- seastar::smp_message_queue::async_work_item::invoke_on_all(seastar::smp_submit_to_options, std::function (service::storage_service&)>)::> Oct 11 12:33:35 NODENAME scylla[18799]: [shard 31: gms] table - Found that storage of group 11 for table 51d04920-86fc-11ef-8be6-7886065e5abf wasn't split correctly, therefore groups cannot be remapped with the new tablet count., at: 0x5e9e22e> -------- seastar::smp_message_queue::async_work_item::invoke_on_all(seastar::smp_submit_to_options, std::function (service::storage_service&)>)::> Oct 11 12:33:35 NODENAME scylla[18799]: [shard 18: gms] table - Found that storage of group 55 for table 51d04920-86fc-11ef-8be6-7886065e5abf wasn't split correctly, therefore groups cannot be remapped with the new tablet count., at: 0x5e9e22e> -------- seastar::smp_message_queue::async_work_item::invoke_on_all(seastar::smp_submit_to_options, std::function (service::storage_service&)>)::> Oct 11 12:33:35 NODENAME scylla[18799]: [shard 4: gms] table - Found that storage of group 37 for table 51d04920-86fc-11ef-8be6-7886065e5abf wasn't split correctly, therefore groups cannot be remapped with the new tablet count., at: 0x5e9e22e> -------- seastar::smp_message_queue::async_work_item::invoke_on_all(seastar::smp_submit_to_options, std::function (service::storage_service&)>)::> Oct 11 12:33:35 NODENAME scylla[18799]: [shard 21: gms] table - Found that storage of group 43 for table 51d04920-86fc-11ef-8be6-7886065e5abf wasn't split correctly, therefore groups cannot be remapped with the new tablet count., at: 0x5e9e22e> -------- seastar::smp_message_queue::async_work_item::invoke_on_all(seastar::smp_submit_to_options, std::function (service::storage_service&)>)::> Oct 11 12:33:35 NODENAME scylla[18799]: [shard 8: gms] table - Found that storage of group 52 for table 51d04920-86fc-11ef-8be6-7886065e5abf wasn't split correctly, therefore groups cannot be remapped with the new tablet count., at: 0x5e9e22e> -------- seastar::smp_message_queue::async_work_item::invoke_on_all(seastar::smp_submit_to_options, std::function (service::storage_service&)>)::> Oct 11 12:33:35 NODENAME scylla[18799]: [shard 17: gms] table - Found that storage of group 56 for table 51d04920-86fc-11ef-8be6-7886065e5abf wasn't split correctly, therefore groups cannot be remapped with the new tablet count., at: 0x5e9e22e> -------- seastar::smp_message_queue::async_work_item::invoke_on_all(seastar::smp_submit_to_options, std::function (service::storage_service&)>)::> Oct 11 12:33:35 NODENAME scylla[18799]: [shard 16: gms] table - Found that storage of group 61 for table 51d04920-86fc-11ef-8be6-7886065e5abf wasn't split correctly, therefore groups cannot be remapped with the new tablet count., at: 0x5e9e22e> -------- seastar::smp_message_queue::async_work_item::invoke_on_all(seastar::smp_submit_to_options, std::function (service::storage_service&)>)::> Oct 11 12:33:35 NODENAME scylla[18799]: [shard 0: gms] storage_service - Failed to apply token_metadata changes: seastar::internal::backtraced (Found that storage of group 1 for table 51d04920-86fc-11ef-8be6-7886065e5abf w> -------- seastar::internal::coroutine_traits_base::promise_type -------- seastar::internal::coroutine_traits_base::promise_type -------- seastar::internal::coroutine_traits_base::promise_type -------- seastar::internal::coroutine_traits_base::promise_type -------- seastar::internal::coroutine_traits_base::promise_type -------- seastar::internal::coroutine_traits_base::promise_type -------- seastar::internal::coroutine_traits_base::promise_type -------- seastar::internal::coroutine_traits_base::promise_type). Aborting. Oct 11 12:37:21 NODENAME systemd[1]: scylla-server.service: Main process exited, code=dumped, status=6/ABRT Oct 11 12:37:21 NODENAME systemd[1]: scylla-server.service: Failed with result 'core-dump'. Oct 11 12:37:22 NODENAME systemd[1]: scylla-server.service: Scheduled restart job, restart counter is at 1. Oct 11 12:37:22 NODENAME systemd[1]: Stopped Scylla Server. Oct 11 12:37:22 NODENAME systemd[1]: Starting Scylla Server... Oct 11 12:37:22 NODENAME scylla[60269]: Scylla version 6.1.2-0.20240915.b60f9ef4c223 with build-id c713ac9e819492d7560aa3ad461c43cf404c977b starting ... Oct 11 12:37:22 NODENAME scylla[60269]: command used: "/usr/bin/scylla --log-to-syslog 1 --log-to-stdout 0 --default-log-level info --network-stack posix --io-properties-file=/etc/scylla.d/io_properties.yaml --lock-memory=1" Oct 11 12:37:22 NODENAME scylla[60269]: pid: 60269 Oct 11 12:37:22 NODENAME scylla[60269]: parsed command line options: [log-to-syslog, (positional) 1, log-to-stdout, (positional) 0, default-log-level, (positional) info, network-stack, (positional) posix, io-properties-file: /etc/scylla.d/io_pr> Oct 11 12:37:22 NODENAME scylla[60269]: seastar - Reactor backend: linux-aio Oct 11 12:37:23 NODENAME scylla[60269]: seastar - Perf-based stall detector creation failed (EACCESS), try setting /proc/sys/kernel/perf_event_paranoid to 1 or less to enable kernel backtraces: falling back to posix timer. Oct 11 12:37:24 NODENAME scylla[60269]: [shard 0:main] seastar - updated: blocked-reactor-notify-ms=36000000000 Oct 11 12:37:24 NODENAME scylla[60269]: [shard 3:main] seastar - updated: blocked-reactor-notify-ms=36000000000 Oct 11 12:37:24 NODENAME scylla[60269]: [shard 10:main] seastar - updated: blocked-reactor-notify-ms=36000000000 Oct 11 12:37:24 NODENAME scylla[60269]: [shard 6:main] seastar - updated: blocked-reactor-notify-ms=36000000000 Oct 11 12:37:24 NODENAME scylla[60269]: [shard 20:main] seastar - updated: blocked-reactor-notify-ms=36000000000 Oct 11 12:37:24 NODENAME scylla[60269]: [shard 29:main] seastar - updated: blocked-reactor-notify-ms=36000000000 Oct 11 12:37:24 NODENAME scylla[60269]: [shard 11:main] seastar - updated: blocked-reactor-notify-ms=36000000000 Oct 11 12:37:24 NODENAME scylla[60269]: [shard 2:main] seastar - updated: blocked-reactor-notify-ms=36000000000 Oct 11 12:37:24 NODENAME scylla[60269]: [shard 27:main] seastar - updated: blocked-reactor-notify-ms=36000000000 Oct 11 12:37:24 NODENAME scylla[60269]: [shard 9:main] seastar - updated: blocked-reactor-notify-ms=36000000000 Oct 11 12:37:24 NODENAME scylla[60269]: [shard 5:main] seastar - updated: blocked-reactor-notify-ms=36000000000 Oct 11 12:37:24 NODENAME scylla[60269]: [shard 4:main] seastar - updated: blocked-reactor-notify-ms=36000000000 Oct 11 12:37:24 NODENAME scylla[60269]: [shard 1:main] seastar - updated: blocked-reactor-notify-ms=36000000000 Oct 11 12:37:24 NODENAME scylla[60269]: [shard 0:main] init - Option is deprecated : force_schema_commit_log Oct 11 12:37:24 NODENAME scylla[60269]: [shard 0:main] init - installing SIGHUP handler Oct 11 12:37:24 NODENAME scylla[60269]: [shard 0:main] init - Scylla version 6.1.2-0.20240915.b60f9ef4c223 with build-id c713ac9e819492d7560aa3ad461c43cf404c977b starting ... Oct 11 12:37:24 NODENAME scylla[60269]: [shard 0:main] init - starting API server Oct 11 12:37:24 NODENAME scylla[60269]: [shard 0:main] init - starting prometheus API server Oct 11 12:37:24 NODENAME scylla[60269]: [shard 0:main] init - creating snitch Oct 11 12:37:24 NODENAME scylla[60269]: [shard 0:main] snitch_logger - GCESnitch using region: asia-southeast1, zone: b. Oct 11 12:37:24 NODENAME scylla[60269]: [shard 0:main] init - starting tokens manager Oct 11 12:37:24 NODENAME scylla[60269]: [shard 0:main] init - starting effective_replication_map factory Oct 11 12:37:24 NODENAME scylla[60269]: [shard 0:main] init - starting migration manager notifier Oct 11 12:37:24 NODENAME scylla[60269]: [shard 0:main] init - starting per-shard database core Oct 11 12:37:24 NODENAME scylla[60269]: [shard 0:main] init - creating and verifying directories Oct 11 12:37:24 NODENAME scylla[60269]: [shard 0:main] init - starting compaction_manager Oct 11 12:37:24 NODENAME scylla[60269]: [shard 0:main] task_manager - Registered module compaction Oct 11 12:37:24 NODENAME scylla[60269]: [shard 7:main] task_manager - Registered module compaction` Tried taking service reboot but didnt helped. Any resolution for such cases where [shard 5: gms] table - Found that storage of group 21 for table 51d04920-86fc-11ef-8be6-7886065e5abf wasn’t split correctly, therefore groups cannot be remapped with the new tablet count. Let us know if anything is needed for debugging ? For the record, this is a duplicate of scylla restart fails after MV created (core dump [shard 11:main] table - Unable to load SSTable) · Issue #20626 · scylladb/scylladb · GitHub which is fixed now in 6.1.3, 6.2.x and later --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-251-2024-10-13/3038 Title: Last week in scylladb.git master (issue #251; 2024-10-13) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 882a3c60e4…18d3a6480d range are covered. There were 111 non-merge commits from 23 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-251-2024-10-13/3038 ## Headings Structure: H1: Last week in scylladb.git master (issue #251; 2024-10-13) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #251; 2024-10-13) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 882a3c60e4…18d3a6480d range are covered. There were 111 non-merge commits from 23 authors in that period. Some notable commits: The alternator /localnodes REST API return nodes local to the current datacenter. It is now enhanced to be able to return other datacenters, and to restrict the returned nodes to a specific rack. This can help reduce networking costs. A crash during shutdown while draining active writes has been fixed. Time Window Compaction Strategy will now return fresher values for its pending compaction count estimate. The sstable scrub/validate commands will now check the digest of the entire sstable data file, in addition to the already-checked per-block checksums. The sstable Scylla.db component now has a copy of the sstable uuid. This is needed for backup deduplication, since the sstable uuid in the file name may change when a tablet is migrated to a different shard. Raft-managed tables used for system metadata now have more eager garbage collection of tombstones, reducing performance problems with many schema or topology changes. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/scylladb-as-a-dynamodb-alternative-frequently-asked-questions/3039 Title: ScyllaDB as a DynamoDB Alternative: Frequently Asked Questions - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: In a recent blog, Felipe Cardeneti (Technical Director at ScyllaDB) wrote… A look at the top questions engineers are asking about moving from DynamoDB to ScyllaDB to reduce cost, avoid throttling, and avoid cloud vendor… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-as-a-dynamodb-alternative-frequently-asked-questions/3039 ## Headings Structure: H1: ScyllaDB as a DynamoDB Alternative: Frequently Asked Questions H2: Why switch from DynamoDB to ScyllaDB? H2: Is ScyllaDB Alternator a DynamoDB drop-in replacement? H2: What are the main differences between ScyllaDB Alternator and AWS DynamoDB? H3: Related topics ## Main Content: H1: ScyllaDB as a DynamoDB Alternative: Frequently Asked Questions H2: Why switch from DynamoDB to ScyllaDB? H2: Is ScyllaDB Alternator a DynamoDB drop-in replacement? H2: What are the main differences between ScyllaDB Alternator and AWS DynamoDB? H3: Related topics In a recent blog, Felipe Cardeneti (Technical Director at ScyllaDB) wrote… A look at the top questions engineers are asking about moving from DynamoDB to ScyllaDB to reduce cost, avoid throttling, and avoid cloud vendor lockin A great thing about working closely with our community is that I get a chance to hear a lot about their needs and – most importantly – listen to and take in their feedback. Lately, we’ve seen a growing interest from organizations considering ScyllaDB as a means to replace their existing DynamoDB deployments and, as happens with any new tech stack, some frequently recurring questions. ScyllaDB provides you with multiple ways to get started: you can choose from its CQL protocol or ScyllaDB Alternator. CQL refers to the Cassandra Query Language, a NoSQL interface that is intentionally similar to SQL. ScyllaDB Alternator is ScyllaDB’s DynamoDB-compatible API, aiming at full compatibility with the DynamoDB protocol. Its goal is to provide a seamless transition from AWS DynamoDB to a cloud-agnostic or on-premise infrastructure while delivering predictable performance at scale. But which protocol should you choose? What are the main differences between ScyllaDB and DynamoDB? And what does a typical migration path look like? I want to answer some of these top questions right here, and right now. If you are here, chances are that you fall under at least one of the following categories: ScyllaDB delivers predictable low latency at scale with less infrastructure required. For DynamoDB specifically, we have an in-depth article covering which pain points we address. In the term’s strict sense, it is not: notable differences across both solutions exist. DynamoDB development is closed source and driven by AWS (which ScyllaDB is not affiliated with), which means that there’s a chance that some specific features launched in DynamoDB may take some time to land in ScyllaDB. A more accurate way to describe it is as an almost drop-in replacement. Whenever you migrate to a different database, some degree of changes will always be required to get started with the new solution. We try to keep the level of changes to a minimum to make the transition as seamless as possible. For example, Digital Turbine easily migrated from DynamoDB to ScyllaDB within just a single two-week sprint, the results showing significant performance improvements and cost savings. Keep reading at ScyllaDB as a DynamoDB Alternative: Frequently Asked Questions - ScyllaDB --- ### Page: https://forum.scylladb.com/t/scylladb-rust-best-practice/3042 Title: ScyllaDB Rust best practice - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: According to ScyllaDB rust documentation session use more resources. So as a developer do I need to keep single separate session for all queries? If yes, how can I do it in rust? Otherwise need to create session for eve… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-rust-best-practice/3042 ## Headings Structure: H1: ScyllaDB Rust best practice H3: Related topics ## Main Content: H1: ScyllaDB Rust best practice H3: Related topics According to ScyllaDB rust documentation session use more resources. So as a developer do I need to keep single separate session for all queries? If yes, how can I do it in rust? Otherwise need to create session for every rust file if it has a database connection. As you said, creating the Session object is expensive, because it requires opening a lot of connections to the cluster. This section in documentation explains all that: Connecting to the cluster | ScyllaDB Docs It also provides a solution to your problem: you should use Arc to share it between threads / tasks / objects etc. Ideally you will create a single Arc object during your application startup and share it with all the code that needs it. Yeah I saw it, Problem is how to share it and when use different ROLE that will be different sessions for each ROLE so it will be much expensive. So that is why I need to check any better solution. You are right, that is the valid use case for multiple Session’s - you will need to create a separate Session for each role. I doubt it will cause performance problems - how many roles do you have? I just keep it as 2 for SELECT and MODIFY. It will be totally fine to have 2 (or some other amount) sessions instead of one. The point of our recommendation in the docs is to not do things like session per user request (in case of e.g. web service) or even worse, session per DB statement. Creating N sessions on startup is fine. --- ### Page: https://forum.scylladb.com/t/how-to-verify-that-a-select-is-using-serial-consistency/3043 Title: How to verify that a select is using SERIAL consistency - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: If i have a gocql query like err = tx.session.Query(fmt.Sprintf("select col, ts, val from \"%s\" where key = ? and ts > ? and col = 'w' order by ts asc limit 1", tx.table), key, tx.readTime.UnixNano()).SerialConsistency… Language: en Canonical URL: https://forum.scylladb.com/t/how-to-verify-that-a-select-is-using-serial-consistency/3043 ## Headings Structure: H1: How to verify that a select is using SERIAL consistency H3: Related topics ## Main Content: H1: How to verify that a select is using SERIAL consistency H3: Related topics If i have a gocql query like that sounds to me like they wont bother applying that consistency level to a select (it will be ignored). Is there any way to debug this so I can verify that select used serial consistency, rather than a non paxos consistency? A subsequent question could be: Does using the .SerialConsistency apply to select statements in gocql? It does not seem like this is simple to verify. You need to set Consistency to SERIAL to run SELECT with SERIAL consistency. You can verify it with LWT metrics in the dashboards and in rest API - specifically, all cas_* metrics will bump once you execute a SELECT with SERIAL consistency, as it runs a full blown Paxos round with EMPTY mutation, so essentially SELECT in SERIAL mode is a write. --- ### Page: https://forum.scylladb.com/t/error-sending-messages-automatically-on-the-scylla-university-website/3046 Title: Error sending messages automatically on the Scylla University website - University and Training - ScyllaDB Community NoSQL Forum Meta Description: The problem has already been sent to university@scylladb.com Good afternoon. I logged in through my Github account and should have received a confirmation email, but it didn’t arrive. Can you resend? Or fix sending emai… Language: en Canonical URL: https://forum.scylladb.com/t/error-sending-messages-automatically-on-the-scylla-university-website/3046 ## Headings Structure: H1: Error sending messages automatically on the Scylla University website H3: Related topics ## Main Content: H1: Error sending messages automatically on the Scylla University website H3: Related topics The problem has already been sent to university@scylladb.com Good afternoon. I logged in through my Github account and should have received a confirmation email, but it didn’t arrive. Can you resend? Or fix sending emails? Thank you in advance) Hi @NikiYani - I asked, and there are many emails tied to it. Please email me at felipe (at) .com and we’ll get in touch with you. I will close this post for now, also feel free to DM me here. --- ### Page: https://forum.scylladb.com/t/scylladb-rust-serializevalue-for-cql-types-and-type-values/3048 Title: Scylladb rust SerializeValue for CQL types and type values - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This is struct use for insert and select data. But SerializeValue given me { user_id: 45b5b8f7-4b5e-4df6-aedc-74327fba0ad5, first_name: “Jane”, last_name: “Smith”, email: “janesmith_example.com”, phone: Some(“234-567-8… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-rust-serializevalue-for-cql-types-and-type-values/3048 ## Headings Structure: H1: Scylladb rust SerializeValue for CQL types and type values H3: Related topics ## Main Content: H1: Scylladb rust SerializeValue for CQL types and type values H3: Related topics This is struct use for insert and select data. But SerializeValue given me { user_id: 45b5b8f7-4b5e-4df6-aedc-74327fba0ad5, first_name: “Jane”, last_name: “Smith”, email: “janesmith_example.com”, phone: Some(“234-567-8901”), address: Some(“456 Oak St, Townsville”), shipping_address: Some(“101 E-Commerce St, Townsville”), date_of_birth: Some(CqlDate(2147489452)), passwordv: “hashed_password_2”, salt: “random_salt_2”, role: Some(“seller”), is_active: true, last_login: CqlTimestamp(1729076400000), created_at: CqlTimestamp(1727254800000), updated_at: CqlTimestamp(1729078200000) } How can i get values only? use scylla::{ frame::value::{CqlDate, CqlTimestamp}, FromRow, SerializeRow, SerializeValue, }; use uuid::Uuid; #[derive(Debug, SerializeRow, FromRow, SerializeValue)] pub struct User { pub user_id: Uuid, pub first_name: String, pub last_name: String, pub email: String, pub phone: Option, pub address: Option, pub shipping_address: Option, pub date_of_birth: Option, pub passwordv: String, pub salt: String, pub role: Option, pub is_active: bool, pub last_login: CqlTimestamp, pub created_at: CqlTimestamp, pub updated_at: CqlTimestamp, } How can i get values only? Remove CqlTimestamp, Some and CqlDate in JSON. Sorry, I don’t really understand the question. Could you share how did you create this JSON? --- ### Page: https://forum.scylladb.com/t/scylladb-mass-insertion-advice-needed/3049 Title: ScyllaDB mass insertion advice needed - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi Scylla Team, So come back after month and my project is up again. TLDR; I wanted to replace my C* cluster by Scylla (hoping for perf gain and cost reduction). Context: I have 3 clusters (different sizing depending… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-mass-insertion-advice-needed/3049 ## Headings Structure: H1: ScyllaDB mass insertion advice needed H3: Related topics ## Main Content: H1: ScyllaDB mass insertion advice needed H3: Related topics So come back after month and my project is up again. TLDR; I wanted to replace my C* cluster by Scylla (hoping for perf gain and cost reduction). I have 3 clusters (different sizing depending on region). The dataset can be big : the most problematic ks/table is 3 billions rows with 1.3 new billions row a day (and then if you do the math quick expiration with TTL) On the biggest region my setup is currently (C* 3.1) : 15 x c6id.8xlarge (32 cores 64g ram local nvme) Peak read/s : 700K requests / s Write : nothing expect when loading (done by sstable loader mostly take 2H) On smalests regions : 8 x c7a.4xlarge (16 cores 32g ram EBS ; I know) Peak read/s : 50K requests / s Write : nothing expect when loading (done by sstable loader mostly take 2H) So my biggest problem is how loading my 1.2 billions row dataset in a reasonable amount of time. On C* sstableloader stream my already made sstables and stream them fast enough to the right shard. My scylla test cluster is currently : 4x i4i.4xlarge (16cpu, 64g ram, NVME) an instance which is advised by the doc. On scylla sstableloader make sstable to cql translation which is slow as hell. I know this is adviced to use a lot of thread and //ize the tool but even I’ve got bad result (20k insert per second, do the math it will take age) What I tried also as insertion method : I really wonder if I missed something : like an obvious settings ? or a way to debug Also I tried nodetool refresh load-and-stream and results are also a bit dissapointing : so again maybe I miss something obvious ? by the way do you know if there is tooling arround nodetool refresh? Where I was surprised is why I am so far from theoretical qps stated on the site. 12.5K per core ; so I my setup 32 cores (64 vcpus) RF=1 I should be around 400k so I am at 15% expectations at best Again I have the feeling that I make something very bad. Maybe something related to the schema of the table (81 columms)? I read this something bad. Does this hit the performance that bad? Anyway I take any advices; without help I will stuck with C* When using load-and-stream, you need to ensure you have good enough concurrency. Upload a batch of sstables to all nodes, and start nodetool refresh -las concurrently on all of them. When the batch finishes on any node, upload the next batch. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-67-2024-10-18/3054 Title: Last week in scylla-cluster-tests.git master (issue #67; 2024-10-18) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the c2a8f32a…99d539bb range are covered. There were 11 non-merge commits from 6 authors in tha… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-67-2024-10-18/3054 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #67; 2024-10-18) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #67; 2024-10-18) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the c2a8f32a…99d539bb range are covered. There were 11 non-merge commits from 6 authors in that period. Some notable commits: When running tests against the same cluster with the ‘reuse_cluster’ feature, some test stages may now be skipped. This allows tests to restart from a specific stage, bypassing intermediary stages (e.g., without reloading the cluster with data, which can be time-consuming). Several fixes have been made around monitoring the decommission topology operation, improving the stability of the decommission_streaming_err nemesis. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-252-2024-10-20/3057 Title: Last week in scylladb.git master (issue #252; 2024-10-20) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 18d3a6480d…29c2d4e7eb range are covered. There were 84 non-merge commits from 16 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-252-2024-10-20/3057 ## Headings Structure: H1: Last week in scylladb.git master (issue #252; 2024-10-20) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #252; 2024-10-20) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 18d3a6480d…29c2d4e7eb range are covered. There were 84 non-merge commits from 16 authors in that period. Some notable commits: The system.sstables table names all the sstables attached to a table, when the sstables are non-local (e.g. on object storage). It now uses the table id rather than a path as its partition key. The CQL server will now wait for the superuser to be created when authentication is enabled. A deprecation noticed as added to some of the Java-based tools. The memtable_flush_period_in_ms option is now implemented. Internode encryption now supports a transitional mode, allowing a cluster to switch to internode encryption without downtime. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/failed-to-refill-emergency-reserve/3058 Title: Failed to refill emergency reserve - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: When I perform removenode operation on a node with permanent DN, other nodes DN are gone. I see the following abnormal information on the DN node: storage_proxy - Exception when communicating with 10.6.0.2, to read from… Language: en Canonical URL: https://forum.scylladb.com/t/failed-to-refill-emergency-reserve/3058 ## Headings Structure: H1: Failed to refill emergency reserve H3: Related topics ## Main Content: H1: Failed to refill emergency reserve H3: Related topics When I perform removenode operation on a node with permanent DN, other nodes DN are gone. I see the following abnormal information on the DN node: Does this result in a node DN? The to read from system.paxos makes me think this is related to LWT and since we don’t use LWT internally, therefore this error shouldn’t cause a node to be DN. This seems like a simple user read failure. I saw that Scylla threw the following message before DN The decode information is as follows: Is it caused by a memory allocation failure during repair? Are you reporting the same as The expansion of the 180-node cluster has failed by any chance? Heavy memory pressure and both in ScyllaDB 5.4. Is it caused by a memory allocation failure during repair? The decoded backtrace definitely points to row_level repair codepath. You may disable it and fallback to old streaming and see if it helps. But as in the aforementioned forum post, do check how your LSA/Non-LSA consumptions were at the time of the crash and take it from there. Same upgrade recommendations and next steps apply. --- ### Page: https://forum.scylladb.com/t/scylla-operator-not-creating-pods/3066 Title: Scylla operator not creating pods - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello, I’m trying to get ScyllaDB up and running in Minikube using the Scylla operator. I’m following the Scylla operator documentation step by step: Everything seems to be going well until I run the following command: … Language: en Canonical URL: https://forum.scylladb.com/t/scylla-operator-not-creating-pods/3066 ## Headings Structure: H1: Scylla operator not creating pods H3: Build software better, together H3: Related topics ## Main Content: H1: Scylla operator not creating pods H3: Build software better, together H3: Related topics I’m trying to get ScyllaDB up and running in Minikube using the Scylla operator. I’m following the Scylla operator documentation step by step: Everything seems to be going well until I run the following command: kubectl -n scylla get pods The output of that command is: No resources found in scylla namespace. That shouldn’t be the case. There should be 3 pods up and running. The thing is I’ve done every check to make sure everything is running well and that seems to be the case. The only command that gives incorrect output is the ‘get pods’ command that I mentioned above. Does anyone know what could be the cause of the pods not spinning up and how to fix it? this is hard to tell without any data. Please file an issue on the GitHub repo and supply must-gather. GitHub is where people build software. More than 100 million people use GitHub to discover, fork, and contribute to over 420 million projects. @davidz , did you find a solution or open a GH issue? Please post here for others who run into a similar issue. --- ### Page: https://forum.scylladb.com/t/lwt-behavior-with-abandoned-lwt-normal-write/3067 Title: LWT behavior with abandoned LWT + normal write - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Since LWT’s get an isolated view the DB, what would the resulting value be in the following scenario: Row A col=1 LWT begins with Update A set col=2 if col=1 LWT fails before completing, currently only a paxos read wou… Language: en Canonical URL: https://forum.scylladb.com/t/lwt-behavior-with-abandoned-lwt-normal-write/3067 ## Headings Structure: H1: LWT behavior with abandoned LWT + normal write H3: Related topics ## Main Content: H1: LWT behavior with abandoned LWT + normal write H3: Related topics Since LWT’s get an isolated view the DB, what would the resulting value be in the following scenario: Will that paxos read see col=2 or col=3? Technically the LWT finished later, but it was based on an old isolated view of the DB. My assumption is that it would see col=3 because even though the LWT would have been applied, there has been a new value since it started and thus it won’t apply the LWT. It depends on the timing of the failure, unfortunately. If the paxos round is finished, it just failed to apply to the base table, the apply will use the original round timestamp. So a (supposedly later) timestamp of step 4 will prevail and the result will be “3”. If, however, the Paxos write fails at accept phase, the SELECT with Storng Consistency, which sees the unfinished accept round, will finish it with a refreshed ballot, so the result is going to be 2. You can see this for yourself in storage_proxy.cc, begin_and_repair_paxos, search for replicas_missing_most_recent_commit, and observe the used mutation timestamp, and then compare with refreshed_in_progress and accept_proposal --- ### Page: https://forum.scylladb.com/t/repair-multiple-node-same-time/3071 Title: Repair multiple node same time - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I have a big cluster around 20 node, and data very big I see doc say best do repair regularly and sequentially current I execute nodetool repair -pr need 1 day per node and I thank my cluster have data issue… there ar… Language: en Canonical URL: https://forum.scylladb.com/t/repair-multiple-node-same-time/3071 ## Headings Structure: H1: Repair multiple node same time H3: Related topics ## Main Content: H1: Repair multiple node same time H3: Related topics I have a big cluster around 20 node, and data very big I see doc say best do repair regularly and sequentially current I execute nodetool repair -pr need 1 day per node and I thank my cluster have data issue… there are many write fail I hope can do repair as soon as possible so I want to ask can I exec repair in multiple node same time ? It would be better if you used ScyllaDB Manager for the job as it knows to handle repair parallelism and intensity with ease, but I see you have a 20-node cluster. That said, you can repair multiple nodes at the same time. You could also simply run nodetool repair once and let it eventually complete. However, you’ve already said you currently are having write failures, so ideally you’d rather check on these before adding more background work to your running cluster. --- ### Page: https://forum.scylladb.com/t/implement-transaction-in-client-side/3082 Title: Implement transaction in client side - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Howdy people! This is my first post on this forum, but I’ve been using ScyllaDB for the better part of this year. My use case requires the use of transactions. This means updating a group of records across partitions wi… Language: en Canonical URL: https://forum.scylladb.com/t/implement-transaction-in-client-side/3082 ## Headings Structure: H1: Implement transaction in client side H3: Related topics ## Main Content: H1: Implement transaction in client side H3: Related topics Howdy people! This is my first post on this forum, but I’ve been using ScyllaDB for the better part of this year. My use case requires the use of transactions. This means updating a group of records across partitions with guaranteed all succeed or rollback, which LWT doesn’t provide. I’m taking inspiration from GitHub - awslabs/dynamodb-transactions with a CQL port to tackle this issue. A simple concrete example – When a user wants to change his email, we need to update both tables, namely: I am aware that we can use GSI, but I’m not doing that for a variety of reasons. Taking a page out of the dynamodb-transactions repo, what I think may be possible (node here means the client program that is interacting with the database): This will require N quorum reads from (2), 1 INSERT with one consistency from (3), and 3N+2 LWT from (4) through (8). Do you suppose this is a good idea? Also a side question: can I regard the result from a LWT query as of quorum consistency? I am aware that we can use GSI, but I’m not doing that for a variety of reasons. Note we do support synchronous GSIs. You may also TTL it, as ideally this tx shouldn’t live for too long anyway. You definitely want to TTL the lock field for all items, in case a worker dies and you don’t hold it for too long (unless you plan to come up with a custom-made unlocking mechanism) Looks fine, could maybe be optimized a few. FWIW, Service Resilience — part 3: Distributed Locking | by Martina Alilovic Rojnic | ReversingLabs Engineering | Medium is a nice write-up on how ReversingLabs did it with ScyllaDB, so it may also help you out. Also a side question: can I regard the result from a LWT query as of quorum consistency? See SERIAL CONSISTENCY in Lightweight Transactions | ScyllaDB Docs, though yea - it requires a majority for consensus. --- ### Page: https://forum.scylladb.com/t/how-to-optimize-scylladb-performance-for-high-throughput-applications/3084 Title: How to Optimize ScyllaDB Performance for High Throughput Applications? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hey guys… :wave: I am currently working on a high throughput application that requires low-latency and high concurrency for handling a large number of requests. I have chosen ScyllaDB due to its performance promises but… Language: en Canonical URL: https://forum.scylladb.com/t/how-to-optimize-scylladb-performance-for-high-throughput-applications/3084 ## Headings Structure: H1: How to Optimize ScyllaDB Performance for High Throughput Applications? H3: Related topics ## Main Content: H1: How to Optimize ScyllaDB Performance for High Throughput Applications? H3: Related topics I am currently working on a high throughput application that requires low-latency and high concurrency for handling a large number of requests. I have chosen ScyllaDB due to its performance promises but I am now looking for advice on optimizing it further for my use case. My application is built on a microservices architecture, where each service performs frequent reads and writes. I am already running a ScyllaDB cluster with 3 nodes, each equipped with NVMe SSDs and 64GB of RAM, but I feel there’s more I can do to fully exploit ScyllaDB’s potential. Here are a few specific points I am looking for guidance on: I also check this: https://forum.scylladb.com/t/scylladb-summit-day-2-continuing-the-high-performance-nosql-conversatlooker But I have not found any solution. Could anyone guide me about this? I have read through the documentation and a few case studies but would love to hear from the community about real-world experiences. Respected community member! In general, there is no one-size-fits-all answer to these questions, as it depends mostly on your use case requirements. That said, I will be as generic as possible and hope it is somehow useful to you: Ensure you walk through our Production Readiness Guidelines | ScyllaDB Docs . If your services are heavy-writers or heavy-readers, then there’s a nice table in https://docs.scylladb.com/stable/architecture/compaction/compaction-strategies.html which should be your guidance. Yes, we recently hosted a Data Modeling for NoSQL Databases Masterclass which should be a good starting point. ScyllaDB University courses are also a great way to get started. Disk utilization, node health are some common KPIs. Get familiar with the Detailed dashboard though, there you can find granular P99 latencies, compaction information, foreground/background queue sizes, cache hit-rate among many other goodies. All that said, our Database Performance at Scale: A Practical Guide covers most of these, and more in a consolidated way if you’d like to read on something. We organized it with recommendations a per-workload basis, driver settings, benchmarking, Monitoring KPIs, among many other topics which are probably somehow related to your question. More recently, ScyllaDB In Action | By Bo Ingram, Staff Engineer at Discord got published, which is very ScyllaDB specific on its own. Lastly, we do offer 1:1 Technical Consultations, so if you’d like to talk to someone, simply book a 30 min slot in ScyllaDB | Technical Strategy Session Hope you find these useful and that you enjoy ScyllaDB! --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-68-2024-10-25/3099 Title: Last week in scylla-cluster-tests.git master (issue #68; 2024-10-25) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 9018c7f2…ee40cd4f range are covered. There were 6 non-merge commits from 5 authors in tha… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-68-2024-10-25/3099 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #68; 2024-10-25) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #68; 2024-10-25) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 9018c7f2…ee40cd4f range are covered. There were 6 non-merge commits from 5 authors in that period. Some notable commits: The hydra utility can now be run on MacOS with Docker Desktop installed. The show-logs command in SCT now includes a --update-argus argument, which updates Argus with log links that were not submitted during the test. The 1to5 ratio test with SLA for 2024.3 was fixed by adjusting loader size, stress command, and setting 1024 initial tablets by default. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/limitations-of-tiny-partitions/3103 Title: Limitations of tiny partitions - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I am having around 150 million rows in a table. Each row is a partition which means 150 million partitions. Size of each partition is 24kb. I know that small partitions work perfectly for even distribution of data witho… Language: en Canonical URL: https://forum.scylladb.com/t/limitations-of-tiny-partitions/3103 ## Headings Structure: H1: Limitations of tiny partitions H3: Related topics ## Main Content: H1: Limitations of tiny partitions H3: Related topics I am having around 150 million rows in a table. Each row is a partition which means 150 million partitions. Size of each partition is 24kb. I know that small partitions work perfectly for even distribution of data without hotspots. But, having partitions as small as 24kb can be a problem in managing them in cache (c++ objects limitations) or using mv’s/secondary index for this table in scylla. Is this true? Is this improved in latest version of scylla? No problem. ScyllaDB is designed to support billions of partitions. Thanks for your reply. Just to cross check, tiny partitions won’t be blocker for You should be fine. Plus, if needed, you can always tune bloom filter, repair intensities, and compaction settings in case you find anything weird later on. --- ### Page: https://forum.scylladb.com/t/the-expansion-of-the-180-node-cluster-has-failed/3111 Title: The expansion of the 180-node cluster has failed - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I currently have a 180-node cluster, distributed across 3 availability zones (AZs), with 60 nodes in each AZ. When I tried to add a new node to one of the AZs, the node failed during the expansion process. The reason for… Language: en Canonical URL: https://forum.scylladb.com/t/the-expansion-of-the-180-node-cluster-has-failed/3111 ## Headings Structure: H1: The expansion of the 180-node cluster has failed H3: Related topics ## Main Content: H1: The expansion of the 180-node cluster has failed H3: Related topics I currently have a 180-node cluster, distributed across 3 availability zones (AZs), with 60 nodes in each AZ. When I tried to add a new node to one of the AZs, the node failed during the expansion process. The reason for the failure was that other nodes in the cluster gradually went down (DN) during the expansion, and the cause of the DN was bad_alloc. Note: I was performing the expansion with business testing, during which a single node had a write load of 1.5k and already contained 1TB of existing data. I want to know why the expansion failed and how to fix it. Here are the logs from the expansion node. Check lsa memory and non-lsa memory per shard in the Detailed dashboard. It’s likely one of the shards is overloaded with metadata. What version are you running? How many shards per node and how much memory per node do you have? Do you have a large amount of keyspaces/tables? Thanks for your reply ! I am now using Scylla version 5.4. Each node has 48 shards and 310 GB of memory. There are more than sixty tables in the cluster, but mainly one table is being written to, which has approximately 170 billion objects, each object is 1 byte in size. The LSA memory of the node was very low before the restart, but it increases after the abort and restart. Here are more information about this expansion: backtrace: The LSA memory of the node was very low before the restart, but it increases after the abort and restart. LSA basically holds cache and memtables and dynamically gets resized. The Non-LSA panel is what’s interesting here: This cyan node (for some reason) had a very high memory consumption, until it eventually OOM’d. One possibility is as @avikivity mentioned, high memory for metadata. Too many small keys under a single shard, perhaps? It is better if you focus on this specific node, check its per-shard view and take it from there. Also upgrade, and if the problem persists, raise a GitHub issue and upload the generated coredump accordingly. Or you may find your way through scylladb/docs/dev/debugging.md at master · scylladb/scylladb · GitHub if you ain’t like upgrading atm (in which case 5.4 is an already EOL release) Right. But the fact the non-LSA memory is low after the restart indicates the problem is not with metadata (since it would then reccover) but with something else. However, since it’s now low, we cannot investigate. Suggest you monitor non-LSA memory and if it starts increasing, we can try to understand why. p.s. 5.4 has reached end-of-life and is no longer supported. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-253-2024-10-27/3112 Title: Last week in scylladb.git master (issue #253; 2024-10-27) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 29c2d4e7eb…e7d6ab576b range are covered. There were 55 non-merge commits from 16 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-253-2024-10-27/3112 ## Headings Structure: H1: Last week in scylladb.git master (issue #253; 2024-10-27) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #253; 2024-10-27) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 29c2d4e7eb…e7d6ab576b range are covered. There were 55 non-merge commits from 16 authors in that period. Some notable commits: The system.peers table was continuously updated even if no change was happening, stressing the disk. This is now fixed. Materialized views perform a read operation on the base table before writing the the view table. This read operation now has better concurrency control: the amount of memory consumed by reads is limited, and when the CPU is the bottleneck, we avoid issuing new reads to avoid flooding the system with competing operations. When using S3, we upload files in chunks. We now recover those chunks on error and delete them. The sstable reader will now consult data in memtable before purging tombstones. This prevents data resurrection in scenarios involving very low write activity, which can lead to data staying in memtables for longer than a repair cycle. Data definition language (DDL) statements are automatically retried in case of an internal race accessing Raft. A crash during this retry, for ALTER KEYSPACE statements, was fixed. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylla-doctor-v1-3/3114 Title: [RELEASE]: Scylla Doctor v1.3 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: Scylla Doctor v1.3 is released. Added Collectors: SystemConfigCollector: collects Scylla actual (in-memory) configuration state. InfrastructureProviderCollector: collect CPU platform where possible. Added Analyzers:… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-doctor-v1-3/3114 ## Headings Structure: H1: [RELEASE]: Scylla Doctor v1.3 H3: Related topics ## Main Content: H1: [RELEASE]: Scylla Doctor v1.3 H3: Related topics Scylla Doctor v1.3 is released. Artifacts can be downloaded from https://downloads.scylladb.com/downloads/scylla-doctor/ or installed from Scylla OSS or Scylla Enterprise repositories. --- ### Page: https://forum.scylladb.com/t/release-scylladb-6-2-0/3115 Title: [RELEASE] ScyllaDB 6.2.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce ScyllaDB Open Source, a production-ready minor release. ScyllaDB 6.2 introduces many Tablets improvements, new zero-token nodes, Alternator RBAC support and many other bug fixes … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-6-2-0/3115 ## Headings Structure: H1: [RELEASE] ScyllaDB 6.2.0 H3: High Availability - zero-token node H3: Alternator RBAC H3: Known Install Issues H3: More updated H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 6.2.0 H3: High Availability - zero-token node H3: Alternator RBAC H3: Known Install Issues H3: More updated H4: Tablets H4: Tracing H4: Stability H4: Admin and Tooling H4: Alternator H4: Performance H4: CQL H4: Materialized view H4: Packaging H4: Config H4: Monitoring H3: Related topics The ScyllaDB team is pleased to announce ScyllaDB Open Source, a production-ready minor release. ScyllaDB 6.2 introduces many Tablets improvements, new zero-token nodes, Alternator RBAC support and many other bug fixes and stabilizations. Only the last two minor releases of the ScyllaDB Open Source project are supported. Once ScyllaDB Open Source 6.2 is officially released, only ScyllaDB Open Source 6.2 and ScyllaDB 6.1 will be supported, and ScyllaDB 6.0 will be retired. Get ScyllaDB Open Source 6.2 as binary packages (RPM/DEB), AWS AMI, GCP Image and Docker image Upgrade from ScyllaDB 6.1 to ScyllaDB 6.2 There is now support for zero-token nodes. Such nodes do not replicate any data, but can participate in query coordination, and in Raft quorum voting. One can use this to create an Arbiter: a tiebreaker node, with no data, that can help maintain quorum in the case of a symmetrical two-datacenter clusters. If one of the data centers fails, the Arbiter, deployed on a 3rd datacenter, keeps quorum on the node alive. Since the Arbiter has zero token, it does not replicate user data, and does not come with network and storage costs. #15360 Authorization: Alternator now supports Role-Based Access Control (RBAC) via CQL commands. #5047 Both issue are expected to be fix in a followup patch release 6.2.x As part of moving to native tooling and away from Java tools, we will deprecate SSTableloader, in future versions of ScyllaDB. You can use the Load and Stream to upload SSTables directly to Scylla, either from Apache Cassandra or other ScyllaDB clusters. We are also deprecating the Java version of nodetool, which was replaced by a compatible native version. A New REST API system/highest_supported_sstable_version, return the sstable format version supported across the cluster #19772 The internal ‘cluster feature’ mechanism now supports suppressing features, enabling simulation of upgrades. This should catch version upgrade problems earlier. #20034 Integrated backup and restore has been merged. A new nodetool backup and restore commands (and corresponding REST API endpoint) will copy a snapshot to and from an S3 compatible endpoint. This is a work in progress aim to replace current external (Manager Agent) backup, with Scylla Core managed backup and restore. #19890 #20305 A new nodetool tasks command can be used to view and manage maintenance tasks running on the node. #19201 ScyllaDB will now tune the number of allowed open files descriptors (LimitNOFILES) for very large nodes, reducing the chance of “Too many files” error. #20443 Tools: Scrub/validate compactions will now verify checksums for uncompressed sstables. #20207 Compaction CLEANUP jobs now run under the maintenance/streaming scheduling/group. #20582 Deprecate IP based node operation REST API. Use host IDs instead. #19218 Scylla Monitoring stack 4.8.1 and later support ScyllaDB 6.2 release. See upgrade docs for Metrics update in ScyllaDB 6.2. --- ### Page: https://forum.scylladb.com/t/indexes-materialized-views-and-disk-usage-in-monitoring/3117 Title: Indexes, Materialized Views and Disk Usage in Monitoring - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/indexes-materialized-views-and-disk-usage-in-monitoring/3117 ## Headings Structure: H1: Indexes, Materialized Views and Disk Usage in Monitoring H3: Related topics ## Main Content: H1: Indexes, Materialized Views and Disk Usage in Monitoring H3: Related topics Originally from the User Slack @Thomas_Foubert: Hi everyone, where should I look into for support regarding indexes? One of my users failed at being mature with the usage of the cluster and I can now see MV operations being run non-stop since 2 days. I want to stop that before it eats the whole disk space and/or crashes the whole cluster. I am not really well informed with indexes/MV on Scylla. This is the query he ran programatically with about 10 different variables as %s : @Piotr_Smaroń: maybe @Nadav can help @Nadav: I think, but not in front of the computer, that it’s allowed to DROP INDEX while an index is still building, and it will stop building. But please test this on some side server before unleashing it on your main deployment. I mean I am not in front of a computer @Thomas_Foubert: ah big problem is, I don’t have the chance to have a dev DB at my client’s can I ensure that MV/indexes are actually being created and it’s not some kind of observability discrepancy? @Nadav: Oh, i thought you said it’s running. Now I see it’s not (disk space not going up in your graph). So what us bothering you, actually? Just that normal writes do a lot of MV work? That’s to be expected if you have many indexes - all of them need to be updated on every write. So what is the problem? If you added indexes by mistake you can drop them. @Thomas_Foubert: My bad, just read the difference between MV and indexes. So this user also ran commands arbitrarily using CQLSH and after he ran commands MV started to show on my graphs, there weren’t being used before and he claims that he didn’t do anything related to MVs. I’m a bit concerned about what’s happening, the grafana reports that MV are being built is this normal behaviour when creating new indexes? Running SELECT * FROM system_schema.views; gives me a list of the indexes he created @Nadav: Scylla’s indexes are implemented using a materializes view for each index. @Thomas_Foubert: Perfect, so this is totally expected right? @Thomas_Foubert: Thank you for your answers and your time, have a great day! --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-12/3119 Title: [RELEASE] ScyllaDB Enterprise 2024.1.12 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2024.1.12, a bug-fix production-ready patch release for ScyllaDB Enterprise 2024.1 LTS Release. Related Links Get ScyllaDB Enterprise 2024.1 (customers only, or 30-day e… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-12/3119 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2024.1.12 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2024.1.12 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2024.1.12, a bug-fix production-ready patch release for ScyllaDB Enterprise 2024.1 LTS Release. The following issues are fixed in this release (with an open-source reference, if available): --- ### Page: https://forum.scylladb.com/t/scylladb-for-centos-7/3123 Title: Scylladb for centos 7 - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, I dont see download for Centos 7 in Scylladb website. Can you any one shed some light here. Thanks Raghav Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-for-centos-7/3123 ## Headings Structure: H1: Scylladb for centos 7 H3: Related topics ## Main Content: H1: Scylladb for centos 7 H3: Related topics Hi, I dont see download for Centos 7 in Scylladb website. Can you any one shed some light here. ScyllaDB 5.2 was the latest OSS release supporting RHEL7 and variants. See OS Support by Linux Distributions and Version | ScyllaDB Docs It is time for you to upgrade. --- ### Page: https://forum.scylladb.com/t/safey-of-deleting-very-large-partitions-to-reclaim-disk-space/3124 Title: Safey of deleting very large partitions to reclaim disk space - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/safey-of-deleting-very-large-partitions-to-reclaim-disk-space/3124 ## Headings Structure: H1: Safey of deleting very large partitions to reclaim disk space H3: Related topics ## Main Content: H1: Safey of deleting very large partitions to reclaim disk space H3: Related topics Originally from the User Slack](https://scylladb-users.slack.com/) @MK: is it safe to do huge partition deletes (like a couple million) in scylla as they are obselete. this is to reclaim space. the entire partition will be deleted. And we dont do any range scans on the table, all the reads are only single partition based. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-69-2024-10-31/3128 Title: Last week in scylla-cluster-tests.git master (issue #69; 2024-10-31) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 6a3ca71c…7843142f range are covered. There were 14 non-merge commits from 6 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-69-2024-10-31/3128 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #69; 2024-10-31) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #69; 2024-10-31) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 6a3ca71c…7843142f range are covered. There were 14 non-merge commits from 6 authors in that period. Some notable commits: We’ve refactored code around Raft helper methods: instead of verifying if Raft is enabled, you can now use node.raft., which should handle the action accordingly. Sending results (metrics) to Argus will raise an Error event if any metric fails validation or has an ERROR status. The Argus client library code is now shipped directly under the SCT root directory, simplifying the backporting process and reducing the need for hydra image rebuilds. The client code is still maintained in the original Argus repository; updates require running the utils/update_argus_client.py script. For performance tests, users can skip the preload_data and steady_state_calc steps when reusing a cluster. SCT now supports testing the Kafka sink connector with limited performance scope due to running Kafka locally in Docker. It uses the confluent_kafka_python library instead of kafka-python for more features and active maintenance. Loader tool updates: Latte upgraded to 0.28.0-scylladb and scylla-bench to v0.1.23, both bringing key fixes and improvements for customer-oriented tests. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/be-featured-at-scylladbs-monster-scale-summit/3129 Title: Be featured at ScyllaDB's Monster Scale Summit - Announcements - ScyllaDB Community NoSQL Forum Meta Description: If you’re working with ScyllaDB, consider applying to be a speaker at Monster Scale Summit. It’s a technical conference that connects the global community of 12K+ professionals working on performance-sensitive data-inten… Language: en Canonical URL: https://forum.scylladb.com/t/be-featured-at-scylladbs-monster-scale-summit/3129 ## Headings Structure: H1: Be featured at ScyllaDB's Monster Scale Summit H3: Related topics ## Main Content: H1: Be featured at ScyllaDB's Monster Scale Summit H3: Related topics If you’re working with ScyllaDB, consider applying to be a speaker at Monster Scale Summit. It’s a technical conference that connects the global community of 12K+ professionals working on performance-sensitive data-intensive applications. The community would love to hear about… If accepted, you’ll share the virtual stage with engineers from Slack, Salesforce, VISA, American Express, ShareChat, Cloudflare, Atlassian, Box, and Disney. Keynotes include Gwen Sharpira, Chris Riccomini, and Martin Kleppmann (Designing Data Intensive Applications). We’ll feature you in our blogs, and you’ll have a nice professionally-produced video you can share with your peers, family, and friends. Plus, our speaker swag is simply amazing. --- ### Page: https://forum.scylladb.com/t/how-to-trace-write-detail-info/3131 Title: How to trace write detail info - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Here we have a 6-node cluster,with replication set to be 5, both read and write consistency level are set to be quorum. But sometimes we put data successfully and request code returns 200, but get the item failed. Then w… Language: en Canonical URL: https://forum.scylladb.com/t/how-to-trace-write-detail-info/3131 ## Headings Structure: H1: How to trace write detail info H3: Related topics ## Main Content: H1: How to trace write detail info H3: Related topics Here we have a 6-node cluster,with replication set to be 5, both read and write consistency level are set to be quorum. But sometimes we put data successfully and request code returns 200, but get the item failed. Then we wanna trace the write procedure to see how data is processed. Addition info: Sometimes we put data successfully and request code returns 200, but get the item failed. And this item will get successfully after serveral minutes or 1 hour also. What do you mean by “get the item failed”? Please share more details, including the exact error message, OS, hardware, which version you’re using etc. Get item failed means that when put item_1 used LWT succeed imediately get item_1, but returned 404 with error info “Empty Item”. But this item_1 could be got by aws or cqlsh after several minutes. OS: Linux scylla version: 5.2.18 consistency: write / read quorum replication: 5 --- ### Page: https://forum.scylladb.com/t/bad-alloc-issue-data-not-accessible-with-large-collections/3133 Title: Bad_alloc issue, data not accessible with large collections - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/bad-alloc-issue-data-not-accessible-with-large-collections/3133 ## Headings Structure: H1: Bad_alloc issue, data not accessible with large collections H2: backtrace that I got by using the provided decoder H2: WARN 2024-09-19 19:37:18,864 [shard 0:strm] seastar_memory - oversized allocation: 8388608 bytes. This is non-fatal, but could lead to latency and/or fragmentation issues H3: Related topics ## Main Content: H1: Bad_alloc issue, data not accessible with large collections H2: backtrace that I got by using the provided decoder H2: WARN 2024-09-19 19:37:18,864 [shard 0:strm] seastar_memory - oversized allocation: 8388608 bytes. This is non-fatal, but could lead to latency and/or fragmentation issues H3: Related topics Originally from the User Slack](https://scylladb-users.slack.com/) @Matheus_Salvia: some rows in my table seem to be unaccessible: Trying to repair the affected range immediately fails with: seems like a bug. I’m on 6.1.0. Any ideas? @Felipe_Cardeneti_Mendes: bad_alloc is often heavy memory pressure. Hard to say its a bug from just these logs. Large partition maybe? This is probably the failure cause (which should be printed to the coordinator node stderr as you receive a server error via CQL) @Matheus_Salvia: large part is a possibility, this table can have some problematic ones. did a rolling restart and it seems like it’s working now. if it happens again I’ll search for the log in the coordinator like you said nvm, this table is tame: Compacted partition maximum bytes: 8409007 @Felipe_Cardeneti_Mendes: Collections by any chance? Either way, best to open an issue. @Matheus_Salvia: yeah it does have arrays @Felipe_Cardeneti_Mendes: that could be a reason. Check if there’s anything absurd on system.large_cells @Matheus_Salvia: sessions | sessions_by_version | me-3gj6_1erf_0xfpc2mbwlidwyx848-big-Data.db | 4186749 is that absurd by scylla standards? 4 megs @Felipe_Cardeneti_Mendes: there should be a column collection_elements with the number of elements in that collection. Collections should be reasonably small and - preferably - infrequently updated. 4 mb is reasonable (not great). @Matheus_Salvia: max elements around 10k @Felipe_Cardeneti_Mendes: Yeah, that’s a lot. @Matheus_Salvia: and I don’t think we do updates on these what’s a rule-of-thumb number I should keep under? @Felipe_Cardeneti_Mendes: > Collections are meant for storing/denormalizing a relatively small amount of data. They work well for things like “the phone numbers of a given user”, “labels applied to an email”, etc. But when items are expected to grow unbounded (“all messages sent by a user”, “events registered by a sensor”…), then collections are not appropriate, and a specific table (with clustering columns) should be used. https://opensource.docs.scylladb.com/stable/cql/types.html#noteworthy-characteristics @Matheus_Salvia: so like a couple hundred? @Felipe_Cardeneti_Mendes: yeah… @avi: Set up metrics, look at LSA and non-LSA memory usage @Matheus_Salvia: dip in the middle is the restart. what should I look for here? am I just OOM and need bigger nodes? @avi: Please decode the trace via https://backtrace.scylladb.com Bigger nodes won’t help, this is a memory fragmentation problem (non-LSA memory is low, which means overall memory consumption is fine) (LSA memory = memtable + cache) oh, the 10k element collections definitely are bad here @Matheus_Salvia: INFO 2024-09-19 19:11:20,531 [shard 5:comp] compaction - [Compact sessions.sessions_by_version f3091cb0-76ba-11ef-ab67-5340763003bd] Compacting of 2 sstables interrupted due to: std::bad_alloc (std::bad_alloc), at 0x5e8296e 0x5e82f80 0x5e83288 0x24b7d9a 0x24b615e 0x5d5d3d6 INFO 2024-09-19 19:25:48,793 [shard 5:stmt] lsa - Standard allocator failure, increasing head-room in section 0x604003b94680 to 8388608 [B]; INFO 2024-09-19 19:25:49,839 [shard 5:stmt] lsa - Standard allocator failure, increasing head-room in section 0x604003b94680 to 16777216 [B]; these were all the traces I could find these are from 6.1.1. we upgraded the cluster yesterday to verify if the problem wasn’t fixed @avi: It seems related to those large collections --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-254-2024-11-03/3134 Title: Last week in scylladb.git master (issue #254; 2024-11-03) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the e7d6ab576b…39b55bd3a0 range are covered. There were 118 non-merge commits from 20 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-254-2024-11-03/3134 ## Headings Structure: H1: Last week in scylladb.git master (issue #254; 2024-11-03) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #254; 2024-11-03) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the e7d6ab576b…39b55bd3a0 range are covered. There were 118 non-merge commits from 20 authors in that period. Some notable commits: nodetool compactionhistory will now report statistics about rows merged during compaction. The bundled node_exporter Prometheus integration now disables temperature monitoring, as it causes bad performance on Azure. nodetool status now shows nodes with no tokens. The efficiency of sstable reads rows within medium or large partitions, when column_index_size_in_kb has been reduced, is now improved. Such reads will generate less I/O. ScyllaDB tracks whether read requests are waiting for CPU or I/O. In one case, a disk read from the primary index was considered to be waiting on CPU, which reduced concurrency. This is now fixed. The CREATE ROLE USING SALTED HASH statement was renamed to CREATE ROLE USING HASHED PASSWORD for improved compatibility with Cassandra. Materialized view building (initiated by CREATE MATERIALIZED VIEW or CREATE INDEX) is now performance-isolated from normal reads and writes. The REST API now reports progress of backup tasks. Repair flushes hints and batchlog in order to reduce the amount of work it has to do, but such flushes also generate work, so these flushes are now batched. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/setting-up-a-3-node-scylladb-cluster-with-k8-using-the-create-command/3148 Title: Setting up a 3 node ScyllaDB cluster with K8, using the create command - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/setting-up-a-3-node-scylladb-cluster-with-k8-using-the-create-command/3148 ## Headings Structure: H1: Setting up a 3 node ScyllaDB cluster with K8, using the create command H3: Related topics ## Main Content: H1: Setting up a 3 node ScyllaDB cluster with K8, using the create command H3: Related topics Originally from the User Slack @Deepak: Hey Guys, I am trying to set up a 3 node scylla cluster on k8s. I can bring individual nodes up and they are bootstrapping without any issues but they are not gossiping and instead behaving as disconnected individual cluster. Is there some property or setting that I should check on my helm chart? @dor: Seems like no connectivity. Get into the container shell and try to see if you can telnet the cql port and change the config accordingly @Maciej_Zimnoch: It’s better to use Scylla Operator for k8s deployments, no need to reinvent the wheel on your own @Deepak: Getting the following error while installing scylla operator in generic K8 cluster Command used kubectl apply -f deploy/operator.yaml My bad works with create kubectl create -f deploy/operator.yaml --- ### Page: https://forum.scylladb.com/t/release-scylla-manager-3-4-0/3149 Title: [RELEASE] Scylla Manager 3.4.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces the release of ScyllaDB Manager 3.4.0, a production-ready patch release of the stable 3.4 branch. This release focuses on controlling the restore procedure and improving its performance. Scy… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-manager-3-4-0/3149 ## Headings Structure: H1: [RELEASE] Scylla Manager 3.4.0 H3: Related topics ## Main Content: H1: [RELEASE] Scylla Manager 3.4.0 H3: Related topics The ScyllaDB team announces the release of ScyllaDB Manager 3.4.0, a production-ready patch release of the stable 3.4 branch. This release focuses on controlling the restore procedure and improving its performance. ScyllaDB Manager is a centralized cluster administration and recurrent tasks automation tool. ScyllaDB Manager is available for ScyllaDB Enterprise customers and ScyllaDB Open Source users. With ScyllaDB Open Source, ScyllaDB Manager is limited to 5 nodes. See the ScyllaDB Manager Proprietary Software License Agreement for details. In version 3.4.0, we reimplemented the core of the restore procedure in order to improve task parallelisation, shard utilization, and overall performance. The chart above compares the restore of identical data on clusters of the same configuration. Scylla Manager 3.4 improved overall CPU utilization during the process, resulting in a shorter time required for a full restore. This second chart shows the results of a restore performed on a 9-node cluster against another 9-node cluster. Once again, Scylla Manager 3.4 improves overall resource utilization during the process. In previous versions of Scylla Manager, a high batch size of 40 improved performance on a 3-node cluster but led to degraded CPU usage during restores on a 9-node cluster. However, with Scylla Manager 3.4, this is no longer an issue, as a higher batch size leads to better resource utilization, ultimately reducing restore time. We also added new restore task flags that will make it easier to configure restore to run as fast as possible: Please see ScyllaDB Manager restore documentation for flag details and recommendations on restore speed. Moreover, we also increased restore performance observability by introducing new metrics that can be used for calculating download and load&stream bandwidths: We also added the per shard download and load&stream bandwidths to the output of the sctool progress command. Apart from restore oriented changes, the 3.4.0 release also contains a few important bug fixes: ScyllaDB Enterprise customers are encouraged to upgrade to ScyllaDB Manager 3.4.0 in coordination with the ScyllaDB support team. The new release includes upgrades of both ScyllaDB Manager Server and Agent. Scylla Manager 3.4.0 supports the following Scylla Enterprise releases: And the following Open Source release (limited to 5 nodes see license 1): You can install and run Scylla Manager on Kubernetes using Scylla Operator. More here. --- ### Page: https://forum.scylladb.com/t/time-to-live-ttl-with-mixed-expiration-times-tombstone-threshold-and-sstable-expiration-issues/3151 Title: Time to Live (TTL) with mixed expiration times, tombstone_threshold and sstable expiration issues - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/time-to-live-ttl-with-mixed-expiration-times-tombstone-threshold-and-sstable-expiration-issues/3151 ## Headings Structure: H1: Time to Live (TTL) with mixed expiration times, tombstone_threshold and sstable expiration issues H3: Related topics ## Main Content: H1: Time to Live (TTL) with mixed expiration times, tombstone_threshold and sstable expiration issues H3: Related topics Originally from the User Slack @Samuel_Hameau: Hi, i’m having sstable expiration issues, i am forced to run manually compactions to gain space. I’m having timeseries workload, with data inserted only, using ttl (of 1 week max), with compaction = TWCS/1/DAY and gc_grace_seconds = 36000. Here is a graph showing data disk usage increase during 13d, where we can see automatic compaction during 10d (minor impact) and then daily manual compaction the last 3 days (stabilisation of disk usage). I expected to get disk usage stagnation at some point, but i only manage to get it with manual compaction. Is it anything I can do to optimize compaction/disk usage without running manual compaction ? (scylla 5.2.17 cluster with 2DC / 8 nodes, RF=3) @avi: Is your workload TWCS compatible? Clustering keys are current time, no overwrites or deletes? @Samuel_Hameau: I believe that yes: we are inserting data on the last 7 days regarding startdate (startdate can be between 7d until 5min before current time; and ajusting ttl accordingly to make it expire 7 days after startdate) PRIMARY KEY is ((userid, devicetype, brand, yearmonth), startdate, deviceid, attrib) There is no delete, there should probably never have any overwrites (but if it occurs, it would overwrite with same expiration date (to make it expires 7d after startdate) @avi: Strange, maybe @raphaelsc can give debugging tips @Samuel_Hameau: using sstabledump $sstable|grep expired.*false, i see that data is stored sorted regarding to insertion date, not regarding expiration: As ttl varies at insertion date, that may explain why sstables are not removed (until one of the records is not expired) ? @raphaelsc: 1180910 = ~13d. yes, twcs waits until sstable is fully expired, so when you mix ttl of different values, the data in a window will only be removed when all data expires. that’s the default behavior, you can force early purge, but with write amplification cost, by setting tombstone_threshold setting to 0.5 in the compaction options. that will allow twcs to remove data incrementally from a window. @Samuel_Hameau: thanks for the hint about huge ttl, and for the tombstone_threshold suggestion --- ### Page: https://forum.scylladb.com/t/release-scylladb-6-1-3/3152 Title: [RELEASE] ScyllaDB 6.1.3 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 6.1.3, a bugfix patch release of the ScyllaDB 6.1 stable branch. ScyllaDB Open Source 6.1.3, like all past and future 6.x.y releases, is backward compatible and supports r… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-6-1-3/3152 ## Headings Structure: H1: [RELEASE] ScyllaDB 6.1.3 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 6.1.3 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 6.1.3, a bugfix patch release of the ScyllaDB 6.1 stable branch. ScyllaDB Open Source 6.1.3, like all past and future 6.x.y releases, is backward compatible and supports rolling upgrades. Note that the latest open-source stable branch is 6.2, and you are encouraged to upgrade to it. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-70-2024-11-08/3153 Title: Last week in scylla-cluster-tests.git master (issue #70; 2024-11-08) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 3f0b2f30…72a67ce9 range are covered. There were 13 non-merge commits from 7 authors in tha… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-70-2024-11-08/3153 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #70; 2024-11-08) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #70; 2024-11-08) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 3f0b2f30…72a67ce9 range are covered. There were 13 non-merge commits from 7 authors in that period. Some notable commits: Improved customer oriented test by adding multi-row partitions queries. Zero-token nodes are now supported in the AWS test backend. Since these nodes don’t contain data, don’t fit many nemesis scenarios, and aren’t counted in replication factors, choose the appropriate node set when adding new test scenario: all nodes (cluster.nodes), data nodes (cluster.data_nodes), or zero nodes (cluster.zero_nodes). A new nemesis set, ZeroTokenSetMonkey , was added, which includes existing StopWaitStartMonkey and NodeTerminateAndReplace nemesis along with the new GrowShrinkZeroTokenNode targeting zero nodes specifically. Scylla-bench has been bumped to v0.1.24, with fixes for host verification and row-thread mapping in sequential workloads. To extend driver CI, we introduced a new 1-hour longevity test with large partitions targeting the gocql driver. This test is based on disruptive nemesis scenarios, including network nemesis with two network interfaces. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/restoring-a-backup-with-scylla-manager-and-sctool-restore-authentication/3155 Title: Restoring a backup with Scylla Manager and sctool restore authentication - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/restoring-a-backup-with-scylla-manager-and-sctool-restore-authentication/3155 ## Headings Structure: H1: Restoring a backup with Scylla Manager and sctool restore authentication H3: Related topics ## Main Content: H1: Restoring a backup with Scylla Manager and sctool restore authentication H3: Related topics Originally from the User Slack @Ishaan_Raghav: Query 1: I am trying to restore data from a backup using scylla-manager and sctool restore command • I restored the schema first and then restarted my db nodes as suggested in the doc. • Next when I run restore with --restore-tables , I get the id of the process • When I check the progress i see this error Can someone please suggest me where I am going wrong? Query2: How to provide authentication in sctool restore? Currently I am testing with AllowAllAuthenticator but please suggest how to do it with password set? @Felipe_Cardeneti_Mendes: 1. Something happened causing load and stream to fail, maybe ENOSPC, best to check the node logs to assess it. 2. You specify it in sctool cluster --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-255-2024-11-10/3156 Title: Last week in scylladb.git master (issue #255; 2024-11-10) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 39b55bd3a0…961a53f716 range are covered. There were 67 non-merge commits from 15 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-255-2024-11-10/3156 ## Headings Structure: H1: Last week in scylladb.git master (issue #255; 2024-11-10) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #255; 2024-11-10) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 39b55bd3a0…961a53f716 range are covered. There were 67 non-merge commits from 15 authors in that period. Some notable commits: Compaction manager stop operations now ignore errors in the compactions it manages, as those errors can affect the caller of the stop operation (such as the shutdown process). Some performance bugs leading to extra I/O when reading the primary index for a large partition are fixed. The bundled Prometheus node_exporter metrics collector now collects fewer metrics by default. CQL DESCRIBE statements for Change Data Capture (CDC) log tables have been improved. The meaning of the enable_tablets configuration has changed. Previously, it controlled whether the cluster supported tablets at all. Now, it specifies whether new keyspaces default to tablets enabled or disabled. The systemd integration now uses KillMode=control-group, as there were reports for systemctl stop not completing. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/safe-way-change-table-compaction-strategy/3163 Title: Safe way change table compaction strategy - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 5.1.19 #Cluster size: 22 #OS: ubuntu Hello everyone, My Scylla cluster version is 5.1. Due to some performance issues, I want to change the compaction strategy of certain ta… Language: en Canonical URL: https://forum.scylladb.com/t/safe-way-change-table-compaction-strategy/3163 ## Headings Structure: H1: Safe way change table compaction strategy H3: Related topics ## Main Content: H1: Safe way change table compaction strategy H3: Related topics Installation details #ScyllaDB version: 5.1.19 #Cluster size: 22 #OS: ubuntu My Scylla cluster version is 5.1. Due to some performance issues, I want to change the compaction strategy of certain tables from LCS (Leveled Compaction Strategy) to STCS (Size-Tiered Compaction Strategy). In my tests, Scylla triggers a major compaction to apply the new compaction strategy, but I want to minimize the impact during this transition. Additionally, I would like a rollback mechanism in case of any issues with Scylla. Here’s the approach I’m considering: If any issues occur, I can revert the table compaction strategy back to LCS. Here are my questions: Q1. Is this process okay, or is there a better approach? Q2. If I disable auto-compaction for an extended period (possibly up to 10 days, due to the large size of my data), could this cause any issues? First off, please update to a more recent version, 5.1 is EOL for a long time now. I think you are overly cautious. Yes, changing the compaction strategy will trigger a major compaction (a reshape actually), but ScyllaDB’s scheduling groups isolation should mean that this doesn’t have adverse affects on your cluster. If you want to be extra sure, you can set compaction_static_shares: 100 in the config, to make sure compaction doesn’t become too aggressive. Note that disabling compaction, or throttling it down to much has its own risks: sstables can start accumulating and reads will start to have higher latencies and take more memory. This is especially true when switching compaction strategies and the sstables are out-of-shape compared to what the new compaction strategy expects. So I don’t recommend disabling auto compaction for any extended period of time. --- ### Page: https://forum.scylladb.com/t/how-the-tombstones-work-in-scylla/3165 Title: How the tombstones work in Scylla - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello everyone, I’m so confused in understanding how tombstones works in scylla and will be really happy if someone could reveal my doubts. I’m using scylla version 6.1.2 Tombstone “validity” As far as I understood t… Language: en Canonical URL: https://forum.scylladb.com/t/how-the-tombstones-work-in-scylla/3165 ## Headings Structure: H1: How the tombstones work in Scylla H3: Related topics ## Main Content: H1: How the tombstones work in Scylla H3: Related topics I’m so confused in understanding how tombstones works in scylla and will be really happy if someone could reveal my doubts. Thanks to Felipe Cardeneti Mendes scylladb-users. slack .com/archives/C2NLNBXLN/p1731423261483859?thread_ts=1731372933.542199&cid=C2NLNBXLN Q1 : Does gc_grace_seconds is using only for tombstone_gc = timeout? Q2 : Does gc_grace_seconds is using for something else? Q3 : but when I described the table I can see “AND tombstone_gc = {‘mode’: ‘timeout’, ‘propagation_delay_in_seconds’: ‘3600’}”. Why mode = timeout? Might be an already fixed bug, it should be repair in this case. Did you enabled tablets in the corresponding keyspace? Either way, 6.2 should have it already changed, otherwise fill an issue. Q4 : what is propagation_delay_in_seconds? I can’t find any documentation about it How long after repair completes compaction is free to evict tombstones. We can’t immediately evict them, due to out of order writes. Q5 : why it’s safe to use for TWCS? TWCS assumes no data deletes and append-only. Thus, as soon as TTL expires compaction is free to purge tombstones. If that’s not the case, you shouldn’t be using TWCS in the first place. Q6 : how it’s differ from the mode = timeout and gc_grace_seconds = 0? There’s a good thelastpickle article which touches on some of the problems involving gc_grace=0, in particular related to hints replay. Q7 : tombstone and underlying data can be removed if they are compacted together to a new SSTable, right? Q8 : Does scylla has tombstone compaction? If so, what the strategy or in other words when this occurs? We do. See opensource.docs. scylladb .com/stable/cql/compaction.html#common-options for the rest of the answer. Q9 : Does scylla supports nodetool garbagecollect? No. You can run a major, tho. Q10 : From TWCS documentation “Tombstone compaction can be enabled to remove data from partially expired SSTables, but this creates additional WA (write amplification).”. How it can be enabled? Q11 : Does tombstone compaction enabled with tombstone_threshold, tombstone_compaction_interval and unchecked_tombstone_compaction options? Also, would like to understand more about these options. The logic is: You configure how often you want compaction to check for SSTable eligible for tombstone compaction, you set a ratio for the single SSTable to be garbage-collected. unchecked_tombstone_compaction disables it altogether. Q12 : if I’m going to delete data with CL=ALL will tombstone created? And how to avoid tombstone creation and force scylla delete data immediately? Tombstone is created irrespective of your CL. Use immediate mode and they will be evicted on flush if you use CL=ALL. Q13 : Using TWCS how to delete data immediately to avoid tombstones? I have a use case when in rare cases I need to remove entire data by partition key or by partition key and clustering range. So, you’ll have tombstone in one windowed SSTable, but actual data in another windowed SSTable. According to the strategy SSTables from different windows never compacted. It means I need tombstone compaction. Currently it leads to the bad performance and it looks like scylla does not have “tombstone compaction” at all or I did something wrong You applied a tombstone to a different compaction window when you deleted. Compaction windows are never compacted together. You must run a major and re-asses the need for TWCS. Issue about default tombstone_gc mode for tables with tablets enabled default tombstone_gc mode is wrong for tables with tablets enabled · Issue #21623 · scylladb/scylladb · GitHub --- ### Page: https://forum.scylladb.com/t/load-imbalance-after-replacing-a-failed-node/3174 Title: Load imbalance after replacing a failed node - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/load-imbalance-after-replacing-a-failed-node/3174 ## Headings Structure: H1: Load imbalance after replacing a failed node H3: Related topics ## Main Content: H1: Load imbalance after replacing a failed node H3: Related topics Originally from the User Slack @Mr._Layup: Hello! I have a cluster with 3 nodes, and yesterday the disk on the DB2 node failed. Today, I replaced the disk and rebuilt the RAID0 (which resulted in data loss). After that, I re-added the node using the replace_node_first_boot option. However, I am now experiencing some imbalance in load: Before performing operations on the production cluster, I tested the process on a similar development cluster using synthetic data, and I did not observe such a significant load imbalance. Is it safe to run a repair on the cluster? I have some concerns about potentially losing data from my cluster. Thank you for your help! @Felipe_Cardeneti_Mendes: Yes, you may repair - though if you are running a reasonably recent version replace should’ve already repaired on its own. Doesn’t seem like a problem, keep in mind the load column refreshes whenever nodes get restarted, etc @Mr._Layup: 5.4.4. Thank yu, didn’t know about refreshing after restart. --- ### Page: https://forum.scylladb.com/t/configmap-in-kind-cluster-not-being-applied/3175 Title: ConfigMap in KinD cluster not being applied - Kubernetes Operator - ScyllaDB Community NoSQL Forum Meta Description: I want to enable authorization and authentication in my ScyllaDB Alternator cluster running in an KinD kubernetes cluster. I’ve created a ConfigMap as described in the docs apiVersion: v1 kind: ConfigMap metadata: na… Language: en Canonical URL: https://forum.scylladb.com/t/configmap-in-kind-cluster-not-being-applied/3175 ## Headings Structure: H1: ConfigMap in KinD cluster not being applied H3: Related topics ## Main Content: H1: ConfigMap in KinD cluster not being applied H3: Related topics I want to enable authorization and authentication in my ScyllaDB Alternator cluster running in an KinD kubernetes cluster. I’ve created a ConfigMap as described in the docs and I’ve added that to the cluster.yaml file I’m using to create the scyllaDb cluster. I can see the configMap via kubectl describe Configmap scylla-config -n scylla so I know it’s there. However, the settings in the configMap do not appear in the /etc/scylla/scylla.yaml file in the created cluster. I’ve even restarted the cluster, but the configMap is there before the scylla cluster is created. Additional information: I am able to directly edit /etc/scylla/scylla.yaml in the pod to add the values in, but when I restart the pod, the values are gone. Had the config map wrong!! Here’s the correct one for anyone else with this problem: --- ### Page: https://forum.scylladb.com/t/role-doesnt-exist-error/3176 Title: Role doesn't exist error - Kubernetes Operator - ScyllaDB Community NoSQL Forum Meta Description: I have a scylladb cluster deployed in a KinD Kubernetes cluster. I have verified the cluster is running as I can connect with cqlsh and create new roles (authentication and authorization have been configured). I have a … Language: en Canonical URL: https://forum.scylladb.com/t/role-doesnt-exist-error/3176 ## Headings Structure: H1: Role doesn't exist error H3: Related topics ## Main Content: H1: Role doesn't exist error H3: Related topics I have a scylladb cluster deployed in a KinD Kubernetes cluster. I have verified the cluster is running as I can connect with cqlsh and create new roles (authentication and authorization have been configured). I have a Java application that uses the AWS DynamoDB SDK (v2) that I’m trying to get to connect to the scylla cluster that I deploy into the KinD cluster. I am able to create the dynamoDbClient, but when I try to use it I get: I’m properly passing the ScyllaDb role I created and its salted hash password: I have created the fap role and given it all permissions to all key spaces. What am I doing wrong? Or what am I missing? Thank you in advance, Chuck Just to make sure, have you followed Using Alternator (DynamoDB) section from Scylla Operator docs? Are you facing the same issues with aws dynamodb cli? --- ### Page: https://forum.scylladb.com/t/upgrade-scylladb-version-to-6-0-elastic-scaling-and-enabling-tablets/3179 Title: Upgrade ScyllaDB version to 6.0, elastic scaling and enabling Tablets - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/upgrade-scylladb-version-to-6-0-elastic-scaling-and-enabling-tablets/3179 ## Headings Structure: H1: Upgrade ScyllaDB version to 6.0, elastic scaling and enabling Tablets H3: Related topics ## Main Content: H1: Upgrade ScyllaDB version to 6.0, elastic scaling and enabling Tablets H3: Related topics Originally from the User Slack @Chaitanya_Tondlekar: @avi Need help in understanding if we are doing a rolling upgrade from 5.4 to 6.0, then how tablets will be enabled on existing keyspace without any data loss ? @Guy @avi: Tablets won’t be enabled on existing keyspaces @Chaitanya_Tondlekar: Main lookout for doing rolling upgrade to 6.0 was elastic scaling which needs tablets. Keyspaces without tablets in 6.0 cluster will not have fast scaling options. Am i correct ? @avi: Correct. Only new keyspaces creates with tablets enabled. @Chaitanya_Tondlekar: So what would be the right way to upgrade from 5.4 to 6.0 with existing keyspace in tablets ? Can we try to do backup restoration with scylla manager from 5.4 to 6.0? Before restoring can we create keyspace in tablets enabled and then do restoration from manager ? @avi: Backup/restore will work, but it’s recommended to test your workload first with tablets to be sure there are no regressions @Chaitanya_Tondlekar: Thanks @avi for confirmation. @avi one last confirmation needed, Scylla 6.0 running comes only with elastic scaling which are done by tablets. Scylla autoscaling comes with operator. Scylla 6.0 running on VM doesn’t comes with autoscaling. Am i correct ? @avi: On VM you are responsible for adding new nodes --- ### Page: https://forum.scylladb.com/t/release-scylladb-6-2-1/3185 Title: [RELEASE] ScyllaDB 6.2.1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 6.2.1, a bugfix patch release of the ScyllaDB 6.2 stable branch. ScyllaDB Open Source 6.2.1, like all past and future 6.x.y releases, is backward compatible and supports r… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-6-2-1/3185 ## Headings Structure: H1: [RELEASE] ScyllaDB 6.2.1 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 6.2.1 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 6.2.1, a bugfix patch release of the ScyllaDB 6.2 stable branch. ScyllaDB Open Source 6.2.1, like all past and future 6.x.y releases, is backward compatible and supports rolling upgrades. Users are encouraged to upgrade to 6.2.1. Issue fixed in this release: CQL and correctness related ScyllaDB can crash during shut down, terminate called after throwing an instance of ‘storage_io_error’ when stopping compaction_manager #21159 Compaction_manager: compaction_disabled: return true if not in compaction_state, result in a failed with ConfigurationException Stability: a potential use after free, make_streaming_consumer passes reference to token_metadata to check_needs_view_update_path without holding ptr #20979 Multishard reader is deadlock prone #21263 raft topology: use-after-free raft_topology_cmd_handler, dereferencing pointer to old topology state after reload #21220 Schema commitlog continuously updated (system.peers writes from storage_service::on_change) #20991 stream-session: can deadlock with db::vew::check_view_update_path() #21264 tasks: task_manager::module::make_task doesn’t set virtual tasks children’s fields #21278 Tooling and monitoring --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-71-2024-11-15/3187 Title: Last week in scylla-cluster-tests.git master (issue #71; 2024-11-15) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 93db67c8…1434c81d range are covered. There were 18 non-merge commits from 10 authors in t… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-71-2024-11-15/3187 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #71; 2024-11-15) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #71; 2024-11-15) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 93db67c8…1434c81d range are covered. There were 18 non-merge commits from 10 authors in that period. Some notable commits: Due to the lack of support for specific features in Amazon Linux 2 and its upcoming EOL, Amazon2 artifact tests were retired. Artifact tests now validate the output of perftune.py, ensuring correct CPU and IRQ masks as well as the default perftune configuration. A new guide has been added for configuring and using full-scan and full-partition-scan background threads. When listing cloud resources with Hydra, users can now limit the scope to a single cloud provider for faster processing. Use --backend where can be one of [aws, gce, azure, eks, docker], e.g., hydra list-resources --backend aws --user . See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/scylla-db-boostrap-issue/3188 Title: Scylla db boostrap issue - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 6.1 #Cluster size: 3 x 2 os (RHEL/CentOS/Ubuntu/AWS AMI): Debian Me again. After finally found a solution to massive insertion (close to the theoritical 10k rps per core) I fo… Language: en Canonical URL: https://forum.scylladb.com/t/scylla-db-boostrap-issue/3188 ## Headings Structure: H1: Scylla db boostrap issue H3: Related topics ## Main Content: H1: Scylla db boostrap issue H3: Related topics Installation details #ScyllaDB version: 6.1 #Cluster size: 3 x 2 os (RHEL/CentOS/Ubuntu/AWS AMI): Debian Me again. After finally found a solution to massive insertion (close to the theoritical 10k rps per core) I found something else. Boostraping a new cluster : 3 nodes on 2 two dc (6 in total) at once I hit this on random node : It was a previous attempt of startup failling with I tried disabling raft (but how?). And btw once a node is in this situation I’m stuck (all I can do is changing IP). The node is nowhere except in gossip. I tried to follow the procedure to delete node or fix raft without success … It looks like your previous attempt failed because of variation of https://github.com/scylladb/scylladb/issues/23536. The new attempt will have to wait until the IP is purged by the gossiper (30 seconds IIRC). --- ### Page: https://forum.scylladb.com/t/recommendations-for-partitioning-imbalanced-data/3191 Title: Recommendations for partitioning imbalanced data - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi all, My team is considering using ScyllaDB (cloud version) to improve cost and efficiency when reading and querying on a couple of our larger datasets. Since we know in advance exactly the kind of queries we want to … Language: en Canonical URL: https://forum.scylladb.com/t/recommendations-for-partitioning-imbalanced-data/3191 ## Headings Structure: H1: Recommendations for partitioning imbalanced data H3: Related topics ## Main Content: H1: Recommendations for partitioning imbalanced data H3: Related topics My team is considering using ScyllaDB (cloud version) to improve cost and efficiency when reading and querying on a couple of our larger datasets. Since we know in advance exactly the kind of queries we want to make it seems like Scylla is a great option but I am having a bit of trouble coming up with an appropriate schema that won’t lead to hot partitions. We deploy a lot of agents that report their health status and other data about where they are installed. One of our customers has hundreds of thousands of agents running, others have tens of thousands, and many have just a few hundred. The most common query we’ll want to do is “information about all running agents health for customer X”. The word “running” there, is highly subjective and my plan was to adjust the TTL of the tables we create to determine what is running vs hasn’t checked in a while. I don’t foresee an scenario where we would ever want to query across different customers (at least at the scylladb level). So a naive schema might look something like: This means we can easily get into a partition to search for all agents for a customer but then there is a hot and large partition on the customer with hundreds of thousands of agents running. Another naive schema might be: Now the partition key is customer + agent id so we no longer have hot and large partitions but we lost the ability to efficiently query all agents for a customer since we need the agent_id to get into the partition. I’ve also considered making separate tables or separate keyspaces per customer but as far as I could find having a dynamic number of keyspaces/tables is an anti-pattern because there is a non-negligible cost to maintaining separate tables (I am not allowed to post links but I’ve seen other discussions in this forum about it as well as in a cassandra blog) Do you all have any advise or suggestions on how best to design a schema for such use cases where one subset of the “same” data is significantly larger than the rest? I wanted to recommend using the second schema (PRIMARY KEY ((customer_name, agent_id))) and creating an materialized view on it, but the I realized the MV would have the same hot partition problem that your first table proposal has and at that point it is better to just go with the first proposal, it will be simpler. Maybe you can tweak it slightly to deal with the hot partition problem: Where agent_group is derived from agent_id. E.g. if agent_id is a monotonically increasing counter, you can make agent_group = agent_id % 1000, dispersing agents into 1000 different partitions. There will still be some which are larger and hotter than others, but the difference will be less drastic. You can think of a more sophisticated method to map agent_id to agent_group, this was just a simple example. --- ### Page: https://forum.scylladb.com/t/release-scylladb-rust-driver-0-15-0/3195 Title: [RELEASE] ScyllaDB Rust Driver 0.15.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce ScyllaDB Rust Driver 0.15.0, an asynchronous CQL driver for Rust, optimized for Scylla, but also compatible with Apache Cassandra! Some interesting statistics: over 2.516k dow… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-rust-driver-0-15-0/3195 ## Headings Structure: H1: [RELEASE] ScyllaDB Rust Driver 0.15.0 H2: Changes H3: Main changes: H3: Other changes by category: H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Rust Driver 0.15.0 H2: Changes H3: Main changes: H4: New deserialization API H3: Other changes by category: H3: Related topics The ScyllaDB team is pleased to announce ScyllaDB Rust Driver 0.15.0, an asynchronous CQL driver for Rust, optimized for Scylla, but also compatible with Apache Cassandra! Some interesting statistics: Beginning with this release, instead of putting all API-breaking changes into one group, now they are in the proper categories, but labeled as [API-breaking]. The main change in this release is the deserialization API refactor. Its primary goal was to reduce overhead caused by all rows being eagerly deserialized to type-erased CqlValue type, only then being converted to end user types. Old traits and structs (FromCqlVal, FromRow, QueryResult - renamed to LegacyQueryResult, RowIterator - renamed to LegacyRowIterator, TypedRowIterator - renamed to LegacyTypedRowIterator) are replaced by new ones (DeserializeValue, DeserializeRow, new QueryResult, QueryPager, TypedRowStream). There are wrappers and helper implementations provided, designed to aid in gradually migrating to new API - see the migration guide in the book for more information. Old traits and structs will be removed in one of future versions. New serialization API has a benefit of increased efficiency - now, rows are deserialized straight to the end user type, without any copying and allocations on the way. Another feature is the ability to deserialize rows to borrowed types (e.g. &str or &[u8]). And the result metadata is now deserialized in the borrowed form, saving even more allocations. The refactor included: New features / enhancements: API cleanups / better types: Internal API cleanups/refactors: CI / developer tool improvements: Congrats to all contributors and thanks everyone for using our driver! The source code of the driver can be found here: The official crates.io registry entry is here: Thank you for your attention, please do not hesitate to contact us if you have any questions, issues, feature requests, or are simply interested in our driver! Contributors since the last release: --- ### Page: https://forum.scylladb.com/t/copying-scylladb-data-to-s3-using-spark-performance-optimization/3196 Title: Copying ScyllaDB data to S3, using Spark, performance optimization - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/copying-scylladb-data-to-s3-using-spark-performance-optimization/3196 ## Headings Structure: H1: Copying ScyllaDB data to S3, using Spark, performance optimization H3: Related topics ## Main Content: H1: Copying ScyllaDB data to S3, using Spark, performance optimization H3: Related topics Originally from the User Slack @Nilesh_Kumar: Hi everybody, I am looking for some way to copy the scylla table partitions key data to s3. Spark is one way but within spark as well is there any optimisation I can do to scan faster with less resource consumption of scylla so it doesn’t impact the running system? Any help or direction will be of great use. Data Info - Table contains billions of partitions and per partition there is just one row. I am trying to take dump of all the partition key available in the table. @Felipe_Cardeneti_Mendes: Use the token() function to scan and limit concurrency as acceptable by your source system, add bypass cache to prevent polluting the cache. See https://www.scylladb.com/2017/03/28/parallel-efficient-full-table-scan-scylla/ --- ### Page: https://forum.scylladb.com/t/release-scylladb-monitoring-stack-4-8-2/3197 Title: [RELEASE] ScyllaDB Monitoring Stack 4.8.2 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce a patch release of ScyllaDB Monitoring Stack 4.8.2. ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB Enterprise and ScyllaDB Open Source, based on Promet… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-monitoring-stack-4-8-2/3197 ## Headings Structure: H1: [RELEASE] ScyllaDB Monitoring Stack 4.8.2 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Monitoring Stack 4.8.2 H3: Related topics The ScyllaDB team is pleased to announce a patch release of ScyllaDB Monitoring Stack 4.8.2. ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB Enterprise and ScyllaDB Open Source, based on Prometheus and Grafana. ScyllaDB Monitoring Stack 4.8.2 supports: This patch release adds the dashboard for ScyllaDB-Manager 3.4.x --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-256-2024-11-17/3198 Title: Last week in scylladb.git master (issue #256; 2024-11-17) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 961a53f716…f23800181a range are covered. There were 71 non-merge commits from 16 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-256-2024-11-17/3198 ## Headings Structure: H1: Last week in scylladb.git master (issue #256; 2024-11-17) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #256; 2024-11-17) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 961a53f716…f23800181a range are covered. There were 71 non-merge commits from 16 authors in that period. Some notable commits: Repair performance in mixed-shard configurations (where different nodes have different shard counts) has been improved. The tarball distribution no longer packages the Java-based tools, just like rpm and deb. The S3 driver is now able to retry failed requests. The DESC TABLE statement will now reject materialized views. The SPLIT compaction type, used to divide tablets into smaller tablets, now uses a better estimate for the partition count of new sstables, leading to better sized bloom filters. The sstable reader now frees memory more quickly, reducing memory requirements. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/does-tombstone-blocks-shadowed-data-sstables-from-being-purged-in-twcs/3205 Title: Does tombstone blocks shadowed data SSTables from being purged in TWCS - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details ScyllaDB version: 6.2.1 From the documentation: Avoid overwriting data and deleting data explicitly at all costs, as this can potentially block an expired SSTable from being purged, due to the che… Language: en Canonical URL: https://forum.scylladb.com/t/does-tombstone-blocks-shadowed-data-sstables-from-being-purged-in-twcs/3205 ## Headings Structure: H1: Does tombstone blocks shadowed data SSTables from being purged in TWCS H3: Compaction | ScyllaDB Docs H3: Related topics ## Main Content: H1: Does tombstone blocks shadowed data SSTables from being purged in TWCS H3: Compaction | ScyllaDB Docs H3: Related topics Installation details ScyllaDB version: 6.2.1 From the documentation: Avoid overwriting data and deleting data explicitly at all costs, as this can potentially block an expired SSTable from being purged, due to the checks that are performed to avoid data resurrection. ScyllaDB is an Apache Cassandra-compatible NoSQL data store that can handle 1 million transactions per second on a single server. Question: I have tombstone in one windowed SSTable A, but actual shadowed data in SSTable B (or others as well). Will tombstone in SSTable A block purging of SSTable B (where is shadowed data is stored) and why if it so? Or documentation means that it will block purging of SSTable A (where is tombstone is stored) because we need to wait tombstone expiration (gc_grace_seconds/repair/etc)? It is the other way around. SSTable B (the one with the shadowed data) will block the purge of SSTable A, in particular it will block the purge of the tombstone in SSTable A and said sstable can no longer be expired as a whole. --- ### Page: https://forum.scylladb.com/t/hinted-handoff-with-gc-mode-immediate/3206 Title: Hinted Handoff with gc_mode = immediate - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details ScyllaDB version: 6.2.1 Question: According to that question How the tombstones work in Scylla, the difference with gc_grace_seconds = 0 and gc_mode = immediate in Hinted Handoff mechanism. I red … Language: en Canonical URL: https://forum.scylladb.com/t/hinted-handoff-with-gc-mode-immediate/3206 ## Headings Structure: H1: Hinted Handoff with gc_mode = immediate H3: Related topics ## Main Content: H1: Hinted Handoff with gc_mode = immediate H3: Related topics Installation details ScyllaDB version: 6.2.1 Question: According to that question How the tombstones work in Scylla, the difference with gc_grace_seconds = 0 and gc_mode = immediate in Hinted Handoff mechanism. I red this article about relations between gc_grace_seconds and Hinted Handoff Hinted Handoff and GC Grace Demystified. My question is how scylla manages Hinted Handoff hints expiration in case of gc_mode = immediate? Hinted handoff considers gc_grace_seconds, expiring hints which are older than this. This mechanism predates the introduction of the new tombstone_gc schema option, so hinted handoff is only safe to use with either timeout or never modes. It is not safe to use with immediate or repair. Actually mode=repair takes care of hints, after each repair, all hints are flushed, to ensure there are no hints remaining, which could be shadowed by tombstones we are about to purge. As for immediate mode, see this issue we just opened: Hinted handoff needs to consider other GC modes than `timeout` · Issue #21662 · scylladb/scylladb · GitHub Shouldn’t the CQL operation DELETE be prohibited for tables with immediate? There were plans to disable it, yes “We are even considering rejecting user deletes if the mode is immediate, to be on the safe side.” Preventing Data Resurrection with Repair Based Tombstone Garbage Collection - ScyllaDB --- ### Page: https://forum.scylladb.com/t/using-scylladb-with-docker-and-performance-optimization/3207 Title: Using ScyllaDB with Docker and performance optimization - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/using-scylladb-with-docker-and-performance-optimization/3207 ## Headings Structure: H1: Using ScyllaDB with Docker and performance optimization H3: Related topics ## Main Content: H1: Using ScyllaDB with Docker and performance optimization H3: Related topics Originally from the User Slack @Ritesh: Hi ScyllaDB Users, I’m running ScyllaDB in Docker using the image scylladb/scylla:5.4.3, and I want to maximize performance by running the scylladb_setup script. Does this Docker image already handle performance optimizations, or are there additional tweaks needed? Thanks! @Felipe_Cardeneti_Mendes: you need to tweak it, in general docker adds lots of overhead alone. https://www.scylladb.com/2018/08/09/cost-containerization-scylla/ Ideally you would be running the Operator and simply running the performance tuning which makes it easier there, and keep using docker simply as a dev environment. Or just install scylla on the host OS instead --- ### Page: https://forum.scylladb.com/t/scylladb-labs-building-high-performance-apps-december-11-2024/3208 Title: ScyllaDB Labs: Building High-Performance Apps - December 11, 2024 - University and Training - ScyllaDB Community NoSQL Forum Meta Description: Join us for an interactive workshop on December 11, where we’ll go hands-on to build and interact with high-performance apps using ScyllaDB. This is a great way to discover the NoSQL strategies used by top teams and app… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-labs-building-high-performance-apps-december-11-2024/3208 ## Headings Structure: H1: ScyllaDB Labs: Building High-Performance Apps - December 11, 2024 H3: Related topics ## Main Content: H1: ScyllaDB Labs: Building High-Performance Apps - December 11, 2024 H3: Related topics Join us for an interactive workshop on December 11, where we’ll go hands-on to build and interact with high-performance apps using ScyllaDB. This is a great way to discover the NoSQL strategies used by top teams and apply them in a guided, supportive environment. As you go live with some sample applications, you’ll learn about the features and best practices that will enable your own applications to get the most out of ScyllaDB. This workshop will cover the following: Hope to see you there! --- ### Page: https://forum.scylladb.com/t/scylladb-enterprise-release-2024-2-0/3210 Title: ScyllaDB Enterprise Release 2024.2.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: ScyllaDB Enterprise Release 2024.2.0 The ScyllaDB team is pleased to announce the release of ScyllaDB Enterprise 2024.2.0 Release, a production-ready ScyllaDB Enterprise Feature (Short Term Support) Release. With 2024.2… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-enterprise-release-2024-2-0/3210 ## Headings Structure: H1: ScyllaDB Enterprise Release 2024.2.0 H1: ScyllaDB Enterprise Release 2024.2.0 H3: Tablets Limited Availability H3: Limited Availability H3: File based streaming for Tablets H3: Strongly Consistent Topology Updates H3: Strongly Consistent Auth Updates H3: Strongly Consistent Service Levels H3: Improved network compression for intra-node RPC H3: Describe Schema with Internals H3: Alternator RBAC H3: Native Nodetool H3: Removing the JMX Server H3: Maintenance Mode H3: Maintenance Socket H3: Deployment H3: Related topics ## Main Content: H1: ScyllaDB Enterprise Release 2024.2.0 H1: ScyllaDB Enterprise Release 2024.2.0 H3: Tablets Limited Availability H3: Limited Availability H4: Using Tablets H4: Procedures H4: Monitor Tablets H4: Driver Support H3: File based streaming for Tablets H3: Strongly Consistent Topology Updates H3: Strongly Consistent Auth Updates H3: Strongly Consistent Service Levels H3: Improved network compression for intra-node RPC H3: Describe Schema with Internals H3: Alternator RBAC H3: Native Nodetool H3: Removing the JMX Server H3: Maintenance Mode H3: Maintenance Socket H3: Deployment H3: Related topics The ScyllaDB team is pleased to announce the release of ScyllaDB Enterprise 2024.2.0 Release, a production-ready ScyllaDB Enterprise Feature (Short Term Support) Release. With 2024.2 out, 2023.1 LTS and 2024.1 LTS are still supported. More information on ScyllaDB Long Term Support (LTS) policy is available here. The ScyllaDB Enterprise 2024.2 release is based on ScyllaDB Open Source 6.0, and includes all the features available in 6.0 like: Note that tablets will not be the default when creating a new Keyspace in ScyllaDB Enterprise (see below). In addition, 2024.2 includes enterprise-only features such as improved network compression (see below), and file-based streaming for tablets, which improves ScyllaDB’s elasticity even further, and a new FIPS enabled Docker Image. ScyllaDB Enterprise customers are encouraged to upgrade to ScyllaDB Enterprise 2024.2, and are welcome to contact our Support Team with questions. In this release, ScyllaDB enabled Tablets, a new data distribution algorithm as a better alternative to the legacy vNodes approach inherited from Apache Cassandra. While the vNodes approach statically distributes all tables across all nodes and shards based on the token ring, the Tablets approach dynamically distributes each table to a subset of nodes and shards based on its size. In the future, distribution will use CPU, OPS, and other information to further optimize the distribution. In particular, Tablets provide the following: Note that you can run a cluster with some of the Keyspaces with Tablets disabled, and some with tablets enabled for as long as you wish. The scaling improvement will be partial, and limited to Keyspaces with Tables enabled. Currently, tablets are ideal for new clusters you frequently scale out or in and have one main large table and potentially many tiny onesScyllaDB Support can help you determine if Tablets in release 2024.2 are a good solution for your use case. In particular, Tablets Keyspaces are not yet enabled for the following features: Alternator support Tablets with the following: Read more about Tablets here. Tablets are disabled by default for new Keyspaces. To use Tablets, create a new keyspace with tablets = { 'enabled': true }. All tables created in this Keyspace will use Tablets by default. You can set the initial number of Tablets per table using the “initial” parameter: CREATE KEYSPACE ... WITH TABLETS = { 'enabled': true, 'initial': 1 } (See CREATE KEYSPACE docs) Note: you can not ALTER an existing Keyspace to switch between Tablets and vNode based table and back. We will remove these restrictions in upcoming releases. Note: Tablets are default in Scylla Open Source 6.0. When Upgrading from ScyllaDB Open Source 6.0 to ScyllaDB Enterprise 2024.2, Keyspaces with Tablets will continue to work with Tablets. You can use DESCRIBE to check if a keyspace or tables are using Tablets. With Tablets, the Replication Factor (RF) cannot be updated to a value higher than the number of nodes per Data Center (DC). This feature protects the Admin from setting an impossible-to-support RF. This affects the following operations: Node Decommission / Remove Starting from 2024.2, you cannot decommission or remove a node if the resulting number of nodes would be smaller than the largest non-zero replication factor (for any keyspace) in this DC 1 DC, 5 nodes, a KS with RF=5 The decommission request will fail The Replication Factor (RF) of Keyspaces must be less than or equal to the number of available nodes per Data Center (DC) Once a tablets-enabled Keyspace has tables, you can not ALTER its Replication Factor to be greater than the number of available nodes per DC. If you create such a Keyspace, you won’t be able to create Tables until you fix the RF or add more nodes. To Monitor Tablets in real time, upgrade ScyllaDB Monitoring Stack to release 4.7, and use the new dynamic Tablet panels, below. The Following Drivers support Tablets Legacy ScyllaDB and Apache Cassandra drivers will continue to work with ScyllaDB but will be less efficient when working with tablet-based Keyspaces. File-based streaming is a ScyllaDB Enterprise-only feature that optimizes tablet migration. In ScyllaDB Open Source, migrating tablets is performed by streaming mutation fragments, which involves deserializing SSTable files into mutation fragments and re-serializing them back into SSTables on the other node. In ScyllaDB Enterprise, migrating tablets is performed by streaming entire SStables, which does not require (de)serializing or processing mutation fragments. As a result, less data is streamed over the network, and less CPU is consumed, especially for data models that contain small cells. File-based streaming is used for tablet migration in all keyspaces created with tablets enabled. With Raft-managed topology enabled, all topology operations are internally sequenced consistently. A centralized coordination process ensures that topology metadata is synchronized across the nodes on each step of a topology change procedure. This makes topology updates fast and safe, as the cluster administrator can trigger many topology operations concurrently, and the coordination process will safely drive all of them to completion. For example, multiple nodes can be bootstrapped concurrently, which couldn’t be done with the previous gossip-based topology. Strongly Consistent Topology Updates is now the default for new clusters, and should be enabled after upgrade for existing clusters. System-auth-2 is a reimplementation of the Authentication and Authorization systems in a strongly consistent way on top of the Raft sub-system. This means that Role-Based Access Control (RBAC) commands like create role or grant permission are safe to run in parallel without a risk of getting out of sync with themselves and other metadata operations, like schema changes. As a result, there is no need to update system_auth RF or run repair when adding a DataCenter. Service Levels allow you to define attributes like timeout per workload. Service levels are now strongly consistent using Raft, like Schema, Topology and Auth. This release adds new Enterprise only RPC compression improvements for node to node communication: Below is a comparison of compressions algorithms on different types of data. Note that dictionary based compression can be used with either lz4 or zstd. Actual compression is very much workload-dependent and can vary between use cases. Until this release, CQL DESCRIBE SCHEMA was not sufficient to do a full schema restore from backup. For example, it lacks information about dropped columns. In 6.0, the DESC SCHEMA WITH INTERNALS command provides more information, streamlining the restore process. Authorization: Alternator supports Role-Based Access Control (RBAC). Control is done via CQL. #5047 The nodetool utility provides simple command-line interface operations and attributes. ScyllaDB inherited the Java based nodetool from Apache Cassandra. In this release, the Java implementation was replaced with a backward-compatible native nodetool. The native nodetool works much faster. Unlike the Java version ,the native nodetool is part of the ScyllaDB repo, and allows easier and faster updates. With the Native Nodetool (above), the JMX server has become redundant and will no longer be part of the default ScyllaDB Installation or image. If you are using the JMX server directly, not via nodetool, you can either: Related issues: #15588 #18566 #18472 #18566 As part of moving to native tooling and away from Java tools, we will deprecate SSTableloader in future versions of ScyllaDB Enterprise. You can use the Load and Stream to upload SSTables directly to Scylla, either from Apache Cassandra or other ScyllaDB clusters. We are also deprecating the Java version of nodetool, which was replaced by a compatible native version (see above). Maintenance mode is a new mode in which the node does not communicate with clients or other nodes and only listens to the local maintenance socket and the REST API. It can be used to fix damaged nodes – for example, by using nodetool compact or nodetool scrub. In maintenance mode, ScyllaDB skips loading tablet metadata if it is corrupted to allow an administrator to fix it. The Maintenance Socket provides a new way to interact with ScyllaDB from within the node it runs on. It is mainly for debugging. You can use CQLSh with the Maintenance Socket as described in the Maintenance Socket docs. #16172 Ubuntu 24.04 is now supported. RHEL / CentOS 7 support is deprecated. Amazon Linux 2 is deprecated and replaced with Amazon Linux 2023 Debian 10 support is deprecated. The setup utility now works with disks that do not have UUIDs, such as those in some virtualized environments #13803 The scylladb-kernel-conf package tunes the Linux kernel scheduler via sysfs to improve latency. These tunings were lost in Linux 5.13+ due to kernel changes. They are now restored. #16077 Docker: can not connect to Scylla 5.4 with CQLSh without providing host IP #16329 On Ubuntu, the installer now handles conflicts between a system process updating apt metadata and the installer itself.#16537 See ScyllaDB Enterprise Release 2024.2.0 - more improvments --- ### Page: https://forum.scylladb.com/t/scylladb-enterprise-release-2024-2-0-more-improvments/3211 Title: ScyllaDB Enterprise Release 2024.2.0 - more improvments - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: See ScyllaDB Enterprise Release 2024.2.0 for main release features. Improvements The following is a list of improvements and bug fixes included in the release, grouped by domain. Bloom Filters Bloom filters are used to… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-enterprise-release-2024-2-0-more-improvments/3211 ## Headings Structure: H1: ScyllaDB Enterprise Release 2024.2.0 - more improvments H2: Improvements H3: Bloom Filters H3: Stability and performance H3: Guardrails H3: Replication strategies Guardrails can be used to deny SimpleStratagy strategy, for example, which is not recommended for production. H3: CQL H3: Alternator H3: Images and packaging H3: Tooling and REST API H3: Configuration H3: Tracing and Auditing H3: Monitoring H3: Deprecated and removed features H3: Related topics ## Main Content: H1: ScyllaDB Enterprise Release 2024.2.0 - more improvments H2: Improvements H3: Bloom Filters H3: Stability and performance H4: Compaction Related H4: Commitlog Related H4: Cluster Operation Related H4: Materialized Views Related H4: Performance Related H4: Edge cases H3: Guardrails H3: Replication strategies Guardrails can be used to deny SimpleStratagy strategy, for example, which is not recommended for production. H3: CQL H3: Alternator H3: Images and packaging H3: Tooling and REST API H3: Configuration H3: Tracing and Auditing H3: Monitoring H3: Deprecated and removed features H3: Related topics See ScyllaDB Enterprise Release 2024.2.0 for main release features. The following is a list of improvements and bug fixes included in the release, grouped by domain. Bloom filters are used to determine which SStables do not contain a partition key, speeding up reads when SStables can be filtered out. Since the Bloom filters are held in memory, and their size depends on the data (small partitions require larger Bloom filters), there is a tradeoff between allocated memory, and the risk of OOM and the filter efficiency. The following improvements were made to Bloom filters in this release: Topology changes, Repairs, etc Guardrails is a framework to protect ScyllaDB users and admins from common mistakes and pitfalls. In this release the following Guardrails are added: Two new configurations replication_strategy_warn_list replication_strategy_fail_list Replace the old restrict_replication_simplestrategy and give more granularity for DB Admin to warn, or block non production strategies. Scylla REST API is now documented! (beta) The sstable validation tools, scylla sstable validate-checksum and scylla sstable validate, now returns output in json format. Admin API: a new API for asynchronous compaction: /tasks/compaction/keyspace_compaction/{keyspace} Similar to the existing synchronos storage_service API. Scylla Monitoring Stack released 4.8.1 and later supports ScyllaDB 2024.2 --- ### Page: https://forum.scylladb.com/t/release-scylladb-monitoring-stack-4-8-3/3212 Title: [RELEASE] ScyllaDB Monitoring Stack 4.8.3 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce a patch release of ScyllaDB Monitoring Stack 4.8.3. ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB Enterprise and ScyllaDB Open Source, based on Promet… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-monitoring-stack-4-8-3/3212 ## Headings Structure: H1: [RELEASE] ScyllaDB Monitoring Stack 4.8.3 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Monitoring Stack 4.8.3 H3: Related topics The ScyllaDB team is pleased to announce a patch release of ScyllaDB Monitoring Stack 4.8.3. ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB Enterprise and ScyllaDB Open Source, based on Prometheus and Grafana. ScyllaDB Monitoring Stack 4.8.3 supports: This patch fixes an issue with the repeated disk pie-chart panels caused by Grafana’s move to Scene-based Dashboards. --- ### Page: https://forum.scylladb.com/t/release-scylla-doctor-v1-4/3213 Title: [RELEASE]: Scylla Doctor v1.4 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: Scylla Doctor v1.4 is released. Fixes: NodeInstanceTypeAnalyzer: add i3.large to the list of supported instance types. SwapAnalyzer: make ram_swap_ratio a floating point value. Analyzers: make sure error messages prin… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-doctor-v1-4/3213 ## Headings Structure: H1: [RELEASE]: Scylla Doctor v1.4 H3: Related topics ## Main Content: H1: [RELEASE]: Scylla Doctor v1.4 H3: Related topics Scylla Doctor v1.4 is released. Artifacts can be downloaded from https://downloads.scylladb.com/downloads/scylla-doctor/ or installed from Scylla OSS or Scylla Enterprise repositories. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-72-2024-11-22/3218 Title: Last week in scylla-cluster-tests.git master (issue #72; 2024-11-22) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 3418fb9c…50020224 range are covered. There were 24 non-merge commits from 10 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-72-2024-11-22/3218 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #72; 2024-11-22) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #72; 2024-11-22) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 3418fb9c…50020224 range are covered. There were 24 non-merge commits from 10 authors in that period. Some notable commits: There’s a new Scylla Manager installation test on the Rocky9 distribution. Download, load&stream bandwidth metrics for Scylla Manager restore tests are now sent to Argus, enhancing the informativeness of restore benchmark graphs. The default Scylla Manager version was updated to 3.4. Instance provisioning fallback to on demand is now enabled by default in the AWS backend, ensuring SCT retries with on-demand instances when spot instance creation fails. The Latte tool version was bumped to 0.28.1-scylladb, allowing “batch” query testing. A new predefined throughput steps with tablets test pipeline was added and scheduled for weekly runs. This test verifies latencies across different throughput steps and the maximum throughput Scylla can achieve. Additionally, write workload test was improved by running c-s directly in the loader OS instead of Docker, enhancing stability. The mgmt_cli_test.py module was refactored to improve organization, naming, and test reusability for Scylla Manager. Dependabot will now track Docker-based loaders. This required updates to configuration files but will not impact their usage in tests. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/manual-bucketing-and-large-partitions-problem/3219 Title: Manual bucketing and large partitions problem - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details ScyllaDB version: 6.2.1 Question: I still can find lot’s of information about large partition problems. And also in scylla monitoring dashboard I could see this information. Is it still a problem … Language: en Canonical URL: https://forum.scylladb.com/t/manual-bucketing-and-large-partitions-problem/3219 ## Headings Structure: H1: Manual bucketing and large partitions problem H3: Related topics ## Main Content: H1: Manual bucketing and large partitions problem H3: Related topics Installation details ScyllaDB version: 6.2.1 Question: I still can find lot’s of information about large partition problems. And also in scylla monitoring dashboard I could see this information. Is it still a problem with large partitions in scylla 6.2.1 and should I manually do time bucketing? Because it seems your team did optimizations in that regards https://youtu.be/n7ljgazbxzA?si=TB2IhDFz-cQT9Uxq&t=37 We recently added a ScyllaDB University topic that deals exactly with Large Partitions, see here. Regarding manual time bucketing, this topic might help you. If you have more specific questions, feel free to ask them here. Thanks for the answer. More specifically I’m interested in time series long term history use case (10 years of data for example). Typical queries is getting the data by sensor_id within a time range. And let’s say we have 10M of sensors in overall I found lot’s of articles about “time bucketing”, e.g.: https://www.scylladb.com/glossary/cassandra-time-series-data-modeling/ https://www.scylladb.com/2019/08/20/best-practices-for-data-modeling/ KariosDB doing also time bucketing (row width argument) I also had watched the video on scylla university and have the following questions: In case of using TWCS compaction will never be done across windows. So, by default after 1 day we’ll have 1 SSTable for that day. Therefore large partitions should not be a problem in case of compaction, since the maximum amount of data that we’re compacting is 1 day. In another words in whole partition we may have a petabytes of data (for 10 years), but within 1 SStable we are storing the data only for 1 day, so it should not be a problem. Is it right? I did not get relations between large partitions and SSTables (https://youtu.be/9HBEDVswLQM?si=3mBmXjJZ3180paRC&t=383). Let’s say I’m using TWCS and data for 1 day will be compacted to one big SSTable. Then let’s have a look at two use cases: big amount of small partitions vs small amount of big partitions. SStable is written not per each partition key, but it contains many partition keys. So, at the end of the day TWCS produce one SSTable which will contains all partitions, but the overall size of SSTable does not depend on the number of partition keys. In another words, if I need to write 100GB of data for 1 day it does not matter if I’ll have 1 partition, 100 partitions or 1M partitions, because anyway the whole data for that 1 day will be stored in 1 SStable file. Also, since your team did an optimization for large partitions(https://youtu.be/n7ljgazbxzA?si=TB2IhDFz-cQT9Uxq&t=37) I’m wondering do I still need to worry about manual bucketing Yes, ScyllaDB has done a lot of work to better handle large partitions, they should not crash ScyllaDB anymore and the performance penalty of working with large partitions was reduced. That said, large partitions above a certain size are still problematic. What this certain size is exactly, I don’t know. Let’s say, if you have partitions in the gigabytes, you should start thinking about bucketing. In general, if you know you have large partitions, watch out for any sign of latency degradation, that is the best way to determine when to take action. As for large partitions vs. SSTables: the problem here is that ScyllaDB currently doesn’t cut SSTables mid-partition. Some compaction strategies want to control the size of the SSTables they create: an example would be LCS or ICS (enterprise-only). These compaction strategies struggle with large partitions, because a single partition will always be in a single SSTable, so they loose control over the size of the SSTables. To my knowledge, TWCS doesn’t suffer from this problem and like you said, partitions are scattered among the windows, further reducing any chance of this problem popping up. In another words in whole partition we may have a petabytes of data (for 10 years), but within 1 SStable we are storing the data only for 1 day, so it should not be a problem. Is it right? I hope this is just a fictitious example. This results in 36510shard windows which is too much. Keep your bucket count reasonable. And just add a corresponding composite key so partitions dont grow indefinitely, it is an append only use case after all so you can do it deterministically anyway… --- ### Page: https://forum.scylladb.com/t/designing-a-notifications-table/3223 Title: Designing a Notifications table - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: i need basically a notification table, app is similar to slack in design, create a notification when a user posts a message to a channel, whenever somebody reacts to a message, follow, etc im having trouble grouping/batc… Language: en Canonical URL: https://forum.scylladb.com/t/designing-a-notifications-table/3223 ## Headings Structure: H1: Designing a Notifications table H3: Related topics ## Main Content: H1: Designing a Notifications table H3: Related topics i need basically a notification table, app is similar to slack in design, create a notification when a user posts a message to a channel, whenever somebody reacts to a message, follow, etc im having trouble grouping/batching notifications. if 100 people like a message, i dont want to create 100 rows, it would be more efficient to just send a number with the context, but im not sure how to do this because you cant use a counter in a non-counter table. any pointers? but im not sure how to do this because you cant use a counter in a non-counter table. any pointers? The answer is to denormalize, you hold the likes for posts under a single/separate table. Although, if I am following you correctly, a “Slack” type of app would allow for multiple different types of reactions. So rather than a likes table, why not a reactions_by_message where you store each reaction separately and then simply count to retrieve the grand total? This could also work with UDTs/collections considering they don’t get too big. wouldn’t this require a batch read though? i read that those are to be avoided generally. so say i pull 8 messages via a group_chat_id, thats just one query, but to now get the reaction counts for each reaction on the message, i have to extract each message_id and do a batch read request also ive gotten conflicting advice on this, ive also been told to just increment a count with LWT’s directly on the message You will probably pull messages within a specific timestamp, as a chat app is really a sliding window of event. That allows you to concurrently and efficiently avoid large partitions when a specific group_chat_id becomes popular or subject to spams. If you use UDTs or collections to store the reactions along each message, then the job is done and you should have everything you need going forward. If you decide to store it under a different table, you would probably also read the same group_chat_id and time-window, which would retrieve you the reactions for each message you stored upfront. also ive gotten conflicting advice on this, ive also been told to just increment a count with LWT’s directly on the message Seems like an overkill for a chatapp which can tolerate eventual consistency - though I’d pick and choose whatever approach works best for the scale you’re after. this was very helpful thank you for replying on a side note what are your thoughts on batch reads in general? the idea of extracting relevant ids after querying one partition, and then using those ids to batch get another set of information from another partition? the most important thing ive noted so far is that ideally you should only every query one partition at a time --- ### Page: https://forum.scylladb.com/t/using-dynamodb-alernator-and-cql-together-data-modeling-and-performance/3224 Title: Using DynamoDB Alernator and CQL together, data modeling and performance - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/using-dynamodb-alernator-and-cql-together-data-modeling-and-performance/3224 ## Headings Structure: H1: Using DynamoDB Alernator and CQL together, data modeling and performance H3: Related topics ## Main Content: H1: Using DynamoDB Alernator and CQL together, data modeling and performance H3: Related topics Originally from the User Slack @Terence_Liu: Hi. Is the choice of CQL and DynamoDB Alternator interface an upfront one, and one can’t switch back and forth between them? @Felipe_Cardeneti_Mendes: You can have both working, but CQL clients will manipulate (read/write) CQL tables, and Dynamo clients will manipulate Alternator tables. There is no inter-op atm. @Terence_Liu: I see. That makes sense. My data modeling is just one primary key that’s evenly distributed (some md5 hash). But a few wide columns that are best modeled with CQL’s UDTs. Sometimes one additional column (like a new version of some ML record) needs to be populated for all billions of keys. This seems to work really well with the native CQL interface, where each column itself is a RocksDB “column family” (key being the primary key, and val being the column data), and UPDATE to a new column doesn’t involve reading other existing columns. Is this a painful operation for the alternator? I understand the underlying implementation uses the CQL Map data type in one single column. Does this mean any update to any DynamoDB field is a read-write for the entire value of that column? @Felipe_Cardeneti_Mendes: > Is this a painful operation for the alternator? I understand the underlying implementation uses the CQL Map data type in one single column. Does this mean any update to any DynamoDB field is a read-write for the entire value of that column? Right, Alternator tables are basically serialized as a CQL table, so you can check how we do it in fact. a DDB hashkey is your partition key, and the rest is serialized as a :attrs column as a map. So you CAN update individual fields for a given key as the idea of this column is to make it work just like DDB’s schemaless idea That said, CQL supports more types, I am biased tho. if anything interesting: https://bsky.app/profile/felipemendes.dev/post/3l4boujjebl2f Bluesky Social: Felipe Mendes (@felipemendes.dev) @Terence_Liu: That was an interesting experiment. How does the storage layer treat the Map data type though? Is the entire Map a single value tracked by an LSM tree (I’m not very accurate with language here because there are partitions, but I just mean RocksDB’s column family), or each key-val pair in the Map is tracked by its own LSM tree? The documentation has these both points. The first suggests my first understanding, but the second seems to suggest my second. I am not able to find further clarification. > • Individual collections are not indexed internally. This means that even to access a single element of a collection, the whole collection has to be read (and reading one is not paged internally). > • While insertion operations on sets and maps never incur a read-before-write internally, some operations on lists do. Further, some list operations are not idempotent by nature (see the section on lists below for details), making their retry in case of timeout problematic. It is thus advised to prefer sets over lists when possible. Data Types | ScyllaDB Docs Ohh, I see, the reality is closer to the first understanding - it’s indeed a big value. But new field inserted is treated as a separate entry, and merged during reads: (From Claude, sounds right to me) @Felipe_Cardeneti_Mendes: ^ this is correct worth to mention that if updates to the same map happen in short fashion, then they should just get merged in memtables and flushed as one to disk but updates to infrequently updated keys will be flushed to a different sstable and merged upon read, compaction will eventually take care of pushing the lowest tiers to higher ones so all in all, you want to avoid a situation where you have to read too many SSTables for a given map. You should also avoid very large maps, because at the end of the day they are all stored under a single cell @Terence_Liu: Cool - this is nice, and should be quite performant. I was only used to the RocksDB’s simple compaction of entire values, but it sure sounds like one can be fancy about compaction, and accommodate maps. And it sounds like modeling a separate field properly as a native column is the least amount of overhead - it’s a brand new column and only queried when requested. @Felipe_Cardeneti_Mendes: sweet, have fun! --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-257-2024-11-24/3225 Title: Last week in scylladb.git master (issue #257; 2024-11-24) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the f23800181a…e2e6f4f441 range are covered. There were 65 non-merge commits from 14 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-257-2024-11-24/3225 ## Headings Structure: H1: Last week in scylladb.git master (issue #257; 2024-11-24) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #257; 2024-11-24) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the f23800181a…e2e6f4f441 range are covered. There were 65 non-merge commits from 14 authors in that period. Some notable commits: Sstable checksums and digests are now checked during compaction, improving overall integrity. Not that checksums for compressed sstables were already checked before. The altternator /localnodes endpoint will not return nodes that are temporarily down. A request to stop all repairs may have missed some ongoing repair operations; this is now fixed. The CQL PER PARTITION LIMIT clause is now respected for aggregating queries. The tablet load balancer is now able to schedule repair operations. This is not yet integrated into nodetool repair or automatic load balancing. During ordinary sstable compaction, we do not purge tombstones if they potentially delete data in commitlog, to avoid data resurrection on restart. However, this is unnecessary for the row cache, so row cache now ignores commitlog when purging tombstones. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/lightweight-transactions-lwt-performance-optimization-client-side-and-server-side/3226 Title: Lightweight Transactions (LWT) performance optimization client side and server side - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/lightweight-transactions-lwt-performance-optimization-client-side-and-server-side/3226 ## Headings Structure: H1: Lightweight Transactions (LWT) performance optimization client side and server side H3: Related topics ## Main Content: H1: Lightweight Transactions (LWT) performance optimization client side and server side H3: Related topics Originally from the User Slack [October 29th, 2024 9:36 PM] career.phuongnguyen: Hi everyone, i have question about lwt optimization done in scylla. Is the cql protocol extension just for better routing in client driver (disable host shuffling, retrying lwt on the same connection) or there is some optimization done on scylla cluster itself once it realizes that client driver knows about SCYLLA_LWT_OPTIMIZATION_META_BIT_MASK, i’m not familiar with C++ so i can not find the answer in scylla source code https://github.com/scylladb/scylladb/commit/6028588148c556151d764f2a987c276cb4a21eb0 @Felipe_Cardeneti_Mendes: For the LWT extension in question, https://groups.google.com/g/scylladb-dev/c/IQsbOpSDcTI/m/43kt-G0fAgAJ answers it: That said, other protocol extensions do exist. For tablet routing (so the client knows understands the new message propagated by the coordinator), and the per-partition-rate limiting error, alowing the driver to know when it needs to backoff and for how long if it hits it. https://github.com/scylladb/scylladb/blob/49d3e281d6a3ebbc23aca8def18408ace06ff9a1/transport/cql_protocol_extension.hh#L36 We have means to know whether a extension was set. For example, in https://github.com/scylladb/scylladb/commit/efc3953c0aafa2a49f610c11273725714a6b4f77#diff-5ddc61dd27612d0e2d48deba0[…]4d0b05de0f3b4126f4632236fR1284 we use: and from https://github.com/scylladb/scylladb/blob/master/docs/dev/protocol-extensions.md the flow is described as: • Client sends the OPTIONS request to the Scylla instance to get a list of protocol extensions that the server understands. • Server sends the SUPPORTED message in reply to the OPTIONS request. The message body is a string multimap, in which keys describe different extensions and possibly one or more additional values specific to a particular extension (specified as distinct values under a feature key in the following form: ARG_NAME=VALUE). • The client determines the set of compatible extensions which it is going to use in the current connection by intersecting known capabilities list with what it has received in SUPPORTED response. • Client driver sends the STARTUP request with additional payload consisting of key-value pairs, each describing a negotiated extension. • Server determines the set of compatible extensions by intersecting known list of protocol extensions with what it has received in STARTUP request. So we have means to know it, but for LWT specifically the optimal routing is responsibility of the client driver. @Shibu_ina: Thank you for the detailed answer! @\Felipe Cardeneti Mendes --- ### Page: https://forum.scylladb.com/t/is-there-a-way-to-perform-arithmetic-operations-in-query/3229 Title: Is there a way to perform arithmetic operations in query - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I need to perform queries which check fields based on some arithmetic ops. For example- select a, b from table_name where a + b == 10 Is there a way for this to be done? Language: en Canonical URL: https://forum.scylladb.com/t/is-there-a-way-to-perform-arithmetic-operations-in-query/3229 ## Headings Structure: H1: Is there a way to perform arithmetic operations in query H3: Related topics ## Main Content: H1: Is there a way to perform arithmetic operations in query H3: Related topics I need to perform queries which check fields based on some arithmetic ops. For example- select a, b from table_name where a + b == 10 Is there a way for this to be done? Im trying this- SELECT bucket, counter FROM used_bucket_tracker WHERE lounge_id = ? AND bucket - counter <= ? LIMIT 10 & I get- :invalid_syntax, message: "line 4:4 no viable alternative at input ‘bucket’", Even enclosing (bucket - counter) with the roud parentheses is useless, in that case it gives error- :invalid_syntax, message: "line 4:4 no viable alternative at input ‘’", As you realized, you can’t do this. Best is to pre-aggregate, so that you could then do WHERE key=? AND c=10 or, as in the latter example, WHERE key=? AND c <= ? LIMIT 10 This works, but of course, it assumes that the other fields (which then would’ve been pre-aggregated) don’t change. Otherwise, perhaps denormalizing to a counter table (if you’re ok w/ eventual consistency) and regular PN-Counter should eventually converge to a common state. thanks again. yeah i figured that out, so opted for somewhat of an architecture change, so yeah…thanks agian --- ### Page: https://forum.scylladb.com/t/how-exactly-are-you-supposed-to-handle-counts-in-scylla/3230 Title: How exactly are you supposed to handle counts in Scylla? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: For example, views_count, likes_count, etc. These are all common counts you need to keep track of but I can’t find a clear answer. 2 options seem clear to me. A separate counter table. Everybody says a counter is archa… Language: en Canonical URL: https://forum.scylladb.com/t/how-exactly-are-you-supposed-to-handle-counts-in-scylla/3230 ## Headings Structure: H1: How exactly are you supposed to handle counts in Scylla? H3: Related topics ## Main Content: H1: How exactly are you supposed to handle counts in Scylla? H3: Related topics For example, views_count, likes_count, etc. These are all common counts you need to keep track of but I can’t find a clear answer. 2 options seem clear to me. A separate counter table. Everybody says a counter is archaic and unreliable to use to keep track of counts. Something about it uses Paxos currently and will use Raft soon? Also it requires an additional query so N+1 problem A bigint count that you increment directly in the table. Requires 3 steps. A read to get the initial value, a write to write the new value, a read to get the final value. So what exactly are you supposed to do here? Everybody says a counter is archaic and unreliable to use to keep track of counts. Who said it? What is exactly the problem you want to solve? If its just count’ing and you do not require idempotency, you should be all set. The alternative really would be to cluster all “likes” “views” for a given entity(key) and then count() on them. --- ### Page: https://forum.scylladb.com/t/how-to-enable-audit-logs-for-scylladb-for-creating-new-functions/3232 Title: How to Enable Audit logs for ScyllaDB for creating new functions - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: ScyllaDB version: 6.0.2 OS: Ubuntu Language: en Canonical URL: https://forum.scylladb.com/t/how-to-enable-audit-logs-for-scylladb-for-creating-new-functions/3232 ## Headings Structure: H1: How to Enable Audit logs for ScyllaDB for creating new functions H3: Related topics ## Main Content: H1: How to Enable Audit logs for ScyllaDB for creating new functions H3: Related topics ScyllaDB version: 6.0.2 OS: Ubuntu What do you mean “for creating new functions”? You can read more about Auditing here. i want to log, for example: CREATE FUNCTION test.add_numbers(a int, b int) RETURNS NULL ON NULL INPUT RETURNS int LANGUAGE lua AS ‘return a + b;’; --- ### Page: https://forum.scylladb.com/t/rust-driver-and-batched-prepared-statements-correct-usage-and-error-message/3322 Title: Rust driver and batched prepared statements, correct usage and error message - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/rust-driver-and-batched-prepared-statements-correct-usage-and-error-message/3322 ## Headings Structure: H1: Rust driver and batched prepared statements, correct usage and error message H3: Related topics ## Main Content: H1: Rust driver and batched prepared statements, correct usage and error message H3: Related topics Originally from the User Slack @Terence_Liu: I am experimenting with batched prepared statements with the Rust driver. Seeing when N statements > N values. However, when N statements < N values, it’s silently going through. Is that the expected behavior? Figure I should post in the main channel, because this should be a server-side question. @Felipe_Cardeneti_Mendes: did you see https://rust-driver.docs.scylladb.com/stable/queries/batch.html#batch-values ? See the boilerplate example on how to specify empty statements Batch statement | ScyllaDB Docs empty values* now I would expect the latter to also fail. Hopefully nothing weird happens, please open an issue (https://github.com/scylladb/scylla-rust-driver/) accordingly @Terence_Liu: Yes. Although my use case is different. I was trying to bulk ingest with the same statement. So I went like this: It’s somewhat awkward, but my understanding is you need to duplicate the prepared statement 100 times in a batch, even if they are identical. And the N of values need to match N of statements in a batch. So I had to take care of the last batch being not rounded to the step size. I found out about the issue when I tried to play around and commented out the initial basically not filling the statements batch. But the code still ran. Not sure if it actually inserted anything. Checking. Yeah, I looked at the count through cqlsh, and the ingested count was 0 If I only have 1 statement in the batch, and 100 values, it would just ingest 1 single entry. So maybe the client is doing an iterator zip, and trimmed the extra values @Felipe_Cardeneti_Mendes: right, that’s what I meant when I asked to open an issue. I just tried it and see the same. Which is “kinda” fine, given that other drivers (eg: python) also require one to append the statement+values in one-shot. But I see where the confusion comes from. I think you raised a good point that silently ingesting just equal the number of bound statements and forgetting the remaining ones is prone to misconceptions. so either way, just append and push values as you go. @Terence_Liu: I’ll open an issue. > so either way, just append and push values as you go. Do you mean create a prepared batch, prepare the batch, and add values every batch cycle? That would repeatedly prepare the same batch of identical statements (other than the last irregular batch), right? @Felipe_Cardeneti_Mendes: start the batch, iterate. push statement+values. when you hit your condition, apply the batch, start a new one. you also have new_with_statements if you know beforehand how much you will need @Terence_Liu: Opened something https://github.com/scylladb/scylla-rust-driver/issues/1114. Hope I wrote it with better clarity. GitHub: When N of batch of statements is smaller than N of values, session.batch() silently drops extra values. · Issue #1114 · scylladb/scylla-rust-driver @Karol_Baryła: We’ll investigate the issue soon. One note regarding the code: it would be better to just prepare a statement once, and use it to create batches. Then you will have no need to call prepare_batch. @Terence_Liu: Ohh, thank you. That’s improves the code a bit! Cloning is OK right? @Karol_Baryła: Yes, we even have a note about this in documentation of PreparedStatement: https://docs.rs/scylla/latest/scylla/statement/prepared_statement/struct.PreparedStatement.html#clone-implementation PreparedStatement in scylla::statement::prepared_statement - Rust @Terence_Liu: cool, thanks --- ### Page: https://forum.scylladb.com/t/raft-upgrade-stuck-waiting-for-ghost-nodes/3369 Title: Raft Upgrade stuck waiting for ghost nodes - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 6.1.3 #Cluster size: 7 nodes (6 nodes + 1 nodes in different DCs) os (RHEL/CentOS/Ubuntu/AWS AMI): Ubuntu 22 Today, I finally started enabling raft, by hitting the API-endpoin… Language: en Canonical URL: https://forum.scylladb.com/t/raft-upgrade-stuck-waiting-for-ghost-nodes/3369 ## Headings Structure: H1: Raft Upgrade stuck waiting for ghost nodes H3: Handling Node Failures | ScyllaDB Docs H3: Related topics ## Main Content: H1: Raft Upgrade stuck waiting for ghost nodes H3: Handling Node Failures | ScyllaDB Docs H3: Related topics Installation details #ScyllaDB version: 6.1.3 #Cluster size: 7 nodes (6 nodes + 1 nodes in different DCs) os (RHEL/CentOS/Ubuntu/AWS AMI): Ubuntu 22 Today, I finally started enabling raft, by hitting the API-endpoint with curl curl -X POST "http://curhost:10000/storage_service/raft_topology/upgrade" Now, as nothing happened, I did a rolling restart, which in the end resulted in Nov 27 10:32:24 o-2 scylla[632625]: [shard 0:strm] raft_topology - waiting for all nodes to finish upgrade to raft schema The last node, however, wasn’t able to finish and instead is spitting out Now, to get rid off those unavailable, not replacable ghost nodes, I tried using removenode --ignore-dead-nodes 8a627941-2f40-47ad-8e5d-6f6e891ab85d,d728fc9d-81ca-4f34-ab5b-3b0858144c61,e16a9c96-d8a0-47fe-8044-37be077f45b9 e16a9c96-d8a0-47fe-8044-37be077f45b9 but that only results in So, the question is: How do I get rid off those ghost nodes? Answering my own question here: For some odd reason, the 2nd DC still had information about already removed nodes, making the Raft Upgrade impossible. I had to follow the following documentation, in particular the section of “Manual Recovery Procedure” which allowed me to cancel the Raft Migration, and remove the ghost nodes, after which everything went the normal (and rather quick) way. ScyllaDB is an Apache Cassandra-compatible NoSQL data store that can handle 1 million transactions per second on a single server. --- ### Page: https://forum.scylladb.com/t/per-partition-local-ordering/3412 Title: Per partition local ordering - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details ScyllaDB version: 6.2.1 Data model: CREATE TABLE data( p TEXT, c TIMESTAMP, v BLOB, PRIMARY KEY(p, c) ) Queries: Give me the first and the last value for each partition. SELECT * FROM … Language: en Canonical URL: https://forum.scylladb.com/t/per-partition-local-ordering/3412 ## Headings Structure: H1: Per partition local ordering H3: Related topics ## Main Content: H1: Per partition local ordering H3: Related topics Installation details ScyllaDB version: 6.2.1 Queries: Give me the first and the last value for each partition. Question: Does scylla supports “per partition ordering” only? And if not, are there any plans to support it? Question: Does scylla supports “per partition ordering” only? And if not, are there any plans to support it? The short answer is that ScyllaDB doesn’t support this and I’m not aware of any plans to support it either. Note that making the ordering restricted to a single partition will still not make it possible to execute this query: a partition can be arbitrarily big and to order it according to any other order than the clustering one, we still have to read it all into memory, which might not be possible. Thanks for the answer, but I still didn’t get why we need read whole result in memory when we are sorting by clustering key, but in reversed order? Let’s have a look at the following queries: We don’t need to read everything in the memory, since we don’t need to sort the final result, right? To follow further examples, let’s introduce new keyword “PER PARTITION ORDER”, so we can write the same query in more explicit way We do need to read everything in memory, because we want the final result to be sorted. But instead I need to do this that introduce overhead and only possible to do with pagging disabled You are right, examples (1) and (3) do not require reading all results into memory, as the end results are partially sorted only. I think this PER PARTITION ORDER would be possible to implement. Thanks, sounds great. Do I need to create feature request on the github for that? Or this feature may be considered only for the Enterprise version? Feel free to create a feature request issue in the open-source repository. --- ### Page: https://forum.scylladb.com/t/performance-issue-throughput-drop-and-latency-increase/3437 Title: Performance issue, throughput drop and latency increase - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/performance-issue-throughput-drop-and-latency-increase/3437 ## Headings Structure: H1: Performance issue, throughput drop and latency increase H2: root@scylladb-node1:/# nodetool cfstats user.segments Total number of tables: 66 H3: Related topics ## Main Content: H1: Performance issue, throughput drop and latency increase H2: root@scylladb-node1:/# nodetool cfstats user.segments Total number of tables: 66 H3: Related topics Originally from the User Slack @Ritesh: Hi, I’m facing an unusual problem where my single-node ScyllaDB instance, with 40 cores and 80GB RAM, is now handling only 1K write QPS with latency of 135 ms. Previously, it was handling around 100K write QPS. Our Spark jobs are writing to this ScyllaDB node, but I don’t see any unusual lines in the logs. Could someone help me troubleshoot this? @Ritesh: This is the htop output- @avi: What storage do you use? @avi: Could you be less specific @Ritesh: @avi Disk & CPU Info- @avi: What’s the ,ale and model of the disk? How many disks are there? How are they organized? @Ritesh: It’s a single 9TB SSD formatted with EXT4 Can you give me some pointers to debug the slow write throughput. I deleted the docker volume twice and it works with 100K Write QPS for a day, then suddenly it drops to 5K. There are no errors in the logs. Below is the screen shoot- @avi I see extremely high disk I/O which I feel is causing the writes to slow down, why isn’t it using the RAM for writes? These are some nodetool command stats- root@scylladb-node1:/# nodetool compactionstats pending tasks: 0 id compaction type keyspace table completed total unit progress 7f181f30-9952-11ef-a52a-ea3f6c0e32ba COMPACTION user segments 10450305 16954752 keys 61.64% e463abc0-9952-11ef-91e6-ea296c0e32ba COMPACTION user segments 16880 536064 keys 3.15% e63e3820-9952-11ef-914e-ea3e6c0e32ba COMPACTION user segments 493 533504 keys 0.09% 9130d7c0-9952-11ef-b05e-ea436c0e32ba COMPACTION user segments 8682107 16182144 keys 53.65% c44e79a0-9952-11ef-9f8d-ea2e6c0e32ba COMPACTION user segments 4336080 7870848 keys 55.09% e4fa1f60-9952-11ef-b6b8-ea306c0e32ba COMPACTION user segments 21118 803200 keys 2.63% 2bbfe8e0-9952-11ef-a7c7-ea3d6c0e32ba COMPACTION user segments 14981297 47618432 keys 31.46% 733bed90-9952-11ef-8887-ea2d6c0e32ba COMPACTION user segments 12171322 15482496 keys 78.61% e4eed4c0-9952-11ef-a9ce-ea326c0e32ba COMPACTION user segments 65605 535936 keys 12.24% e4bffc90-9952-11ef-8317-ea316c0e32ba COMPACTION user segments 13424 534400 keys 2.51% e5a07180-9952-11ef-a1e8-ea346c0e32ba COMPACTION user segments 6845 1066112 keys 0.64% e461aff0-9952-11ef-8622-ea256c0e32ba COMPACTION user segments 234639 536320 keys 43.75% e418c010-9952-11ef-940b-ea336c0e32ba COMPACTION user segments 53084 803712 keys 6.60% e4edea60-9952-11ef-9c7d-ea406c0e32ba COMPACTION user segments 59928 802432 keys 7.47% e3cb1540-9952-11ef-acf7-ea376c0e32ba COMPACTION user segments 69324 535296 keys 12.95% e5e9af80-9952-11ef-ae29-ea476c0e32ba COMPACTION user segments 12129 1065984 keys 1.14% a2bb5a10-9952-11ef-877b-ea286c0e32ba COMPACTION user segments 5516920 35002240 keys 15.76% 8834e2b0-9952-11ef-9a27-ea2b6c0e32ba COMPACTION user segments 5432425 7618816 keys 71.30% e463f9e0-9952-11ef-bf63-ea416c0e32ba COMPACTION user segments 24961 1066240 keys 2.34% Active compaction remaining time : n/a Keyspace : user Read Count: 18546287 Read Latency: 2.0165853143542964E-4 ms Write Count: 58488507 Write Latency: 4.2926142737751884E-5 ms Pending Flushes: 0 Table: segments SSTable count: 227 SSTables in each level: [227/4] Space used (live): 255941509511 Space used (total): 255941509511 Space used by snapshots (total): 0 Off heap memory used (total): 11382524916 SSTable Compression Ratio: 0.5989474298522396 Number of partitions (estimate): 2778003890 Memtable cell count: 4803792 Memtable data size: 2806399012 Memtable off heap memory used: 5588910080 Memtable switch count: 198 Local read count: 18553639 Local read latency: 0.202 ms Local write count: 58565454 Local write latency: 0.043 ms Pending flushes: 0 Percent repaired: 0.0 Bloom filter false positives: 58849 Bloom filter false ratio: 0.11260 Bloom filter space used: 5599935004 Bloom filter off heap memory used: 5599935004 Index summary off heap memory used: 193679832 Compression metadata off heap memory used: 0 Compacted partition minimum bytes: 43 Compacted partition maximum bytes: 446 Compacted partition mean bytes: 111 Average live cells per slice (last five minutes): 0.0 Maximum live cells per slice (last five minutes): 0 Average tombstones per slice (last five minutes): 0.0 Maximum tombstones per slice (last five minutes): 0 Dropped Mutations: 0 root@scylladb-node1:/# cqlsh 10.0.7.135 -e “DESCRIBE user.segments;” CREATE TABLE user.segments ( maid uuid PRIMARY KEY, segment_map map ) WITH bloom_filter_fp_chance = 0.01 AND caching = {‘keys’: ‘ALL’, ‘rows_per_partition’: ‘ALL’} AND comment = ‘’ AND compaction = {‘class’: ‘SizeTieredCompactionStrategy’} AND compression = {‘sstable_compression’: ‘org.apache.cassandra.io.compress.LZ4Compressor’} AND crc_check_chance = 1.0 AND dclocal_read_repair_chance = 0.0 AND default_time_to_live = 0 AND gc_grace_seconds = 864000 AND max_index_interval = 2048 AND memtable_flush_period_in_ms = 0 AND min_index_interval = 128 AND read_repair_chance = 0.0 AND speculative_retry = ‘99.0PERCENTILE’; @avi: What’s the make and model of the disk? Check scylla-monitoring, Advanced dashboard, commitlog and compaction I/O panels. Look at the disk latency and bandwidth charts. @Robert: And use xfs instead ext4 and I’m not sure if You are not overload a single partitions by map operations - data model doesn’t looks so efficient Btw You wrote 80GB memory peer 40 cores, but htop shows 96 cores and 128GB… but even that it’s around 2GB peer Scylla shard, there is maybe not enough space for a memtable @Ritesh: @Robert Thanks for the suggestion on using XFS instead of EXT4! I’m concerned about potential issue you specified about partition overload due to the map operations. Could you recommend an alternative data model that would be more efficient for handling high-cardinality segments in this setup? > And use xfs instead ext4 and I’m not sure if You are not overload a single partitions by map operations - data model doesn’t looks so efficient In our use case, each UUID is associated with multiple integer segments with high cardinality, and we frequently update specific UUIDs with their associated segments. For our query pattern, we primarily perform lookups using UUIDs in the WHERE clause. Based on this, I thought the below data model would be well-suited for our queries: CREATE TABLE user.segments ( maid uuid PRIMARY KEY, segment_map map ) > Btw You wrote 80GB memory peer 40 cores, but htop shows 96 cores and 128GB… but even that it’s around 2GB peer Scylla shard, there is maybe not enough space for a memtable The difference in cores and RAM is due to resources allocated for running Apache Spark, which handles writing partitions to the ScyllaDB table > What’s the make and model of the disk? > > Check scylla-monitoring, Advanced dashboard, commitlog and compaction I/O panels. Look at the disk latency and bandwidth charts. @avi Thanks for the suggestion! I’ll check the scylla-monitoring, Advanced dashboard, commitlog and compaction I/O panels you have mentioned. This is info about the disk- ``Model: PERC H745 Front Firmware: 51.16.0-4076 Product Firmware Size Diskbay RPM SAMSUNG RAID0 1920 GB 134:2 SSD SAMSUNG RAID0 1920 GB 134:3 SSD SAMSUNG RAID0 1920 GB 134:4 SSD SAMSUNG RAID0 1920 GB 134:5 SSD SAMSUNG RAID0 1920 GB 134:6 SSD @Robert: Imo instead of map should be just a simple table with partition, clustering column as map key and normal column as map value: CREATE TABLE user.segments ( maid uuid, segment_key int, segment_value int, Primary key(maid, segment_key) ); Looks like a multi row (per uuid) instead just one in Your model, but from Scylla perspective it’s a single partition (per uuid) which contains some rows. So You will directly stream a whole partition at once. But maybe currently Scylla handle both model in the same way. But in past iirc operations on the big collections has some troubles. Additional with that data model You are capable to sort (asc/desc - worth to thing about that when table is created) and additionally keys could be filtered and range filtered (for example: select * from segments where maid = x and segment_key < 10;) --- ### Page: https://forum.scylladb.com/t/release-scylladb-cpp-over-rust-driver-0-3-0/3469 Title: [RELEASE] ScyllaDB CPP-over-Rust Driver 0.3.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce ScyllaDB CPP Rust Driver 0.3.0, an API-compatible rewrite of GitHub - scylladb/cpp-driver: Scylla C/C++ Driver as a wrapper for the Rust driver. It will fully replace the CPP driv… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-cpp-over-rust-driver-0-3-0/3469 ## Headings Structure: H1: [RELEASE] ScyllaDB CPP-over-Rust Driver 0.3.0 H2: Changes H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB CPP-over-Rust Driver 0.3.0 H2: Changes H3: Related topics The ScyllaDB team is pleased to announce ScyllaDB CPP Rust Driver 0.3.0, an API-compatible rewrite of GitHub - scylladb/cpp-driver: Scylla C/C++ Driver as a wrapper for the Rust driver. It will fully replace the CPP driver, which is already reaching its End of Life. The current driver version should be considered Alpha. Some minor features still need to be included. See Limitations section in README.md. The underlying Rust driver used version: 0.15.0. Currently, the cpp-rust-driver does not fully utilize lazy deserialization mechanism introduced in rust-driver 0.15.0. The support for lazy deserialization, and all features that depend on it are going to be introduced in some future release. Starting from this release, features that are not part of original cpp-driver, but are implemented in cpp-rust-driver as an extension to original API, will be labeled as [Extension]. Implemented API functions: New features / enhancements CI / developer tool improvements: Congrats to all contributors and thanks everyone for using our driver! The source code of the driver can be found here: Thank you for your attention, please do not hesitate to contact us if you have any questions, issues, feature requests, or are simply interested in our driver! Contributors since the last release: --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-73-2024-11-29/3865 Title: Last week in scylla-cluster-tests.git master (issue #73; 2024-11-29) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 6e300864…63b10fb8 range are covered. There were 26 non-merge commits from 14 authors in t… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-73-2024-11-29/3865 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #73; 2024-11-29) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #73; 2024-11-29) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 6e300864…63b10fb8 range are covered. There were 26 non-merge commits from 14 authors in that period. Some notable commits: Added support for configuration with the ‘latte schema’ command, enabling users to specify custom schema parameters via a config file or environment variables (e.g., SCT_LATTE_SCHEMA_PARAMETERS.tombstone_gc_mode=repair). This was applied for issue reproduction. Improved Scylla Manager restore benchmarks by ensuring comparability and determinism across tests with different data sizes, achieved by disabling autocompaction. Cassandra-stress was updated to version 3.17.0, addressing an issue with unbalanced node loads during cluster growth in tablet tests. Introduced a new performance test for Scylla Manager backup time under read stress conditions. A new nemesis simulating multiple restarts of the Raft coordinator node in a row was added. Scylla upgrade tests were sped up by parallelizing system upgrades across nodes before upgrading Scylla itself. We now collect schema information during test teardown and store it in the SCT logs archive. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/out-scaling-a-cluster-restoring-a-backup-into-a-new-cluster-to-avoid-streaming/4011 Title: Out scaling a cluster, restoring a backup into a new cluster to avoid streaming - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/out-scaling-a-cluster-restoring-a-backup-into-a-new-cluster-to-avoid-streaming/4011 ## Headings Structure: H1: Out scaling a cluster, restoring a backup into a new cluster to avoid streaming H3: Related topics ## Main Content: H1: Out scaling a cluster, restoring a backup into a new cluster to avoid streaming H3: Related topics Originally from the User Slack @Terence_Liu: Have been reading this doc https://opensource.docs.scylladb.com/stable/operating-scylla/procedures/cluster-management/add-node-to-cluster.html. By my understanding, if I just ingested to a big single-node cluster, and backed up the relevant keyspace/tables to the cloud, I should restore this backup to a node prod space, add two more nodes to make it RF=3, and wait for streaming from the first node to the other two. Can I do this instead - restore the same backup to all three nodes, and boot them up together to avoid the streaming process? Assuming all three nodes have identical data, this should be possible? I understand it’s harder when the original backup is more than one host, because it’s a lot harder to know how the key ranges are distributed. I assume nodetool refresh will help in this case? ScyllaDB node will ignore the partitions in the sstables which are not assigned to this node. For example, if sstable are copied from a different node. Adding a New Node Into an Existing ScyllaDB Cluster (Out Scale) | ScyllaDB Docs @Pete_Aven: Hi Terence, you should be able to boot up an empty 3 node cluster, then follow the restore process for a table. It consists of: • Create table schema on empty cluster • Copy all the table data into the table directory - usually /var/lib/scylla/data/keyspacename/tablename-uuid/upload/ ( copy the table data from the single node to each node in the new cluster - a 1:3 copy • Then use nodetool refresh -- keyspacename tablenamehttps://opensource.docs.scylladb.com/stable/operating-scylla/nodetool-commands/refresh.html This will ingest the backup into a running cluster, which could also be serving traffic (especially writes) @Terence_Liu: Thank you Pete. Do I need to execute nodetool refresh 1 time on each of the nodes? Or doing it once on any node will cause every node to take in the upload folder sstables? Hi @Pete_Aven. If I do a 1:3 copy (RF=3), will this speed up ingestion by bypassing the load-and-stream process? Or will it actually create more load because each node needs to duplicate the streams to other nodes over essentially the same data? @Pete_Aven: @Terence_Liu Following the suggestion above, there should be no streaming. You copy the table data over to each node . You run nodetool refresh on each node. All refresh operations will be node local. Then you run repair when all that’s done and the cluster will be in sync. @Terence_Liu: thank you! --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-258-2024-12-01/4017 Title: Last week in scylladb.git master (issue #258; 2024-12-01) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the e2e6f4f441b…65949ce6078 range are covered. There were 85 non-merge commits from 18 authors in that per… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-258-2024-12-01/4017 ## Headings Structure: H1: Last week in scylladb.git master (issue #258; 2024-12-01) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #258; 2024-12-01) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the e2e6f4f441b…65949ce6078 range are covered. There were 85 non-merge commits from 18 authors in that period. Some notable commits: ScyllaDB computes a schema version in order to see if it needs to perform a transparent upgrade during a query. For system tables, it now uses a hash based algorithm which is more robust compared to the manual annotation applied by developers which was used earlier. The materialized view update process updates views when the base tables are updated by an UPDATE or INSERT query. It is now able to avoid unnecessary updates in some circumstances. Alternator, ScyllaDB’s implementation of the DynamoDB API, now reports consumed RCU and WCU to the user, and adds metrics to measure overall consumption. Bootstrap and decommission now enable the small-table repair optimization. This speeds up bootstrap in large clusters when small or empty system tables have to be migrated to other nodes. We no longer take snapshots of materialized views, as their contents can be regenerated from the base table. Internal tasks are now kept for an hour after completion so their status can be queried with nodetool. The ScyllaDB process now handles the SIGQUIT signal by dumping memory diagnostics into the log. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-6-1-4/4035 Title: [RELEASE] ScyllaDB 6.1.4 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 6.1.4, a bugfix patch release of the ScyllaDB 6.1 stable branch. ScyllaDB Open Source 6.1.4, like all past and future 6.x.y releases, is backward compatible and supports r… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-6-1-4/4035 ## Headings Structure: H1: [RELEASE] ScyllaDB 6.1.4 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 6.1.4 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 6.1.4, a bugfix patch release of the ScyllaDB 6.1 stable branch. ScyllaDB Open Source 6.1.4, like all past and future 6.x.y releases, is backward compatible and supports rolling upgrades. Note that the latest open-source stable branch is 6.2, and you are encouraged to upgrade to it. Issue fixed in this release: Tooling and Monitoring --- ### Page: https://forum.scylladb.com/t/tombstones-gc-grace-seconds-propagation-delay-in-seconds-twcs-compaction-and-related-questions/4064 Title: Tombstones, gc_grace_seconds, propagation_delay_in_seconds, TWCS, compaction and related questions - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/tombstones-gc-grace-seconds-propagation-delay-in-seconds-twcs-compaction-and-related-questions/4064 ## Headings Structure: H1: Tombstones, gc_grace_seconds, propagation_delay_in_seconds, TWCS, compaction and related questions H3: Related topics ## Main Content: H1: Tombstones, gc_grace_seconds, propagation_delay_in_seconds, TWCS, compaction and related questions H3: Related topics Originally from the User Slack @Artem_Golovko: Hello everyone, I’m so confused in understanding how tombstones works in scylla and will be really happy if someone could reveal my doubts. c) disabled - not interested d) immediate Q5: why it’s safe to use for TWCS? Q6: how it’s differ from the mode = timeout and gc_grace_seconds = 0? 2. When tombstone and underlying data can be removed? a) During tables compaction Q7: tombstone and underlying data can be removed if they are compacted together to a new SSTable, right? b) During tombstone compaction when only one SSTable itself got compacted to remove data. Q8: Does scylla has tombstone compaction? If so, what the strategy or in other words when this occurs? Q9: Does scylla supports nodetool garbagecollect? Q10: From TWCS documentation “Tombstone compaction can be enabled to remove data from partially expired SSTables, but this creates additional WA (write amplification).”. How it can be enabled? Q11: Does tombstone compaction enabled with tombstone_threshold, tombstone_compaction_interval and unchecked_tombstone_compaction options? Also, would like to understand more about these options. Q12: if I’m going to delete data with CL=ALL will tombstone created? And how to avoid tombstone creation and force scylla delete data immediately? Q13: Using TWCS how to delete data immediately to avoid tombstones? I have a use case when in rare cases I need to remove entire data by partition key or by partition key and clustering range. So, you’ll have tombstone in one windowed SSTable, but actual data in another windowed SSTable. According to the strategy SSTables from different windows never compacted. It means I need tombstone compaction. Currently it leads to the bad performance and it looks like scylla does not have “tombstone compaction” at all or I did something wrong, @Felipe_Cardeneti_Mendes: Q1: Does gc_grace_seconds is using only for tombstone_gc = timeout? Yes Q2: Does gc_grace_seconds is using for something else? No. Q3: but when I described the table I can see “AND tombstone_gc = {‘mode’: ‘timeout’, ‘propagation_delay_in_seconds’: ‘3600’}”. Why mode = timeout? Might be an already fixed bug, it should be repair in this case. Did you enabled tablets in the corresponding keyspace? Either way, 6.2 should have it already changed, otherwise fill an issue. Q4: what is propagation_delay_in_seconds? I can’t find any documentation about it How long after repair completes compaction is free to evict tombstones. We can’t immediately evict them, due to out of order writes. Q5: why it’s safe to use for TWCS? TWCS assumes no data deletes and append-only. Thus, as soon as TTL expires compaction is free to purge tombstones. If that’s not the case, you shouldn’t be using TWCS in the first place. Q6: how it’s differ from the mode = timeout and gc_grace_seconds = 0? There’s a good thelastpickle article which touches on some of the problems involving gc_grace=0, in particular related to hints replay. Q7: tombstone and underlying data can be removed if they are compacted together to a new SSTable, right? Right. Q8: Does scylla has tombstone compaction? If so, what the strategy or in other words when this occurs? We do. See https://opensource.docs.scylladb.com/stable/cql/compaction.html#common-options for the rest of the answer. Q9: Does scylla supports nodetool garbagecollect? No. You can run a major, tho. Q10: From TWCS documentation “Tombstone compaction can be enabled to remove data from partially expired SSTables, but this creates additional WA (write amplification).”. How it can be enabled? See link from Q8 Q11: Does tombstone compaction enabled with tombstone_threshold, tombstone_compaction_interval and unchecked_tombstone_compaction options? Also, would like to understand more about these options. The logic is: You configure how often you want compaction to check for SSTable eligible for tombstone compaction, you set a ratio for the single SSTable to be garbage-collected. unchecked_tombstone_compaction disables it altogether. Q12: if I’m going to delete data with CL=ALL will tombstone created? And how to avoid tombstone creation and force scylla delete data immediately? Tombstone is created irrespective of your CL. Use immediate mode and they will be evicted on flush if you use CL=ALL. Q13: Using TWCS how to delete data immediately to avoid tombstones? I have a use case when in rare cases I need to remove entire data by partition key or by partition key and clustering range. So, you’ll have tombstone in one windowed SSTable, but actual data in another windowed SSTable. According to the strategy SSTables from different windows never compacted. It means I need tombstone compaction. Currently it leads to the bad performance and it looks like scylla does not have “tombstone compaction” at all or I did something wrong, You applied a tombstone to a different compaction window when you deleted. Compaction windows are never compacted together. You must run a major and re-asses the need for TWCS. @Artem_Golovko: @Felipe_Cardeneti_Mendes so many thanks for the answers! --- ### Page: https://forum.scylladb.com/t/making-outh-app-using-django/4082 Title: Making outh app using django - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: latest #Cluster size:1 os (RHEL/CentOS/Ubuntu/AWS AMI): docker from django_cassandra_engine.models import DjangoCassandraModel from cassandra.cqlengine import columns from dj… Language: en Canonical URL: https://forum.scylladb.com/t/making-outh-app-using-django/4082 ## Headings Structure: H1: Making outh app using django H1: Primary key for Cassandra H3: Related topics ## Main Content: H1: Making outh app using django H1: Primary key for Cassandra H3: Related topics Installation details #ScyllaDB version: latest #Cluster size:1 os (RHEL/CentOS/Ubuntu/AWS AMI): docker from django_cassandra_engine.models import DjangoCassandraModel from cassandra.cqlengine import columns from django.contrib.auth.models import AbstractBaseUser, BaseUserManager from django.utils import timezone import uuid class UserManager(BaseUserManager): def create_user(self, username, email, password=None, **extra_fields): if not email: raise ValueError(‘The Email field must be set’) class User(AbstractBaseUser, DjangoCassandraModel): user_id = columns.UUID(primary_key=True, default=uuid.uuid4) PS D:\codes\Galileo\galileo> python manage.py sync_cassandra Traceback (most recent call last): File “D:\codes\Galileo\galileo\manage.py”, line 22, in main() File “D:\codes\Galileo\galileo\manage.py”, line 18, in main execute_from_command_line(sys.argv) File “C:\Users\hp\AppData\Local\Programs\Python\Python310\lib\site-packages\django\core\management_init.py”, line 442, in execute_from_command_line utility.execute() File “C:\Users\hp\AppData\Local\Programs\Python\Python310\lib\site-packages\django\core\managementinit_.py”, line 436, in execute self.fetch_command(subcommand).run_from_argv(self.argv) File “C:\Users\hp\AppData\Local\Programs\Python\Python310\lib\site-packages\django\core\management\base.py”, line 412, in run_from_argv self.execute(*args, **cmd_options) File “C:\Users\hp\AppData\Local\Programs\Python\Python310\lib\site-packages\django\core\management\base.py”, line 453, in execute self.check() File “C:\Users\hp\AppData\Local\Programs\Python\Python310\lib\site-packages\django\core\management\base.py”, line 485, in check all_issues = checks.run_checks( File “C:\Users\hp\AppData\Local\Programs\Python\Python310\lib\site-packages\django\core\checks\registry.py”, line 88, in run_checks new_errors = check(app_configs=app_configs, databases=databases) File “C:\Users\hp\AppData\Local\Programs\Python\Python310\lib\site-packages\django\contrib\auth\checks.py”, line 84, in check_user_model if isinstance(cls().is_anonymous, MethodType): File “C:\Users\hp\AppData\Local\Programs\Python\Python310\lib\site-packages\django\db\models\base.py”, line 511, in init if field.attname not in kwargs and field.column is None or field.generated: AttributeError: ‘UUID’ object has no attribute ‘column’. Did you mean: ‘db_column’? Is the model synced to the database? Make sure ./manage.py sync_cassandra is successful. If you still have issues, please post a complete example - snippets above are a bit broken in formatting. --- ### Page: https://forum.scylladb.com/t/node-stuck-in-none-state-after-topology-changes-raft-recovery-mode/4085 Title: Node stuck in "none" state after topology changes, raft recovery mode - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/node-stuck-in-none-state-after-topology-changes-raft-recovery-mode/4085 ## Headings Structure: H1: Node stuck in "none" state after topology changes, raft recovery mode H3: Related topics ## Main Content: H1: Node stuck in "none" state after topology changes, raft recovery mode H3: Related topics Originally from the User Slack @Thomas_Foubert: Hi, I just encountered a bootstrap failure and now a node is stuck in none state : can’t be replaced at boot time : and cant be nodetool remove : Here’s a small dump of the raft tables : This failure happened while I was adding 2 new nodes to an existing cluster after a cluster upgrade from 6.1 to 6.2 The other node is still joining and seems to be doing it well EDIT: I remember manipulating raft and system tables to forcefully remove a node when I was first experimenting with Scylla, but I can’t find the doc anymore @Kamil_Braun: please open an issue and post logs from the topology coordinator node and the node that failed to bootstrap, and ping me on github (@kbr-scylla) > The other node is still joining and seems to be doing it well I didn’t notice this sentence before. Given that, I suspect that the none node will most likely be removed automatically once the other joining node finishes. If that doesn’t happen, then you should open the issue. @Thomas_Foubert: It didn’t happen so I put the cluster into raft recovery mode, wiped the topology tables and performed a rolling restart as per : https://opensource.docs.scylladb.com/stable/troubleshooting/handling-node-failures.html#manual-recovery-procedure Handling Node Failures | ScyllaDB Docs --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-2-1/4101 Title: [RELEASE] ScyllaDB Enterprise 2024.2.1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2024.2.1, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2024.2 Feature (Short Term Support) Release. 2024.2.1 patch release includes… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-2-1/4101 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2024.2.1 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2024.2.1 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2024.2.1, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2024.2 Feature (Short Term Support) Release. 2024.2.1 patch release includes multiple minor bug fixes. The following issues are fixed in this release (with an open-source reference, if available): CQL and correctness related Alternator - Amazon DynamoDB-compatible API --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-13/4102 Title: [RELEASE] ScyllaDB Enterprise 2024.1.13 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2024.1.13, a bug-fix production-ready patch release for ScyllaDB Enterprise 2024.1 Long Term Support (LTS) Release. Note there is a later Feature (Short term Support) Rele… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-13/4102 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2024.1.13 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2024.1.13 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2024.1.13, a bug-fix production-ready patch release for ScyllaDB Enterprise 2024.1 Long Term Support (LTS) Release. Note there is a later Feature (Short term Support) Release 2024.2. The following issues are fixed in this release (with an open-source reference, if available): Tooling and Monitoring --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-259-2024-12-08/4103 Title: Last week in scylladb.git master (issue #259; 2024-12-08) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 65949ce6078…f744007e136 range are covered. There were 150 non-merge commits from 19 authors in that pe… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-259-2024-12-08/4103 ## Headings Structure: H1: Last week in scylladb.git master (issue #259; 2024-12-08) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #259; 2024-12-08) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 65949ce6078…f744007e136 range are covered. There were 150 non-merge commits from 19 authors in that period. Some notable commits: The topology coordinator will now detect tables that have shrunk, and merge adjacent tablets in order to meet the average tablet replica size goal. An sstable scrub results in rewriting sstables, but did not remove the original sstable from metadata, causing a later reshape compaction to fail. This is now fixed. Tablet repair operations are now tracked via the task manager. The topology coordinator reads the request table, but if a new request arrives while the read is in progress, it could be missed. This is now fixed. The decimal data type comparison operations have been improved not to cause stalls when comparing numbers with vastly different scales. The data plane coordination code (“storage_proxy”) now uses host UUIDs to track hosts rather than network addresses. This simplifies the code and brings a nice performance improvement. The scylla sstable dump-summary command now displays the tokens of the first and last keys. This helps associating an sstable with a node or a tablet. Validation of sstable data checksums has been improved. Alternator, ScyllaDB’s implementation of the DynamoDB API, now reports consumed WCU for item deletion operations. The topology coordinator no longer waits for a replaced node to appear in gossip, since it may be gone already. The gossip code will now clean up nodes that died before they could join Raft. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/how-to-model-my-data-when-having-subgropus-of-data-with-vastly-different-volumes/4104 Title: How to model my data when having subgropus of data with vastly different volumes - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/how-to-model-my-data-when-having-subgropus-of-data-with-vastly-different-volumes/4104 ## Headings Structure: H1: How to model my data when having subgropus of data with vastly different volumes H3: Related topics ## Main Content: H1: How to model my data when having subgropus of data with vastly different volumes H3: Related topics Originally from the User Slack @Andres: I have an issue where subgroups of my data have vastly different volumes and this volume disparity isn’t something we can change. A common query I need to do is “what is the latest state for all active ‘agents’ in a subgroup” so I need to design my tables such that searching with just the subgroup is efficient, meaning have it be the partition key: However, this leads to hot/large partitions for the larger subgroups that have magnitudes more number of agents than our smaller subgroups. If I make the agent_id part the partition key, then I no longer have a hot partition (each agent writes at about the same rate) but I then lose the ability to quickly query for all agents in a subgroup (queries for smaller subgroups still have to go sift through the data of the larger subgroups). I considered maybe making separate tables or keyspaces per subgroup (since I don’t need to ever search across subgroups) but it seems like having a dynamic number of tables is an anti-pattern (https://forum.scylladb.com/t/any-restriction-on-number-of-tables-in-a-keyspace/482). Do you all have any ideas or alternatives? @Felipe_Cardeneti_Mendes: Thinking out loud — maybe hash the agent_id UUID (a 16 byte array) and split it to corresponding buckets When you need to query a specific agent you already know which bucket it lives When you want to scan all agents, simply read from all buckets concurrently. Find the good balance where ((subgroup, bucket), agent_id) is optimal @Andres: So keep a static set of bucket identifiers so if I need to query for all agents I iterate through all the buckets client-side, doing a query per bucket? @Felipe_Cardeneti_Mendes: Yeap— this should spread the load accordingly. The problem is that you cant be too optimistic on the number of buckets as you otherwise risk the smaller subgroups from becoming expensive to read — if you are too conservative you wont resolve the problem as it is. Another option would be to maybe only apply this to outliers (under a corresponding table) and transition subgroups as needed. Scanning should be fast as you query all buckets in parallel @Andres: It’s an interesting idea, unfortunate that unless I only do it to outliers then even the small subgroups also get the “cost” of buckets but it may be worth it. I’d have to think further on how painful, if at all, it would be to do it to only outliers and then transition as needed. I was also considering maintaining a redis ordered set of subgroup to agent uuids, so that way I can have the agent uuids in the partition key because I would know the uuids via the redis cache @Felipe_Cardeneti_Mendes: Yeap, a reverse lookup table also works. Scanning an entire subgroup with too many agents may be slower though. Instead of scanning each bucket individually you will now query each agent— you will need to decide on what works and scales best for your use case. @Andres: Since it will be a query using the IN operation, instead of a query per agent, it shouldn’t be that much slower right? I was imagining a query for every few thousand agents so similarly to paging. @Felipe_Cardeneti_Mendes: > Since it will be a query using the IN operation, instead of a query per agent, it shouldn’t be that much slower right? I was imagining a query for every few thousand agents so similarly to paging. It works, but may create shard contention under some circumstances, beware of it. A query (IN or not) is handled by a single coordinator shard. When you ask for many other keys on potentially different shards, this coordinator needs to coordinate with other replicas. Best to test and ensure it matches your P99 latencies requirements. --- ### Page: https://forum.scylladb.com/t/codec-not-found-exception-enum-codecs-issue-when-using-data-access-objects/4109 Title: Codec not found exception - enum codecs issue when using data access objects - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/codec-not-found-exception-enum-codecs-issue-when-using-data-access-objects/4109 ## Headings Structure: H1: Codec not found exception - enum codecs issue when using data access objects H3: Related topics ## Main Content: H1: Codec not found exception - enum codecs issue when using data access objects H3: Related topics Originally from the User Slack @Leo_Urbina: I’m running into an issue with enum codecs. I have a Dao that has a method like the following The issue is specifically with statuses in the WHERE clause. I’ve registered the EnumCodec like so: This seems to work fine as I’m able to insert data into the table. However when I try tu run getTasksForTriggerWithStatus I get the following error: I assumed that the codec would take care of that, no? Since I’m using the Enum names, I’m able to just map over the enums in the caller and search directly for the names, which works - but seems somewhat brittle Any insight here would be helpful Ok, figured it out - this is an issue with Kotlin and how it handles generics. I needed to use @JvmSuppressWildcards to avoid it thinking of the List If it’s incorrect and replace is not ongoing, to clean it up you can patch CRD subresource status using kubectl patch with --subresource=status parameter. It works really well, thank you very much! --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-260-2024-12-15/4183 Title: Last week in scylladb.git master (issue #260; 2024-12-15) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the f744007e136…5880a1b90b4 range are covered. There were 67 non-merge commits from 16 authors in that per… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-260-2024-12-15/4183 ## Headings Structure: H1: Last week in scylladb.git master (issue #260; 2024-12-15) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #260; 2024-12-15) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the f744007e136…5880a1b90b4 range are covered. There were 67 non-merge commits from 16 authors in that period. Some notable commits: The TRUNCATE operation has been promoted to a topology level operation. This allows the topology change coordinator to drive the operation to completion in the face of node failures and concurrent tablet migrations. ScyllaDB now leaves some Linux aio control blocks for use of native tooling such as the scylla sstable command. Linux aio control blocks are a limited resource and can prevent tool start up when exhausted. As part of that change the io_uring Seastar backend was enabled (as a non-default backend). Commitlog replay fixes a corner case where a shard that received no mutations before crashing had its commitlog interpreted incorrectly. When decommissioning a node in a cluster that has different shard counts per node, we are now more careful to preserve tablet balance among the remaining nodes. Tablet migrations are now surfaced as tasks that can be observed and controlled by the nodetool tasks command family. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-6-2-2/4184 Title: [RELEASE] ScyllaDB 6.2.2 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 6.2.2, a bugfix patch release of the ScyllaDB 6.2 stable branch. ScyllaDB Open Source 6.2.2, like all past and future 6.x.y releases, is backward compatible and supports r… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-6-2-2/4184 ## Headings Structure: H1: [RELEASE] ScyllaDB 6.2.2 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 6.2.2 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 6.2.2, a bugfix patch release of the ScyllaDB 6.2 stable branch. ScyllaDB Open Source 6.2.2, like all past and future 6.x.y releases, is backward compatible and supports rolling upgrades. Users are encouraged to upgrade to 6.2.2. Issue fixed in this release: --- ### Page: https://forum.scylladb.com/t/how-to-optimize-tombstone-issues-caused-by-updating-primary-keys-in-scylladb-materialized-views/4189 Title: How to Optimize Tombstone Issues Caused by Updating Primary Keys in ScyllaDB Materialized Views? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I’m using ScyllaDB to store subscription relationships in the following business scenario: Requirements: The data model needs to support these queries: app_id = ? AND user_id = ? AND ts >= <lastTs> LIMIT 10 for increm… Language: en Canonical URL: https://forum.scylladb.com/t/how-to-optimize-tombstone-issues-caused-by-updating-primary-keys-in-scylladb-materialized-views/4189 ## Headings Structure: H1: How to Optimize Tombstone Issues Caused by Updating Primary Keys in ScyllaDB Materialized Views? H3: Related topics ## Main Content: H1: How to Optimize Tombstone Issues Caused by Updating Primary Keys in ScyllaDB Materialized Views? H3: Related topics I’m using ScyllaDB to store subscription relationships in the following business scenario: Requirements: The data model needs to support these queries: Is there a data structure that satisfies the functional requirements while avoiding tombstone issues? Hello @qijia_wang, as you correctly noted, modifying any column that is part of the materialized view primary key generates a tombstone for the previous value of the row and inserts a row with the new value. There was a recent fix in this area: Improve timestamp heuristics for tombstone garbage collection by bhalevy · Pull Request #20446 · scylladb/scylladb · GitHub that improves the garbage collection of those tombstones so they don’t accumulate over time. It is available since scylla-6.2. What version are you running? --- ### Page: https://forum.scylladb.com/t/pagination-with-concurrent-inserts-stale-queries-and-caching/4195 Title: Pagination with concurrent inserts, stale queries and caching - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/pagination-with-concurrent-inserts-stale-queries-and-caching/4195 ## Headings Structure: H1: Pagination with concurrent inserts, stale queries and caching H3: Related topics ## Main Content: H1: Pagination with concurrent inserts, stale queries and caching H3: Related topics Originally from the User Slack @Andres: How does pagination state account for concurrent inserts, if at all? I am assuming it doesn’t and can thus cause the pagination state to be “stale” and queries relying on pagination state can then return duplicate and/or skip data, is that correct? @Felipe_Cardeneti_Mendes: > ScyllaDB aims to provide partition-level write isolation, which means that reads must not see only parts of a write made to a given partition, but either all or nothing. To support this, the cache and memtables use MVCC internally. Multiple versions of partition data, each holding incremental writes, may exist in memory and later get merged. https://www.scylladb.com/2018/07/26/how-scylla-data-cache-works/ @Andres: So as long as the partition is still in the cache layer then it will be fine? But if it has been evicted before the next page is grabbed that’s when issues may occur? @Felipe_Cardeneti_Mendes: no, due to MVCC. The read will retrieve whatever existed the moment it gets executed @Andres: I am having a hard time understanding how it solves paging, but perhaps I am just not seeing something obvious. @Felipe_Cardeneti_Mendes: we can’t move backwards on a clustering slice. What happens is that the paging state will hold the last position and continue the read as the application requests for the next page. In other words, if you have a paging state of 1, read 10 non-consecutive clustering rows, inserted a new clustering row which is before what had already been provided, it will just be skipped - but will be available upon a subsequent read. On the other hand, if you insert data on a clustering slice on a clustering slice which is yet to be paged back, and this insert happens prior to when the client requests for the next page, then you’ll see the data. If this doesn’t addresses your question, you should be able to mimic it fairly easily with a paging size of 1 and just playing with another cqlsh opened while tracing your initial read. @avi: Pagination is one of many things that break partition isolation An insert on an already-passed row will obviously not be seen. An insert on a page that hasn’t been fetched yet may or may not be seen. --- ### Page: https://forum.scylladb.com/t/regarding-selecting-map-index-from-a-table/4197 Title: Regarding selecting map index from a table - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I am using this version : [cqlsh 6.1.0 | Cassandra 3.0.8 | CQL spec 3.3.1 | Native protocol v4] and have this following udt and table schema : CREATE TYPE notificationstore.fcm_details ( app_version_code int, … Language: en Canonical URL: https://forum.scylladb.com/t/regarding-selecting-map-index-from-a-table/4197 ## Headings Structure: H1: Regarding selecting map index from a table H3: Related topics ## Main Content: H1: Regarding selecting map index from a table H3: Related topics I am using this version : [cqlsh 6.1.0 | Cassandra 3.0.8 | CQL spec 3.3.1 | Native protocol v4] and have this following udt and table schema : I want to use this CQL query in ScyllaDB Java driver (fetch the fcm_details for the provided userId and fcmToken) : By creating this Select query and SimpleStatement : I also tried using Selector in-place of CqlIdentifier. But in both the scenarios I was getting this error : no viable alternative at input 'fcm_detail_map' Is it possible to serve the use-case with the existing version or any further releases or it’s not yet supported in ScyllaDB? I am using com.datastax.oss.driver.api.core driver in java. You should have been using Selector.column instead of CqlIdentifier.fromCql: --- ### Page: https://forum.scylladb.com/t/why-am-i-seeing-more-than-expected-gossip-writes-in-a-6-node-scylla-cluster-26-writes-instead-of-18/4201 Title: Why am I seeing more than expected gossip writes in a 6-node Scylla cluster? (26 writes instead of 18) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 6.0.2 #Cluster size: 6 os (RHEL/CentOS/Ubuntu/AWS AMI): AWS AMI ami-0d0131dc0789683cc Hello everyone, I have a 6-node Scylla cluster, and according to the official Scylla do… Language: en Canonical URL: https://forum.scylladb.com/t/why-am-i-seeing-more-than-expected-gossip-writes-in-a-6-node-scylla-cluster-26-writes-instead-of-18/4201 ## Headings Structure: H1: Why am I seeing more than expected gossip writes in a 6-node Scylla cluster? (26 writes instead of 18) H3: Related topics ## Main Content: H1: Why am I seeing more than expected gossip writes in a 6-node Scylla cluster? (26 writes instead of 18) H3: Related topics Installation details #ScyllaDB version: 6.0.2 #Cluster size: 6 os (RHEL/CentOS/Ubuntu/AWS AMI): AWS AMI ami-0d0131dc0789683cc I have a 6-node Scylla cluster, and according to the official Scylla documentation, each node communicates with 1 to a maximum of 3 other nodes using the gossip protocol. However, when I check the gossip writes, I’m seeing about 26 writes instead of the expected 18, which corresponds to the 3x number of nodes in my cluster. Has the number of gossip writes increased recently, or could there be another reason for this behavior? Here are some additional details: #ScyllaDB version: 6.0.2 #Cluster size: 6 os (RHEL/CentOS/Ubuntu/AWS AMI): AWS AMI ami-0d0131dc0789683cc Any insights or pointers on why this is happening would be greatly appreciated! However, when I check the gossip writes, I’m seeing about 26 writes How do you count gossip writes? metric, log, other? I am able to see the gossip writes on scylla monitor in advanced tab. Gossiper exchange is not one way, it exchanges information both ways, so theoretically, if each exchange result in a new information to both parties the max amount of writes should be 36. But gossiper does not suppose to write on each exchange, only if there is a new information that has to persisted (which is not usually the case), so the number in steady state should be 0. The gossiper is not the only thing that runs is gossiper scheduling group though, so it is hard to say where the writes are coming from without digging much deeper. --- ### Page: https://forum.scylladb.com/t/select-element-from-collection-map-set-like-map-key/4206 Title: Select element from collection map/set like map[key] - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I want to select specific keys from a map/set in scylla, it seems this has been an open issue since a long time. Is this functionality available in the latest version, if not when can we expect this? Query : select map[… Language: en Canonical URL: https://forum.scylladb.com/t/select-element-from-collection-map-set-like-map-key/4206 ## Headings Structure: H1: Select element from collection map/set like map[key] H3: Related topics ## Main Content: H1: Select element from collection map/set like map[key] H3: Related topics I want to select specific keys from a map/set in scylla, it seems this has been an open issue since a long time. Is this functionality available in the latest version, if not when can we expect this? Query : select map[‘key’] from table where id = ? Git issue : Allow selecting map values and set elements, like in Cassandra 4.0 #7751 Take a look at this PR# https://github.com/scylladb/scylladb/pull/22051 If you grab our latest 2025.1.2 release, you should have the functionally as described in the PR. Best regards, Gabriel --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-2-1-is-now-available-in-scylladb-cloud/4207 Title: [RELEASE] ScyllaDB Enterprise 2024.2.1 is now available in ScyllaDB Cloud! - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: For more information on this version, see here Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-2-1-is-now-available-in-scylladb-cloud/4207 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2024.2.1 is now available in ScyllaDB Cloud! H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2024.2.1 is now available in ScyllaDB Cloud! H3: Related topics For more information on this version, see here --- ### Page: https://forum.scylladb.com/t/enabling-tablets-for-keyspace-after-version-upgrade-sstable-size-limit-when-importing/4209 Title: Enabling Tablets for keyspace after version upgrade, SSTable size limit when importing - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/enabling-tablets-for-keyspace-after-version-upgrade-sstable-size-limit-when-importing/4209 ## Headings Structure: H1: Enabling Tablets for keyspace after version upgrade, SSTable size limit when importing H3: Related topics ## Main Content: H1: Enabling Tablets for keyspace after version upgrade, SSTable size limit when importing H3: Related topics Originally from the User Slack @Ahmed: can anyone guide me for enabling tablet for my existing keyspace upgraded from 5 to 6.2 tried sstableloader it’s give unsupported SSTable format or version @avi: It’s not possible to help with such a vague description of the problem @Ahmed: Hi Avi, Thanks for your response. Let me clarify: • I’ve migrated my cluster to version 6.2 from 5.x. • I have a keyspace with tablet disabled, and its size is approximately 3TB. • My goal is to enable tablet for this keyspace. However, as far as I understand, I can’t simply use an ALTER command to achieve this. • I’ve attempted to use sstableloader for this migration, but I encountered an issue where it reports “unsupported SSTable format or version.” I’ve searched for documentation on enabling tablet for existing keyspaces post-migration but haven’t found anything relevant. Could you guide me on the proper procedure for enabling tablet in this scenario? Thanks in advance! @avi: First, you need to create an empty keyspace with tablets enabled. To migrate data, it’s best to use nodetool refresh --load-and-stream , see https://opensource.docs.scylladb.com/stable/operating-scylla/nodetool-commands/refresh.html Nodetool refresh | ScyllaDB Docs @Ahmed: version 6.2 enable_tablets: true in scylla.yaml CREATE KEYSPACE IF NOT EXISTS test_new WITH replication = {'class': 'org.apache.cassandra.locator.NetworkTopologyStrategy', 'replication_factor': '2'} AND durable_writes = true AND tablets = {'enabled': true}; error ConfigurationException: Tablet replication is not enabled @avi: Set enable_tablets in scylla.yaml and perform a rolling restart @Ahmed: Issue with nodetool refresh Command I am encountering an issue when running the following command: The error message is as follows: @avi: Check the logs for the actual failure. It’s better to load batches of a smaller number of sstables at a time. @Ahmed: i tried importing in multiple chunks and it’s worked thanks I have been encountering a persistent issue where the logs frequently show errors like the following: [shard 0:main] storage_proxy - exception during mutation write to 192.168.100.1: std::runtime_error (Key size too large: 70754 > 65535) @avi: Keys are limited to 64k. This would usually be caught at the CQL layer. But maybe the sstables you imported contained illegal keys. You can use the scylla sstable dump-data command to search for these large keys --- ### Page: https://forum.scylladb.com/t/scylladb-source-available-licensing/4214 Title: ScyllaDB Source Available Licensing - Announcements - ScyllaDB Community NoSQL Forum Meta Description: As announced in this blog, ScyllaDB has decided to focus on a single release stream: ScyllaDB Enterprise. ScyllaDB OSS AGPL 6.2 will stand as the final OSS AGPL release. A free tier of the full-featured ScyllaDB Enterpr… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-source-available-licensing/4214 ## Headings Structure: H1: ScyllaDB Source Available Licensing H3: Related topics ## Main Content: H1: ScyllaDB Source Available Licensing H3: Related topics As announced in this blog, ScyllaDB has decided to focus on a single release stream: ScyllaDB Enterprise. ScyllaDB OSS AGPL 6.2 will stand as the final OSS AGPL release. A free tier of the full-featured ScyllaDB Enterprise will be available to the community. This includes all the performance, efficiency, and security features previously reserved for ScyllaDB Enterprise. Under the ScyllaDB Source Available License, users can download the Enterprise 2024.2 version binaries and run production workloads free up to 50 vCPU and 10 TB (total storage space of all ScyllaDB servers/instances per organization). --- ### Page: https://forum.scylladb.com/t/release-scylla-operator-v1-15-0/4221 Title: [RELEASE] Scylla Operator v1.15.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of Scylla Operator 1.15.0. Scylla Operator is an open-source project that helps ScyllaDB Open Source and ScyllaDB Enterprise users run ScyllaDB on Kubernetes. The … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-operator-v1-15-0/4221 ## Headings Structure: H1: [RELEASE] Scylla Operator v1.15.0 H1: Notable changes H1: Supported versions H1: Upgrade instructions H2: Getting started with Scylla Operator H2: Related Links H3: Related topics ## Main Content: H1: [RELEASE] Scylla Operator v1.15.0 H1: Notable changes H1: Supported versions H1: Upgrade instructions H2: Getting started with Scylla Operator H2: Related Links H3: Related topics The ScyllaDB team is pleased to announce the release of Scylla Operator 1.15.0. Scylla Operator is an open-source project that helps ScyllaDB Open Source and ScyllaDB Enterprise users run ScyllaDB on Kubernetes. The Scylla Operator manages ScyllaDB clusters deployed to Kubernetes and automates tasks related to operating a ScyllaDB cluster, like installation, vertical and horizontal scaling, as well as rolling upgrades. Scylla Operator 1.15.0 improves stability and brings new features. As with all of our releases, all API changes are backward compatible. For more changes and details check out the GitHub release notes. Upgrading from v1.14.x with kubectl apply doesn’t require any extra action, just take the manifest from v1.15.0 tag and substitute the released image. Using helm requires a mandatory manual step for every release because helm can’t handle CRDs updates. For details, see our upgrade documentation. We’ll welcome your feedback! Feel free to open an issue or reach out on the #scylla-operator channel in Scylla User Slack. --- ### Page: https://forum.scylladb.com/t/last-3-weeks-in-scylla-cluster-tests-git-master-issue-74-2024-12-20/4239 Title: Last 3 weeks in scylla-cluster-tests.git master (issue #74; 2024-12-20) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last 3 weeks. Commits in the f21011c6…78c864c9 range are covered. There were 50 non-merge commits from 16 authors i… Language: en Canonical URL: https://forum.scylladb.com/t/last-3-weeks-in-scylla-cluster-tests-git-master-issue-74-2024-12-20/4239 ## Headings Structure: H1: Last 3 weeks in scylla-cluster-tests.git master (issue #74; 2024-12-20) H3: Related topics ## Main Content: H1: Last 3 weeks in scylla-cluster-tests.git master (issue #74; 2024-12-20) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last 3 weeks. Commits in the f21011c6…78c864c9 range are covered. There were 50 non-merge commits from 16 authors in that period. Some notable commits: Collect logs step collects SSL configuration from SCT runner and specific certificates from DB/loader nodes. This will facilitate root cause analysis of SCT failures caused by certificate related issues. To collect performance or other metrics, node-exporter command-line arguments can now be set, e.g., append_scylla_node_exporter_args: '--collector.perf'. A new performance regression test measures latency during operations with disabled RBNO. EBS throughput for gp3 volumes can be configured via the data_volume_disk_throughput parameter. This supports testing scenarios where users favor gp3 volumes. Per-folder job files can now be created to store additional Jenkins job metadata. These are consumed by Argus to determine job types/subtypes. Additionally, users can add a description annotation to the pipeline file, allowing jobs to have descriptions specified. A new test for backup snapshot preparation was added. Backup size can be managed via the SCT_MGMT_PREPARE_SNAPSHOT_SIZE environment variable. docs/configuration_options.md has been reformatted for improved readability and automatic linking. The reliable_replication_factor method no longer returns the highest possible number when tablets are enabled. It is now capped at 3 nodes to address issues in large cluster tests. A mixed keyspaces (vnodes + tablets) Scylla Manager test for Enterprise 2024.2 has been added. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/scylladb-vs-postgresql-citus-cluster-for-real-estate-portal/4248 Title: Scylladb vs postgresql citus cluster for real estate portal - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Thinking if scylladb is the best choice or citus cluster (postgresql) will be easier and also very fast/scalable option for database. It will be for portal in real estate niche so “geo tools” are also important. Any tip… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-vs-postgresql-citus-cluster-for-real-estate-portal/4248 ## Headings Structure: H1: Scylladb vs postgresql citus cluster for real estate portal H3: Related topics ## Main Content: H1: Scylladb vs postgresql citus cluster for real estate portal H3: Related topics Thinking if scylladb is the best choice or citus cluster (postgresql) will be easier and also very fast/scalable option for database. It will be for portal in real estate niche so “geo tools” are also important. It’s hard to say without any information about your use case, such as required performance, ops per second, data size, availability requirements, etc. it is too early to know it but what if it will be something similar in size of zillow after some time? You can get a general idea of what some other users are doing. See a migration use case from Postgres to ScyllaDB by Coralogix and a comparison of PostgreSQL to ScyllaDB (and to MongoDB). --- ### Page: https://forum.scylladb.com/t/error-when-adding-new-nodes-to-cluster-and-repair-based-node-operations-rbno/4254 Title: Error when adding new nodes to cluster, and repair based node operations (RBNO) - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/error-when-adding-new-nodes-to-cluster-and-repair-based-node-operations-rbno/4254 ## Headings Structure: H1: Error when adding new nodes to cluster, and repair based node operations (RBNO) H3: Related topics ## Main Content: H1: Error when adding new nodes to cluster, and repair based node operations (RBNO) H3: Related topics Originally from the User Slack @Joakim_Lindqvist: Hey I am trying to add new larger nodes to my cluster (I intend to eventually replace the old ones with this new larger nodes as we need more storage space for Scylla). Unfortunately I seem to be unable to create new nodes as when they finish bootstrapping (which usually takes about 15 hrs or so) they end up with errors and the node then aborts. This is the error we see. This is on a older version (5.4.5). More Error logs in thread. Some more context around that error: Ater about 15 minutes the shutdown process has completed and it outputs this error, seems to just be a repeat of the reason why it aborted. @dor: It’s not an ‘official’ answer, just a guess - if you didn’t run repair, it’s the streaming that uses RBNO (repair based node operations). You can disable it in the config and add a node in the other way @Joakim_Lindqvist: I can try that. I was not running a repair at the time. @dor: RBNO is repair under the hood @Joakim_Lindqvist: Thank you for your suggestion Dor, I can confirm that this was an issue when using RBNO, when I disabled using repair during bootstrap my nodes were able to bootstrap and join successfully. Thanks! --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-261-2024-12-22/4259 Title: Last week in scylladb.git master (issue #261; 2024-12-22) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 5880a1b90b…10c79a4d47 range are covered. There were 87 non-merge commits from 20 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-261-2024-12-22/4259 ## Headings Structure: H1: Last week in scylladb.git master (issue #261; 2024-12-22) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #261; 2024-12-22) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 5880a1b90b…10c79a4d47 range are covered. There were 87 non-merge commits from 20 authors in that period. Some notable commits: The master branch has been relicensed from AGPL-3.0-or-later to a proprietary source-available license. Read more here. The frozen toolchain used to build ScyllaDB was updated with new build dependencies, in anticipation of porting features from ScyllaDB Enterprise. The container image can now be based on Ubuntu Pro for FIPS support. This feature was ported from ScyllaDB Enterprise. A use-after-free bug due to a race between tablet split and cleanup has been fixed. The materialized view flow control algorithm parameters can now be configured. The container image no longer contains the rsyslog package. The read-repair code is more careful to avoid stalls when reconciling reads with many deleted rows. Restore operations (nodetool restore) can now be aborted. Materialized view updates are now more careful to avoid stalls when calculating affected clustering keys during a view update. The sstable management code now avoids treating sstables that have been unlinked from the filesystem but are still in use as deleted. Node rebuilds that use repair-based node operations now apply the small-table optimization when beneficial. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/latency-issue-and-data-retrieval-page-size-and-number-of-threads-java-driver/4260 Title: Latency issue and data retrieval, page size and number of threads (Java Driver) - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/latency-issue-and-data-retrieval-page-size-and-number-of-threads-java-driver/4260 ## Headings Structure: H1: Latency issue and data retrieval, page size and number of threads (Java Driver) H3: Related topics ## Main Content: H1: Latency issue and data retrieval, page size and number of threads (Java Driver) H3: Related topics Originally from the User Slack @Shresth_Jain: Hi, is there a size limit for partitions cached in Scylla version 2024.1.12-0.20241023.6140bb5b2d0a? I have a partition with around 24k records, and despite running 100-200 RPS queries for the same key, it seems like data is still being fetched from disk. The cache usage in Scylla Monitor isn’t spiking, and this is causing latency to reach several seconds. Any suggestions? @dor: There is no limit. Do you retrieve the entire partitions, or just several rows in each query? Anyway, 24k rows is more work. Best is to work closely with our solution architect @Shresth_Jain: I retrieve an aggregate of 60 columns ( min, max, sum, count). So, all i Retreive is a single row which is the aggregation for a particular partition key. FYi: I am running the exact same query each time. Also, 24k is the max number of rows for a partition. The average is less than 2-3k. @dor: We cache rows but not the aggregation values. You can turn own cql trace to see what’s happening under the hood @Shresth_Jain: Will surely check this. Thanks. Also, is 24k rows too much for a single partition? Say, If I do need to read all this data for a query then would it still be a better decision to try to look for ways to separate out this data across multiple partitions ( Maybe by changing the table schema ) @dor: Well, not necessary, we do have parallel aggregations that read and count partitions in parallel. Again, best to work with the SA who support you to figure out the cons/pros of your data model @Shresth_Jain: Got it. I checked the trace for the query and found that data is indeed being fetched from the cache but it is happening in pages. Each time about 1.1k records are being fetched either from disk or cache and then the next range is being fetched. I hope this page size is configurable. eg: This also seems to me to reason for why the request served by coordinator were around 6ops but the read requests were around 150ops. As there are total of 24k records so it is approx making around 24 read requests per request received at coordinator giving total reads = 24* 6 ~150. Am I correct? Also, Is this good/expected? @dor: Regarding the page size, it’s configurable, on the client side @Shresth_Jain: Could you please help me regarding this? I tried .setPageSize(2000) in java driver but still the page details in trace is same. Reference: @avi: ScyllaDB will limit itself to 1MB pages, so increasing the page size won’t help. If I understand correctly, you’re reading the entire partition at a rate of 6 partitions/sec, and since it has 24k rows you’re reading 144k rows/sec. That’s fine for a single-threaded workload. If you add more threads, reading other partitions, you’ll get a much higher row rate. From the tracing, the partition is cached. You can also check monitoring, there’s cache statistics in the Detailed dashboard. @Shresth_Jain: For reading more threads, the only option would be horizontal or vertical scaling. RIght? Can we configure the 1MB limit? @avi: More threads = change the application to have more threads that read data independently What problem are you trying to solve? @Shresth_Jain: I have an API created using java spring. This API will receive a userId as its parameter and then run a range aggregate query on scyllaDB and then return that aggregated row response. I am using the scyllaDB driver for java. @avi: It’s working fine. Changing parameters won’t help. If your application has a single thread of execution, it won’t be able to utilize all the nodes and all the CPUs that ScyllaDB runs on. @Shresth_Jain: So, you are saying that using more number of thread at the application level will improve this performance? Is there any particular configuration in scylladb Java driver for the same? @avi: You need to make the application parallel. More threads, each consuming a different partition. @Shresth_Jain: The spring api creates a new thread for every new api request so would’t this create the application parallel? @Guy: Hey @Shresth_Jain, were you able to solve this? --- ### Page: https://forum.scylladb.com/t/how-to-stop-scylla-gracefully/4267 Title: How to stop scylla gracefully - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, I installed scylla 5.2.9 on centos 7 as non root and custom directory. I was trying to find a way to stop scylla gracefully instead of using kill command. Also i want to send the logs to a custom directory. Is ther… Language: en Canonical URL: https://forum.scylladb.com/t/how-to-stop-scylla-gracefully/4267 ## Headings Structure: H1: How to stop scylla gracefully H3: Related topics ## Main Content: H1: How to stop scylla gracefully H3: Related topics Hi, I installed scylla 5.2.9 on centos 7 as non root and custom directory. I was trying to find a way to stop scylla gracefully instead of using kill command. Also i want to send the logs to a custom directory. Is there a way or parameter i can do that. started scylla as below /opt/scylladb/bin/scylla --io-properties-file=/opt/scylladb/etc/scylla.d/io_properties.yaml --options-file=/opt/scylladb/conf/scylla.yaml I am working on poc where there are restrictions not to use systemd or systemctl commands. Since you explicitly asked not to use systemd or systemctl (which are the “normal” way to start and stop services), you are left with the kill command. However, before you kill the node you should consider running nodetool drain on this node. This command (see its documentation) will tell the running Scylla to flush all the data to disk, to finish all ongoing request and stop accepting new requests. Once nodetool drain is done, it is safe to kill the scylla process with whichever signal you want (SIGTERM or SIGKILL), and there is no risk of losing any unsaved data or interrupting a request. --- ### Page: https://forum.scylladb.com/t/anyone-writing-to-the-alternator-interface-in-rust/4273 Title: Anyone writing to the Alternator interface in Rust? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I’m working in Rust, and am interested in talking to ScyllaDB via their Alternator API. I gather I can just use the AWS S3 SDK, configured appropriately. Not seeing a lot of examples laying around… anyone here done this? … Language: en Canonical URL: https://forum.scylladb.com/t/anyone-writing-to-the-alternator-interface-in-rust/4273 ## Headings Structure: H1: Anyone writing to the Alternator interface in Rust? H3: Related topics ## Main Content: H1: Anyone writing to the Alternator interface in Rust? H3: Related topics I’m working in Rust, and am interested in talking to ScyllaDB via their Alternator API. I gather I can just use the AWS S3 SDK, configured appropriately. Not seeing a lot of examples laying around… anyone here done this? Indeed, Alternator is API-compatible with DynamoDB, so you can just use Amazon’s regular AWS SDK for Rust - which you called it “S3 SDK” but they are actually SDKs for all of AWS, including S3 but also DynamoDB and other things. There’s a lot of Amazon documentation on how to do this - e.g. DynamoDB examples using SDK for Rust - AWS SDK for Rust. One thing you should be aware of, however, is that Amazon’s SDKs normally accept one “endpoint URL” for the DynamoDB service. An example endpoint URL is https://dynamodb.us-east-1.amazonaws.com. But Alternator is actually a cluster of multiple nodes - so which of them will you send the requests to? You don’t want to just pick one of them, you want to balance the load of requests between all nodes. So you’ll probably need some sort of load-balancing solution. You can have a server-side load balancer (e.g., an HTTP load balancer, or a DNS load balancer) in front of your cluster. Or you can have a client-side load balancer - a small “hack” to the AWS SDK that will allow it to pick a different ScyllaDB node for each request instead of always the same node. We have examples of such client-side load balancers in GitHub - scylladb/alternator-load-balancing: Various tricks, scripts, and libraries, for load balancing multiple Alternator nodes - but unfortunately not yet in Rust. Thanks for getting back to me. I of course meant to say “DynamoDB SDK” (not S3) and have been happily using it to code against the Alternator API for the past few months. Also aware of the load-balancing issue. To anyone who finds this via search, you can just use the DDB SDK. Connect with aws_config::defaults(BehaviorVersion::latest()).endpoint_url(/* Alternator endpoint */) & off you go. --- ### Page: https://forum.scylladb.com/t/change-column-type-of-existing-table-change-data-capture-cdc-and-collection-support-for-analytics/4274 Title: Change column type of existing table, Change Data Capture (CDC) and collection support for analytics - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/change-column-type-of-existing-table-change-data-capture-cdc-and-collection-support-for-analytics/4274 ## Headings Structure: H1: Change column type of existing table, Change Data Capture (CDC) and collection support for analytics H3: Related topics ## Main Content: H1: Change column type of existing table, Change Data Capture (CDC) and collection support for analytics H3: Related topics Originally from the User Slack @Nishant_Kumar: Hii, I have a table, in which for few column type is UDT, for some reason I can’t use UDT and want to change it to text. Is it supported in scylla? I tried with alter type query locally Scylla 6.0.1, it didn’t work. Any way to do it? other than creating a new table and migrating the data? Your help highly appreciated. @Botond_Dénes: In ScyllaDB you can only alter the type of columns to compatible types. So e.g. you can change from int to bigint, but not the other way around. There are no compatible types that you can convert a UDT to, so the only option is to migrate to a new table. What is the reason why you can’t use UDT? Is it some ScyllaDB problem, or is it something on the app/business side? @Nishant_Kumar: @Botond_Dénes We need to put scylla db data to athena for analytics purposes, currently it’s executed BI_HOURLY, and at this time we are seeing spike in latency. bcz data-platform team is saying Scylla Debezium CDC doesn’t have support for collection types (LIST, SET, MAP) and UDT - columns. These columns are omitted before Debezium CDC processing. For handling complex data types, we currently scan the entire CDC table, which retains two days of data, and subsequently process it to extract newly updated or inserted records. This full table scan is necessary because Scylla CDC tables store timestamps as UUIDs, which cannot directly track newer records. As a result, we bring in all the data, and the UUIDs are then converted into DateTime format for further processing. is there any way to solve this? For ref: https://opensource.docs.scylladb.com/stable/using-scylla/integrations/scylla-cdc-source-connector.html https://opensource.docs.scylladb.com/stable/features/cdc/cdc-log-table.html#time-column ScyllaDB CDC Source Connector | ScyllaDB Docs The CDC Log Table | ScyllaDB Docs @Md_Anees: @Botond_Dénes https://github.com/scylladb/scylla-cdc-source-connector/issues/9 GitHub: feature request: support for collection types (LIST, SET, MAP) and UDT · Issue #9 · scylladb/scylla-cdc-source-connector We are heavily using UDT (frozen) in our tables, which is causing issues while adding them to CDC table. Any alternates would help. @Botond_Dénes: Sorry, I’m not aware of a workaround for this. @Karol_Baryła: https://github.com/scylladb/scylla-cdc-source-connector/pull/21 adds collections support to scylla-cdc-connector. I’ve raised the topic of merging it and releasing a new version recently - cc @Wojciech_Bączkowski @Dmitry_Kropachev GitHub: Support for collections by Lorak-mmk · Pull Request #21 · scylladb/scylla-cdc-source-connector --- ### Page: https://forum.scylladb.com/t/how-to-update-the-tombstone-warn-threshold-value-rolling-restart/4276 Title: How to update the tombstone_warn_threshold value, rolling restart - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/how-to-update-the-tombstone-warn-threshold-value-rolling-restart/4276 ## Headings Structure: H1: How to update the tombstone_warn_threshold value, rolling restart H3: Related topics ## Main Content: H1: How to update the tombstone_warn_threshold value, rolling restart H3: Related topics Originally from the User Slack @Piette_Dylan: Hello everyone, Is it possible to update tombstone_warn_threshold by just updating its value in the system.config table via a CQL query ? If yes and if I want all my nodes to share the same value do I need to execute the query on every node of the cluster or just one ? https://opensource.docs.scylladb.com/stable/reference/configuration-parameters.html#tombstone-settings Configuration Parameters | ScyllaDB Docs @Felipe_Cardeneti_Mendes: no, it is not a live-updateable value - maybe it should be. cc @Botond_Dénes From an operational perspective - for liveupdate config changes you’d be better off updating scylla.yaml and SIGHUP the scylla pid on all nodes. @Piette_Dylan: Thanks for you answer @Felipe_Cardeneti_Mendes So I don’t need to restart Scylla with that method then ? And sorry for my (maybe dumb) question but how do you SIGHUP ? @Botond_Dénes: You do have to do a rolling restart, because this config item is not live-update. kill -SIGHUP $(pidof scylla) @Piette_Dylan: But the method from Felipe works right ? @Botond_Dénes: SIGHUP only works with live-update config items. For non live-update, you have to update every scylla.yaml then do rolling restart @Piette_Dylan: Ha ok, thanks a lot ! --- ### Page: https://forum.scylladb.com/t/scylla-manager-using-helm-and-configure-s3-storage/4282 Title: Scylla Manager using Helm and configure S3 storage - Kubernetes Operator - ScyllaDB Community NoSQL Forum Meta Description: Hi everyone, I’m trying to deploy Scylla Manager using Helm and configure S3 storage with Wasabi. Below is my values.manager.yaml: image: tag: 3.3.0 resources: limits: cpu: 200m memory: 256Mi requests: … Language: en Canonical URL: https://forum.scylladb.com/t/scylla-manager-using-helm-and-configure-s3-storage/4282 ## Headings Structure: H1: Scylla Manager using Helm and configure S3 storage H3: Related topics ## Main Content: H1: Scylla Manager using Helm and configure S3 storage H3: Related topics I’m trying to deploy Scylla Manager using Helm and configure S3 storage with Wasabi. Below is my values.manager.yaml: I also created a scylla-agent-config ConfigMap with the following content: When I attempt to take a backup using: What configurations or steps am I missing to successfully configure Scylla Manager with Wasabi S3 storage so my backups work? Overall scyllaAgentConfig needs to reference a Secret, not a ConfigMap, and needs to be set up for the target ScyllaDB cluster, not the one used by Scylla Manager. With just these snippets it’s hard to say if these are all the changes required to wire it correctly, so please create an issue with must-gather archive attached. --- ### Page: https://forum.scylladb.com/t/trying-to-configure-alternator-authorization-in-scylladb-cloud/4283 Title: Trying to configure alternator authorization in ScyllaDB cloud - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I’m experimenting with the Alternator interface in a cluster deployed to ScyllaDB cloud. I’m using OpenTofu with the ScyllaDB provider, scylladbcloud_cluster resource, but it has no option to configure alternator_enforce… Language: en Canonical URL: https://forum.scylladb.com/t/trying-to-configure-alternator-authorization-in-scylladb-cloud/4283 ## Headings Structure: H1: Trying to configure alternator authorization in ScyllaDB cloud H3: Related topics ## Main Content: H1: Trying to configure alternator authorization in ScyllaDB cloud H3: Related topics I’m experimenting with the Alternator interface in a cluster deployed to ScyllaDB cloud. I’m using OpenTofu with the ScyllaDB provider, scylladbcloud_cluster resource, but it has no option to configure alternator_enforce_authorization. What am I missing? Hey sp1ff, thanks for reaching out. Currently, the alternator_enforce_authorization feature is not supported as self-service for Alternator in ScyllaDB Cloud. This can be enabled for your cluster, by reaching to our support. However, this will be added in an upcoming release, as a full Role-Based Access Control (RBAC) feature to be applied to all Alternator tables. Stay tuned for updates! --- ### Page: https://forum.scylladb.com/t/java-driver-configuration/4285 Title: Java driver configuration - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: When i try to use application.conf file in my spring boot application for customize e.g contact-points there is no effect. It seems that application.conf doesnt load…But if i use application.yaml with contact-points its … Language: en Canonical URL: https://forum.scylladb.com/t/java-driver-configuration/4285 ## Headings Structure: H1: Java driver configuration H3: Related topics ## Main Content: H1: Java driver configuration H3: Related topics When i try to use application.conf file in my spring boot application for customize e.g contact-points there is no effect. It seems that application.conf doesnt load…But if i use application.yaml with contact-points its work java driver version 4.15.0 spring boot v3.1.2 application.conf in src/main/resources: datastax-java-driver { basic.contact-points = [10.72.63.81,10.72.63.82,10.72.63.83] } I don’t have experience with building Spring applications, so it’s hard for me to say what could be going wrong here without seeing the application. If my understanding is correct application.yml is a configuration file for the spring application and application.conf is the usual Typesafe config that should be picked up by the driver. It should be possible to pass driver configuration both ways, but using different formats. First, it seems that the format you’re using for your application.conf is slightly wrong. Perhaps this is why the driver rejects those contact points. Each address should be of the “host:port” format. Basically this: I tried setting up example application with such application.conf and I’ve had no issues customizing the contact points. I’ve put it in src/main/resources too. Another issue you may be facing is that one of the configs overwrites what another says. Make sure that you define contact points only in one of them and also not in any extra environment variables. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-75-2024-12-28/4286 Title: Last week in scylla-cluster-tests.git master (issue #75; 2024-12-28) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 03cf4e54…f8670793 range are covered. There were 8 non-merge commits from 5 authors in tha… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-75-2024-12-28/4286 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #75; 2024-12-28) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #75; 2024-12-28) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 03cf4e54…f8670793 range are covered. There were 8 non-merge commits from 5 authors in that period. Some notable commits: Sending results to Argus with the latency_calculator_decorator now supports validation rules. Default error thresholds are set to 5ms for P90 and 10ms for P99 latencies, but these can be overridden for each workload and result type (e.g., nemesis or predefined step). Throughput validation is also supported. The int_or_list config parameter type has been renamed to int_or_space_separated_ints, making it more precise and descriptive. To reduce noise from development images, Scylla image listings (excluding branches) are now limited to those with the environment=production label. A new option for appending string or list configuration values has been introduced. Strings can be appended by prefixing with ++: export SCT_APPEND_SCYLLA_ARGS="++ --overprovisioned 1". Similarly, lists can be appended by including ++ as the first item: export SCT_SCYLLA_D_OVERRIDES_FILES='["++", "extra_file/scylla.d/io.conf"]'. This feature applies only to str_or_list_or_eval parameters and is disabled for certain ones. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/query-timeouts-using-timeout-and-read-request-timeout-in-ms-behavior/4295 Title: Query timeouts - Using Timeout and read_request_timeout_in_ms behavior - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/query-timeouts-using-timeout-and-read-request-timeout-in-ms-behavior/4295 ## Headings Structure: H1: Query timeouts - Using Timeout and read_request_timeout_in_ms behavior H3: Related topics ## Main Content: H1: Query timeouts - Using Timeout and read_request_timeout_in_ms behavior H3: Related topics Originally from the User Slack @Lina_Sharifi: basic question: USING TIMEOUT - https://enterprise.docs.scylladb.com/stable/cql/cql-extensions.html#using-timeout for a given query how does the server reacts when the the query does not finish by that timeout for these scenarios : 1- USING TIMEOUT value is less than server timeout read_request_timeout_in_ms 2-USING TIMEOUT value is more than server timeout read_request_timeout_in_ms @Botond_Dénes: There is no difference. When a query has USING TIMEOUT, the specified timeout will be used instead of read_request_timeout_in_ms for that query. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-262-2024-12-29/4299 Title: Last week in scylladb.git master (issue #262; 2024-12-29) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 10c79a4d47…3e22998dc1 range are covered. There were 28 non-merge commits from 8 authors in that period… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-262-2024-12-29/4299 ## Headings Structure: H1: Last week in scylladb.git master (issue #262; 2024-12-29) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #262; 2024-12-29) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 10c79a4d47…3e22998dc1 range are covered. There were 28 non-merge commits from 8 authors in that period. Some notable commits: The CQL NOT IN operator has been implemented. It can be used in the WHERE clause and the IF clause. Many unit tests are now built into a single executable, reducing build footprint and time. When selecting replicas for a read, we randomize the replica list order to get better distribution with tablets. The replica list selection itself is faster as well. The load-and-stream facility has a new scope parameter that can limit data copying to within the same node, the same rack, or the same datacenter. This can improve restore performance when restoring sstables on all racks simultaneously. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/drivers-and-consistency-with-multiple-datacenters-and-local-quorum/4305 Title: Drivers and consistency with multiple datacenters and local quorum - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/drivers-and-consistency-with-multiple-datacenters-and-local-quorum/4305 ## Headings Structure: H1: Drivers and consistency with multiple datacenters and local quorum H3: Related topics ## Main Content: H1: Drivers and consistency with multiple datacenters and local quorum H3: Related topics Originally from the User Slack @Some_Random_Guy: quick ? re: clients and consistency. if you have a cluster (let’s call the nodes n1-dc1 and n2-dc2). n1-dc1 is in dc1 rack1, n2-dc2 is in dc2 rack1. In my app, I have the consistency configured to be local_quorum and the host set to n1-dc1. Should my app be trying to connect to n2-dc2? @Felipe_Cardeneti_Mendes: Connecting yes, using dc2 nodes as coordinators no If you want to prevent connecting most drivers allow you to either blacklist a DC @Some_Random_Guy: oh ok thanks @Felipe_Cardeneti_Mendes --- ### Page: https://forum.scylladb.com/t/allocation-problems-6-1/4314 Title: Allocation problems 6.1 - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, ScyllaDB version 6.1.4 seems to have less memory allocation warnings than 6.0. Nevertheless we are sporadically still seeing some allocator failures: Dec 30 08:44:14 o-p-L309-7 scylla[1843673]: [shard 0:stmt] lsa… Language: en Canonical URL: https://forum.scylladb.com/t/allocation-problems-6-1/4314 ## Headings Structure: H1: Allocation problems 6.1 H3: Related topics ## Main Content: H1: Allocation problems 6.1 H3: Related topics ScyllaDB version 6.1.4 seems to have less memory allocation warnings than 6.0. Nevertheless we are sporadically still seeing some allocator failures: An allocation of 128MB seems rather large. Might this be an issue with scylla or are we doing something wrong in our queries? Today I saw a strange one, related to compactions: A alloc error during compaction seems worrisome to me. Do these mean that its loading 128MB of mutations in a query? The allocation warnings and failures you’re seeing can be due to large partitions and large cells, as well as internal memory management details of Scylla’s allocator. Here are some points to consider: Allocations of 128MB or more can happen if a single partition is very large (you mentioned some partitions up to 1GB with 30 million rows), which puts high pressure on memory allocation and can trigger allocator headroom increases or failures. Large cells (up to 3MB) are generally acceptable but combined with large partitions can exacerbate memory pressure. Scylla’s internal allocator (logalloc) may increase allocation granularity under load, which is normal but can cause warnings during peak demand. Frequent memory allocation failures might also correlate with query patterns or long-running queries that require large buffers. It is recommended to review data model to avoid extremely large partitions if possible (e.g., split data into smaller partitions). Monitoring and tuning JVM heap, page cache, and memory-related settings in Scylla may help, but the fundamental issue here is the workload profile and partition size. If patterns of frequent bad_alloc or allocator failures persist and impact stability, consider reaching out to ScyllaDB support with detailed logs and metrics to investigate potential improvements or fixes. In summary, this is likely not a bug but a symptom of heavy memory pressure caused by very large partitions and data sizes, combined with high query demands. The best path is to monitor partition sizes, avoid extremely large partitions if possible, and tune memory settings accordingly. I would also recommend upgrading to a recent version - ScyllaDB 2025.3.0 and see if the issues persist. Thanks for the feedback! sorry, I forgot to update the thread: Since we have upgraded to 2024.1.8 our issues are gone and the logs are also so much cleaner (due to `managed_bytes` violates preferred contiguous allocation size · Issue #23781 · scylladb/scylladb · GitHub). I think these few large partitions we have are basically never read and therefore dont cause any/much harm. As since that fix our logs are pretty clean. Nevertheless we are trying to remove the large partitions. --- ### Page: https://forum.scylladb.com/t/aio-max-nr-issue-could-not-setup-async-i-o-resource-temporarily-unavailable/4315 Title: Aio-max-nr issue - Could not setup Async I/O: Resource temporarily unavailable - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/aio-max-nr-issue-could-not-setup-async-i-o-resource-temporarily-unavailable/4315 ## Headings Structure: H1: Aio-max-nr issue - Could not setup Async I/O: Resource temporarily unavailable H3: Related topics ## Main Content: H1: Aio-max-nr issue - Could not setup Async I/O: Resource temporarily unavailable H3: Related topics Originally from the User Slack @Johannes_Gilger: I am getting an error when trying to run nodetool on a running node: what(): Could not setup Async I/O: Resource temporarily unavailable. The required nr_event (1) exceeds the limit of request capacity in /proc/sys/fs/aio-max-nr (1048576) I have a 32-vcpu node and aix-max-nr set to 1048576. I don’t know why I’m getting this error, it works fine on other nodes with the same configuration. It’s just two out of six nodes that throw this error. Oh, hold on, another node has 2097152. Hmm @Johannes_Gilger: Mystery solved, don’t mind me @Robert_Bindar: If you’d like to share what was wrong, it’d be highly appreciated, others might search the chat in the future and find this helpful, cheers! @Johannes_Gilger: Well the aix-max-nr was too low on three out of six nodes Set it to 2097152 now and it works --- ### Page: https://forum.scylladb.com/t/iosetup-error-aio-max-nr-error-root-did-not-pass-validation-tests-it-may-not-be-on-xfs-and-or-has-limited-disk-space/4319 Title: IOsetup error, aio-max-nr, ERROR: root: ... did not pass validation tests, it may not be on XFS and/or has limited disk space - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/iosetup-error-aio-max-nr-error-root-did-not-pass-validation-tests-it-may-not-be-on-xfs-and-or-has-limited-disk-space/4319 ## Headings Structure: H1: IOsetup error, aio-max-nr, ERROR: root: ... did not pass validation tests, it may not be on XFS and/or has limited disk space H3: Related topics ## Main Content: H1: IOsetup error, aio-max-nr, ERROR: root: ... did not pass validation tests, it may not be on XFS and/or has limited disk space H3: Related topics Originally from the User Slack @Keith_Mciff: Hey folks. I am trying to run iosetup and I get this error and if but the directories are there and I even used chmod 777 to see if it was some kind of permissions issue Any suggestions where I can check next for the problem? @Jethro: What is the output of the following commands; mount | grep '/data' and: df -h /data @Jethro: hmm thats look good to me. And ls -lah / | grep data just to check if /data has the correct user/group? @Keith_Mciff: yeah. we were able to solve the issue with sudo sysctl -w fs.aio-max-nr=1048576 It’s not a very good error message but after lots of google we found something on seastar with that error and solution --- ### Page: https://forum.scylladb.com/t/meaning-of-reader-concurrency-semaphore-error-latency-spikes-and-hot-partitions/4320 Title: Meaning of reader_concurrency_semaphore error, latency spikes and hot partitions - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/meaning-of-reader-concurrency-semaphore-error-latency-spikes-and-hot-partitions/4320 ## Headings Structure: H1: Meaning of reader_concurrency_semaphore error, latency spikes and hot partitions H3: Related topics ## Main Content: H1: Meaning of reader_concurrency_semaphore error, latency spikes and hot partitions H3: Related topics Originally from the User Slack @Chaitanya_Tondlekar: what does the below error means ? [shard 3] reader_concurrency_semaphore @Guy ^^ Scylla version : 5.2.19 Compaction strategy : SizeTieredCompactionStrategy Latencies are shooting up in seconds whereas reads/writes are having same. Also cache HITS are being too high. @avi can you help here ? @Guy: Hey @Chaitanya_Tondlekar, did you figure this out? @Chaitanya_Tondlekar: nope we are not able to understand @avi: You probably have a hot partition. Try the nodetool toppartitions command to find out which. @Chaitanya_Tondlekar: @avi: Add -s 2000 to get better discrimination @Chaitanya_Tondlekar: @avi: Looks like you don’t have a hot partition Look at the advanced dashboard, per shard, CPU panel @Chaitanya_Tondlekar: when issue was happening, we saw that cache hits went high for long time and it came down after sometime on its own. @avi: This is typical of a hot partition. Maybe the call to nodetool was after the event passed. @Chaitanya_Tondlekar: cool --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-2-2/4321 Title: [RELEASE] ScyllaDB Enterprise 2024.2.2 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2024.2.2, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2024.2 Feature (Short Term Support) Release. 2024.2.2 patch release includes… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-2-2/4321 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2024.2.2 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2024.2.2 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2024.2.2, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2024.2 Feature (Short Term Support) Release. 2024.2.2 patch release includes multiple bug fixes. The following issues are fixed in this release (with an open-source reference, if available): CQL and correctness related --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-76-2025-01-03/4324 Title: Last week in scylla-cluster-tests.git master (issue #76; 2025-01-03) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 2bfc7917…ceb441b9 range are covered. There were 11 non-merge commits from 5 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-76-2025-01-03/4324 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #76; 2025-01-03) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #76; 2025-01-03) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 2bfc7917…ceb441b9 range are covered. There were 11 non-merge commits from 5 authors in that period. Some notable commits: latency_calculator_decorator, which measures latencies, duration, and throughput during a decorated function and automatically sends results to Argus, can now be easily used in all longevity tests. To use it, specify workload_name (as read, write, or mixed) and set use_hdr_cs_histogram: true in test config. The Java version on sct-runners has been updated to resolve a version mismatch issue between Jenkins master and the SCT builder. Scylla Manager tests will now run on Scylla 2024.2 by default. Nemesis flow is aborted when a stress command within nemesis fails, as continuing without data or in a corrupted state is pointless. The performance regression Grow-Shrink cluster test now runs with hinted handoff enabled. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/replaced-one-dead-node-and-new-node-came-up-same-hostid-in-scylla-4-6/4325 Title: Replaced one dead node and new node came up same hostid in scylla 4.6 - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Replaced one dead node and new node came up with same hostid in scylla 4.6. We have replaced one dead node and the new node came up with same hostid.we are getting below Host ID collision in log.Could you please help on… Language: en Canonical URL: https://forum.scylladb.com/t/replaced-one-dead-node-and-new-node-came-up-same-hostid-in-scylla-4-6/4325 ## Headings Structure: H1: Replaced one dead node and new node came up same hostid in scylla 4.6 H3: Related topics ## Main Content: H1: Replaced one dead node and new node came up same hostid in scylla 4.6 H3: Related topics Replaced one dead node and new node came up with same hostid in scylla 4.6. We have replaced one dead node and the new node came up with same hostid.we are getting below Host ID collision in log.Could you please help on it.Please find below log Jan 3 18:10:14 ip-10-233-65-246 scylla: [shard 0] gossip - 60000 ms elapsed, gossip quarantine over Jan 3 18:10:15 ip-10-233-65-246 scylla: [shard 0] storage_service - Host ID collision for between and ; ignored Jan 3 18:10:15 ip-10-233-65-246 scylla: [shard 0] storage_service - handle_state_normal: Nodes and have the same token -9212246330932197951. Ignoring Jan 3 18:10:15 ip-10-233-65-246 scylla: [shard 0] storage_service - handle_state_normal: Nodes and have the same token -2023138897313104219. Ignoring Jan 3 18:10:15 ip-10-233-65-246 scylla: [shard 0] storage_service - handle_state_normal: Nodes and have the same token 4660978197628450497. Ignoring Jan 3 18:10:15 ip-10-233-65-246 scylla: [shard 0] storage_service - handle_state_normal: Nodes and have the same token -9091405894166869642. Ignoring Jan 3 18:10:15 ip-10-233-65-246 scylla: [shard 0] storage_service - handle_state_normal: Nodes and have the same token 1716034972238154026. Ignoring Jan 3 18:10:15 ip-10-233-65-246 scylla: [shard 0] storage_service - handle_state_normal: Nodes and have the same token 202759629005760746. Ignoring Hi Ranjeet, 4.6 is old and no longer supported. Please upgrade to the latest version and let me know if the problem still occurs. Hi Guy, As the cluster is unstable so we can not proceed with upgrade.As one node was terminated and trying to replace that node but cluster topology change is not working.(replace node and removenode ) is not working. even try to remove the dead node but not able to remove that node and getting below error Jan 3 20:03:22 ip-10-233-64-213 scylla: [shard 0] storage_service - removenode[0694a7e0-361f-4d07-94a9-e97b962de6aa]: Added node= as leaving node, coordinator= Jan 3 20:04:27 ip-10-233-64-213 scylla: [shard 0] storage_service - removenode[0694a7e0-361f-4d07-94a9-e97b962de6aa]: Node is down for node_ops_cmd verb Jan 3 20:04:27 ip-10-233-64-213 scylla: [shard 0] storage_service - removenode[0694a7e0-361f-4d07-94a9-e97b962de6aa]: Node is down for node_ops_cmd verb Jan 3 20:04:27 ip-10-233-64-213 scylla: [shard 0] storage_service - removenode[0694a7e0-361f-4d07-94a9-e97b962de6aa]: Node is down for node_ops_cmd verb Jan 3 20:08:16 ip-10-233-64-213 scylla: [shard 0] storage_service - Operation removenode is in progress, wait for it to complete Jan 3 20:08:17 ip-10-233-64-213 scylla: [shard 0] storage_service - No tokens to force removal on, call ‘removenode’ --- ### Page: https://forum.scylladb.com/t/if-the-synchronous-write-of-the-materialized-view-fails-will-the-base-table-also-fail-to-be-written/4329 Title: If the synchronous write of the materialized view fails, will the base table also fail to be written? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: After synchronous write is enabled for the materialized view, if a replica of the corresponding materialized view fails to be written during the alternator put_item, will put_item fail? From the code, I can see that eve… Language: en Canonical URL: https://forum.scylladb.com/t/if-the-synchronous-write-of-the-materialized-view-fails-will-the-base-table-also-fail-to-be-written/4329 ## Headings Structure: H1: If the synchronous write of the materialized view fails, will the base table also fail to be written? H3: Related topics ## Main Content: H1: If the synchronous write of the materialized view fails, will the base table also fail to be written? H3: Related topics After synchronous write is enabled for the materialized view, if a replica of the corresponding materialized view fails to be written during the alternator put_item, will put_item fail? From the code, I can see that even if an exception occurs in function mutate_MV, it will not be handled in function generate_and_propagate_view_updates. With synchronous materialized views, the view replica == base replica, so the view-write is local and will immediately succeed/fail. The base write will wait up the view update to complete and failures are propagated, so a failed view update will fail the base write too. --- ### Page: https://forum.scylladb.com/t/how-to-migrate-data-from-one-cluster-to-another/4330 Title: How to migrate data from one cluster to another - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/how-to-migrate-data-from-one-cluster-to-another/4330 ## Headings Structure: H1: How to migrate data from one cluster to another H3: Related topics ## Main Content: H1: How to migrate data from one cluster to another H3: Related topics Originally from the User Slack @Sujay_KS: hey, I want to migrate my entire data which is in scylla cluster 1 to cluster 2, is there any way by which we can do this, I have explored backup and restore way but realised for each table I should do the restoration process this will take lot of time doc: https://opensource.docs.scylladb.com/stable/operating-scylla/procedures/backup-restore/restore.html Restore from a Backup and Incremental Backup | ScyllaDB Docs @Lubos: why not making second cluster a DC of old cluster and then decomm old DCs? @Sujay_KS: yes, will try this approach --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-263-2025-01-05/4334 Title: Last week in scylladb.git master (issue #263; 2025-01-05) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 3e22998dc1…c973254362 range are covered. There were 81 non-merge commits from 15 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-263-2025-01-05/4334 ## Headings Structure: H1: Last week in scylladb.git master (issue #263; 2025-01-05) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #263; 2025-01-05) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 3e22998dc1…c973254362 range are covered. There were 81 non-merge commits from 15 authors in that period. Some notable commits: Release builds now use link-time optimization (LTO) which considers all source files together rather than independently, and profile-guided optimization (PGO) which use training workloads to guide the optimizer. Performance improvements of around 30% have been observed. This feature was contributed from ScyllaDB Enterprise. Intra-node remote procedure call (RPC) learned dictionary-based compression, which gathers a dictionary to get better compression ratio. After the dictionary is gathered and distributed, compression is able to refer to byte strings from the dictionary, rather than just each message as it is compressed. This feature was contributed from ScyllaDB Enterprise. The CQL language can now select a collection element rather than the entire collection. It is now possible to assign CPU consumption and disk bandwidth shares to different workloads on the same cluster. Assignment is based on login role via the service level mechanism. This feature was contributed from ScyllaDB Enterprise. A new compaction strategy, Incremental Compaction Strategy, is available. It splits sstables into one gigabyte fragments. The principal benefit is that a 50% space reserve is no longer necessary. This feature was contributed from ScyllaDB Enterprise. The system.clients lists the connection stage of each connection. A bug where a ready connection was listed as still authenticating was fixed. The reader concurrency semaphore manages query concurrency on the replica. It received a small performance optimization for workloads that typically miss the cache. The effective replication map caches the result of applying the replication strategy to the current topology. A race condition in updating it after a tablet split was fixed. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/maximal-limit-for-the-number-of-partitions-and-other-partition-recommendations/4338 Title: Maximal limit for the number of partitions and other partition recommendations - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/maximal-limit-for-the-number-of-partitions-and-other-partition-recommendations/4338 ## Headings Structure: H1: Maximal limit for the number of partitions and other partition recommendations H3: Related topics ## Main Content: H1: Maximal limit for the number of partitions and other partition recommendations H3: Related topics Originally from the User Slack @hamonica: Hello, I have a question about partitions. There are 100 WAS servers, and each server stores 100 data points per second in ScyllaDB. On average, each server processes 100 TPS, which means 10,000 records are stored per second, and 600,000 records are stored per minute. I plan to use a partition key based on minute + WAS ID (min_timestamp, instance_id). Would this partition be too large and negatively impact performance? @hamonica: um… I’d like to know the recommended number of partitions. Would it be better to split the partitions into smaller sizes rather than having larger partitions? @Botond_Dénes: In general, the more partitions the better. @hamonica: Is there a limit on the number of partitions that can be created in a day? Or is there a recommended number of partitions? @Botond_Dénes: In ScyllaDB there are not such thing as too many partitions. I have never seen any production problem that resulted from the number of partitions. When there is a problem with partitions, it is always related to their size. Both too large and too small (just a few bytes) can be a problem. Ok, having very few partitions can cause problems, because it is impossible to spread the load across the cluster. But on the other end of the scale, I have not seen problems yet. @hamonica: @Botond_Dénes Thank you very much, it is very helpful. I will think about how to further subdivide the partitions without causing problems in querying. --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-14/4339 Title: [RELEASE] ScyllaDB Enterprise 2024.1.14 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2024.1.14, a bug-fix production-ready patch release for ScyllaDB Enterprise 2024.1 Long Term Support (LTS) Release. Note there is a later Feature (Short term Support) Rele… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-14/4339 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2024.1.14 H3: Related Links H2: Fixed Issues H3: Performance Related H3: Stability related H3: Materialized Views Related H3: Documentation H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2024.1.14 H3: Related Links H2: Fixed Issues H3: Performance Related H3: Stability related H3: Materialized Views Related H3: Documentation H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2024.1.14, a bug-fix production-ready patch release for ScyllaDB Enterprise 2024.1 Long Term Support (LTS) Release. Note there is a later Feature (Short term Support) Release 2024.2. The following issues are fixed in this release (with an open-source reference, if available): In this release, the hard-coded default is replaced with a live-updateable configuration parameter, view_flow_control_delay_limit_in_ms, which defaults to 1000ms as before. #18187 Add metrics for MV base_writes throttling: --- ### Page: https://forum.scylladb.com/t/unexpected-behavior-with-timewindowcompactionstrategy-in-scylladb-6-2-open-source/4345 Title: Unexpected Behavior with TimeWindowCompactionStrategy in ScyllaDB 6.2 Open Source - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi Community, I’m using ScyllaDB 6.1 Open Source and have a table configured to store 30 days of data with the following compaction strategy: compaction = { ** ‘class’: ‘TimeWindowCompactionStrategy’,** ** ‘compact… Language: en Canonical URL: https://forum.scylladb.com/t/unexpected-behavior-with-timewindowcompactionstrategy-in-scylladb-6-2-open-source/4345 ## Headings Structure: H1: Unexpected Behavior with TimeWindowCompactionStrategy in ScyllaDB 6.2 Open Source H3: Related topics ## Main Content: H1: Unexpected Behavior with TimeWindowCompactionStrategy in ScyllaDB 6.2 Open Source H3: Related topics I’m using ScyllaDB 6.1 Open Source and have a table configured to store 30 days of data with the following compaction strategy: compaction = { ** ‘class’: ‘TimeWindowCompactionStrategy’,** ** ‘compaction_window_size’: ‘3’,** ** ‘compaction_window_unit’: ‘DAYS’** } Currently, no reads are performed on this table—only data inserts. Here’s what I observed: I haven’t manually triggered any compaction apart from observing this autocompaction behavior. Does anyone have insights into why this might have happened? Could it be related to some internal behavior or a specific setting I might have missed? Any help or suggestions would be greatly appreciated! I posted this question a few days ago, but I haven’t received a response yet. I wanted to provide some additional details about the issue I’m facing. I’m using ScyllaDB 6.1 Open Source and have a table configured to store 30 days of data with the following compaction strategy: compaction = { ** ‘class’: ‘TimeWindowCompactionStrategy’,** ** ‘compaction_window_size’: ‘3’,** ** ‘compaction_window_unit’: ‘DAYS’,** ** ‘max_threshold’: ‘32’,** ** ‘min_threshold’: ‘4’** } Upon further investigation, I noticed that instead of forming a single large SSTable for each 3-day window, there are multiple small SSTables within each window. These smaller SSTables are not being compacted into a single SSTable as I expected, even though the compaction strategy specifies min_threshold = 4 and max_threshold = 32. I’d appreciate any insights or suggestions to troubleshoot and resolve this issue. Thank you in advance! TWCS only compacts expired windows. The current/active 3-day window is not compacted by default, your original post is from Jan 8th and you reference files from Jan 7th, so it seems it’s an expected behaviour . if you’re seeing multiple SSTables in the current window, and you’re expecting them to compact: they won’t be compacted yet. The compaction is delayed until the window is considered “cold” (i.e., new writes are no longer hitting it). If your write volume is too low to generate 4 SSTables per window (or it’s spread across the window unevenly), compaction won’t be triggered. --- ### Page: https://forum.scylladb.com/t/is-it-possible-to-retrieve-a-value-after-an-update-without-using-lightweight-transactions/4348 Title: Is it possible to retrieve a value after an update without using Lightweight Transactions? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/is-it-possible-to-retrieve-a-value-after-an-update-without-using-lightweight-transactions/4348 ## Headings Structure: H1: Is it possible to retrieve a value after an update without using Lightweight Transactions? H3: Related topics ## Main Content: H1: Is it possible to retrieve a value after an update without using Lightweight Transactions? H3: Related topics Originally from the User Slack is here any way to do do something like without lightweight transactions and needing to do another operation with “SELECT x FROM val…”? @avi: No, you will need to SELECT after an update @Celal: Is there any wisdom behind it? Or is it just not implemented If there would be an option to retrieve the value right after an option, it would speed up my API x1.5 overall and one specific part even x100 throughput wise since i can then leave out the LWT @avi: We’re following CQL/SQL, and it’s not how the API flows --- ### Page: https://forum.scylladb.com/t/release-scylla-manager-3-4-1/4349 Title: [RELEASE] Scylla Manager 3.4.1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces the release of ScyllaDB Manager 3.4.1, a production-ready patch release of the stable 3.4 branch. ScyllaDB Manager is a centralized cluster administration and recurrent tasks automation tool. … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-manager-3-4-1/4349 ## Headings Structure: H1: [RELEASE] Scylla Manager 3.4.1 H3: Related topics ## Main Content: H1: [RELEASE] Scylla Manager 3.4.1 H3: Related topics The ScyllaDB team announces the release of ScyllaDB Manager 3.4.1, a production-ready patch release of the stable 3.4 branch. ScyllaDB Manager is a centralized cluster administration and recurrent tasks automation tool. ScyllaDB Manager is available for ScyllaDB Enterprise customers and ScyllaDB Open Source users. With ScyllaDB Open Source, ScyllaDB Manager is limited to 5 nodes. See the ScyllaDB Manager Proprietary Software License Agreement for details. The full list of issues fixed by the 3.4.1 release can be found here, but most notably: ScyllaDB Enterprise customers are encouraged to upgrade to ScyllaDB Manager 3.4.1 in coordination with the ScyllaDB support team. The new release includes upgrades of both ScyllaDB Manager Server and Agent. Scylla Manager 3.4.1 supports the following Scylla Enterprise releases: And the following Open Source release (limited to 5 nodes see license 1): You can install and run Scylla Manager on Kubernetes using Scylla Operator. More here. --- ### Page: https://forum.scylladb.com/t/in-clause-composite-partition-key-and-range-quereis/4356 Title: IN clause, composite partition key and range quereis - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/in-clause-composite-partition-key-and-range-quereis/4356 ## Headings Structure: H1: IN clause, composite partition key and range quereis H3: Related topics ## Main Content: H1: IN clause, composite partition key and range quereis H3: Related topics Originally from the User Slack @hamonica: I’m using a composite partition key. If I have two composite partition keys, can I use the IN clause for each key? Or can I only use the IN clause for one key in a composite partition key? PRIMARY KEY ((created_at, service_name), name, end_time_unix_nano) What I want is not this: Instead, I want to retrieve the results with a query like this: SELECT * FROM sample WHERE create_at IN ('c1', 'c2') AND service_name IN ('s1', 's2'); Additionally, is there a limit to the number of values that can be included in the IN clause? @hamonica: ah… um…, is range query only possible on the first partition key? @avi: No, use WHERE token(pk) >= ? AND WHERE token(pk) < ? to delimit the ranges @hamonica: @avi Thank you. I already found that information and am applying it. I’m trying to use it like a time-series database, but it seems like scylladb a lot of thought is needed to apply it to our scenario. Anyway, thank you so much for the response. --- ### Page: https://forum.scylladb.com/t/cant-connect-to-the-node/4357 Title: Can't connect to the node - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: hello, I am new to scylla (and docker so maybe the problem is with my docker command ) I am trying to connect from my rust code to the scylla node that I run from docker my docker command is docker run --name some… Language: en Canonical URL: https://forum.scylladb.com/t/cant-connect-to-the-node/4357 ## Headings Structure: H1: Can't connect to the node H3: Related topics ## Main Content: H1: Can't connect to the node H3: Related topics hello, I am new to scylla (and docker so maybe the problem is with my docker command ) I am trying to connect from my rust code to the scylla node that I run from docker my docker command is docker run --name some-scylla -d scylladb/scylla --smp 1 and I simply have this fn to connect and it giving me always this error I forgot to mention I pass this socket_addr so I seems to fix the connection by running this docker command and changed the socket_addr to localhost but now I am wondering why I couldn’t connect to the private ip I got from the docker vm ? it in my network . it seems I need use the docker port mapping . What operating system do you use? Do you use docker or podman? I am using macOS, and docker This is a known behavior of docker on MacOS and Windows - you can’t connect to internal docker IPs from outside on those systems. so even if it private ip I have to map it to my address to communicate with the container? Alternatively you could also run your application in the docker container. that an option too, thank you for helping me out ! so I seems to fix the connection by running this docker command and changed the socket_addr to localhost but now I am wondering why I couldn’t connect to the private ip I got from the docker vm ? it in my network . it seems I need use the docker port mapping . Private IP likely failed due to Docker’s networking isolation. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-77-2025-01-10/4360 Title: Last week in scylla-cluster-tests.git master (issue #77; 2025-01-10) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the ea1a40b5…63550823 range are covered. There were 32 non-merge commits from 9 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-77-2025-01-10/4360 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #77; 2025-01-10) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #77; 2025-01-10) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the ea1a40b5…63550823 range are covered. There were 32 non-merge commits from 9 authors in that period. Some notable commits: SCT now sends RackAwarePolicy cassandra-stress log events when running with RackAwareRoundRobinPolicy and provided rack name. This information appears in report emails similarly to shard awareness info. For easier Azure VM access, hydra ssh now supports Azure backend. Renovate was introduced for dependency tracking, running weekly on requirements files and stress-tools docker images. Several updates were already made including jinja, Azure identity and Scylla python driver. Mergify was replaced with auto-backport.py script for backport automation, similar to other Scylla repositories. For first-time PR owners, an invite to scylladbbot/scylla-cluster-tests fork will be sent due to GitHub permission limitations. Node-exporter was upgraded to v1.8.2 with removal of unneeded default collectors (matching Scylla’s configuration but without interrupts). See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/how-to-backup-a-scylladb-to-local-storage/4366 Title: How to backup a ScyllaDB to local storage? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Scylla seems perfect from a performance standpoint and it seems easy enough to get it running and configure tables. CONFIGURATION I believe configuration of a 3 node 3 zone system is as follows. Start a Ubuntu EC2 with … Language: en Canonical URL: https://forum.scylladb.com/t/how-to-backup-a-scylladb-to-local-storage/4366 ## Headings Structure: H1: How to backup a ScyllaDB to local storage? H1: CONFIGURATION H1: BACKUP H1: CHATGPT H3: Related topics ## Main Content: H1: How to backup a ScyllaDB to local storage? H1: CONFIGURATION H1: BACKUP H1: CHATGPT H3: Related topics Scylla seems perfect from a performance standpoint and it seems easy enough to get it running and configure tables. I believe configuration of a 3 node 3 zone system is as follows. Start a Ubuntu EC2 with docker installed and then: Then I believe on my next EC2 (in same zone z1), I run: Is this correct so far? If I am connecting 9 nodes (3 zones of 3 nodes each = 9 nodes) do I just keep going like this or what? Let’s say hypothetically I have this system now running. I do not want to back it up to S3. I just want to copy the database to my local disk. How can I do this? I have read the documentation but I see no clear explanation. I need to install and run scylla manager, right? Can I just run that on my local machine? If my local machine has no firewall obstructions to the 9 EC2 nodes, can I just run it locally via docker WSL Ubuntu and connect to them that way? Chat GPT says I can run Then edit Scylla Manager’s configuration to point to your ScyllaDB cluster. “The Scylla Manager configuration file (scylla-manager.conf) typically resides in /etc/scylla-manager”: Then it says I can run things like: But I suspect this is all hallucinations. What is the actual method to copy the database to my local drive if it is running say on 9 EC2’s as described above with a replication factor of 3? Can I do this with a single command? Can I do it with Scylla Manager docker just running locally? Do I need to still take 9 snapshots and then deal with a mess of trying to reconstitute them if there is a problem later? Or can I get a singular backup file of some kind? I saw another post like this but the only reply just said something like “you need to configure MinIO” which is not actionable advice if I have not used MinIO or done any of this before. According to the Scylla Manager Backup documentation, Scylla Manager only supports S3 (or a compatible API) or Google Cloud as backup targets. If you want to backup to a local disk with Scylla Manager, MinIO is your best option, as you already mentioned. MinIO is an S3 compatible object store that you can run locally and thus achieve local backup with Scylla Manager. --- ### Page: https://forum.scylladb.com/t/imds-v2-support-in-scylladb/4367 Title: IMDS (v2) support in ScyllaDB - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/imds-v2-support-in-scylladb/4367 ## Headings Structure: H1: IMDS (v2) support in ScyllaDB H3: Related topics ## Main Content: H1: IMDS (v2) support in ScyllaDB H3: Related topics Originally from the User Slack @Sujay_KS: hey team, we have scylla 4.6.3 running as a docker container, in ec2-server with IMDSV2 enabled, but scylla is unable to access metadata and giving error, but it works if i disable IMDS v2 has anyone faced this ? , is there any fix for this https://stackoverflow.com/questions/71884350/using-imds-v2-with-token-inside-docker-on-ec2-or-ecs I have also configured hop as 2 in ec2-server but still facing the error Stack Overflow: Using IMDS (v2) with token inside docker on EC2 or ECS @Felipe_Cardeneti_Mendes: IMDsv2 support started in 5.2 — latest Scylla is 6.2 you must upgrade @Sujay_KS: cool got it --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-264-2025-01-12/4368 Title: Last week in scylladb.git master (issue #264; 2025-01-12) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the c973254362…2a8ff478f0 range are covered. There were 47 non-merge commits from 17 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-264-2025-01-12/4368 ## Headings Structure: H1: Last week in scylladb.git master (issue #264; 2025-01-12) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #264; 2025-01-12) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the c973254362…2a8ff478f0 range are covered. There were 47 non-merge commits from 17 authors in that period. Some notable commits: Some edge cases related to tablet draining were fixed. The cluster will now reject decommissioning a node if that results in failing to satisfy the specified replication factor. We now create XFS filesystems with reduced metadata overhead. The commitlog hard limit, introduced in ScyllaDB 4.6, is now mandatory. The hard limit prevents the commit log from expanding and instead restricts the write rate. The small-table optimization for repair-based node operations is now enabled by default. This speeds up bootstrap and decommission operations for clusters with small amounts of data. There is a new experimental feature for enabling materialized views with tablets, for testing. When writing to a table that has a materialized view, the coordinator checks the backlog of the participating replicas in order to apply back-pressure to the client. Due to a bug, some replicas were not considered in this calculation. This is now fixed. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-15/4382 Title: [RELEASE] ScyllaDB Enterprise 2024.1.15 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2024.1.15, a bug-fix production-ready patch release for ScyllaDB Enterprise 2024.1 Long Term Support (LTS) Release. Note there is a later Feature (Short term Support) Rele… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-15/4382 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2024.1.15 H3: Related Links H2: Fixed Issue with an open-source reference: H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2024.1.15 H3: Related Links H2: Fixed Issue with an open-source reference: H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2024.1.15, a bug-fix production-ready patch release for ScyllaDB Enterprise 2024.1 Long Term Support (LTS) Release. Note there is a later Feature (Short term Support) Release 2024.2. --- ### Page: https://forum.scylladb.com/t/batch-data-inserts-number-of-rows-and-parallel-processing/4385 Title: Batch data inserts, number of rows and parallel processing - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/batch-data-inserts-number-of-rows-and-parallel-processing/4385 ## Headings Structure: H1: Batch data inserts, number of rows and parallel processing H3: Related topics ## Main Content: H1: Batch data inserts, number of rows and parallel processing H3: Related topics Originally from the User Slack @hamonica: hello~, I’m currently developing in Java and attempting to perform batch data inserts into ScyllaDB. I receive warning messages when attempting to insert approximately 500 to 2,000 rows into ScyllaDB. Upon reviewing the batch-related settings, I found that the recommended values are: batch_size_warn_threshold_in_kb: 128 KB batch_size_fail_threshold_in_kb: 1,024 KB However, for my batch operations, I require significantly larger sizes. Therefore, I have increased these thresholds to: batch_size_warn_threshold_in_kb: 512 KB batch_size_fail_threshold_in_kb: 5,120 KB Given these adjustments, I am concerned about the potential impact on system stability. Could you please advise if increasing these thresholds to the specified values could lead to any stability issues? @hamonica: Is it a better approach to split batch processing into 100 rows and perform parallel processing (insert)? @avi: yes, better to have smaller batches @hamonica: @avi Currently, I’m performing batch inserts with only 500 rows.there don’t seem to be any major issues based on the tests conducted. If problems arise in the future, I will experiment by dividing it into smaller units. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-78-2025-01-17/4386 Title: Last week in scylla-cluster-tests.git master (issue #78; 2025-01-17) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 5fe7ebd2…c22fba4d range are covered. There were 17 non-merge commits from 9 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-78-2025-01-17/4386 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #78; 2025-01-17) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #78; 2025-01-17) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 5fe7ebd2…c22fba4d range are covered. There were 17 non-merge commits from 9 authors in that period. Some notable commits: When sending perf_simple_query benchmark results, numbers will be validated based on all history submitted to Argus instead of only 10. Peer verification is now enabled by default for cassandra-stress, scylla-bench and latte stress tools when client encryption is configured in Scylla. Scylla logs will include start and stop markers for nemesis operations containing nemesis name, target node and completion status for improved debugging. Due to recent builder issues, journalctl logs are now collected and linked in Argus. Gemini stress tool docker image moved to dedicated project with added CQL statement logging. Scylla Enterprise features test pipelines were reorganized under jenkins-pipelines/oss/features directory. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/testing-scylladb-performance-issue-with-kubernetes-clusters-using-containers/4390 Title: Testing ScyllaDB performance issue with Kubernetes clusters, using containers - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/testing-scylladb-performance-issue-with-kubernetes-clusters-using-containers/4390 ## Headings Structure: H1: Testing ScyllaDB performance issue with Kubernetes clusters, using containers H1: @Serhii_DevOps: Hello everyone, I’m not sure if it’s worth opening an issue, so I wanted to ask here first. I deployed Scylla in Kubernetes clusters, and it’s performing very poorly compared to the local database. I wanted to ask if this is normal, but looking at the numbers, it seems like it’s not. Can anyone suggest what might be the problem or if I should go ahead and create an issue on GitHub? H3: Related topics ## Main Content: H1: Testing ScyllaDB performance issue with Kubernetes clusters, using containers H1: @Serhii_DevOps: Hello everyone, I’m not sure if it’s worth opening an issue, so I wanted to ask here first. I deployed Scylla in Kubernetes clusters, and it’s performing very poorly compared to the local database. I wanted to ask if this is normal, but looking at the numbers, it seems like it’s not. Can anyone suggest what might be the problem or if I should go ahead and create an issue on GitHub? H3: Related topics Originally from the User Slack Summary of Scylla Testing in Kubernetes Server Specifications: • Processor: Intel Core i9 14900K (8 cores x 3.2GHz up to 6.0GHz & 16 cores x 2.4GHz up to 4.4GHz) • Memory: 128GB RAM DDR5 • Storage: 2x 2TB NVMe • Software RAID: RAID 0 • File System: ext4 Kubernetes Infrastructure: • System: Vanilla Kubernetes, Istio Service Mesh, Flannel, OpenEBS Storage Provider • Tested Modes: Local tests, Kubernetes tests, Kubernetes tests with local databases. • Testing showed that the results on the DigitalOcean Kubernetes cluster align with those on bare metal Kubernetes. Test Description: A Go application was run to clear and populate the database by processing data from another source. Test Results: • Test 1: ◦ Insert Service: Minimum resources ◦ Scylla: 5 racks in two different data centers, each rack with 2 CPUs and 18GB RAM, 200GB storage ◦ Execution Time: 6 minutes 46 seconds • Test 2: ◦ Insert Service: Minimum resources ◦ Scylla: 5 racks in two different data centers, each rack with 6 CPUs and 18GB RAM, 200GB storage ◦ Execution Time: 6 minutes 41 seconds • Test 3: ◦ Insert Service: 4 CPUs, 4GB RAM ◦ Scylla: 5 racks in two different data centers, each rack with 6 CPUs and 18GB RAM, 200GB storage ◦ Execution Time: 5 minutes 45 seconds • Test 4: ◦ Insert Service: 4 CPUs, 4GB RAM ◦ Scylla: 1 rack in a single data center, 6 CPUs and 18GB RAM, 200GB storage ◦ Execution Time: 6 minutes 20 seconds • Test 5 (Local Installation): ◦ Insert Service: All services in Docker on a server with the same specifications (Single Scylla) ◦ Execution Time: 2.4 seconds • Test 6: ◦ Insert Service: 4 CPUs, 4GB RAM ◦ Scylla: 1 rack in a single data center, 6 CPUs and 18GB RAM, 200GB storage ◦ Istio Service Mesh disabled ◦ Execution Time: 6 minutes 48 seconds • Test 7: ◦ Insert Service in Kubernetes: 4 CPUs, 4GB RAM ◦ Scylla: One Scylla instance deployed in Docker on a server ◦ Execution Time: 10.5 seconds Summary and Comparison: • The local Docker installation (Test 5) produced the best result, taking only 2.4 seconds, significantly faster than all other tests. • Kubernetes tests with Istio (Tests 1-4, 6) had longer execution times, ranging from 5 to 7 minutes. • Disabling Istio Service Mesh (Test 6) did not significantly affect the result, as the execution time remained close to previous tests. • The last test (Test 7), where Scylla was running in Docker on the server, showed a time of 10.5 seconds, which was an order of magnitude faster than Kubernetes. Percentage Comparison: • Local Docker Installation vs Kubernetes: 2.4 seconds vs 6 minutes 46 seconds = 99.4% faster • Kubernetes without Istio vs with Istio: 6 minutes 48 seconds vs 6 minutes 41 seconds = 1.7% difference • Docker (Single Scylla) vs Kubernetes (5 racks): 10.5 seconds vs 6 minutes 46 seconds = 97.4% faster @dor: ScyllaDB expect not to over commit the disk and the cpu. Ideally, no container/process would run on the cpus that scylla will be scheduled on. Otherwise, huge latency will occur. So you should pin your containers. You can avoid that by providing a --overcommit flag to the scylla executable In a similar way, Scylla expects to know the disk capacity, that’s why we have iotune script. The result is io.conf, the max capacity of a scylla nodde. If you don’t use it, or used data which doesn’t match the hardware, huge latencies will occur. So this setup is almost the worst case scenario for scylla. Best would be to just run a single test and optimize it with the above suggestions, share the results and continue from there lstio also makes it worse… @Serhii_DevOps: Thank you very much for your response. Everything you mentioned sounds quite reasonable. However, I’m a bit unclear about your comments regarding Istio. Based on the tests, the results are as follows: Kubernetes without Istio vs. with Istio: 6 minutes 48 seconds vs. 6 minutes 41 seconds = 1.7% difference. Disabling the Istio Service Mesh (Test 6) did not significantly affect the results, as the execution time remained close to previous tests. So, did you mean that it causes some slight delays, or are you suggesting not to use it at all? Additionally, I’d like to ask for your opinion as a Scylla expert. In our case, where we need multiple clusters across different regions, would you recommend deploying Scylla on dedicated servers, or should we continue trying to optimize it in Kubernetes to achieve comparable performance? @dor: I misread the lstio results, I agree that 1.7% is a neglectable difference. It will effect latency more but let’s ignore it for the time being You can run Scylla in k8s and we have an operator project too. As long as you follow the tunning it’s fine. Ideally, you either run a single container per host or you run multiple containers, just make sure they are pinned and configured accordingly. @Serhii_DevOps: Hello @dor, thank you very much for your help. Unfortunately, the Kubernetes cluster configuration didn’t help much. So, I set up a Scylla node on a separate server. Processor: Single Intel E-2276G (6c x 3.8GHz) Memory: 64GB RAM DDR4 Storage: 1x 1TB NVMe I split the disk into two 300GB partitions for Scylla. The Ansible role successfully reformatted them for Scylla’s data. Unfortunately, as a result, we only gained a 2-minute faster processing time. So, if previously it took over 6 minutes, now it takes over 4 minutes. Tomorrow, I will continue testing and further investigating this issue. But if you have any suggestions, I would appreciate it. I just don’t think the latency between servers could increase the processing time this much. @dor: Did you run scylla setup? Why did you split the disk? @Serhii_DevOps: I’m installing Scylla with ansible role Because this server has 1tb nvme disk like 1 device But for Ansible role require to have two devices which formatted to xfs and merged to one md0 device hey @dor After numerous tests and experiments, I can conclude that our current issue does not with Scylla. It is configured and working perfectly. Thank you very much for your help! @dor: Thanks for the report. It’s hard to figure out what is the problem. Yes, k8s can be complex but I don’t know whether this is a high op/s use case, real time or batch and why it didn’t work @Serhii_DevOps: Yes, I completely agree, especially since the issue itself is quite strange. What takes 5 seconds locally took over 6 minutes in multi-region clusters. There were many potential culprits, ranging from latency and Scylla (that’s our first experience with scylla) to Redis and Citus. After numerous test cases, I managed to achieve a 5-second response time with Scylla across multiple regions. This suggests that the issue is not so much with latency or Scylla but rather with Redis or Citus. So now I’ll focus on those. Once again, thank you so much! Also, moving Scylla from Kubernetes to a standard installation will be beneficial in the future. @dor: Got it, thanks for the answer. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-265-2025-01-20/4391 Title: Last week in scylladb.git master (issue #265; 2025-01-20) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 2a8ff478f0…1ef2d9d076 range are covered. There were 176 non-merge commits from 25 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-265-2025-01-20/4391 ## Headings Structure: H1: Last week in scylladb.git master (issue #265; 2025-01-20) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #265; 2025-01-20) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 2a8ff478f0…1ef2d9d076 range are covered. There were 176 non-merge commits from 25 authors in that period. Some notable commits: ScyllaDB can now use Lightweight Directory Access Protocol (LDAP) for authentication and authorization. This feature was ported from ScyllaDB Enterprise. ScyllaDB can now encrypt sstables and the commitlog on disk (encryption-at-rest) using a variety of methods to obtain encryption keys. This feature was ported from ScyllaDB Enterprise. ScyllaDB can now audit CQL operations. Audit can be configures by statement category and keyspace/table. Audit logs can be stored in a table or forwarded via rsyslog. This feature was ported from ScyllaDB Enterprise. A disk space monitor was added, to aid the load balancer in its decision. A possible data corruption during read repair of different partitions hashing to the same token was fixed. Additional bandwidth limit configuration on compaction and streaming were implemented. The next version was renamed to 2025.1, indicating the merging of the open-source and enterprise release streams into a single source-available release stream. Tablet split and merge operations are now reflected via the task manager. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/i-want-to-know-best-practices-for-scaling-scylladb-in-a-high-traffic-environment/4394 Title: I want to know Best Practices for Scaling ScyllaDB in a High-Traffic Environment - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hey everyone, I am new to ScyllaDB and have been exploring its potential for handling high-traffic workloads. I am impressed with it is performance but I wanted to tap into the collective wisdom of this community for … Language: en Canonical URL: https://forum.scylladb.com/t/i-want-to-know-best-practices-for-scaling-scylladb-in-a-high-traffic-environment/4394 ## Headings Structure: H1: I want to know Best Practices for Scaling ScyllaDB in a High-Traffic Environment H3: Related topics ## Main Content: H1: I want to know Best Practices for Scaling ScyllaDB in a High-Traffic Environment H3: Related topics I am new to ScyllaDB and have been exploring its potential for handling high-traffic workloads. I am impressed with it is performance but I wanted to tap into the collective wisdom of this community for some practical advice. I am working on an application that is expected to scale rapidly, with a mix of heavy reads and writes. While I have read the documentation and a few case studies, I want to hear real-world experiences or tips on Optimal cluster configurations for scaling efficiently. Any gotchas or common mistakes to avoid when setting up for high traffic. Monitoring tools or practices you have found essential for maintaining performance. Tips for designing a data model that can handle growth without causing bottlenecks. Also i want to know if anyone here has experience running ScyllaDB on aws and could share insights into balancing cost vs. performance as the cluster grows. --- ### Page: https://forum.scylladb.com/t/why-do-tombstones-exist-in-the-scylla-cdc-log-table/4395 Title: Why do tombstones exist in the scylla_cdc_log table? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 5.4.10 #Cluster size: 6 os (RHEL/CentOS/Ubuntu/AWS AMI): CentOS Hi ! After I turned on cdc for the table, I found tombstones when scanning the cdc_log table. The log is as fo… Language: en Canonical URL: https://forum.scylladb.com/t/why-do-tombstones-exist-in-the-scylla-cdc-log-table/4395 ## Headings Structure: H1: Why do tombstones exist in the scylla_cdc_log table? H3: Related topics ## Main Content: H1: Why do tombstones exist in the scylla_cdc_log table? H3: Related topics Installation details #ScyllaDB version: 5.4.10 #Cluster size: 6 os (RHEL/CentOS/Ubuntu/AWS AMI): CentOS Hi ! After I turned on cdc for the table, I found tombstones when scanning the cdc_log table. The log is as follows: Isn’t the cdc_log table only for writes? Why do tombstones exist? Is it related to TTL? Yes it is related to TTL. Expired cells are identical to dead cells (tombstones) in the internal representation and are reported as such when calculating this tombstone statistic during read. --- ### Page: https://forum.scylladb.com/t/error-when-running-scylladb-with-docker-in-macos/4397 Title: Error when running ScyllaDB with Docker in macOS - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/error-when-running-scylladb-with-docker-in-macos/4397 ## Headings Structure: H1: Error when running ScyllaDB with Docker in macOS H3: Related topics ## Main Content: H1: Error when running ScyllaDB with Docker in macOS H3: Related topics Originally from the User Slack @Pritam_Nagar: Dear @avi and team, I am currently running ScyllaDB on my local macOS system within Docker. However, I encountered an error during startup: @avi: Try increasing /proc/sys/fs/aio-max-nr or reducing the number of shards you run scylla with @Pritam_Nagar: syclla container or host machine ? @avi: The VM that runs the container @Pritam_Nagar: my docker setup is on mac i don’t think so it can be configured @Josh_Mengerink: I encountered this error in kubernetes context, reducing the number of cpus for the deployment (and by extension IO requirements) solved the issue for me. Perhaps that points you in a direction @Patrick_Bossman: add: --reactor-backend=epoll https://forum.scylladb.com/t/running-3-node-scylladb-in-docker/1057/7 @Pritam_Nagar ^^ @Pritam_Nagar: Thanks @Patrick_Bossman --- ### Page: https://forum.scylladb.com/t/restore-data-initialize-versioned-sstables-not-a-scylla-manager-snapshot/4404 Title: Restore data: initialize versioned SSTables: not a Scylla Manager snapshot - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: *Installation details #ScyllaDB version: 5.0.13 #Cluster size: 3 nodes os (RHEL/CentOS/Ubuntu/AWS AMI): Ubuntu restore data: initialize versioned SSTables: not a Scylla Manager snapshot despite the snapshot tag is … Language: en Canonical URL: https://forum.scylladb.com/t/restore-data-initialize-versioned-sstables-not-a-scylla-manager-snapshot/4404 ## Headings Structure: H1: Restore data: initialize versioned SSTables: not a Scylla Manager snapshot H3: Related topics ## Main Content: H1: Restore data: initialize versioned SSTables: not a Scylla Manager snapshot H3: Related topics *Installation details #ScyllaDB version: 5.0.13 #Cluster size: 3 nodes os (RHEL/CentOS/Ubuntu/AWS AMI): Ubuntu restore data: initialize versioned SSTables: not a Scylla Manager snapshot despite the snapshot tag is exact same. What am I missing? How to resolve this? Please share more details, the exact steps you took, the full error message, and any other details you have. Hello Guy, Thank you for your reply and help - Here is the summary what we did - Run: 818902-de0b-11ef-91h8-02d6b68b54cd Status: ERROR Cause: restore data: initialize versioned SSTables: not a Scylla Manager snapshot tag, expected format is sm_20060102150405UTC Start time: 24 Jan 25 04:27:21 UTC End time: 24 Jan 25 04:27:22 UTC Duration: 0s Progress: 0% | 0% Snapshot Tag: sm_20250123110656UTC From the information pasted above it looks like sm_20250123110656UTC has a correct snapshot tag format. Please provide Scylla Manager logs from the time of running the restore. Hello @Michal_Trzonkowski, Thank you for your reply, Below are the logs - Jan 23 10:37:18 scylla-manager[2100920]: {“L”:“INFO”,“T”:“2025-01-23T10:37:18.052Z”,“N”:“scheduler”,“M”:“PutTask”,“task”:“restore/25132a36-9f31-40c9-b7e5-fa5d020ff0fe”,“schedule”:{“cron”:“”,“window”:null,“timezone”:“Etc/UTC”,“start_date”:“0001-01-01T00:00:00Z”,“interval”:“”,“num_retries”:3,“retry_wait”:“10m”},“properties”:{“keyspace”:[“myredactedkeyspace”],“location”:[“s3:my-reducted-s3-bucket”],“restore_tables”:true,“snapshot_tag”:“sm_20250121065957UTC”},“create”:true,“_trace_id”:“jwRC1_rlT4OgDEFa8h4scg”} Jan 23 10:37:18 scylla-manager[2100920]: {“L”:“INFO”,“T”:“2025-01-23T10:37:18.055Z”,“N”:“scheduler.136499a1”,“M”:“Schedule”,“task”:“restore/25132a36-9f31-40c9-b7e5-fa5d020ff0fe”,“in”:“0s”,“begin”:“2025-01-23T10:37:18.055Z”,“retry”:0,“_trace_id”:“jwRC1_rlT4OgDEFa8h4scg”} Jan 23 10:37:18 scylla-manager[2100920]: {“L”:“INFO”,“T”:“2025-01-23T10:37:18.056Z”,“N”:“scheduler.136499a1”,“M”:“Run started”,“task”:“restore/25132a36-9f31-40c9-b7e5-fa5d020ff0fe”,“retry”:0,“_trace_id”:“93PMHTMNRLClxnK6fQG5xA”} Jan 23 10:37:18 scylla-manager[2100920]: {“L”:“INFO”,“T”:“2025-01-23T10:37:18.265Z”,“N”:“backup”,“M”:“Restore”,“cluster_id”:“136499a1-313a-4d2c-b090-e9945efe8ad8”,“task_id”:“25132a36-9f31-40c9-b7e5-fa5d020ff0fe”,“run_id”:“04bbd675-d976-11ef-804e-0237f91a5b9b”,“target”:{“location”:[“s3:my-reducted-s3-bucket”],“keyspace”:[“myredactedkeyspace”,“!system”,“!system_schema”,“!system_distributed_everywhere.cdc_generation_descriptions_v2”,“!system_distributed.cdc_streams_descriptions_v2”,“!system_distributed.cdc_generation_timestamps”,“!._scylla_cdc_log”],“snapshot_tag”:“sm_20250121065957UTC”,“batch_size”:2,“restore_tables”:true,“continue”:true},“_trace_id”:“93PMHTMNRLClxnK6fQG5xA”} Jan 23 10:37:18 scylla-manager[2100920]: {“L”:“INFO”,“T”:“2025-01-23T10:37:18.434Z”,“N”:“backup.restore”,“M”:“Created restore units”,“units”:[{“Keyspace”:“myredactedkeyspace”,“Size”:711636847007,“Tables”:[{“Table”:“myreductedtable”,“Size”:388144379774},{“Table”:“myreductedtable”,“Size”:323492467233}]}],“_trace_id”:“93PMHTMNRLClxnK6fQG5xA”} Jan 23 10:37:18 scylla-manager[2100920]: {“L”:“INFO”,“T”:“2025-01-23T10:37:18.436Z”,“N”:“backup.restore”,“M”:“Awaiting schema agreement…”,“_trace_id”:“93PMHTMNRLClxnK6fQG5xA”} Jan 23 10:37:18 scylla-manager[2100920]: {“L”:“INFO”,“T”:“2025-01-23T10:37:18.443Z”,“N”:“backup.restore”,“M”:“Done awaiting schema agreement”,“duration”:“6.811967ms”,“_trace_id”:“93PMHTMNRLClxnK6fQG5xA”} Jan 23 10:37:18 scylla-manager[2100920]: {“L”:“INFO”,“T”:“2025-01-23T10:37:18.542Z”,“N”:“backup.restore”,“M”:“Started restoring tables”,“_trace_id”:“93PMHTMNRLClxnK6fQG5xA”} Jan 23 10:37:18 scylla-manager[2100920]: {“L”:“INFO”,“T”:“2025-01-23T10:37:18.542Z”,“N”:“backup.restore”,“M”:“Disabling table’s gc_grace_seconds”,“keyspace”:“myredactedkeyspace”,“table”:“myreductedtable”,“_trace_id”:“93PMHTMNRLClxnK6fQG5xA”} Jan 23 10:37:19 scylla-manager[2100920]: {“L”:“INFO”,“T”:“2025-01-23T10:37:19.239Z”,“N”:“backup.restore”,“M”:“Disabling table’s gc_grace_seconds”,“keyspace”:“myredactedkeyspace”,“table”:“myreductedtable”,“_trace_id”:“93PMHTMNRLClxnK6fQG5xA”} Jan 23 10:37:20 scylla-manager[2100920]: {“L”:“INFO”,“T”:“2025-01-23T10:37:20.128Z”,“N”:“backup.restore”,“M”:“Awaiting schema agreement…”,“_trace_id”:“93PMHTMNRLClxnK6fQG5xA”} Jan 23 10:37:20 scylla-manager[2100920]: {“L”:“INFO”,“T”:“2025-01-23T10:37:20.136Z”,“N”:“backup.restore”,“M”:“Done awaiting schema agreement”,“duration”:“7.631435ms”,“_trace_id”:“93PMHTMNRLClxnK6fQG5xA”} Jan 23 10:37:20 scylla-manager[2100920]: {“L”:“INFO”,“T”:“2025-01-23T10:37:20.136Z”,“N”:“backup.restore”,“M”:“Restoring location”,“location”:“s3:my-reducted-s3-bucket”,“_trace_id”:“93PMHTMNRLClxnK6fQG5xA”} Jan 23 10:37:20 scylla-manager[2100920]: {“L”:“INFO”,“T”:“2025-01-23T10:37:20.394Z”,“N”:“backup.restore”,“M”:“Initialized restore hosts”,“hosts”:[{“Host”:“172.31.20.203”,“OngoingRunProgress”:null},{“Host”:“172.31.20.72”,“OngoingRunProgress”:null},{“Host”:“172.31.27.215”,“OngoingRunProgress”:null}],“_trace_id”:“93PMHTMNRLClxnK6fQG5xA”} Jan 23 10:37:20 scylla-manager[2100920]: {“L”:“INFO”,“T”:“2025-01-23T10:37:20.444Z”,“N”:“backup.restore”,“M”:“Restoring manifest”,“manifest”:{“Location”:“s3:my-reducted-s3-bucket”,“DC”:“datacenter1”,“ClusterID”:“46s32eg7-9013-45b7-a59b-f071cot3jpae”,“NodeID”:“80ba459f-2500-4f96-a8b0-69ea500b0848”,“TaskID”:“7e1d88ad-0dc8-455c-ab94-42fd3e6619ae”,“SnapshotTag”:“sm_20250121065957UTC”,“Temporary”:false},“_trace_id”:“93PMHTMNRLClxnK6fQG5xA”} Jan 23 10:37:20 scylla-manager[2100920]: {“L”:“INFO”,“T”:“2025-01-23T10:37:20.445Z”,“N”:“backup.restore”,“M”:“Restoring table”,“keyspace”:“myredactedkeyspace”,“table”:“myreductedtable”,“_trace_id”:“93PMHTMNRLClxnK6fQG5xA”} Jan 23 10:37:20 scylla-manager[2100920]: {“L”:“INFO”,“T”:“2025-01-23T10:37:20.447Z”,“N”:“backup.restore”,“M”:“Initialized SSTable bundle pool”,“sstable_ids”:[“3631”,“4413”,“4478”,“719”,“4452”,“1017”,“4030”,“4171”,“4356”,“4249”,“4387”,“4417”,“4355”,“4430”,“4486”,“652”,“2555”,“3173”,“4383”,“4456”,“673”,“718”,“4373”,“4424”,“4466”,“4482”,“4224”,“4295”,“4401”,“4421”],“_trace_id”:“93PMHTMNRLClxnK6fQG5xA”} Jan 23 10:37:20 scylla-manager[2100920]: {“L”:“INFO”,“T”:“2025-01-23T10:37:20.450Z”,“N”:“backup.restore”,“M”:“Received table’s version”,“keyspace”:“myredactedkeyspace”,“table”:“myreductedtable”,“version”:“efaebc80d97211efb40d5d352ab1f974”,“_trace_id”:“93PMHTMNRLClxnK6fQG5xA”} Jan 23 10:37:20 scylla-manager[2100920]: {“L”:“INFO”,“T”:“2025-01-23T10:37:20.450Z”,“N”:“backup.restore”,“M”:“Found table’s source and destination directory”,“keyspace”:“myredactedkeyspace”,“table”:“myreductedtable”,“src_dir”:“s3:my-reducted-s3-bucket/backup/sst/cluster/46s32eg7-9013-45b7-a59b-f071cot3jpae/dc/datacenter1/node/80ba459f-2500-4f96-a8b0-69ea500b0848/keyspace/myredactedkeyspace/table/myreductedtable/1a600df0fa2511ed9a140dd3a1cdb004”,“dst_dir”:“data:myredactedkeyspace/myreductedtable-efaebc80d97211efb40d5d352ab1f974/upload”,“_trace_id”:“93PMHTMNRLClxnK6fQG5xA”} Jan 23 10:37:20 scylla-manager[2100920]: {“L”:“INFO”,“T”:“2025-01-23T10:37:20.500Z”,“N”:“backup.restore”,“M”:“Restoring table finished”,“keyspace”:“myredactedkeyspace”,“table”:“myreductedtable”,“_trace_id”:“93PMHTMNRLClxnK6fQG5xA”} Jan 23 10:37:20 scylla-manager[2100920]: {“L”:“INFO”,“T”:“2025-01-23T10:37:20.500Z”,“N”:“backup.restore”,“M”:“Restoring manifest finished”,“manifest”:{“Location”:“s3:my-reducted-s3-bucket”,“DC”:“datacenter1”,“ClusterID”:“46s32eg7-9013-45b7-a59b-f071cot3jpae”,“NodeID”:“80ba459f-2500-4f96-a8b0-69ea500b0848”,“TaskID”:“7e1d88ad-0dc8-455c-ab94-42fd3e6619ae”,“SnapshotTag”:“sm_20250121065957UTC”,“Temporary”:false},“_trace_id”:“93PMHTMNRLClxnK6fQG5xA”} Jan 23 10:37:20 scylla-manager[2100920]: {“L”:“INFO”,“T”:“2025-01-23T10:37:20.500Z”,“N”:“backup.restore”,“M”:“Restoring location finished”,“location”:“s3:my-reducted-s3-bucket”,“_trace_id”:“93PMHTMNRLClxnK6fQG5xA”} Jan 23 10:37:20 scylla-manager[2100920]: {“L”:“INFO”,“T”:“2025-01-23T10:37:20.500Z”,“N”:“backup.restore”,“M”:“Restoring tables finished”,“_trace_id”:“93PMHTMNRLClxnK6fQG5xA”} Jan 23 10:37:20 scylla-manager[2100920]: {“L”:“INFO”,“T”:“2025-01-23T10:37:20.510Z”,“N”:“scheduler.136499a1”,“M”:“Run ended with ERROR”,“task”:“restore/25132a36-9f31-40c9-b7e5-fa5d020ff0fe”,“status”:“ERROR”,“cause”:“restore data: not restored bundles [3631 4413 4478 719 4452 1017 4030 4171 4356 4249 4387 4417 4355 4430 4486 652 2555 3173 4383 4456 673 718 4373 4424 4466 4482 4224 4295 4401 4421]: initialize versioned SSTables: not a Scylla Manager snapshot tag, expected format is sm_20060102150405UTC”,“duration”:“2.444916447s”,“_trace_id”:“93PMHTMNRLClxnK6fQG5xA”} Jan 23 10:37:20 scylla-manager[2100920]: {“L”:“INFO”,“T”:“2025-01-23T10:37:20.510Z”,“N”:“scheduler.136499a1”,“M”:“Retry backoff”,“task”:“restore/25132a36-9f31-40c9-b7e5-fa5d020ff0fe”,“backoff”:“10m0s”,“retry”:1} Jan 23 10:37:20 scylla-manager[2100920]: {“L”:“INFO”,“T”:“2025-01-23T10:37:20.510Z”,“N”:“scheduler.136499a1”,“M”:“Schedule”,“task”:“restore/25132a36-9f31-40c9-b7e5-fa5d020ff0fe”,“in”:“9m59s”,“begin”:“2025-01-23T10:47:20.510Z”,“retry”:1} Jan 23 10:37:34 scylla-manager[2100920]: {“L”:“INFO”,“T”:“2025-01-23T10:37:34.394Z”,“N”:“http”,“M”:“GET /api/v1/cluster/test_indhan/task/restore/25132a36-9f31-40c9-b7e5-fa5d020ff0fe”,“from”:“127.0.0.1:45462”,“status”:200,“bytes”:553,“duration”:“26ms”,“_trace_id”:“fqBU2r3oR0isqU1VSDWoXw”} Jan 23 10:37:34 scylla-manager[2100920]: {“L”:“INFO”,“T”:“2025-01-23T10:37:34.442Z”,“N”:“http”,“M”:“GET /api/v1/cluster/test_indhan/task/restore/25132a36-9f31-40c9-b7e5-fa5d020ff0fe/latest”,“from”:“127.0.0.1:45462”,“status”:200,“bytes”:1164,“duration”:“47ms”,“_trace_id”:“Gj8_DhBhQyevo8oukRNz0Q”} Jan 23 10:38:18 scylla-manager[2100920]: {“L”:“INFO”,“T”:“2025-01-23T10:38:18.115Z”,“N”:“scheduler”,“M”:“Task deleted”,“cluster_id”:“136499a1-313a-4d2c-b090-e9945efe8ad8”,“task_type”:“restore”,“task_id”:“25132a36-9f31-40c9-b7e5-fa5d020ff0fe”,“_trace_id”:“kymEDapqRIWArT5iijM0sw”} Let’s see if @Michal_Leszczynski can help Hi and sorry for late response, Could you verify that scylla-manager and scylla-manager-agents have the same versions? Forgetting to update scylla-manager-agents when updating scylla-manager might have caused this issue. If the versions are equal, please also write them in the comment. In such case, we would need to verify the file names in the s3:my-reducted-s3-bucket/backup/sst/cluster/46s32eg7-9013-45b7-a59b-f071cot3jpae/dc/datacenter1/node/80ba459f-2500-4f96-a8b0-69ea500b0848/keyspace/myredactedkeyspace/table/myreductedtable/1a600df0fa2511ed9a140dd3a1cdb004 directory. Please also send them (just the file names) in the comment. The version of scylla manager version is 3.1.2-0.20230704.bd349aa4 and the scylla manager agent version is 2.6.5-0.20221128.1ced9336 Hi, I will update the agent version and try restoring. Will keep you updated. Thank you. --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-2-3/4405 Title: [RELEASE] ScyllaDB Enterprise 2024.2.3 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2024.2.3, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2024.2 Feature (Short Term Support) Release. 2024.2.3 patch release includes… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-2-3/4405 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2024.2.3 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2024.2.3 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2024.2.3, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2024.2 Feature (Short Term Support) Release. 2024.2.3 patch release includes two CQL extensions and multiple bug fixes. The following issues are fixed in this release (with an open-source reference, if available): Allow selecting map values and set elements, compatible to Cassandra 4.0 SELECT map['key'] FROM table SELECT map['key1']['key2'] FROM table cql3: implement NOT IN #21992 select * from TBL where v NOT IN (5,7) ALLOW FILTERING Until this release, the materialized view flow-control algorithm used a constant delay_limit_us hard-coded to one second, which means that when the size of view-update backlog reached the maximum (10% of memory), we delay every request by an additional second - while smaller amounts of backlog will result in smaller delays. So this patch replaces the hard-coded default with a live-updateable configuration parameter, view_flow_control_delay_limit_in_ms, which defaults to 1000ms as before. #18187 --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-79-2025-01-24/4407 Title: Last week in scylla-cluster-tests.git master (issue #79; 2025-01-24) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 0c7fa60a…f64203d2 range are covered. There were 20 non-merge commits from 9 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-79-2025-01-24/4407 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #79; 2025-01-24) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #79; 2025-01-24) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 0c7fa60a…f64203d2 range are covered. There were 20 non-merge commits from 9 authors in that period. Some notable commits: Test case configuration for 5000 tables was updated to make it more robust, reduce duration and cost. Jenkins jobs will now automatically add Argus link to job description. New documentation page explains SCT Docker backend specifics and limitations. Added new parameter enable_views_with_tablets_on_upgrade to enable experimental views-with-tablets feature for Scylla 2025.1 and upgrade testing. Log collection switched to zstd compression instead of gzip. Since counters are not supported with tablets, stress commands with counters will now error out and continue operation when tablets are used. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/multi-data-center-cluster-setup-latencies-and-data-replication/4415 Title: Multi data-center cluster setup, latencies and data replication - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/multi-data-center-cluster-setup-latencies-and-data-replication/4415 ## Headings Structure: H1: Multi data-center cluster setup, latencies and data replication H3: Related topics ## Main Content: H1: Multi data-center cluster setup, latencies and data replication H3: Related topics Originally from the User Slack @Chaitanya_Tondlekar: Wanted to understand what would be pros and cons if we setup multi DC cluster setup. We will be changing cassandrarackdc.properties file but all nodes belong to same region. In application connection string, IPs of DC 1 will be provided. In case of issue, application will start pointing to DC2. What are the main things to consider while setting this kind of setup ? How latencies can be managed properly while it replicates data from DC1 to DC2 ? @Guy @avi If any doc is available for this , please do let us know. @dor: You can access every DC locally in a local quorum. So there is not latency problem It works well if you need > 1 DCs @Chaitanya_Tondlekar: writes are also showing latencies more than 1 second. @dor What should be consistency level for writes ? --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-266-2025-01-26/4417 Title: Last week in scylladb.git master (issue #266; 2025-01-26) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 1ef2d9d076…f4b1ad43d4 range are covered. There were 34 non-merge commits from 11 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-266-2025-01-26/4417 ## Headings Structure: H1: Last week in scylladb.git master (issue #266; 2025-01-26) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #266; 2025-01-26) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 1ef2d9d076…f4b1ad43d4 range are covered. There were 34 non-merge commits from 11 authors in that period. Some notable commits: Tablets will now use file streaming rather than mutation streaming. File streaming takes advantage of the fact that all of a tablet’s mutations are segregated in their own log-structured merge tree to copy and then erase those files, resulting in higher bandwidth and lower CPU utilization. The repair time for a tablet is now recorded in the system.tablets table. This is more reliable than storing repair time in local tables. New COMPACT STORAGE tables can no longer be created. They have been deprecated for a long while. A bug where materialized views lost track of the base table schema during reverse queries was fixed. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-2-4/4418 Title: [RELEASE] ScyllaDB Enterprise 2024.2.4 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2024.2.4, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2024.2 Feature (Short Term Support) Release. 2024.2.4 patch release includes… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-2-4/4418 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2024.2.4 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2024.2.4 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2024.2.4, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2024.2 Feature (Short Term Support) Release. 2024.2.4 patch release includes two CQL extensions and multiple bug fixes. The following issues are fixed in this release (with an open-source reference, if available): --- ### Page: https://forum.scylladb.com/t/how-long-does-it-take-to-add-a-new-node-to-an-existing-cluster/4419 Title: How long does it take to add a new node to an existing cluster? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/how-long-does-it-take-to-add-a-new-node-to-an-existing-cluster/4419 ## Headings Structure: H1: How long does it take to add a new node to an existing cluster? H3: Related topics ## Main Content: H1: How long does it take to add a new node to an existing cluster? H3: Related topics Originally from the User Slack @Hyunwoo_Kim: Hi, I’m planning to add a new node to a cluster with the following setup: • 5 nodes, each holding 500GB of data • RF is 3 Is there any benchmark or reference data on how long it typically takes to add a new node under these conditions? @dor: In the old architecture, it depends on the schema, can take 1-2 hours but it varies. The new tablet code, in the source available version - 2024.2 will do it in 15 minutes @Hyunwoo_Kim: Oh, I got it. I appreciate --- ### Page: https://forum.scylladb.com/t/release-added-support-for-aws-i7ie-instance-family-28-january-2025/4421 Title: [RELEASE] Added support for AWS i7ie instance family - 28 January 2025 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: We are happy to announce a new update to the ScyllaDB Cloud service. New features and enhancements: We’ve added support for the AWS i7ie instance family. This is to support intermediate workloads, offering a cost-effe… Language: en Canonical URL: https://forum.scylladb.com/t/release-added-support-for-aws-i7ie-instance-family-28-january-2025/4421 ## Headings Structure: H1: [RELEASE] Added support for AWS i7ie instance family - 28 January 2025 H3: Related topics ## Main Content: H1: [RELEASE] Added support for AWS i7ie instance family - 28 January 2025 H3: Related topics We are happy to announce a new update to the ScyllaDB Cloud service. New features and enhancements: We’ve added support for the AWS i7ie instance family. This is to support intermediate workloads, offering a cost-effective midpoint between compute and storage performance. The following regions are currently supported: Is the support limited to only cloud service ? What’s the ETA for the enterprise and opensource ? This is supported in Enterprise in 2024.2.4 - see here. --- ### Page: https://forum.scylladb.com/t/what-are-the-options-for-replacing-disks-on-nodes-of-a-running-cluster/4422 Title: What are the options for replacing disks on nodes of a running cluster? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/what-are-the-options-for-replacing-disks-on-nodes-of-a-running-cluster/4422 ## Headings Structure: H1: What are the options for replacing disks on nodes of a running cluster? H3: Related topics ## Main Content: H1: What are the options for replacing disks on nodes of a running cluster? H3: Related topics Originally from the User Slack @Hyunwoo_Kim: -------- one more question When there is 6-node cluster, and I need to replace the disks on 3 of the nodes. There seem to be two possible approaches. @avi: Decommission + bootstrap is safer, since you have 3 replicas at all points in the process. Drain + copy will be faster, but puts you at some risk in case another node fails during the process. @Hyunwoo_Kim: Sounds interesting… Thank you for answering. I’m gonna find more info about Drain + Copy --- ### Page: https://forum.scylladb.com/t/how-to-use-sstableloader-in-scylladb-docker-container/4426 Title: How to use sstableloader in ScyllaDB docker container - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: scylladb/scylla:latest #Cluster size: os (RHEL/CentOS/Ubuntu/AWS AMI): Hello I guess I’m just to dumb to see it. How can I use/start sstableloader inside the scylla-Container… Language: en Canonical URL: https://forum.scylladb.com/t/how-to-use-sstableloader-in-scylladb-docker-container/4426 ## Headings Structure: H1: How to use sstableloader in ScyllaDB docker container H3: Related topics ## Main Content: H1: How to use sstableloader in ScyllaDB docker container H3: Related topics Installation details #ScyllaDB version: scylladb/scylla:latest #Cluster size: os (RHEL/CentOS/Ubuntu/AWS AMI): I guess I’m just to dumb to see it. How can I use/start sstableloader inside the scylla-Container? Can someone please help me out? Thanks in advance! daprodigy Which version of ScyllaDB are you using? Note that we have deprecated and removed sstableloader in ScyllaDB 6.1. We recommend that you use nodetool refresh --load-and-stream to load SSTables into the cluster instead. Thank you very much for the answer. So what are the basic steps for a cold migration from Cassandra to ScyllaDB? I put all the .db-Files from the Cassandra-Snapshot to the import-Folder of the appropriate Table in ScyllaDB. After running nodetool refresh --load-and-stream the files in the import-Folder were removed but select count() still returns 0. Is there anything in the ScyllaDB logs? Does the table you import the data to have sstables after the import? Where can I find the logs? I tried the following: Afterwards I check the data like 'Select count(Id) from keyspace.table; and it returns 0 in every of the described tries. I have the following setup: 1 Cassandra node with … Now i want to cold-migrate to scylladb running in a docker container. The setup stays the same. Just one simple node, one keyspace and 2 tables. What are the basic steps? Which image should I use? scylladb or scylladb-enterprise? Appreciate any help! Thank you. Ok, i tried a simple test and get an error. Here the steps: In cassandra (5.0.4) container: 5: Stoping cassandra container and starting scylladb (2025.1.2) container 6: Reexecute steps 1 and 2 inside ScyllaDB 7: Copy the following files from cassandra/data/edm/zeitreihe-uuid/snapshots/foobar/ to scylla/data/edm/zeitreihe-uuid/upload 8: Import data via nodetool … returns the following error: I get the same error if I copy only the file nb-1-big-Data.db into the upload-Folder. I hope this helps to identify the problem or my misunderstanding. Many thanks in advance! --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-80-2025-01-31/4431 Title: Last week in scylla-cluster-tests.git master (issue #80; 2025-01-31) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the ac2ac49d…9518705b range are covered. There were 20 non-merge commits from 8 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-80-2025-01-31/4431 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #80; 2025-01-31) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #80; 2025-01-31) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the ac2ac49d…9518705b range are covered. There were 20 non-merge commits from 8 authors in that period. Some notable commits: cql-stress benchmarking tool has been updated to the latest version, adding insert operations for user profiles and upgrading the Rust driver to 0.14. The issue with loader’s rack name in Argus has been fixed (the Resources tab). This will improve rack-aware feature testing by making rack information more accessible. SCT is now prepared for upgrades to 2025.1, with fixes ensuring the correct product (enterprise) is used from 2025.1 onward. Additionally, upgrades from enterprise to 2025.1 and beyond will require a reinstall. Since the enterprise version no longer exists and the master branch is now 2025.x, all enterprise performance tests will run with the master:latest version, and OSS performance test triggers have been disabled. Cloud Usage Report now includes active Capacity Reservations to help detect any unused ones. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/cpuset-conf-setup/4433 Title: Cpuset.conf - setup - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Upgrade from 5.0 to 5.1 #ScyllaDB version: 5.1 #Cluster size: 3 nodes os amazon Linux 2023 Hello team, I am following documentation Upgrade Guide - ScyllaDB 5.0 to 5.1 | ScyllaDB Docs and to complete upgrade, it see… Language: en Canonical URL: https://forum.scylladb.com/t/cpuset-conf-setup/4433 ## Headings Structure: H1: Cpuset.conf - setup H3: Related topics ## Main Content: H1: Cpuset.conf - setup H3: Related topics Upgrade from 5.0 to 5.1 #ScyllaDB version: 5.1 #Cluster size: 3 nodes os amazon Linux 2023 Hello team, I am following documentation Upgrade Guide - ScyllaDB 5.0 to 5.1 | ScyllaDB Docs and to complete upgrade, it seems I have to change Mode in perftune.yaml. When I execute command “sudo scylla_sysconfig_setup --nic ens5 --homedir /scylla --confdir /scylla/config” file /etc/scylla.d/cpuset.conf is generated. After that I restart scylla service and /etc/scylla.d/perftune.yaml is generated. it suppose to be mode sq_split but instead it is mq. wonder if there is something wrong with command and/or current config? btw, cluster is up and running but trying to complete scylladb documentation and tune best as possible ec2 instance type i4g.large. You are running an old and non-supported version of ScyllaDB, I suggest you gradually upgrade to at least ScyllaDB 6.2 and execute the command again. --- ### Page: https://forum.scylladb.com/t/how-to-create-a-local-snapshot-of-the-data/4435 Title: How to create a local snapshot of the data - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/how-to-create-a-local-snapshot-of-the-data/4435 ## Headings Structure: H1: How to create a local snapshot of the data H3: Related topics ## Main Content: H1: How to create a local snapshot of the data H3: Related topics Originally from the User Slack @Some_Random_Guy: I want to provide local snapshots of our data for the developers, what’s the best way to go about this? @avi: nodetool snapshot and copy the snapshot directories @Some_Random_Guy: I thought you had to restore the schema and copy the data to the newly created directories? cuz they are table_name-uuid did i misunderstand something? @avi: You can create the tables with CQL, and copy the data to the /upload directory of the new tables, then nodetool refresh @Some_Random_Guy: ok, ill look into this, thanks --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-267-2025-02-02/4439 Title: Last week in scylladb.git master (issue #267; 2025-02-02) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the f4b1ad43d4…e1b1a2068a range are covered. There were 94 non-merge commits from 23 authors in that perio… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-267-2025-02-02/4439 ## Headings Structure: H1: Last week in scylladb.git master (issue #267; 2025-02-02) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #267; 2025-02-02) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the f4b1ad43d4…e1b1a2068a range are covered. There were 94 non-merge commits from 23 authors in that period. Some notable commits: Keyspaces that use tablets now move data by streaming files instead of mutations. This reduces CPU usage and network bandwidth. This feature was ported from ScyllaDB Enterprise. ScyllaDB can now run on FIPS-compliant operating systems. This feature was ported from ScyllaDB Enterprise. The version mark in the master branch was updated to 2025.2, indicating the start of the 2025.1 stabilization cycle. An initialization order problem that could make zstd unavailable for compressing sstables has been fixed. The nodetool removenode force command was removed. Forcing a removal leaves the cluster in a worse state than it started. Transport layer security (TLS) uses Linux inotify to watch for certificates changing on disk. It will now consume fewer inotify resources and so have a lesser chance of failing due to that. Backup will now clean up snapshots as components are uploaded to reduce disk space consumption. A bug in reading encrypted sstables when their size slightly crosses over the buffer alignment has been fixed. Audit configuration can now be updated without restarting the server. Materialized views pair each view replica with a base replica. This pairing is now rack-aware - the database will prefer to pair a base replica and a view replica on the same rack. This reduces rack crossings which can be expensive on public clouds, and generally have lower bandwidth and higher latency. We now validate PERCENTILE values in speculative_retry schema attributes. Materialized views with tablets are disabled by default, until the combination is known to work well. A rare case of materialized view builds never completing was fixed. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/c-net-native-driver/4444 Title: C#/.NET native driver - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: During recent ScyllaDB University LIVE event it was suggested to me to post this request here. I’m extensively working with .NET codebases and we are planning to switch to ScyllaDB. So the request is simple - add .NET “n… Language: en Canonical URL: https://forum.scylladb.com/t/c-net-native-driver/4444 ## Headings Structure: H1: C#/.NET native driver H3: Related topics ## Main Content: H1: C#/.NET native driver H3: Related topics During recent ScyllaDB University LIVE event it was suggested to me to post this request here. I’m extensively working with .NET codebases and we are planning to switch to ScyllaDB. So the request is simple - add .NET “native” driver. I’m open contributing to it =) There is no ScyllaDB C# Driver yet, we plan to create one this year. Till we do, an Apache Cassandra compatible driver will work, even if as fast as the future native version. I will update here once the work starts. Help in code or reviews will be appreciated! @tzach thank you for the reply! Looking forward to contributing to the driver. Is there any ETA on when it will start? --- ### Page: https://forum.scylladb.com/t/graceful-way-to-change-seed-node/4445 Title: Graceful way to change seed node - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 5.1 #Cluster size: 23 node os : Ubuntu hello I have question about to change seed node if I have 9 node node1 ~ node9 the node 1-3 is seed node if I want change seed node… Language: en Canonical URL: https://forum.scylladb.com/t/graceful-way-to-change-seed-node/4445 ## Headings Structure: H1: Graceful way to change seed node H3: Remove a Seed Node from Seed List | ScyllaDB Docs H3: Related topics ## Main Content: H1: Graceful way to change seed node H3: Remove a Seed Node from Seed List | ScyllaDB Docs H3: Related topics Installation details #ScyllaDB version: 5.1 #Cluster size: 23 node os : Ubuntu hello I have question about to change seed node if I have 9 node node1 ~ node9 the node 1-3 is seed node if I want change seed node to node 7-9 first change node 7-9 scylla.yaml seed node config use node 1-3,7-9 then change node 1-6 scylla.yaml seed node config use node 7-9 last change node 7-9 scylla.yaml seed node config use node 7-9 Look good. The key is to validate there at least one seed node available when you add a new node. See here for more ScyllaDB is an Apache Cassandra-compatible NoSQL data store that can handle 1 million transactions per second on a single server. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-81-2025-02-07/4457 Title: Last week in scylla-cluster-tests.git master (issue #81; 2025-02-07) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the c03df5c5…80eb8032 range are covered. There were 5 non-merge commits from 5 authors in tha… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-81-2025-02-07/4457 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #81; 2025-02-07) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #81; 2025-02-07) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the c03df5c5…80eb8032 range are covered. There were 5 non-merge commits from 5 authors in that period. Some notable commits: FullScanAggregate now follows a common error handling pattern, converting ERROR to WARNING based on error messages or active nemesis. Additionally, the severity of mapreduce metric validation has been reduced due to an existing issue, until it is resolved. scylla-driver has been upgraded from 3.28.0 to 3.28.2. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-268-2025-02-09/4462 Title: Last week in scylladb.git master (issue #268; 2025-02-09) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the e1b1a2068a4…93f53f4eb82 range are covered. There were 94 non-merge commits from 22 authors in that per… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-268-2025-02-09/4462 ## Headings Structure: H1: Last week in scylladb.git master (issue #268; 2025-02-09) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #268; 2025-02-09) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the e1b1a2068a4…93f53f4eb82 range are covered. There were 94 non-merge commits from 22 authors in that period. Some notable commits: There is now support for vector types in CQL. Vectors are fixed-size arrays of another data type, commonly used for AI. Note nearest-neighbor vector search is not yet supported. This project was contributed by Warsaw University students. The reader concurrency semaphore controls concurrent reads at the replica level. It could lose track of reads in some circumstances, resulting in those zombie reads leaking memory. This is now fixed. ScyllaDB breaks long query results into pages to reduce transient memory consumption and latency. When it does so, it caches the query running on the replica and resumes it on the next page. This resuming broke when a paging decision was made due to a large number of tombstones, requiring the query to be restarted on the next page instead of resumed. This is now fixed. We no longer crash when creating a table while there is a rack that has no nodes in the NORMAL state. The reader concurrency semaphore tries to restrict the number of reads competing for CPU, since the competition delays all those reads. We now allow up to two reads to compete for the CPU in the default configuration. This allows common fast reads to compete and bypass rare long reads, reducing head-of-line blocking. The reader concurrency semaphore could accidentally use the the query timeout for evicting cached queries, resulting in reduced performance. This is now fixed. There are now more ways to specify AWS credentials for S3. Some configuration parameters can be live-updated on a running server by sending SIGUP. We now prevent parameters that are not designed to be live updated from being updated in the same manner, as it can cause unpredictable behavior. Authentication now ensures the default superuser password is set before serving CQL, reducing authentication problems. The memtable_flush_period_in_ms option now works for system tables. The S3 driver was tuned for improved backup performance. The toolchain used to build ScyllaDB is now based on Fedora 41 with clang 19. The compiler upgrade results in a small performance improvement. There are now per-table tablet configuration options used to control how many tablets are created for a table, allowing planning ahead for performance. Seamless upgrades from enterprise versions of ScyllaDB to new source-available versions are now supported. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-6-2-3/4466 Title: [RELEASE] ScyllaDB 6.2.3 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 6.2.3, a bugfix patch release of the ScyllaDB 6.2 stable branch. ScyllaDB Open Source 6.2.3, like all past and future 6.x.y releases, is backward compatible and supports r… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-6-2-3/4466 ## Headings Structure: H1: [RELEASE] ScyllaDB 6.2.3 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 6.2.3 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 6.2.3, a bugfix patch release of the ScyllaDB 6.2 stable branch. ScyllaDB Open Source 6.2.3, like all past and future 6.x.y releases, is backward compatible and supports rolling upgrades. Users are encouraged to upgrade to 6.2.3. Issue fixed in this release: Related issues: #21534 --- ### Page: https://forum.scylladb.com/t/affect-of-num-tokens-when-all-keyspaces-have-tablets-enabled/4469 Title: Affect of num_tokens when all keyspaces have Tablets enabled - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/affect-of-num-tokens-when-all-keyspaces-have-tablets-enabled/4469 ## Headings Structure: H1: Affect of num_tokens when all keyspaces have Tablets enabled H3: Related topics ## Main Content: H1: Affect of num_tokens when all keyspaces have Tablets enabled H3: Related topics Originally from the User Slack @Chaitanya_Tondlekar: what is the best value to set num_tokens on scylla 6.1 (Tablets) ? Is it correct to use(num_tokens=1) single-token mode + tablets for better performance and scalability in tablets ? @dor: I might be wrong but IIUC num_tokens is the amount of vnodes. It’s irrelevant if all of your tables are tablets. Otherwise, just use the default (I think it’s 256) @Chaitanya_Tondlekar: Yes considering scylla 6.1 , all keyspaces have tablets enabled. @avi: num_tokens is immaterial for tablets @Chaitanya_Tondlekar: @avi so what should we set ? @dor: If it’s all tablets, it doesn’t matter. Otherwise, use the default (256) --- ### Page: https://forum.scylladb.com/t/release-scylladb-6-1-5/4475 Title: [RELEASE] ScyllaDB 6.1.5 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Open Source 6.1.5, a bugfix patch release of the ScyllaDB 6.1 stable branch. ScyllaDB Open Source 6.1.5, like all past and future 6.x.y releases, is backward compatible and supports r… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-6-1-5/4475 ## Headings Structure: H1: [RELEASE] ScyllaDB 6.1.5 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 6.1.5 H3: Related topics The ScyllaDB team announces ScyllaDB Open Source 6.1.5, a bugfix patch release of the ScyllaDB 6.1 stable branch. ScyllaDB Open Source 6.1.5, like all past and future 6.x.y releases, is backward compatible and supports rolling upgrades. Note that the latest open-source stable branch is 6.2, and you are encouraged to upgrade to it. Issue fixed in this release: raft: improve logs for abort while waiting for apply #22159 --- ### Page: https://forum.scylladb.com/t/check-scylladb-configuration-in-docker-container-is-developer-mode-enabled/4477 Title: Check ScyllaDB configuration in Docker container, is developer mode enabled? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/check-scylladb-configuration-in-docker-container-is-developer-mode-enabled/4477 ## Headings Structure: H1: Check ScyllaDB configuration in Docker container, is developer mode enabled? H3: Related topics ## Main Content: H1: Check ScyllaDB configuration in Docker container, is developer mode enabled? H3: Related topics Originally from the User Slack @Sujay_KS: hey, is there any command to check in which mode scylla is running in docker comntainer like developer mode disabled or in enabled mode, also if i had not passed this flag during docker run will developer mode will be disabled by default ? @Botond_Dénes: One way is to check the logs, ScyllaDB will log its command-line on startup. You can also check with CQL: This flag defaults to false, if you don’t pass it it will be false. @Sujay_KS: got it will check this out, Thank you very much @Botond_Dénes if I change it to true or false using cql will it work I have not passed developer_mode field, but still its set to “true” @Botond_Dénes: Changing via CQL only works for live-update confguration items. Developer-mode is not one of those. It takes effect at startup. @Sujay_KS: okay, Thanks for the clarification @Botond_Dénes: > I have not passed developer_mode field, but still its set to “true” I seem to recall that our docker start scripts passes it auomatically --- ### Page: https://forum.scylladb.com/t/release-gocql-v1-14-5/4479 Title: [RELEASE]: GoCQL v1.14.5 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: Driver Release Summary: v1.14.5 Release link: Release v1.14.5 · scylladb/gocql · GitHub New Features: Compression Update: Replaced Snappy compressor with Snappy-compatible S2 algorithm. Pooling Improvements: Debouncin… Language: en Canonical URL: https://forum.scylladb.com/t/release-gocql-v1-14-5/4479 ## Headings Structure: H1: [RELEASE]: GoCQL v1.14.5 H3: Related topics ## Main Content: H1: [RELEASE]: GoCQL v1.14.5 H4: New Features: H4: Zero Token Nodes Support H4: Tablets Invalidation H4: Better Error Handling H4: Marshalling & Unmarshalling Refactor H4: Bug Fixes H4: Documentation Updates H4: Testing & CI Improvements H4: New Contributors H3: Related topics Driver Release Summary: v1.14.5 Release link: Release v1.14.5 · scylladb/gocql · GitHub --- ### Page: https://forum.scylladb.com/t/release-python-driver-3-28-2/4481 Title: [RELEASE] Python Driver 3.28.2 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: Driver Release Summary: 3.28.2 Release Link: Release 3.28.2 · scylladb/python-driver · GitHub Driver Release Notes Summary Improvements & Fixes No longer throw warnings for zero-token nodes. Prevent pytest from picki… Language: en Canonical URL: https://forum.scylladb.com/t/release-python-driver-3-28-2/4481 ## Headings Structure: H1: [RELEASE] Python Driver 3.28.2 H3: Driver Release Notes Summary H3: Related topics ## Main Content: H1: [RELEASE] Python Driver 3.28.2 H3: Driver Release Notes Summary H3: Related topics Driver Release Summary: 3.28.2 Release Link: Release 3.28.2 · scylladb/python-driver · GitHub Environment & Compatibility Updates --- ### Page: https://forum.scylladb.com/t/6-0-4-major-crashes-due-to-memory-gossip-failures/4483 Title: 6.0.4 Major Crashes Due to Memory/Gossip Failures - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 6.0.4 #Cluster size: 39 i3en.6xlarge EC2 instances os (RHEL/CentOS/Ubuntu/AWS AMI): Customized based on the scylla-6.0.4-x86_64 AMI builder - Ubuntu 22.04 We upgraded to 6.0.4… Language: en Canonical URL: https://forum.scylladb.com/t/6-0-4-major-crashes-due-to-memory-gossip-failures/4483 ## Headings Structure: H1: 6.0.4 Major Crashes Due to Memory/Gossip Failures H3: Related topics ## Main Content: H1: 6.0.4 Major Crashes Due to Memory/Gossip Failures H3: Related topics Installation details #ScyllaDB version: 6.0.4 #Cluster size: 39 i3en.6xlarge EC2 instances os (RHEL/CentOS/Ubuntu/AWS AMI): Customized based on the scylla-6.0.4-x86_64 AMI builder - Ubuntu 22.04 We upgraded to 6.0.4 a few weeks ago, and since then on our largest cluster have seen restart issues where a single node will crash/restart, then large chunks of the cluster will restart together. 10+ nodes will fail, restart, then another 10 or so will restart, then 1 or 2 to finish off, then we are fine for another few days. Checking logs doesn’t really leave us with much, plenty of bad_alloc errors around the failure. This error seems fairly common though. Today the 1st node that failed had this error, then immediately crashed (core dump) and restarted. These crashes mean the cluster cannot be read from or written to for around 30 minutes until it recovers. We’re looking to upgrade to 6.1, but I couldn’t find anything related in issues / release notes, so am not sure if this specific issue will be resolved. I’m curious if this was resolved in some way or if Raft/Gossip has improved and the upgrade should help. Another day and more failures. This time it was during compactions. Initial failure: Secondary error, right when the first node is coming back up and there are log messages about it rejoining the token ring: we are also having memory issues since 6.0. Could you check the scylla_memory_regular_dirty_bytes metric? Perhaps you are having the same problem as we have: We noticed that for us scylla_memory_regular_dirty_bytes goes up over time, at some point not going down with flushes any more. Disclaimer: I am just a Scylla user, not a dev. But I am interested in understanding allocation problems better. @horschi Thanks for the input. I checked scylla_memory_regular_dirty_bytes and don’t really see any large spikes before the restarts. It builds up over an hour or so after restarts then stays within ~20% of where it peaks. Are you suggesting it should drop down fairly often? We have constant activity on the cluster, so we never really have downtime unless we plan for it (or the crashes). If the nodes are under memory pressure, the problems can manifest in many different places. The code causing the crash might be completely innocent, it was just the one which attempted a badly failing allocation. Any time there is memory pressure, the metric to look at is non-LSA memory usage in the “Detailed” dashboard. Almost always, you will see this metric being elevated on the sick node. The next step is to find something which correlates with this, lately we often found the memory used by bloom filter to be the culprit, but it can be something else as well. Not in our cases @Botond_Denes, we do have 1 node that is always higher on Non-LSA, but I think it just has more “load” than the others. It reports more on nodetool status than the others, and is our oldest node, so I assume that has something to do with it. Here’s an example of one of our failures this week (of 4), there isn’t a change in Non-LSA at all before the crash (even when I isolate to that node it has no change). any large spikes before the restarts. It builds up over an hour or so after restarts then stays within ~20% of where it peaks. Are you suggesting it should drop down fairly often? No, I think what you describe sounds like normal behaviour. For us it keeps building up over time, at some point being permanently stuck at 100%. How does it behave when you manually call nodetool flush? If its a properly functioning node, then the metric goes close to 0. Another question: Are you doing a lot of “in” queries on larger partitions? It feels like Scylla does not like them very much. I find your post quite interesting, as your traces look similar to what we are seeing. @horschi, @GarrettPoore , thanks for reporting this. Notice that version 6.0.x is End-of-life. Please open an issue for it, with as many details as possible, so that it would be easier to investigate. @horschi Hm, I haven’t manually flushed any of these nodes yet, but will give it a try tonight. As for what queries we use, I’m not sure, I’ll have to check with app developers. @Guy We found this issue on GitHub last week and have been talking there a bit as well. Though we have had trouble analyzing the coredumps on our end. We are unable to send the coredumps due to compliance constraints. We also tried upgraded this specific cluster to 6.1.5 over the weekend, hoping it helps. Though since the above issue is still open we’re not convinced we’re safe yet. @horschi I tried a nodetool flush and that metric does drop very low, so I think we’re seeing different issues. Or at least different ways to get to the same problem. --- ### Page: https://forum.scylladb.com/t/scylladb-rust-driver-how-do-i-create-a-vec-of-batch-values/4484 Title: ScyllaDB Rust Driver: How do I create a Vec of batch values? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-rust-driver-how-do-i-create-a-vec-of-batch-values/4484 ## Headings Structure: H1: ScyllaDB Rust Driver: How do I create a Vec of batch values? H3: Related topics ## Main Content: H1: ScyllaDB Rust Driver: How do I create a Vec of batch values? H3: Related topics Originally from the User Slack @Terry_Davis: How can you create a Vec of batch values in v0.15 like you used to be able to do in scylla rust driver v0.14? The examples only show how to do it using a tuple with hardcoded values: https://rust-driver.docs.scylladb.com/stable/queries/batch.html#batch-statement I can’t create something like Vec or Vec or Vec> or anything like that. Do I have to type out every single field manually like this? batch_values.push((player.id, player.account_id)); actually even that won’t work because tuples differ in shape. @Karol_Baryła: 1. You should not use LegacySerializedValues - it is a not type safe remnant of previous serialization framework and will be removed in the next version. In your case you should probably use Player (or whatever is the type of player var). 2. fn generate_populate_queries(json_values: &JsonValues) -> (Batch, impl BatchValues) should work. @Terry_Davis: 1. This is actually my objective to move away from LegacySerializedValues (0.14 -> 0.15 update) 2. The issue arises when I try to push another value of a different shape to batch_values Previously, the "polymorphism" was achieved thanks toValueList`, but now I don’t know how to type that vector. @Karol_Baryła: This would be a great use case for “BoundStatement” which we currently lack, but will be added in the future. For now you can: • Use 2 batches instead of one, so that the type of values in the batch is constant. • Box your values and use Vec> I think you could also make your own enum, with variants for your types, and implement SerializeRow for it (and then use Vec) but I’m not sure because I didn’t try to do this before. I can try in a week (I am going for a leave now), or maybe @Mikołaj_Uzarski or @Wojciech_Przytuła will be able to help. @Terry_Davis: The enum route will work for now. I had to implement the BatchValues trait which is required when calling batch. If someone has a better solution please lmk. This is horrid but works: @Karol_Baryła: Hi @Terry_Davis, I’m back now so I can help again. There should be no need to implement BatchValues in your case - there is already an impl for Vec. I wrote a quick example of using enum for batch with multiple row types: --- ### Page: https://forum.scylladb.com/t/could-sstable-dump-data-by-api/4485 Title: Could sstable dump-data by api - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Is there any api to dump sstable files ? Language: en Canonical URL: https://forum.scylladb.com/t/could-sstable-dump-data-by-api/4485 ## Headings Structure: H1: Could sstable dump-data by api H3: Related topics ## Main Content: H1: Could sstable dump-data by api H3: Related topics Is there any api to dump sstable files ? We have offline tools to dump sstables, see ScyllaDB SStable | ScyllaDB Docs. --- ### Page: https://forum.scylladb.com/t/release-scylladb-java-driver-3-11-5-5/4486 Title: [RELEASE] ScyllaDB Java Driver 3.11.5.5 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces release of ScyllaDB Java Driver 3.11.5.5. This marks the first release post for the 3.x series of the driver, despite being around for a long while. Here is a quick overview of recent notable… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-java-driver-3-11-5-5/4486 ## Headings Structure: H1: [RELEASE] ScyllaDB Java Driver 3.11.5.5 H3: 3.11.5.3 H3: 3.11.5.2 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Java Driver 3.11.5.5 H4: 3.11.5.5 H4: 3.11.5.4 H3: 3.11.5.3 H3: 3.11.5.2 H3: Related topics The ScyllaDB team announces release of ScyllaDB Java Driver 3.11.5.5. This marks the first release post for the 3.x series of the driver, despite being around for a long while. Here is a quick overview of recent notable changes: side note: this version is unavailable on maven repository. Please use 3.11.5.5 instead. Latest code, tagged versions and full history can be checked out on GitHub: https://github.com/scylladb/java-driver/tree/scylla-3.x Check out “Getting the driver” section for how to incorporate the driver in your project: --- ### Page: https://forum.scylladb.com/t/cluster-overloaded-getting-error-message-reader-concurrency-semaphore-rate-limiting-dropped-49-similar-messages-semaphore-read-concurrency-sem-with-55-100-count-and-13255581-147597557-memory-resources-timed-out-dumping-permit-diagnostics/4487 Title: Cluster overloaded? Getting error message: reader_concurrency_semaphore - (rate limiting dropped 49 similar messages) Semaphore _read_concurrency_sem with 55/100 count and 13255581/147597557 memory resources: timed out, dumping permit diagnostics: - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/cluster-overloaded-getting-error-message-reader-concurrency-semaphore-rate-limiting-dropped-49-similar-messages-semaphore-read-concurrency-sem-with-55-100-count-and-13255581-147597557-memory-resources-timed-out-dumping-permit-diagnostics/4487 ## Headings Structure: H1: Cluster overloaded? Getting error message: reader_concurrency_semaphore - (rate limiting dropped 49 similar messages) Semaphore _read_concurrency_sem with 55/100 count and 13255581/147597557 memory resources: timed out, dumping permit diagnostics: H3: Related topics ## Main Content: H1: Cluster overloaded? Getting error message: reader_concurrency_semaphore - (rate limiting dropped 49 similar messages) Semaphore _read_concurrency_sem with 55/100 count and 13255581/147597557 memory resources: timed out, dumping permit diagnostics: H3: Related topics Originally from the User Slack @Sujay_KS: hey, i’m facing frequent issues in 3 node scylla cluster setup (here we are running as docker container) this happens very frequently, is there any resolution for this issue ? @Botond_Dénes: Look like your cluster is overloaded. Have a look at monitoring, you will probably see high reactor load. @Sujay_KS: yeah memory usage was pretty high, will check this from application side row_cache_size_in_mb can this flag help in reducing cache overload ? @Botond_Dénes: No, this option is not implemented (it is there for Cassandra backwards-compatibility). ScyllaDB always tries to use all of the memory, that is not the problem here. The cluster seems to be overloaded on the CPU and possibly on the I/O level. @Sujay_KS: can we limit the memory usage and also looks like it’s using memory a lot without flushing ( 32 Gb Fully used ) @Botond_Dénes: Memory usage can be limited by the --memory command-line parameter. By default ScyllaDB will use all available memory, minus the reserve (set by --reserve-memory). can you pls review this and suggest any abnormality which can cause this issue? @Botond_Dénes: I don’t see anything obviously wrong here. To resolve the issue, you will probably either have to reduce load or scale out/up the cluster. @avi: If it’s always shard 2, then it can be a hot partition. Use nodetool toppartition to track which one. @Sujay_KS: is there any command or query using which i can check which is overloading scylla and spiking cpu usage @Botond_Dénes: Unfortunately no (although we should definitely have such a command). The above command (nodetool toppartitions ) is your best bet, it will tell you which partition is queried the most and it will show you which partition is hot (if any). --- ### Page: https://forum.scylladb.com/t/release-scylladb-java-driver-4-18-0-2/4490 Title: [RELEASE] ScyllaDB Java Driver 4.18.0.2 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces release of ScyllaDB Java Driver 4.18.0.2. Previous release post: 4.18.0.1 This version brings the following notable changes: Improvements to tablet handling. Checking tablet map only for … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-java-driver-4-18-0-2/4490 ## Headings Structure: H1: [RELEASE] ScyllaDB Java Driver 4.18.0.2 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Java Driver 4.18.0.2 H3: Related topics The ScyllaDB team announces release of ScyllaDB Java Driver 4.18.0.2. Previous release post: 4.18.0.1 This version brings the following notable changes: You can check out the source code and full history on GitHub https://github.com/scylladb/java-driver/tree/scylla-4.x For how to use the driver please refer to “Getting the driver” section. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-82-2025-02-14/4497 Title: Last week in scylla-cluster-tests.git master (issue #82; 2025-02-14) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 29744f38…baa8624a range are covered. There were 15 non-merge commits from 5 authors in tha… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-82-2025-02-14/4497 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #82; 2025-02-14) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #82; 2025-02-14) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 29744f38…baa8624a range are covered. There were 15 non-merge commits from 5 authors in that period. Some notable commits: Latte tool was upgraded to 0.28.2-scylladb, introducing real-time streaming of HDR histogram data and improving latency calculations by addressing coordinated omission. It can now run similarly to cassandra-stress by mimicking its schema and queries. cassandra-stress, cql-stress, and latte metrics are now combined in monitoring panels based on metric types. Scylla Manager testing saw several improvements, including a more flexible snapshot preparation process, additional parameterization, support for cloud clusters, and sending snapshot details to Argus Results for easier access. To prepare for broader HDR histogram support in benchmarking tools, the use_hdr_cs_histogram SCT config option was renamed to use_hdrhistogram, removing the cassandra-stress reference. Related modules were also renamed and moved to a common location. is_enterprise node property was fixed for 2025.1, resolving several issues with enterprise feature testing in source-available Scylla. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/using-elasticsearch-for-full-text-search-on-a-specific-column-with-scylladb-performance/4501 Title: Using Elasticsearch for full-text search on a specific column with ScyllaDB - performance - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/using-elasticsearch-for-full-text-search-on-a-specific-column-with-scylladb-performance/4501 ## Headings Structure: H1: Using Elasticsearch for full-text search on a specific column with ScyllaDB - performance H3: Related topics ## Main Content: H1: Using Elasticsearch for full-text search on a specific column with ScyllaDB - performance H3: Related topics Originally from the User Slack @Ahmed: We’ve set up Elasticsearch for full-text search on a specific column. The challenge is retrieving full records from ScyllaDB in bulk. After finding results in Elasticsearch, querying ScyllaDB with multiple keys using the IN operator requires ALLOW FILTERING and doesn’t use the secondary index. Searching one by one is too slow. Is there a better way to fetch bulk data efficiently from ScyllaDB based on Elasticsearch results? @avi: You should store the primary key in Elasticsearch along with the column you’re indexing @Ahmed: Even with the primary key stored in Elasticsearch, retrieving thousands of records from ScyllaDB is still challenging due to the need for multiple queries. @avi: If you fire them off in parallel (not using IN) the entire cluster bandwidth can be utilized. @Ahmed: that’s what i’m doing right now using multiple threads to query prepared statement it’s work but i was looking for proper solution @avi: Launching thousands of threads will be slow, better to use async @Ahmed: i’m using threadpool and like 10 max threads @avi: That’s not enough for thousands of keys. Better to use async. @Ahmed: thanks it’s works great --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-269-2025-02-16/4502 Title: Last week in scylladb.git master (issue #269; 2025-02-16) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 93f53f4eb82…0d5f5e6c9da range are covered. There were 89 non-merge commits from 25 authors in that per… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-269-2025-02-16/4502 ## Headings Structure: H1: Last week in scylladb.git master (issue #269; 2025-02-16) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #269; 2025-02-16) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 93f53f4eb82…0d5f5e6c9da range are covered. There were 89 non-merge commits from 25 authors in that period. Some notable commits: A materialized view created using AS SELECT * will show as such in the DESCRIBE statement, instead of expanding to the table columns. Error handling while streaming mutations was improved. ScyllaDB will automatically parallelize some aggregation queries, such as SELECT count(*) FROM table. Such parallelized queries are now cancelled if a node is being shut down. Alternator, ScyllaDB’s implementation of the DynamoDB API, can now add or delete global secondary indices after the table is creates. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/scylladb-monitoring-stack-how-do-i-generate-the-grafana-json-files/4506 Title: ScyllaDB Monitoring stack: how do I generate the Grafana JSON files? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-monitoring-stack-how-do-i-generate-the-grafana-json-files/4506 ## Headings Structure: H1: ScyllaDB Monitoring stack: how do I generate the Grafana JSON files? H3: Related topics ## Main Content: H1: ScyllaDB Monitoring stack: how do I generate the Grafana JSON files? H3: Related topics Originally from the User Slack @Nathan_Li: for the scylla monitoring stack’s grafana json files (from https://github.com/scylladb/scylla-monitoring/tree/master/grafana), does anyone know how to generate the grafana json files from the templates to be imported into their own self hosted grafana? GitHub: scylla-monitoring/grafana at master · scylladb/scylla-monitoring @amnon: if you take a release (not master) the dashboards are already there - no need to generate them. You can then load the json to a grafana server (from the server or from the api) if you want to create the dashboards from template by yourself, use ./generate-dashboards.sh script --- ### Page: https://forum.scylladb.com/t/explanation-of-each-widget-in-monitoring-stack/4507 Title: Explanation of each widget in monitoring stack - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 6.2 #Cluster size: 3 os (RHEL/CentOS/Ubuntu/AWS AMI): AWS AMI Hi, We have ScyallDB monitoring stack installed. There are so many widgets. Is there a cheat sheet to look for s… Language: en Canonical URL: https://forum.scylladb.com/t/explanation-of-each-widget-in-monitoring-stack/4507 ## Headings Structure: H1: Explanation of each widget in monitoring stack H3: Related topics ## Main Content: H1: Explanation of each widget in monitoring stack H3: Related topics Installation details #ScyllaDB version: 6.2 #Cluster size: 3 os (RHEL/CentOS/Ubuntu/AWS AMI): AWS AMI We have ScyallDB monitoring stack installed. There are so many widgets. Is there a cheat sheet to look for specific widgets which indicates cluster health? Is there a doc which explains each of the widgets in details. You can find the Monitoring user guide here. I also recommend the Monitoring lesson on ScyllaDB University, it covers different widgets, how to troubleshoot issues and things to look at for general cluster health. If you have a question about a specific dashboard/widget, you can ask here. --- ### Page: https://forum.scylladb.com/t/release-scylla-doctor-v1-5/4508 Title: [RELEASE]: Scylla Doctor v1.5 - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Scylla Doctor v1.5 is released. Added Collectors: NTP status collector. ScyllaConfigurationFileNoParsingCollector (needed for ScyllaConfigurationFileFormatAnalyzer). Added Analyzers: NTP status analyzer. Scylla… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-doctor-v1-5/4508 ## Headings Structure: H1: [RELEASE]: Scylla Doctor v1.5 H3: Related topics ## Main Content: H1: [RELEASE]: Scylla Doctor v1.5 H3: Related topics Scylla Doctor v1.5 is released. Artifacts can be downloaded from https://downloads.scylladb.com/downloads/scylla-doctor/ or installed from Scylla OSS or Scylla Enterprise repositories. --- ### Page: https://forum.scylladb.com/t/besides-writing-commitlog-during-i-o-are-there-any-other-ways-to-generate-commitlog/4509 Title: Besides writing commitlog during I/O, are there any other ways to generate commitlog? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 5.2.13 #Cluster size: 15 os (RHEL/CentOS/Ubuntu/AWS AMI): CentOS Hi! My cluster uses alternator and does not have any calls to cqlsh. I found an error log in the log of one o… Language: en Canonical URL: https://forum.scylladb.com/t/besides-writing-commitlog-during-i-o-are-there-any-other-ways-to-generate-commitlog/4509 ## Headings Structure: H1: Besides writing commitlog during I/O, are there any other ways to generate commitlog? H3: Related topics ## Main Content: H1: Besides writing commitlog during I/O, are there any other ways to generate commitlog? H3: Related topics Installation details #ScyllaDB version: 5.2.13 #Cluster size: 15 os (RHEL/CentOS/Ubuntu/AWS AMI): CentOS Hi! My cluster uses alternator and does not have any calls to cqlsh. I found an error log in the log of one of the nodes in the cluster as follows: I found from the code that max_mutation_size = 1/2 of commitlog segment size, which is 16M. But what makes me wonder is how this 36M commitlog entry was generated? Writing an item that is too large directly from Alternator will return a 413 Request Body Too Large error. So besides writing commitlog during I/O, are there any other ways to generate commitlog? Commitlog is only written to on the write path. It is possible to generate a too large mutation, even if the client protocol (CQL or Alternator) itself tries to guard against it: via read-modify-write. I don’t know if Alternator has any such write-paths, @nyh do you know of any Alternator command does such read-modify-write? Another explanation is the discrepancy between JSON (the serialization format used by Alternator) and the mutation format used internally – namely that even though some item is still small enough as JSON but when re-serialized to frozen_mutation, it is now too large. This is hard to imagine because JSON itself is a very bloated text-based serialization format, while our frozen_mutation format is a binary representation (although a bloated one to be sure). @bo_li you’re right that Alternator prevents you from creating a huge item via PutItem, because as you noted the size of the request itself is limited to content_length_limit = 16 * MB. However, as @Botond_Denes noted you can still create larger items incrementally, e.g., by setting additional attributes on an existing item, or by enlarging an existing attribute. For example consider a read-modify-write UpdateItem which asks to append an element to a list attribute. It’s a tiny request, but each time you run it the attribute grows a little bit longer, and the mutation that writes the new full value of the attribute (yes, that’s how it is done) grows larger and larger. Do you know if your workload has such things - like appending things to existing lists, sets or maps or strings? We have an issue alternator: add a test for large items, and document limits · Issue #11880 · scylladb/scylladb · GitHub about testing these schenarios. Scylla doesn’t necessarily need to support huge items or attributes (DynamoDB does not), but it needs to behave reasonably in this case - it must not crash or hang or break the database in a way that no future writes will succeed. Yes, I used the read-modify-write UpdateItem. It should be related to him. --- ### Page: https://forum.scylladb.com/t/refershing-data-expiring-old-records-when-uploading-batch-data-daily-aggregates-full-table-scan-ttl-and-secondary-indexes/4514 Title: Refershing data, expiring old records when uploading batch data (daily aggregates), full table scan, TTL and Secondary Indexes - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/refershing-data-expiring-old-records-when-uploading-batch-data-daily-aggregates-full-table-scan-ttl-and-secondary-indexes/4514 ## Headings Structure: H1: Refershing data, expiring old records when uploading batch data (daily aggregates), full table scan, TTL and Secondary Indexes H3: Related topics ## Main Content: H1: Refershing data, expiring old records when uploading batch data (daily aggregates), full table scan, TTL and Secondary Indexes H3: Related topics Originally from the User Slack @Igor_Q: hello, everyone i have a running scylla cluster storing some features in a table (id, field, data, updated_at) with primary key (id, field) this is great for saving realtime updates by (id, field) and querying the whole wide row by (id) now what i want to do is to also upload batch data (daily aggregates). upsert by (id, field) works fine here, too, but the problem is, these data are meant to be fully refreshed, so i need some way to expire old records from the table i have considered the following solutions: • implement full scan as described in https://www.scylladb.com/2017/03/28/parallel-efficient-full-table-scan-scylla/ with bypass caching. this can be used to expire records from all aggregate uploads by (field, updated_at) and also to collect some data metrics at the same time. but i am concerned about performance penalty. there’s probably non-negligible limitation on how often i can run full scans • use secondary index to filter by field , but essentially run full scan on all records with matching field • set short TTL (3-6 hours) for records updated from aggregates and constantly reupload them to refresh the TTL. the downside is, if constant reuploading breaks, we lose the data from scylla is there something i am missing? what would you suggest? @avi: Full scan is a good solution here, especially with workload prioritization pushing it to only use idle time --- ### Page: https://forum.scylladb.com/t/after-node-goes-down-getting-error-cannot-achieve-consistency-level-for-cl-local-one-requires-1-alive-0/4517 Title: After node goes down, getting error Cannot achieve consistency level for cl LOCAL_ONE. Requires 1, alive 0 - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/after-node-goes-down-getting-error-cannot-achieve-consistency-level-for-cl-local-one-requires-1-alive-0/4517 ## Headings Structure: H1: After node goes down, getting error Cannot achieve consistency level for cl LOCAL_ONE. Requires 1, alive 0 H3: Related topics ## Main Content: H1: After node goes down, getting error Cannot achieve consistency level for cl LOCAL_ONE. Requires 1, alive 0 H3: Related topics Originally from the User Slack @Marcondes_Viana_de_Oliveira_Junior: Hey, one question, I’m connecting to a 3 nodes cluster with a key space using replica of 3 and CL of local quorum. I intentionally killed one node to test and I got some logs: Cannot achieve consistency level for cl LOCAL_ONE. Requires 1, alive 0 , but the other two nodes seems alive. When I connect to the cluster I pass all ips. @Karol_Baryła: How are nodes distributed between DCs and what is your exact replication strategy? @Marcondes_Viana_de_Oliveira_Junior: Only one DC NetworkTopologyStrategy I’m using golang and I set this to the client: Can it lock a host? So it don’t try other hosts? Also, I just tried to recreate the connection and it fails. Fails to start even I pass 3 hosts, one is down It do not connects @avi: Check if you misspelled the datacenter name at the client side @Marcondes_Viana_de_Oliveira_Junior: Is the IP that I killed, even I pass 3 ips to the client connection. The other 2 IPS are live DC is correct we have 2 alive and one dead, and the ips are correct. also dc I added some logs in the lib and it fails to get the protocol version It only logs the latest error, that’s why I was seeing different errors, now It’s more clear, but I still don’t understand @avi: Double-check the replication factor @Marcondes_Viana_de_Oliveira_Junior: Is 3 @avi: I don’t have an explanation then, if a node was alive enough to respond, it is alive enough to be a replica. Are you using tablets? @Marcondes_Viana_de_Oliveira_Junior: Not sure. I asked the responsible for the deployment. I can perform queries using cqlsh But the go client can’t even connect, fails in this protocol version request @avi: Then the problem is somewhere in the client or networking @Marcondes_Viana_de_Oliveira_Junior: Ok, I will keep digging @avi: wireshark may help @Marcondes_Viana_de_Oliveira_Junior: One user can login other don’t, when I try to list the roles from cqlsh I get: the only diff between the user is that one can only read and the other can read/modify @avi: What version are you running? @Marcondes_Viana_de_Oliveira_Junior: I guess this is the problem, auth is spread Version: 6.0.2-0.20240703.c9cd171f426e @avi: But in 6.0 auth was moved to raft @Marcondes_Viana_de_Oliveira_Junior: i got version from nodetool status --version wrong place? @avi: no, it’s the right place you can fix the problem by increasing the replication factor of system_auth and running repair (this is documented), but it should be in raft in 6.0 @Marcondes_Viana_de_Oliveira_Junior: Is there a place where raft is configurable so I can check if is it on/off? @avi: It’s automatic https://github.com/scylladb/scylladb/commit/19b816bb68292b2a5ff7d8e8ec374ceb0d5ed85e GitHub: Merge ‘Migrate system_auth to raft group0’ from Marcin Maliszkiewicz · scylladb/scylladb@19b816b @Marcondes_Viana_de_Oliveira_Junior: This is the current configuration @avi: Was this cluster upgraded, or did it begin life as 6.0? @Marcondes_Viana_de_Oliveira_Junior: This I need to ask the devops guy, not available now. But my bet it was not upgraded. @avi: Well you can fix it with ALTER KEYSPACE + repair @Marcondes_Viana_de_Oliveira_Junior: Sure Appreciate you. Tomorrow I will try the fix and let you know! nodetool repair system_auth -full Right? @Marcondes_Viana_de_Oliveira_Junior: Thanks again You were right. It was born in 5 and migrated to 6. Is there a tool/tutorial to migrate auth from system_auth to raft? We found the docs, Thanks! Worked, thank you!! --- ### Page: https://forum.scylladb.com/t/nodetool-clearsnapshot-error/4521 Title: Nodetool clearsnapshot - error - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 6.0.3 #Cluster size: 3 i4g.large #OS: AWS linux 2023 Hi team, got following error in recent versions of scyllaDB, when I execute command “nodetool clearsnapshot” or “nodetool … Language: en Canonical URL: https://forum.scylladb.com/t/nodetool-clearsnapshot-error/4521 ## Headings Structure: H1: Nodetool clearsnapshot - error H3: Related topics ## Main Content: H1: Nodetool clearsnapshot - error H3: Related topics Installation details #ScyllaDB version: 6.0.3 #Cluster size: 3 i4g.large #OS: AWS linux 2023 Hi team, got following error in recent versions of scyllaDB, when I execute command “nodetool clearsnapshot” or “nodetool clearsnapshot -t myname”. btw, not issues if we include keyspace name. seastar::internal::backtraced (table directory entry name '1' is invalid: no '-' separator found at pos -32 Looks like you have a directory/file named 1 within one of the tables’ snapshot dir. Seastar reconizes it does not match the expected format (which is snapshot name, a dash, and a timestamp) and raises the error. You might need to manually delete that file/dir and only let objects created by ScyllaDB to live on the snapshots directories. --- ### Page: https://forum.scylladb.com/t/serialize-and-deserialize-decimal-in-rust-driver-open-source-contribution/4523 Title: Serialize and deserialize Decimal in Rust Driver, open source contribution - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/serialize-and-deserialize-decimal-in-rust-driver-open-source-contribution/4523 ## Headings Structure: H1: Serialize and deserialize Decimal in Rust Driver, open source contribution H3: Related topics ## Main Content: H1: Serialize and deserialize Decimal in Rust Driver, open source contribution H3: Related topics Originally from the User Slack @erik404: Hi all! New user here! I’ve been using the ScyllaDB Rust driver since a couple of months now in a project that does algorithm optimizations. Before I had build my own storage system that managed data but eventually this failed after we reached 10 million+ simulations/hour. So far ScyllaDB seems to hold up very well! And I am pleased with the state of the driver so far! Thanks for all the hard work! I had some issues with (de)serializing rust_decimal::Decimal to ScyllaDB and back without having to write boiler code all the time so I made a Crate that wraps that Decimal in a DecimalCql that implements (De)SerializeValue so I can easily use (De)SerializeRow on my structs. I am not a frequent opensource contributor, or someone who works in large teams of programmers at all. So all feedback is welcome! https://crates.io/crates/rust_decimal_cql/ @Felipe_Cardeneti_Mendes: oh nice! @Wojciech_Przytuła maybe we could make its consumption easier as a feature somehow? @Guy: Interesting, thanks! --- ### Page: https://forum.scylladb.com/t/i-want-to-know-best-practices-for-designing-an-efficient-data-model-in-scylladb/4526 Title: I want to know Best Practices for Designing an Efficient Data Model in ScyllaDB? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hey everyone, I am working on designing a data model for a new project using ScyllaDB & want to get some insights from the community. I know that ScyllaDB is built for high performance & scalability but I want to make s… Language: en Canonical URL: https://forum.scylladb.com/t/i-want-to-know-best-practices-for-designing-an-efficient-data-model-in-scylladb/4526 ## Headings Structure: H1: I want to know Best Practices for Designing an Efficient Data Model in ScyllaDB? H3: Related topics ## Main Content: H1: I want to know Best Practices for Designing an Efficient Data Model in ScyllaDB? H3: Related topics I am working on designing a data model for a new project using ScyllaDB & want to get some insights from the community. I know that ScyllaDB is built for high performance & scalability but I want to make sure I am structuring my tables and queries in the best possible way. From what I understand, avoiding too many partitions and making good use of clustering keys is key But I am still a bit unsure about things like Would love to hear from folks who have worked with ScyllaDB in production. if anyone have any resources, tutorials or personal experiences please share with me. There are many tutorials and lessons, covering the topics you mentioned in ScyllaDB University. The How to Write Better Apps lesson covers best practices and tips, including common pitfalls, performance, and tombstones. Materialized Views Indexing and Filtering, including performance impact, is covered in [this lesson](https://Materialized Views, Secondary Indexes, and Filtering). If you have specific questions, you can ask them here. Provide as many details as possible about your use case and the data model. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-83-2025-02-21/4527 Title: Last week in scylla-cluster-tests.git master (issue #83; 2025-02-21) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 9c63d6d6…0be52b00 range are covered. There were 33 non-merge commits from 13 authors in t… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-83-2025-02-21/4527 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #83; 2025-02-21) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #83; 2025-02-21) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 9c63d6d6…0be52b00 range are covered. There were 33 non-merge commits from 13 authors in that period. Some notable commits: Refactored HDR histogram support to work with multiple benchmarking tools, not just cassandra-stress. Stress threads now automatically detect HDR support, enabling future integration with tools like latte, scylla-bench, and cql-stress-cassandra-stress. Enabled a tombstone GC verification thread for large-partitions tests, requiring an adjustment to tombstone lookup to match the new output format of scylla sstable dump-data. Due to a change in the semantics of the enable_tablets Scylla config parameter, is_tablets_feature_enabled now returns the default tablets setting for new keyspaces. Nemesis can now be flagged with supports_high_disk_utilization to indicate compatibility with 90% disk utilization tests, and a 90% utilization test has been added. Fixed Gemini command arguments when no oracle cluster is provided, resolving an issue in upgrade tests. Enabled parallel node operations by default, improving cluster provisioning speed. This enhances test scenarios that involve adding multiple nodes (nemesis_add_node_cnt > 1), aligning with recent Scylla rapid scaling improvements. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/how-to-perform-schema-changes-in-scylladb/4531 Title: How to perform schema changes in ScyllaDB? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/how-to-perform-schema-changes-in-scylladb/4531 ## Headings Structure: H1: How to perform schema changes in ScyllaDB? H3: Related topics ## Main Content: H1: How to perform schema changes in ScyllaDB? H3: Related topics Originally from the User Slack @Roman_Hargrave: maybe we’re missing something obvious or just haven’t tested this by hand, but what is the best way to handle schema changes in a service with multiple instances, e.g. do we need to find some way for the instances to elect a migration runner? in most of our other systems, we are able to start a single pilot/setup/config instance before rolling out the whole deployment, but that’s not an option in this case. @dor: You just send the CQL statement. Schema changes are transactional and Scylla uses raft and leader election underneath --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-270-2025-02-23/4534 Title: Last week in scylladb.git master (issue #270; 2025-02-23) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 0d5f5e6c9da…a6c882e4e37 range are covered. There were 137 non-merge commits from 23 authors in that pe… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-270-2025-02-23/4534 ## Headings Structure: H1: Last week in scylladb.git master (issue #270; 2025-02-23) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #270; 2025-02-23) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 0d5f5e6c9da…a6c882e4e37 range are covered. There were 137 non-merge commits from 23 authors in that period. Some notable commits: A new CQL function, set_intersection, is now available to calculate the intersection of two set values. Tablet repair can now filter by host or datacenter. The Raft implementation now limits consumption of memory for replication. Handling of TRUNCATE statements while previous TRUNCATE statements are still processing was improved. A race condition between splitting tablets of a table, and a DROP of the same table, was fixed. A case where Raft initialization loads incorrect values from disk was fixed. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/full-scan-performance-impact-on-cluster-impact-on-other-queries-and-workloads/4536 Title: Full scan performance impact on cluster, impact on other queries and workloads - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/full-scan-performance-impact-on-cluster-impact-on-other-queries-and-workloads/4536 ## Headings Structure: H1: Full scan performance impact on cluster, impact on other queries and workloads H3: Related topics ## Main Content: H1: Full scan performance impact on cluster, impact on other queries and workloads H3: Related topics Originally from the User Slack @Hyunwoo_Kim: Question about Full Scan Count Query Impact on Cluster [Situation] • Our cluster consists of 3 nodes each in DC1, DC2. • One user ran the following query to count rows in a specific table ◦ SELECT count(*) FROM keyspace.table_inference • At the same time, another user exerienced a failure when running a read query. ◦ Read from keyspace2.table_2 • ScyllaDB Version : 5.4.0 I understand that a full scan query can take a long time to return results, but I’m unsure why it caused another read query to fail. Could anyone provide insights into why this might have happend? ps. I’ve attached ScyllaDB server error logs from when the read query failed. [Follow up Question] When managing Scylladb, would decreasing range_request_timeout_in_ms or read_request_timeout_in_ms help prevent full scan queries from impacting other users? @Botond_Dénes: Full scans have impact on other queries because they generate a lot of work for the cluster, increasing the load on the nodes. Decreasing timeouts is counter-productive and leads to thrown away work. @Hyunwoo_Kim: Thank you for answering Denes! So… from an operational perspective, how can we prevent this kind of failure? I think it’s not enough to tell users that count queries are not appropriate. Or add one more API layer to read or write data into scylladb? I expect only count (full scan) query to fail, and the other queries will not be affected by full scan query. @Botond_Dénes: You can achieve that, but you need to use workload prioritization, which is available in enterprise or the upcoming source-available release. With workload prioritization you can isolate workloads from each other, so a count() query running in one priority group will not hurt other queries in another priority group. @Hyunwoo_Kim: Oh I see… Too bad we’re using version 5.4 @Robert: in 5.4.x that feature also exists @Botond_Dénes: No, system table schema was adjusted to be compatible with that of enterprise, but the feature is not implemented in open-source ScyllaDB. @avi: Once you migrate to 2025.1, you can isolate the parallel scan and give it fewer resources The table will have an additional column shares denoting how much resources to allocate to these queries --- ### Page: https://forum.scylladb.com/t/how-to-debug-possible-consistency-issues-with-a-specific-table-using-mutation-fragments/4538 Title: How to debug possible consistency issues with a specific table? Using mutation_fragments - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/how-to-debug-possible-consistency-issues-with-a-specific-table-using-mutation-fragments/4538 ## Headings Structure: H1: How to debug possible consistency issues with a specific table? Using mutation_fragments H3: Related topics ## Main Content: H1: How to debug possible consistency issues with a specific table? Using mutation_fragments H3: Related topics Originally from the User Slack @Matheus_Salvia: I’m finding that sometimes I have some consistency issues with a particular table. the operations are as follow: • 1. calculate data locally and build a struct • 2. write calculated data (single row) with CL=QUORUM • 3. read entire partition where the row was written (they are small partitions with at most 100 rows or something - rows are also light) also with CL=QUORUM I expect to always see the newly inserted row at step 3, but I find that sometimes that’s not the case, the row is not there. Both the read and the write are made from the same client, serially (read is only fired after write has returned). What steps can I take to help debug this? @avi: You can try SELECT FROM mutation_fragments() on each replica to see what it did with your data Reading mutation fragments | ScyllaDB Docs It basically reads from cache/memtable/sstables and shows you if some node missed the update @Matheus_Salvia: this is gonna be hard to pull off since the problem is only very sporadic and in production I have read in a couple of places that clock sync could be a culprit. does that make sense? if nodes and/or the client’s time drifts @avi: Check that server clocks on the clients and servers are synced yes @Matheus_Salvia: will do, thanks @Matheus_Salvia: I figured this out, wrong assumption. I trusted that queries were being made with CL=QUORUM but instead they were being made with CL=ONE. If you’re experiencing the same situation I suggest you make absolutely sure that your queries have the correct CL and then try and sync your clocks. --- ### Page: https://forum.scylladb.com/t/release-scylladb-monitoring-stack-4-9-0/4539 Title: [RELEASE] ScyllaDB Monitoring Stack 4.9.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of ScyllaDB Monitoring Stack 4.9.0 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB Enterprise and ScyllaDB Open Source, based on Prometheus … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-monitoring-stack-4-9-0/4539 ## Headings Structure: H1: [RELEASE] ScyllaDB Monitoring Stack 4.9.0 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Monitoring Stack 4.9.0 H3: Related topics The ScyllaDB team is pleased to announce the release of ScyllaDB Monitoring Stack 4.9.0 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB Enterprise and ScyllaDB Open Source, based on Prometheus and Grafana. ScyllaDB Monitoring Stack 4.9.0 supports: Versions updates for ScyllaDB Monitoring Stack 4.9.0 New Information in ScyllaDB Dashboards Keyspace Dashboard Change The new table gives a summary of all tables in the keyspace. It shows disk space usage and reads and writes information. Alternator Dashboard Change Removed the live, compaction, and cache misses from the table. The P50 is the average of get, put, update, and delete operations. There is now a P99 for each of the get, put, update, and delete operations. The new panel shows all the consistency level requests together and makes it easier to compare. General Dashboard Changes The dashboard limits the number of series for better performance, the original filter was hard to understand. The current implementation lets you choose what kind of filter you want (top-k, bottom-k or limit-k) and how many results. The introduced experimental limit-k function in Prometheus filters a given number of results, but without specifying a specific rule (like top or bottom), it makes it easier to understand the overall distribution of the values. ScyllaDB Monitoring ships with the ScyllaDB plugin. By default, it shipped with all plugin implementations, which made the repository too big. Now, only the needed implementation is shipped. For multi-cluster monitoring, it is sometimes better to use a centralized Thanos query to read from a local per-cluster Prometheus with a sidecar. To split a query calculation between, so each cluster will do its calculation and the centralized query will combine the results, we need to run a thanos query locally. As part of the work towards making user-facing dashboards clearer, while keeping the Support dashboards verbose, a new option was added that splits the dashboards into user-facing and support. To make the startup script run quicker, two changes were made. First, the regular wait interval when testing a container start was shortened, which reduced the start time of a clean installation by 80%. Second, there is an option to start the monitoring stack without waiting for application validation. While usually unnecessary, it will reduce the startup time of even a non-clean installation to a few seconds. --- ### Page: https://forum.scylladb.com/t/replacing-disks-in-running-nodes-cleanup-and-what-is-the-correct-process/4540 Title: Replacing disks in running nodes, cleanup and what is the correct process? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/replacing-disks-in-running-nodes-cleanup-and-what-is-the-correct-process/4540 ## Headings Structure: H1: Replacing disks in running nodes, cleanup and what is the correct process? H3: Related topics ## Main Content: H1: Replacing disks in running nodes, cleanup and what is the correct process? H3: Related topics Originally from the User Slack @Hans: Hey, I have a question about nodetool cleanup. We have 3 nodes in each of DC1, DC2, and we need to replace a disk on every nodes in DC1, my understanding is that we have to • Decommission the node, replace the disk, recreate the scylla dir, and join it as a new node. • Once the new node is added to the cluster, run nodetool cleanup on all existing nodes. My question is: Since cleanup just removes outdated data from SSTables, is it sufficient to run it once on each node after all rolling replacements are completed? Or do we need to run cleanup on existing nodes after each individual node replacement? Would appreciate any insights! Thanks @stewart: Replacing a disk and doing a nodetool cleanup are different operations. Which one are you trying to do? @Samuel_Hameau: you may just stop scylla on your node ; copy the disk content to your new disks ; replace disks ; restart scylla with the new disks @Hans: Thanks, everyone! Stewart gave me the similiar advice as Samuel said. Just drain → turn off scylladb server → replace disk → turn the scylla server back on. This approach is much better than decommissioning and rejoining the server to the cluster. --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-2-5/4543 Title: [RELEASE] ScyllaDB Enterprise 2024.2.5 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2024.2.5, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2024.2 Feature (Short Term Support) Release. Related Links Get ScyllaDB En… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-2-5/4543 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2024.2.5 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2024.2.5 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2024.2.5, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2024.2 Feature (Short Term Support) Release. The following issues are fixed in this release (with an open-source reference, if available): --- ### Page: https://forum.scylladb.com/t/why-use-update-entry-in-view-updates/4545 Title: Why use update_entry in view_updates? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: master #Cluster size: 1 os (RHEL/CentOS/Ubuntu/AWS AMI): centos I noticed that when a view update is generated, it determines whether to use create_entry or update_entry based… Language: en Canonical URL: https://forum.scylladb.com/t/why-use-update-entry-in-view-updates/4545 ## Headings Structure: H1: Why use update_entry in view_updates? H3: Related topics ## Main Content: H1: Why use update_entry in view_updates? H3: Related topics Installation details #ScyllaDB version: master #Cluster size: 1 os (RHEL/CentOS/Ubuntu/AWS AMI): centos I noticed that when a view update is generated, it determines whether to use create_entry or update_entry based on the existing and update data of the base. When both existing and update exist and are alive, update_entry is called. update_entry calculates the difference between existing and update. And create_entry will directly create the entire update data. Suppose there is a 3-node 2-replica cluster, where nodes A and B are two replica nodes of a certain data. The data of replica A node is 1 and 2, and the data of replica B node is 1 and 3. Then executing repair on node A will generate a new data of 1, 2 and 3. Then update_entry will be used when generating a view. update_entry will calculate the diff and get 3. At this time, the view table of node A previously had data 1 and 2, plus the newly generated data 3, so the data read is 1, 2 and 3. But assume that the base table data 1 and 2 of node A fail to propagate the view for some reason, then we will read 3 when reading the view table. Assume that the index table has not executed the repair of this token during this period. Because the consistency level of reading the index table is ONE in Alternator. Read repair will not be triggered at consistency level ONE. Since the index table of the scylladb database only supports eventual consistency in Alternator, this is not a problem. However, can we replace update_entry with create_entry? This can reduce the generation of these intermediate data. What is the purpose of using update_entry ? To save some space overhead? Hello @nyh and @Botond_Denes, can you help me answer this? Thanks! There are several separate issues here, I’m not sure if this forum is the best place for it (the mailing list, scylla-dev@googlegroups.com, is probably a better venue for prolonged discussion threads). First of all, the main reason why update_entry exists separately from a delete_entry/create_entry pair is because a delete and then create with the same timestamp will not work - the delete would win when the timestamp ties, and the data will disappear instead of being updated. This is why we need need this separate update_entry() function. The second issue is why update_entry() needs to have this “optimization” where if we believe that the view already contains some data, we don’t write it again. In the specific case you mentioned, update_entry() (view row key is known and hasn’t changed), I think you’re right and the optimization isn’t necessary, although to be honest I don’t remember every detail so please use “git blame” on the relevant line of code and see if comes from a commit that explained why this optimization was added. A third question is why update_entry() and other code can assume it begins with the view and base replicas having matching data - so we can read from the base table to decide what to do to the view table, assuming we know what’s there. Well, as I already noted elsewhere, we often don’t have any way NOT to make this assumption. Consider the case where the view’s key is not the same as the base’s - for example, the base key is p and the view key is p, x. Now imagine a write setting SET x=3 WHERE p=7. We need to not only write the new row in the view (7,3) we also need to delete the previous row, say (7,2) (if the previous value of x was 2) - but how will we know the previous value of x was 2? We need to read it from the base table, and then assume that the view table matches it and has this row (7.2) and delete that row - not any other row. So Scylla needs to assume that “paired” base and view replicas match up. The fact this assumption can become wrong on unrepaired tables and because of other problems is not lost on us, but we don’t have good documentation of when exactly this can happen, or whether repairs of base and/or view tables can fix some of these problems - or whether it can fix all of them. This definitely needs more work, but I’m afraid that we might not be able to fix all these more and more obscure bugs without completely changing the materialized views algorithm and dropping the “paired replicas” approach. Finally, you are right that reading an unrepaired table with a consistency of ONE - which is the only way to read a GSI in Alternator - is a problem and we will never enjoy read-repair. I don’t know what to do about this. In the past, we actually believed that the lack of read-repair was a good thing (see Disable read-repair in materialized-view tables · Issue #3933 · scylladb/scylladb · GitHub) but I no longer believe this to be the case. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-84-2025-02-28/4546 Title: Last week in scylla-cluster-tests.git master (issue #84; 2025-02-28) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 08dba227…8287d6e1 range are covered. There were 27 non-merge commits from 12 authors in t… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-84-2025-02-28/4546 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #84; 2025-02-28) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #84; 2025-02-28) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 08dba227…8287d6e1 range are covered. There were 27 non-merge commits from 12 authors in that period. Some notable commits: SCT now allows specifying different seeds for parallel nemesis, useful for testing scenarios with multiple parallel nemesis running. A new test with 90% disk utilization and 500 small tables has been added. HDR histograms are now supported in the Latte tool, and an example job has been added. Client encryption was disabled for all Manager tests and has now been enabled for all Manager tests except ubuntu24-manager-sanity and ubuntu24-manager-upgrade. To address resource leftovers, the post behavior of instances is now set to destroy by default. If you need to keep them, set keep or keep-on-failure in the pipeline configuration. The monitoring branch has been switched to 4.9 and new images have been set for AWS and GCE configs. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/error-with-where-in-x-y-z-query-with-a-cartesian-product-over-a-maximum-size/4554 Title: Error with WHERE…IN(x,y,z) query with a cartesian product over a maximum size - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/error-with-where-in-x-y-z-query-with-a-cartesian-product-over-a-maximum-size/4554 ## Headings Structure: H1: Error with WHERE…IN(x,y,z) query with a cartesian product over a maximum size H3: Related topics ## Main Content: H1: Error with WHERE…IN(x,y,z) query with a cartesian product over a maximum size H3: Related topics Originally from the User Slack @Chris_Spence: Hello, I have a question regarding a limit to the parameters for a WHERE…IN(x,y,z) I am getting an error regarding a cartesian product being over a limit of 100. It is a very simple query and a simple table Table is similar to and the query SELECT id, ColumnA, ColumnB, ColumnC FROM table WHERE id IN (x, y, z ….) There are 105 strings in the array for the where clause so that matches the errors message above. I’ve not come across any published limit to an IN clause for scylla or cassandra and compared to SQL it seems very low I would expect to be able to pass a lot more values in an IN clause. An example of the value passed is E7E5079CC9444E4B8AC6FD00981B3DDB strings all of uniform size. Can you help point me in the right direction to resolve this or explain what is causing the issue or if this is expected behaviour please? Thank you @Marko_Ćorić: you can set for example max_clustering_key_restrictions_per_query : 250 via configuration, but it’s better to review your database scheme. ofc, that required full cluster restart @Chris_Spence: Thanks @Marko_Ćorić if it is an expected limit then we will review the schema and re-organise data. Thanks --- ### Page: https://forum.scylladb.com/t/scylladb-prometheus-graphana-dashboard-using-our-own-setup/4555 Title: ScyllaDB Prometheus Graphana dashboard, using our own setup - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-prometheus-graphana-dashboard-using-our-own-setup/4555 ## Headings Structure: H1: ScyllaDB Prometheus Graphana dashboard, using our own setup H3: Related topics ## Main Content: H1: ScyllaDB Prometheus Graphana dashboard, using our own setup H3: Related topics Originally from the User Slack @Naveed_Khan: Hello, We are looking for basic simple Graphana dashboard for ScyllaDB. Any pointers? Note that we are not using the Scylla monitoring stack. We have a Prometheus Graphana setup of our own and we’d prefer to just hook it up with the exporter. @Felipe_Cardeneti_Mendes: Can’t you simply deploy the monitoring stack somewhere else and grab the panels and metrics you’d want? @tzach: https://monitoring.docs.scylladb.com/stable/ ScyllaDB Monitoring Stack | ScyllaDB Docs --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-271-2025-03-02/4556 Title: Last week in scylladb.git master (issue #271; 2025-03-02) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the a6c882e4e37…0343235aa26 range are covered. There were 92 non-merge commits from 22 authors in that per… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-271-2025-03-02/4556 ## Headings Structure: H1: Last week in scylladb.git master (issue #271; 2025-03-02) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #271; 2025-03-02) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the a6c882e4e37…0343235aa26 range are covered. There were 92 non-merge commits from 22 authors in that period. Some notable commits: The bundled node_exporter package was updated to version 1.9.0 to fix some vulnerabilities. The SELECT DISTINCT statement now rejects the PER PARTITION LIMIT clause, since it is redundant. The tablet-mon.py, used to visualize tablet operations, improves merge and split support. Tables creates with tablets now have topology that is better prepared for immediate ingestion. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/how-do-i-solve-node-stuck-in-dj-state-it-was-destroyed-during-join/4557 Title: How do I solve node stuck in DJ state (it was destroyed during join)? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/how-do-i-solve-node-stuck-in-dj-state-it-was-destroyed-during-join/4557 ## Headings Structure: H1: How do I solve node stuck in DJ state (it was destroyed during join)? H3: Related topics ## Main Content: H1: How do I solve node stuck in DJ state (it was destroyed during join)? H3: Related topics Originally from the User Slack @Andrey_Kojevnikov: Hi. I’ve got a node stuck in DJ state as it was destroyed during join. There is no way to conveniently replace or remove it, as it is both “exists” and “bootstrapping”. Is there a way to fix this? Seems like the only relevant stuff in docs is under “Manual Recovery Procedure”, which I really would like to avoid. The answer is “upgrade to 6.2.3” https://github.com/scylladb/scylladb/issues/20082 GitHub: node could stay in gossip forever if node start gossiping but topology coordinator didn’t get request to bootstrap · Issue #20082 · scylladb/scylladb --- ### Page: https://forum.scylladb.com/t/scylla-is-killed-due-to-space-issue/4564 Title: Scylla is killed due to space issue - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 6.2 #Cluster size: 3 nodes os (RHEL/CentOS/Ubuntu/AWS AMI): amzn2 ec2 Hi Folks, We have a scylla operator that is deployed to k8s using helm, but we keep getting disk pressure… Language: en Canonical URL: https://forum.scylladb.com/t/scylla-is-killed-due-to-space-issue/4564 ## Headings Structure: H1: Scylla is killed due to space issue H3: Related topics ## Main Content: H1: Scylla is killed due to space issue H3: Related topics Installation details #ScyllaDB version: 6.2 #Cluster size: 3 nodes os (RHEL/CentOS/Ubuntu/AWS AMI): amzn2 ec2 Hi Folks, We have a scylla operator that is deployed to k8s using helm, but we keep getting disk pressure tain and pods are evicted due to following error: Warning Evicted 7m9s kubelet The node was low on resource: ephemeral-storage. Threshold quantity: 857733132, available: 4252Ki. Container scylla was using 152416Ki, request is 0, has larger consumption of ephemeral-storage. Container scylla-manager-agent was using 88Ki, request is 0, has larger consumption of ephemeral-storage. Container scylladb-api-status-probe was using 8Ki, request is 0, has larger consumption of ephemeral-storage. Container scylladb-ignition was using 384Ki, request is 0, has larger consumption of ephemeral-storage. Normal Killing 7m9s kubelet Stopping container scylla I have given 256mb ephemeral space to the pods, what I am missing, thank you for your help! Your ScyllaDB pods are being evicted because they are exceeding the ephemeral storage allocation you’ve set (256MB). The logs clearly state: 256MB ephemeral storage is extremely limited for running ScyllaDB, even for just logs and basic container overhead. You have two practical solutions here: Adjust based on your environment, number of nodes, and expected log verbosity. --- ### Page: https://forum.scylladb.com/t/warning-message-with-java-driver-error-while-opening-new-channel/4567 Title: Warning message with Java driver - Error while opening new channel - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/warning-message-with-java-driver-error-while-opening-new-channel/4567 ## Headings Structure: H1: Warning message with Java driver - Error while opening new channel H3: Related topics ## Main Content: H1: Warning message with Java driver - Error while opening new channel H3: Related topics Originally from the User Slack @stewart: Hi Guys, Every time we replace a ScyllaDB node, the client generates the following warning log. In this case, it doesn’t seem to cause any major issues with the service, and restarting the service makes the warning log disappear. However, we would like the warning log to disappear immediately when the node is replaced. How should we change the ScyllaDB Datastax Java driver settings? we are using scylladb 4.6.3. @Felipe_Cardeneti_Mendes: Apparently could be due to https://datastax-oss.atlassian.net/browse/JAVA-3005 and its ramifications. https://github.com/apache/cassandra-java-driver/pull/1604/files seems to be in 4.15 - I believe (though no java expert here) you should be able to force a metadata refresh, if this solves it for you then that would help. @stewart: I just found that feature as well, but I’m not sure if it has been bumped in the ScyllaDB Java driver. How does the ScyllaDB Java driver import the functionality from the DataStax Cassandra Java driver? @Felipe_Cardeneti_Mendes: we should be at 4.18.x https://forum.scylladb.com/t/release-scylladb-java-driver-4-18-0-2/4490 - so latest should it, considering upstream never dropped it in between @stewart: Then, should I upgrade the driver version to 4.18.0.2? @Felipe_Cardeneti_Mendes: Yes but which driver version are you currently at? @stewart: so many microservices using scylla, so I need to check it one by one lol @Felipe_Cardeneti_Mendes: right, and always remember 3.x->4.x is a breaking change though AFAICT 3.x shouldn’t be affected @stewart: thx felipe! --- ### Page: https://forum.scylladb.com/t/release-scylladb-rust-driver-1-0-0/4568 Title: [RELEASE] ScyllaDB Rust Driver 1.0.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce ScyllaDB Rust Driver 1.0.0, an asynchronous CQL driver for Rust, optimized for Scylla, but also compatible with Apache Cassandra! Some interesting statistics: over 3.162k down… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-rust-driver-1-0-0/4568 ## Headings Structure: H1: [RELEASE] ScyllaDB Rust Driver 1.0.0 H2: Stabilization H2: Changes H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Rust Driver 1.0.0 H2: Stabilization H2: Changes H3: Related topics The ScyllaDB team is pleased to announce ScyllaDB Rust Driver 1.0.0, an asynchronous CQL driver for Rust, optimized for Scylla, but also compatible with Apache Cassandra! Some interesting statistics: 1.0 has a special meaning for Rust crates - it means that the crate is considered “API-stable” - however there is no specific definition of what it means to be “API-stable”. So what does it mean to us, and why did we release this version? Up until now, we introduced breaking changes whenever we felt we could improve the API this way (sometimes we kept the old API for a bit, and provided a migration guide). This caused the vast majority of our releases to have to bump the major version number - we didn’t reach minor version greater than “2” for a few years now. This is not a great situation for our users, because it means it is basically never possible to update the driver without adjusting their codebases. At the same time we know that we can’t (and don’t want to) freeze the API forever - there are still many improvements we can make, and the databases (Scylla / Cassandra) are always changing, which sometimes requires breaking changes on our part (an example of this is Tablets feature in Scylla, which required us to change our load balancing APIs). This means we won’t stay on 1.x forever. We decided to stabilize the API by providing a guarantee regarding how long a breaking change won’t occur, which will allows users to better allocate time for dealing with breakage. For the 1.0, we won’t release 2.0 earlier than 9 months after 1.0 (but we may release it later). Until then we will release minor (1.x) versions. In 1.x versions we may of course introduce new APIs, and deprecate old ones, but there will be no breaking changes. Two exceptions that may force us to release major versions quicker are: After releasing new major version (e.g. 2.0) we will continue to provide bugfixes (but no new features) for the previous major version (e.g. 1.0). The exact duration of such bugfix support will be provided after the new major version is released. Additionally, 1.0 release signifies that we view the driver as production ready. It has been for a long time now, but the 0.x version number may have been suggesting otherwise to some people, and we don’t want to send mixed signals. New features / enhancements: API cleanups / better types: Internal API cleanups/refactors: CI / developer tool improvements: Congrats to all contributors and thanks everyone for using our driver! The source code of the driver can be found here: The official crates.io registry entry is here: Thank you for your attention, please do not hesitate to contact us if you have any questions, issues, feature requests, or are simply interested in our driver! Contributors since the last release: --- ### Page: https://forum.scylladb.com/t/whats-the-best-way-to-split-a-cluster-thats-on-two-data-centers-into-two-clusters/4570 Title: What's the best way to split a cluster that's on two data centers into two clusters? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/whats-the-best-way-to-split-a-cluster-thats-on-two-data-centers-into-two-clusters/4570 ## Headings Structure: H1: What's the best way to split a cluster that's on two data centers into two clusters? H3: Related topics ## Main Content: H1: What's the best way to split a cluster that's on two data centers into two clusters? H3: Related topics Originally from the User Slack @Dmitriy_Karpman: I want to split a scylla cluster with two datacenters into two clusters (where each datacenter becomes its own cluster and no longer knows about each other) — what’s the best way to do this? @avi: There’s no way to do it. Create a new cluster and copy the data. --- ### Page: https://forum.scylladb.com/t/release-scylla-operator-v1-16-0/4572 Title: [RELEASE] Scylla Operator v1.16.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of Scylla Operator 1.16.0. Scylla Operator is an open-source project that helps users run ScyllaDB on Kubernetes. The Scylla Operator manages ScyllaDB clusters dep… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-operator-v1-16-0/4572 ## Headings Structure: H1: [RELEASE] Scylla Operator v1.16.0 H1: Notable changes H2: Coming soon: Multi-Datacenter Deployments H1: Supported versions H1: Upgrade instructions H2: Getting started with Scylla Operator H2: Related Links H3: Related topics ## Main Content: H1: [RELEASE] Scylla Operator v1.16.0 H1: Notable changes H2: Coming soon: Multi-Datacenter Deployments H1: Supported versions H1: Upgrade instructions H2: Getting started with Scylla Operator H2: Related Links H3: Related topics The ScyllaDB team is pleased to announce the release of Scylla Operator 1.16.0. Scylla Operator is an open-source project that helps users run ScyllaDB on Kubernetes. The Scylla Operator manages ScyllaDB clusters deployed to Kubernetes and automates tasks related to operating a ScyllaDB cluster, like installation, vertical and horizontal scaling, as well as rolling upgrades. Scylla Operator 1.16.0 improves stability and brings new features. As with all of our releases, all API changes are backward compatible. In preparation for future introduction of automated multi-datacenter deployments, we’ve introduced a bunch of new v1alpha APIs subject to future development. (#2215, #2278, #2271, #2272): Stay tuned for future updates on readiness of multi-datacenter deployments for general use, including quickstarts/examples and documentation. For more changes and details, check out the GitHub release notes. Upgrading from v1.15.x with kubectl apply doesn’t require any extra action, just take the manifest from v1.16.0 tag and substitute the released image. Using helm requires a mandatory manual step for every release because helm can’t handle CRDs updates. For details, see our upgrade documentation. We’ll welcome your feedback! Feel free to open an issue or reach out on the #scylla-operator channel in Scylla User Slack. The Scylla Operator Team --- ### Page: https://forum.scylladb.com/t/is-delete-insert-to-update-immutable-data-in-rows-a-recommended-approach/4576 Title: Is DELETE + INSERT to update immutable data in rows a recommended approach? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: In the leaderboard modeling blog, it’s mentioned that updating fields that are part of the primary key (whether partition or clustering keys) requires deleting the existing row and inserting a modified version. However,… Language: en Canonical URL: https://forum.scylladb.com/t/is-delete-insert-to-update-immutable-data-in-rows-a-recommended-approach/4576 ## Headings Structure: H1: Is DELETE + INSERT to update immutable data in rows a recommended approach? H3: Related topics ## Main Content: H1: Is DELETE + INSERT to update immutable data in rows a recommended approach? H3: Related topics In the leaderboard modeling blog, it’s mentioned that updating fields that are part of the primary key (whether partition or clustering keys) requires deleting the existing row and inserting a modified version. However, since DELETE operations introduce tombstones, they can contribute to read latency issues, increased compaction overhead, and potential performance degradation over time. Additionally, frequent DELETE + INSERT operations could lead to bloated SSTables, impacting disk space and query efficiency. Given these factors, in a large-scale environment is it advisable to rely on an approach which uses delete + insert very commonly? leaderboard modeling blog - scylladb .com/2024/09/03/model-game-leaderboards-scylladb/ You are right that delete+insert gives quite some work to the database, but so does any workload which has a roughly constant-size working-set. If your dataset is not growing, it means that your writes are either overwrites or delete+insert. Both are challenging for the database, just in different ways. All that said, the database is designed to cope with this so such a workload on its own will not be a problem. --- ### Page: https://forum.scylladb.com/t/release-python-driver-3-29-0/4577 Title: [RELEASE] Python Driver 3.29.0 - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Driver Release Summary: 3.29.0 Release Link: Release 3.29.0 · scylladb/python-driver · GitHub Release Summary: Query Optimization: All queries to system.local now use a WHERE clause for improved efficiency. CI/CD Up… Language: en Canonical URL: https://forum.scylladb.com/t/release-python-driver-3-29-0/4577 ## Headings Structure: H1: [RELEASE] Python Driver 3.29.0 H3: Release Summary: H3: Related topics ## Main Content: H1: [RELEASE] Python Driver 3.29.0 H3: Release Summary: H3: Related topics Driver Release Summary: 3.29.0 Release Link: Release 3.29.0 · scylladb/python-driver · GitHub Python 3.12 Compatibility Fixes: Error Handling Improvement: Documentation Updates: --- ### Page: https://forum.scylladb.com/t/release-python-driver-3-29-2/4578 Title: [RELEASE] Python Driver 3.29.2 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: Driver Release Summary: 3.29.2 Release Link: Release 3.29.2 · scylladb/python-driver · GitHub Build & CI/CD Improvements Simplified the build pipeline to reduce confusion. (#425) Modified build-push workflow to push c… Language: en Canonical URL: https://forum.scylladb.com/t/release-python-driver-3-29-2/4578 ## Headings Structure: H1: [RELEASE] Python Driver 3.29.2 H3: Related topics ## Main Content: H1: [RELEASE] Python Driver 3.29.2 H4: Build & CI/CD Improvements H4: CQL Engine Fixes H4: Environment & Compatibility Updates H4: Testing & Extension Updates H3: Related topics Driver Release Summary: 3.29.2 Release Link: Release 3.29.2 · scylladb/python-driver · GitHub --- ### Page: https://forum.scylladb.com/t/arm-vs-intel-instance-types-on-aws-and-performance-differences-when-benchmarking/4579 Title: ARM vs. Intel instance types on AWS and performance differences when benchmarking - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/arm-vs-intel-instance-types-on-aws-and-performance-differences-when-benchmarking/4579 ## Headings Structure: H1: ARM vs. Intel instance types on AWS and performance differences when benchmarking H3: Related topics ## Main Content: H1: ARM vs. Intel instance types on AWS and performance differences when benchmarking H3: Related topics Originally from the User Slack @stewart: Hi all, I discovered an interesting phenomenon while benchmarking ScyllaDB. I provisioned separate ScyllaDB clusters using two different instance types: r6i.2xlarge and r6g.2xlarge, and ran benchmarks on each. Interestingly, the r6i (Intel) instance type showed better performance. (throughput, latency) I’m not sure why the Intel-based instance outperforms the ARM-based one. Could someone help explain the possible reasons for this? and also tested with r6gd.2xlarge and i4i.2xlarge, i4i.2xlarge got better performance @dor: Yes, it’s consistent with what we know. Intel is stronger than arm @avi: The ARM instances used by AWS have relatively small caches @stewart: you mean small L1 cache? Can you explain the difference in more detail? @Robert: which distribution and version of linux did You take for benchmarking? @stewart: scylladb version : 5.4.9 distribution : 5.10.224-212.876.amzn2.x86_64 tested on kubernetes v1.29. cc. @Robert @Robert: uch, need to figure out which kernel version it uses, because on ubuntu20 vs ubuntu22 there is a cache issue on graviton. U20 doesn’t use all available L1… if You have a time to play with, simple execute lscpu on Amazon linux vs Ubuntu22 vs Ubuntu20. On the other hand it will be double weird if amazon distro doesn’t fully operate with amazon cpus @stewart: i’ll check it on i4i.4xlarge @Robert: and how about graviton? --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-85-2025-03-07/4583 Title: Last week in scylla-cluster-tests.git master (issue #85; 2025-03-07) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 528b0a52…df2b91cf range are covered. There were 38 non-merge commits from 15 authors in t… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-85-2025-03-07/4583 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #85; 2025-03-07) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #85; 2025-03-07) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 528b0a52…df2b91cf range are covered. There were 38 non-merge commits from 15 authors in that period. Some notable commits: Nemeses target only data nodes by default. Since most also apply to zero-token nodes (including config and schema changes), both data and zero-token nodes are now explicitly set as nemesis targets. A new job was added with four zero-token nodes to run all supported nemeses. All pip dependencies were updated. A templating mechanism for monitoring dashboards was introduced, keeping SCT in sync with changes in the Scylla Monitoring repo. Gemini now supports log compression using ZSTD, allowing longer runs without running out of disk space. Sanity jobs and 90% utilization tests now use simulated racks for a more realistic testing environment. Due to a Grafana renderer issue, monitoring was reverted to version 4.8. Performance test step durations can now be configured using the perf_gradual_step_duration parameter, making it easier to modify predefined step lengths. This was used in a new Java driver performance test. The modify_table_twcs_window_size nemesis was refactored and now tests TWCS borders across wider ranges. Performance regression tests now respect the pre_create_keyspace parameter, allowing easier configuration of the initial number of tablets. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/caching-for-performance-masterclass-slides-and-recording/4587 Title: Caching for Performance Masterclass slides and recording - University and Training - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/caching-for-performance-masterclass-slides-and-recording/4587 ## Headings Structure: H1: Caching for Performance Masterclass slides and recording H3: Related topics ## Main Content: H1: Caching for Performance Masterclass slides and recording H3: Related topics Originally from the User Slack @Jenagan_Sivakumar: Hi I attended a caching masterclass session last week. It was great, but I had to run to another meeting so I didn’t catch the last of it. I was told it would be uploaded, is there anywhere I can find it? Also, I did the practice test, and I am waiting for the certification and to know if I am eligible to win the goodie bag . I was wondering what happens after the mini exam? Do I get notified by email? @Ellen_Trieu: Hi @Jenagan_Sivakumar You can view the record and slides for the Caching Masterclass here. We’ll send an email about the certificate the swag after March 18th. Thank you for your participation! --- ### Page: https://forum.scylladb.com/t/read-fails-for-one-node-in-the-cluster-no-response-received-on-gcp/4589 Title: Read fails for one node in the cluster, no response received on GCP - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/read-fails-for-one-node-in-the-cluster-no-response-received-on-gcp/4589 ## Headings Structure: H1: Read fails for one node in the cluster, no response received on GCP H3: Related topics ## Main Content: H1: Read fails for one node in the cluster, no response received on GCP H3: Related topics Originally from the User Slack @Chaitanya_Tondlekar: Hello All, Seeing issues on scylla one of the node where queued reads & reads failed happen for only one node in cluster. This got auto resolved but didnt understand why did it happened. Checked disk level IOPS as well but didnt find any. We are getting errors on application level at same time : logs of the node. Also there are spikes at local read error & C ++ exception at same time @avi: It looks like you have a large+hot partition. File an issue, slack isn’t a good place to debug this. @Chaitanya_Tondlekar: https://github.com/scylladb/scylladb/issues/22506 GitHub: Seeing spikes in queued reads & reads failed on one of the node. · Issue #22506 · scylladb/scylladb @avi this command is ran when we started facing this issue, @avi: Need to increase sampling capacity (try -k 1000) until the “+/-” column is small compared to the Count column @Chaitanya_Tondlekar: even fro -k 300 was giving same error. TopK count (-k) option must be smaller then the summary capacity (-s) @avi: Increase -s, not -k. My mistake. @Chaitanya_Tondlekar: will check. @Guy: @Chaitanya_Tondlekar, did this solve the issue? @Chaitanya_Tondlekar: It seems the maintenance is getting triggered for nodes from GCP level. @avi: I think you can tell them not to do it @Chaitanya_Tondlekar: @avi @Guy Need your help on this. We initially thought that the timings were matching with the maintenance which was happening but even after stopping maintenance we are seeing queues reads and reads failed on couple of nodes. Logs of the node where reads failed are printing memory resources: timed out with semaphore. @Botond_Dénes is it possible to help here ? @avi: Did you check for hot partitions? @Chaitanya_Tondlekar: its happening now @avi. It happens majorly for two nodes only. @avi what can be done over here ? tried with @Botond_Dénes: The disk on the node is not keeping up with requests, so you have a lots of queued reads and timeouts. @Chaitanya_Tondlekar: yes, but it keeps on happening all of sudden. and only on 2 nodes. It spikes on queued reads and reads failed. @Botond_Dénes: You possibly have a spike in load (check requests count for replica metrics on the dashboard) or possibly a spike in request for certain partitions @Chaitanya_Tondlekar: partition hits seems same on all nodes. Even on the nodes where queued reads and reads failed are there. There’s no sudden increased in QPS. I have shared output for top paritions from the node which are affected. Can we get any help from there ? @Botond_Dénes @Botond_Dénes: partition load seems pretty uniform do partitions vary in size? @Chaitanya_Tondlekar: how to get the partition sizes ? @Botond_Dénes: is there anything in the logs about large partitions how about system.large_partitions? @Chaitanya_Tondlekar: Logs are printing only those which we have posted here. logs are only printing for semaphores @Botond_Dénes anything else we can check now ? As we are seeing this and can check things activiely @Botond_Dénes: check CPU load metric with shard resolution do you see some shards pegged at 100% CPU load? @Chaitanya_Tondlekar: check CPU load metric with shard resolution – > where we can get this ? we are checking load metrics from detailed dashboard and seem load is even. @Botond_Dénes: Change by Instance to by Shard @Chaitanya_Tondlekar: shard 18 is having 100% @Botond_Dénes: Is this the shard that is logging the semaphore timeotus? I see both reports from above are from shard 18 So you do have some partitions which are hotter they live on shard 18 @Chaitanya_Tondlekar: Got it @Botond_Dénes how do we solve hot partition issues? @Botond_Dénes: Two ways to solve it: • change data model to split the partitions • scale up/out so hopefully the hot partitions are spread between multiple shards/nodes there is a 3rd way: chance clients so they don’t hammer a subset of partiitons – this is not always possible @Chaitanya_Tondlekar: chance clients didnt understand @Botond_Dénes: change clients – change application code @Chaitanya_Tondlekar: Oh got it @avi: The capacity for nodetool toppartitions is too low, use the -s parameter. @Chaitanya_Tondlekar: Yes we have used -s only @Botond_Dénes Is there any way we can increase the value for reader_concurrency_semaphore in scylla.yaml ? @Botond_Dénes: No way to increase the limit but it would be pointless even if it would be possible. If the shard is overloaded, any limit will be reached eventually, because client sends more requests than what the node can retire, so the queue keeps on growing forever. @Chaitanya_Tondlekar: Can you elaborate the third approach of changing clients ? How we can do that ? @Botond_Dénes: I cannot give you concrete advice, I don’t know anything about your application. I just mentioned the possibility – sometimes it may be possible to change the application to not request some partition much more often than other partitions. This may not always be possible e.g. if the requests are generated by users who you can’t control and some partitions are just more popular than others. --- ### Page: https://forum.scylladb.com/t/some-bad-alloc-on-6-0-6-1/4592 Title: Some bad_alloc on 6.0/6.1 - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 6.1.5-0.20250119.c84780618297 #Cluster size: 150 To on 10 nodes os (RHEL/CentOS/Ubuntu/AWS AMI): Ubuntu 22.04 Hi, since I have upgrade from scylla 5.4 to 6.0, i have ~20/min b… Language: en Canonical URL: https://forum.scylladb.com/t/some-bad-alloc-on-6-0-6-1/4592 ## Headings Structure: H1: Some bad_alloc on 6.0/6.1 H3: Related topics ## Main Content: H1: Some bad_alloc on 6.0/6.1 H3: Related topics Installation details #ScyllaDB version: 6.1.5-0.20250119.c84780618297 #Cluster size: 150 To on 10 nodes os (RHEL/CentOS/Ubuntu/AWS AMI): Ubuntu 22.04 Hi, since I have upgrade from scylla 5.4 to 6.0, i have ~20/min batch insert queries failing with “bad_alloc” reason (while between 5k and 15k queries beside are ok). Upgrading to 6.1 did not helped, neither reducing the size of batches. I did not encounter this problem with v5.4, and I have no idea of which node is generating this kind of error. I have tried to add a few gb of memory whithout any luck. How much memory do you have per node? We usually recommend 8GB per shard - so if you have 150TB on 10 nodes, thus 15TB per node, the minimum memory on each node should be ~160GB. The recommended number of vCPUs in that case is ~20 per node. Thanks for your reply Initial setup was 125 Go with 4 vcpu, that went fine with 5.4 ; i tried to raise it up to 175G on 3 nodes but without any significant win. Payload is now ~10To per node, still 10 nodes. --- ### Page: https://forum.scylladb.com/t/time-window-compaction-strategy-ttl-number-of-windows-and-performance/4594 Title: Time Window Compaction Strategy, TTL, number of windows and performance - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/time-window-compaction-strategy-ttl-number-of-windows-and-performance/4594 ## Headings Structure: H1: Time Window Compaction Strategy, TTL, number of windows and performance H3: Related topics ## Main Content: H1: Time Window Compaction Strategy, TTL, number of windows and performance H3: Related topics Originally from the User Slack @Daria_Fedorova: I have a question about twcs_max_window_count config. How will scylla be affected if we raise this limit ? We used to have no TTL at all with TWC strategy and now I want to set TTL = 3 years with window size = 1 week resulting in 156 windows. @Felipe_Cardeneti_Mendes: See https://github.com/scylladb/scylladb/issues/6923 The more windows you have, the more SSTables you have, adding more memory pressure. 156 may be fine, but read in particular https://github.com/scylladb/scylladb/issues/6923#issuecomment-1016592469 from @raphaelsc, which is the reason why it is currently hardcoded at 50. GitHub: Too many windows with default Time Window Compaction Strategy configuration · Issue #6923 · scylladb/scylladb @raphaelsc: if your queries touch many windows, you’d rather stay with a small number of windows for better perf. we have had improvements around single-key reads, where only one window is consumed at a time, reducing significantly memory usage. but more windows still can yield higher read amplification, decreasing throughput. if you do full scans (or think you may need at some point), I’d stay with a small number of windows, around 20. example of queries that might benefit from optimization aforementioned: select * from partition_key=X LIMIT 1; select * from partition_key=X where timestamp_clustering_key > some_timestamp; @Daria_Fedorova: Thank you for reply we try to read in weeks , year&week number are in in the key for convenience --- ### Page: https://forum.scylladb.com/t/introducing-demo-ui-simple-way-to-run-scylladb-demos-and-pocs-in-the-cloud/4595 Title: Introducing Demo UI: Simple way to run ScyllaDB DEMOs and POCs in the cloud - Announcements - ScyllaDB Community NoSQL Forum Meta Description: We’ve just published Demo UI, a user interface for our existing 1M ops/sec GitHub repository! Originally, this repository focused on spinning up pre-defined high-performance ScyllaDB workloads using Terraform. Now, with… Language: en Canonical URL: https://forum.scylladb.com/t/introducing-demo-ui-simple-way-to-run-scylladb-demos-and-pocs-in-the-cloud/4595 ## Headings Structure: H1: Introducing Demo UI: Simple way to run ScyllaDB DEMOs and POCs in the cloud H3: Related topics ## Main Content: H1: Introducing Demo UI: Simple way to run ScyllaDB DEMOs and POCs in the cloud H3: Related topics We’ve just published Demo UI, a user interface for our existing 1M ops/sec GitHub repository! Originally, this repository focused on spinning up pre-defined high-performance ScyllaDB workloads using Terraform. Now, with Demo UI, it’s easier and customizable. You can configure your cluster size (e.g., 6 nodes), hardware type (e.g., i4i.2xlarge), and workload (read/write ops per second) for quick PoCs and demos. All you need is Docker installed on your machine and an AWS account to get started. Try it out on ScyllaDB Cloud or host yourself in AWS and see how ScyllaDB scales in real-time! Follow the instructions in readme: GitHub - scylladb/1m-ops-demo: Set up custom DEMOs and PoCs with ScyllaDB in the cloud --- ### Page: https://forum.scylladb.com/t/disk-space-usage-after-deleting-tables-how-can-i-verify-data-was-actually-removed/4599 Title: Disk space usage after deleting tables, how can I verify data was actually removed? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/disk-space-usage-after-deleting-tables-how-can-i-verify-data-was-actually-removed/4599 ## Headings Structure: H1: Disk space usage after deleting tables, how can I verify data was actually removed? H3: Related topics ## Main Content: H1: Disk space usage after deleting tables, how can I verify data was actually removed? H3: Related topics Originally from the User Slack @Matheus_Salvia: Hi, nodetool status is reporting around 3.3TB of data in a node, but df -h shows around 5TB used in the disk. I suspect there are some tables that I deleted last year that didn’t get actually removed. how can I figure this out? @Botond_Dénes: You can list all the tables scylla knows about by querying system_schema.tables, then compare this with what you have on disk. In data directory, we have the following structure: /${keyspace-name}/${table-name}-${table-id} Look for orphan directories, they may still hold data in the snapshot subdir, we don’t remove snapshots when removing a table. Maybe nodetool listsnapshots is also able to list these. @Matheus_Salvia: thanks --- ### Page: https://forum.scylladb.com/t/release-scylladb-monitoring-stack-4-9-2/4600 Title: [RELEASE] ScyllaDB Monitoring Stack 4.9.2 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce a patch release of ScyllaDB Monitoring Stack 4.9.2 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB Enterprise and ScyllaDB Open Source, based on Prometh… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-monitoring-stack-4-9-2/4600 ## Headings Structure: H1: [RELEASE] ScyllaDB Monitoring Stack 4.9.2 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Monitoring Stack 4.9.2 H3: Related topics The ScyllaDB team is pleased to announce a patch release of ScyllaDB Monitoring Stack 4.9.2 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB Enterprise and ScyllaDB Open Source, based on Prometheus and Grafana. ScyllaDB Monitoring Stack 4.9.2 supports: --- ### Page: https://forum.scylladb.com/t/release-scylla-manager-3-4-2/4601 Title: [RELEASE] Scylla Manager 3.4.2 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces the release of ScyllaDB Manager 3.4.2, a production-ready patch release of the stable 3.4 branch. ScyllaDB Manager is a centralized cluster administration and recurrent tasks automation tool. … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-manager-3-4-2/4601 ## Headings Structure: H1: [RELEASE] Scylla Manager 3.4.2 H3: Related topics ## Main Content: H1: [RELEASE] Scylla Manager 3.4.2 H3: Related topics The ScyllaDB team announces the release of ScyllaDB Manager 3.4.2, a production-ready patch release of the stable 3.4 branch. ScyllaDB Manager is a centralized cluster administration and recurrent tasks automation tool. ScyllaDB Manager is available for ScyllaDB Enterprise customers and ScyllaDB Open Source users. With ScyllaDB Open Source, ScyllaDB Manager is limited to 5 nodes. See the ScyllaDB Manager Proprietary Software License Agreement for details. The list of issues fixed by the 3.4.2 release can be found here. It includes: ScyllaDB Enterprise customers are encouraged to upgrade to ScyllaDB Manager 3.4.2 in coordination with the ScyllaDB support team. The new release includes upgrades of both ScyllaDB Manager Server and Agent. Scylla Manager 3.4.2 supports the following Scylla Enterprise releases: And the following Open Source release (limited to 5 nodes see license 1): You can install and run Scylla Manager on Kubernetes using Scylla Operator. More here. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-86-2025-03-14/4606 Title: Last week in scylla-cluster-tests.git master (issue #86; 2025-03-14) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 0c569b88…31cb6202 range are covered. There were 54 non-merge commits from 17 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-86-2025-03-14/4606 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #86; 2025-03-14) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #86; 2025-03-14) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 0c569b88…31cb6202 range are covered. There were 54 non-merge commits from 17 authors in that period. Some notable commits: A new nemesis was added for the Raft feature consistent-topology-changes, verifying that a node can’t connect to the cluster after being removed. A new Severity.SUPPRESS level stops generating events for specified log lines, improving test stability when excessive events could clog SCT. A cluster target size config parameter can now be defined for multi-DC test cases, allowing scale testing with multi-DC configurations. Scylla driver was updated to v3.29.2. When testing Scylla AMI, we now validate the preconfigured IOTune config by comparing it to actual machine values and showing deviations in Argus. Monitoring is back to version 4.9. The Docker backend now supports the new UBI9-based image. Users can still use the older image by setting SCT_SCYLLA_LINUX_DISTRO to Ubuntu. nemesis_add_node_cnt now defaults to 3, meaning three nodes will be added to the cluster by default in supported nemesis (e.g., grow_shrink_cluster). The cql-stress tool was updated to use the rust driver v1.0.0. PRs that change Python dependencies now trigger an automatic SCT Docker image build via a GitHub action, pushing the new image info back into the PR. If not desired, the PR author can remove the New Hydra Version label. We now test perf simple query write path on a weekly basis. To improve unit testing, we now use the moto library to mock AWS services, eliminating the need to spin up a mock service. A new unit test module for AWS services has been added. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/adding-a-cluster-using-sctool-and-using-backup-getting-error-deduplicating-the-snapshot/4608 Title: Adding a cluster using sctool and using backup, getting "ERROR (deduplicating the snapshot)" - Kubernetes Operator - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/adding-a-cluster-using-sctool-and-using-backup-getting-error-deduplicating-the-snapshot/4608 ## Headings Structure: H1: Adding a cluster using sctool and using backup, getting "ERROR (deduplicating the snapshot)" H3: Related topics ## Main Content: H1: Adding a cluster using sctool and using backup, getting "ERROR (deduplicating the snapshot)" H3: Related topics Originally from the User Slack @Pritam_Nagar: @Kacper_Rzetelski hey, i’,m using scylla in gke after adding a cluster using when i schedule a backup the. very first backup is goes successfully but second one gives error i’m stuck over here please let me know if any one of you know the solution @Kacper_Rzetelski: Hi, clusters added to SM directly through sctool won’t be controlled by the operator. Any clusters created with operator should automatically be registered with SM. For backup scheduling see https://operator.docs.scylladb.com/stable/architecture/manager.html. If you have issues with how to register a cluster with SM or schedule a backup through the operator, feel free to create an issue with must-gather attached so we can verify it. As for the deduplication error from SM - I’m afraid I won’t be able to help here, try asking on scylla-manager channel. Hi, could you share logs and version (both for both scylla-manager and scylla-manager-agents)? My suspicion is that you updated scylla-manager server, but you forgot to update scylla-manager-agents, so scylla-manager is trying to query endpoints, which were not exposed in the previous versions, and that causes the error. I got fixed when i update the versions of scylla and scylla manager(5.2.x -->6.2.0). And schedule the backup it works fine for some days and got error of file indexing just like this issue: --- ### Page: https://forum.scylladb.com/t/ntp-configuration-in-scylladb-can-the-operator-be-used-to-configure/4609 Title: NTP configuration in ScyllaDB, can the Operator be used to configure? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/ntp-configuration-in-scylladb-can-the-operator-be-used-to-configure/4609 ## Headings Structure: H1: NTP configuration in ScyllaDB, can the Operator be used to configure? H3: Related topics ## Main Content: H1: NTP configuration in ScyllaDB, can the Operator be used to configure? H3: Related topics Originally from the User Slack @Matheus_Salvia: does the operator provide any api to configure NTP? if not how should it be handled? @Maciej_Zimnoch: Containers share time set on the host, your Kubernetes nodes should be configured with NTP. Scylla Operator can’t help with that. You can learn more about the NTP configuration for ScyllaDB here. --- ### Page: https://forum.scylladb.com/t/not-able-to-create-schema-while-restoration-in-version-6-2-0/4611 Title: Not able to create schema while restoration in version 6.2.0 - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details Running a 3 node scylla cluster in GKE. #ScyllaDB version: 6.2.0 #Cluster size: 3 node os (RHEL/CentOS/Ubuntu/AWS AMI): Ubuntu Hi all, I have upgraded my Scylla cluster from 5.2.4 to 6.2.0. no… Language: en Canonical URL: https://forum.scylladb.com/t/not-able-to-create-schema-while-restoration-in-version-6-2-0/4611 ## Headings Structure: H1: Not able to create schema while restoration in version 6.2.0 H3: Related topics ## Main Content: H1: Not able to create schema while restoration in version 6.2.0 H3: Related topics Installation details Running a 3 node scylla cluster in GKE. #ScyllaDB version: 6.2.0 #Cluster size: 3 node os (RHEL/CentOS/Ubuntu/AWS AMI): Ubuntu Hi all, I have upgraded my Scylla cluster from 5.2.4 to 6.2.0. now when I try to restore a schema from already running scylla cluster to newly created scylla cluster, using command: sctool restore -c scylla-backup -L gcs:scylla-backup-prod -T sm_20250214180010UTC --restore-schema I’m getting issue that scylla-manager keyspace already exists. but when I delete scylla-manager namespace from new cluster, I can’t even run the above command. Earlier scylla used to store schema in .cql format but now it’s storing in json format. Any help on this is appreciated. Thanks! Hi, could you share the scylla-manager logs and version? Are you perhaps using the same cluster for your data and as the scylla-manager back-end? Have you created a new scylla-manager instance for this new cluster (and also using this new cluster as its backend)? If that’s the case, then backup contains the schema definition of scylla-manager keyspace and the new cluster also contains such definition (created by the new scylla-manager on its start-up, as it is its back-end). A simple fix is to not create a new scylla-manager, but use the old one for schema restoration in the new, fresh cluster. After schema and data restoration, for whatever reason, you can create a new scylla-manager instance with the newly restored cluster as its back-end, and all should work fine (although using the same cluster for your data and scylla-manager back-end is problematic as in this case). Another option would be to restore schema by taking the .cql file from the backup and applying it manually (see backup spec for more info). --- ### Page: https://forum.scylladb.com/t/backup-fails-with-indexing-files-error/4612 Title: Backup fails with indexing files error - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details Running a 3 node scylla cluster in GKE. #ScyllaDB version: 6.2.0 #Cluster size: 3 node os (RHEL/CentOS/Ubuntu/AWS AMI): Ubuntu Hi guys, I’m getting indexing files error in scylla backup. Language: en Canonical URL: https://forum.scylladb.com/t/backup-fails-with-indexing-files-error/4612 ## Headings Structure: H1: Backup fails with indexing files error H3: Related topics ## Main Content: H1: Backup fails with indexing files error H3: Related topics Installation details Running a 3 node scylla cluster in GKE. #ScyllaDB version: 6.2.0 #Cluster size: 3 node os (RHEL/CentOS/Ubuntu/AWS AMI): Ubuntu Hi guys, I’m getting indexing files error in scylla backup. Hi, could you share the Scylla Manager logs from the backup execution? Was this backup paused/resumed/retried or is this the first fresh backup execution? Were there any non empty tables being backed up? This error occured after taking couple (depends on cronjob how it is set in my case it is hourly) of backups and then failed giving this error. when we stop this task and schedule a new one it works perfectly fine. Hey team, kindly reply to this question either it is a bug or we had configured our scylla-manager incorrect. Hi, could you share the Scylla Manager logs from the backup execution? Were there any non empty tables being backed up? --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-272-2025-03-16/4613 Title: Last week in scylladb.git master (issue #272; 2025-03-16) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 0343235aa26…d84da3dc11c range are covered. There were 169 non-merge commits from 28 authors in that pe… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-272-2025-03-16/4613 ## Headings Structure: H1: Last week in scylladb.git master (issue #272; 2025-03-16) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #272; 2025-03-16) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 0343235aa26…d84da3dc11c range are covered. There were 169 non-merge commits from 28 authors in that period. Some notable commits: Tablet load-balancing is now aware of each node’s capacity. Different nodes can have different ratios between storage size and shard count. This change prevents some nodes from reaching 100% utilization while others have free space. There is now configuration for enabling TLS session cookies (unfortunately defaulting to false). TLS session cookies reduce TLS handshake cost. This is important when a node restarts, and especially for alternator that requires many connections as it uses HTTP for transport. The scylla sstable command, used for inspecting sstables outside ScyllaDB itself, can now run CQL queries against individual sstables. This is useful for debugging problems. The scylla sstable command can now access sstables stored on S3 rather than local disk. The container image is now based on Red Hat Universal Base Image 9. This is necessary for OpenShift certification. The values provided to the LIMIT and PER PARTITION LIMIT CQL clauses are now required to be strictly positive. CQL PER PARTITION LIMIT queries are now rejected if aggregate functions are present. Querying via a secondary index is now careful not to fetch too many rows from the index, as this can cause allocation related stalls. Basic metrics are now labeled so it is possible to fetch only those metrics. A race condition between the cleanup operation and shapshot operation was fixed. The S3 client now retries failed instance metadata operations and credential operations. Node shutdown now cancels draining hints. This reduces problems shutting down a node if the rest of the cluster is not healthy. Most of the gossip code now addresses nodes using host IDs rather than IP addresses, reducing problems in environments that have variable IP addresses such as Kubernetes. See you in the next issue of last week in scylladb.git master! Uploading LFS objects: 100% (2/2), 12 MB | 0 B/s, done. --- ### Page: https://forum.scylladb.com/t/can-i-update-configuration-variables-using-help/4614 Title: Can I update configuration variables using help? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/can-i-update-configuration-variables-using-help/4614 ## Headings Structure: H1: Can I update configuration variables using help? H3: Related topics ## Main Content: H1: Can I update configuration variables using help? H3: Related topics Originally from the User Slack @kam_ka: Hi FOlks, how do we update certain config variables using helm chart, for example compaction_large_row_warning_threshold_mb @Maciej_Zimnoch: not entirely in the helm. You can manage Scylla configuration through ConfigMap (having scylla.yaml key) and supply it’s name in the ScyllaCluster rack spec --- ### Page: https://forum.scylladb.com/t/listening-on-0-0-0-0/4616 Title: Listening on 0.0.0.0 - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 6.2.3 #Cluster size: 3 os (RHEL/CentOS/Ubuntu/AWS AMI): Debian I want port 9042 to listen on all IP addresses. I am trying to configure networking as described in scylla.yaml… Language: en Canonical URL: https://forum.scylladb.com/t/listening-on-0-0-0-0/4616 ## Headings Structure: H1: Listening on 0.0.0.0 H3: Related topics ## Main Content: H1: Listening on 0.0.0.0 H3: Related topics Installation details #ScyllaDB version: 6.2.3 #Cluster size: 3 os (RHEL/CentOS/Ubuntu/AWS AMI): Debian I want port 9042 to listen on all IP addresses. I am trying to configure networking as described in scylla.yaml: # If you leave broadcast_address (below) empty, then setting listen_address # to 0.0.0.0 is wrong as other nodes will not know how to reach this node. # If you set broadcast_address, then you can set listen_address to 0.0.0.0. For example: listen_address: 0.0.0.0 broadcast_address: 10.10.10.10 After I do so, port 9042 is not brought up and there is nothing in debug logs to say why. Check if port 9042 is being listened on by any other service. --- ### Page: https://forum.scylladb.com/t/scylla-constantly-flushes-memtables-and-runs-huge-number-of-compactions/4621 Title: Scylla constantly flushes memtables and runs huge number of compactions - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 6.1.4 #Cluster size: 6 os (RHEL/CentOS/Ubuntu/AWS AMI): Ubuntu 22.04/24.04 (on some nodes), ScyllaDB is running in Docker containers (image: https://hub.docker.com/layers/scyll… Language: en Canonical URL: https://forum.scylladb.com/t/scylla-constantly-flushes-memtables-and-runs-huge-number-of-compactions/4621 ## Headings Structure: H1: Scylla constantly flushes memtables and runs huge number of compactions H3: Related topics ## Main Content: H1: Scylla constantly flushes memtables and runs huge number of compactions H3: Related topics #ScyllaDB version: 6.1.4 #Cluster size: 6 os (RHEL/CentOS/Ubuntu/AWS AMI): Ubuntu 22.04/24.04 (on some nodes), ScyllaDB is running in Docker containers (image: https://hub.docker.com/layers/scylladb/scylla/6.1.4/images/sha256-a507e50f703662580230d54269876d491c9bb7f6110a1a95af6341b74e86154c) Limits: memory - 16 GiB, CPU - 12 cores I’m facing quite a strange Scylla’s behavior: it seems like Scylla doesn’t keep any data in memtables and instantly flushes any new data into new SSTable causing too frequent compaction runs as soon as there are at least 2 new equally-sized SSTables. This leads to performance issues (~20% of queries are timed out), resource shortage (quite frequent OOM’s, CPU throttling), and overall cluster overload. On the other hand we have several production Scylla clusters in our company running in the same or very close environments - and there are no issues with them at all. Data models, configuration, workload patterns, RPSs are very similar too. So there is something here completely going wrong, or maybe I’m just missing something obvious. Anyway I need your help - any hints on causes of such behavior would be very appreciated, details are following Data model: there are 8 tables in the keyspace, general schema: Compaction strategy used is default SizeTieredCompactionStrategy with default options. Earlier I’ve tried to increase the tombstone_compaction_interval option but it didn’t improve anything Workload: typical workload - read some entity, update some fields and save it, so numbers of reads and writes are expected to be approximately equal Symptoms: first signs of problems came with alerts of high memory and CPU consumption from monitoring system. After that random nodes started to crash periodically with 139 exit code (segfaults I guess). Searching through Scylla’s logs I discovered only the following types of errors (some info is replaced with placeholders): Further research showed that there is always quite a large scheduler’s task queue (something about 50-100 tasks all the time) and Scylla is constantly running some compactions (there are 50-100 running compactions all the time according to compaction manager’s data). This data agrees with Scylla’s log records on compactions - there are 200k (!) successful compactions per table per day on average according to the logs Running nodetool tablestats gives the following data (I’m showing only one table - other tables’ stats are very similar): And here come first oddities: nodetool sstableinfo on the same table reports something quite similar to the following data for the vast majority of SSTables: And here comes the next chunk of oddities: And finally nodetool cfhistograms reports the following for this table: Taking into account all of the collected data I supposed that there might be some issue with instant memtables flushing and keeping SSTable size very small but I’m not quite sure Please let me know if you have any clues on such Scylla’s behavior that can help me to elucidate the root cause of cluster’s problems or if there’s some additional data I could collect that can help in this investigation. Thank you in advance! --- ### Page: https://forum.scylladb.com/t/compaction-task-occurred-bad-alloc-when-there-is-large-partition/4631 Title: Compaction task occurred bad_alloc when there is large partition - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 5.2.13 #Cluster size: os (RHEL/CentOS/Ubuntu/AWS AMI): CentOS Is there any potential risk in large partitions or rows or cells if do compaction, or others? Actually, a node i… Language: en Canonical URL: https://forum.scylladb.com/t/compaction-task-occurred-bad-alloc-when-there-is-large-partition/4631 ## Headings Structure: H1: Compaction task occurred bad_alloc when there is large partition H3: Related topics ## Main Content: H1: Compaction task occurred bad_alloc when there is large partition H3: Related topics Installation details #ScyllaDB version: 5.2.13 #Cluster size: os (RHEL/CentOS/Ubuntu/AWS AMI): CentOS Is there any potential risk in large partitions or rows or cells if do compaction, or others? Actually, a node is restarted with log “compaction_manager - compaction task … failed : std::bad_alloc(std::bad_alloc), will retry in 300 seconds” Look for log lines for the same shard just before the bad_alloc. ScyllaDB will log when it’s starting a compaction task on a SSTable set, so you might have a pointer to which table/SSTables the problem is related to. Also make sure you follow best practices in terms of vCPU/memory/storage to avoid bad_allocs: System Requirements | ScyllaDB Docs --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-2-6/4633 Title: [RELEASE] ScyllaDB Enterprise 2024.2.6 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2024.2.6, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2024.2 Feature (Short Term Support) Release. Related Links Get ScyllaDB En… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-2-6/4633 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2024.2.6 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2024.2.6 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2024.2.6, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2024.2 Feature (Short Term Support) Release. The following issues are fixed in this release (with an open-source reference, if available): --- ### Page: https://forum.scylladb.com/t/getting-logalloc-bad-alloc-on-3-node-cluster/4636 Title: Getting logalloc::bad_alloc on 3 node cluster - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 6.1.2-0.20240915.b60f9ef4c223 #Cluster size: 3 os (RHEL/CentOS/Ubuntu/AWS AMI): Ubuntu 22.04.5 LTS 2 of the 3 nodes are down with the following error. Any idea how I could re… Language: en Canonical URL: https://forum.scylladb.com/t/getting-logalloc-bad-alloc-on-3-node-cluster/4636 ## Headings Structure: H1: Getting logalloc::bad_alloc on 3 node cluster H3: Related topics ## Main Content: H1: Getting logalloc::bad_alloc on 3 node cluster H3: Related topics Installation details #ScyllaDB version: 6.1.2-0.20240915.b60f9ef4c223 #Cluster size: 3 os (RHEL/CentOS/Ubuntu/AWS AMI): Ubuntu 22.04.5 LTS 2 of the 3 nodes are down with the following error. Any idea how I could recover these 2 nodes. It keeps restarting. When you restart the nodes, they replay the Commitlog, which restores the content of the Memtable, the flushing of which seems to be causing this error. To restore your nodes, move aside the content of /var/lib/data/commitlog/, then start the node. Note!!! This will result in the loss of the writes contained in said Commitlogs. You can try to move aside only some of the files, start the node, and if it succeeds, stop it and copy back some of the files. Repeat until all Commitlogs are restored. In any case, make sure to repair the cluster after doing this. Does your schema has collections by any chance? --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-87-2025-03-21/4638 Title: Last week in scylla-cluster-tests.git master (issue #87; 2025-03-21) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 6e74ce13…8d10481f range are covered. There were 21 non-merge commits from 10 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-87-2025-03-21/4638 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #87; 2025-03-21) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #87; 2025-03-21) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 6e74ce13…8d10481f range are covered. There were 21 non-merge commits from 10 authors in that period. Some notable commits: Latte tool was bumped to 0.28.3 with a switch to Rust driver 1.0.0 and support for the Decimal CQL data type. System logs of cluster nodes will now include SCT disruptive operations executed on those nodes. Scylla-doctor installation in artifact tests was refactored to use the official package from the repo, or a package from S3 for offline installation. Scylla Python driver was updated to v3.29.3. The nemesis_selector config parameter now supports full logical expressions, allowing the use of not , and , and or operators instead of complex list elements, e.g., config_changes and topology_changes , topology_changes or disruptive , not disruptive . Cloud usage reports for GCE will now handle and destroy stopped instances. To prevent termination, use the keep_action=stop label. Following the migration of stress tools from the hydra-loaders Docker repository, cql-stress and scylla-bench now use their official images, and SCT no longer includes their build processes. Latte and Gemini had already been migrated in previous commits. Loader rack awareness was always enabled for multi-AZ, multi-DC cases. To avoid issues when too few loaders are used after switching some tests to simulated racks, a new “rack_aware_loader” boolean parameter was introduced in the test config, allowing this feature to be enabled only in specific tests. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/scylladb-monitoring-how-can-i-reuse-an-existing-prometheus-installation-with-a-kubernetes-cluster/4641 Title: ScyllaDB Monitoring, how can I reuse an existing Prometheus installation with a Kubernetes cluster? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-monitoring-how-can-i-reuse-an-existing-prometheus-installation-with-a-kubernetes-cluster/4641 ## Headings Structure: H1: ScyllaDB Monitoring, how can I reuse an existing Prometheus installation with a Kubernetes cluster? H3: Related topics ## Main Content: H1: ScyllaDB Monitoring, how can I reuse an existing Prometheus installation with a Kubernetes cluster? H3: Related topics Originally from the User Slack @Tom_Pester: For the scylla monitoring stack, when using the CRD to anable it (more details in thread), is there a way to What did I already do? Thanks for any help. BTW love the CRD that was created for this @Maciej_Zimnoch: so far there’s no support for external prometheus. We plan to add it sooner than later though. If you want to disable both prometheus and grafana why this CRD is useful? for ServiceMonitors? @Tom_Pester: Indeed, for the • ServiceMonitor • and the alert rules which is also very valuable Come to think of it, is there a way to let scylla monitoring add and maintain its rules to a pre-existing prometheus? If there is an update of scylla that fixes some bug to an alert rule, can it even apply this to the pre-existing prometheus? If not, than spinning up a new prometheus, even not resource friendly, is a more clean solution. @Maciej_Zimnoch: I see. is there a way to let scylla monitoring add and maintain its rules to a pre-existing prometheus? we create PrometheusRule resources which contain those alerts, whether they are installed in external prometheus it depends on operator managing it. But currenly we create those only when we manage Prometheus as well. So in theory you could spawn little instance of Prometheus using ScyllaDBMonitoring to get those PrometheusRule’s and ServiceMonitor’s, and configure your Prometheus to pick them up. @Tom_Pester: This is what we ended up doing if I understand you correctly@Maciej_Zimnoch https://github.com/scylladb/scylla-operator/issues/2490#issuecomment-2725600233 ^^ Is this also what you had in mind and can you improve on it? This exercise was a bit daunting for me as I had to learn about prometheus managed alert rules, external alert managers and k8s operators. At some point I saw 5 possible solutions and could zoom in on the best one. So the only thing left for us to do is disable the grafana and prometheus that the ScyllaDBMonitoring CR spins up. Would it be possible add this little functionality so we can disable them from the CR and not have a patch or hack that deletes them afterwards? This would be of great help! I believe it wasn’t possible to configure the prometheus created by ScyllaDBMonitoring to use an external alertmanager because the ScyllaDBMonitoring CRD doesn’t expose it. Or is there a way that you can reach into the underlying objects? But in the end we pointed to our pre-existing Prometheus that we had full control over. @amnon We have disabled the prometheus and grafana with a bit of a hack for the moment. Can we influence the ScyllaDBMonitoring CRD in another way to not let he pods exist in the first place? Thanks for this CRD and the existing monitoring solution. It already set us on the right path. @Maciej_Zimnoch: No, that’s the only option atm @Tom_Pester: Thanks for confirming Maciej. Would it be possible to add an enable:true|false key for Prometheus an Grafana? @Maciej_Zimnoch: it’s in our backlog @Tom_Pester: Thank you Maciej --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-273-2025-03-23/4644 Title: Last week in scylladb.git master (issue #273; 2025-03-23) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the d84da3dc11c..7646e1448a4 range are covered. There were 81 non-merge commits from 16 authors in that pe… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-273-2025-03-23/4644 ## Headings Structure: H1: Last week in scylladb.git master (issue #273; 2025-03-23) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #273; 2025-03-23) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the d84da3dc11c..7646e1448a4 range are covered. There were 81 non-merge commits from 16 authors in that period. Some notable commits: Tablet merges happen when the load balancer wants to reduce the number of tablets in a table. To merge tablets, the load balancer performs a “colocation migration” to move one tablet of the pair to the same node and shard as the other. It will now prefer migrating within the same rack. This is important for materialized view consistency. A Seastar bug triggered a gdb bug which caused core dumps not to be processed correctly. Seastar now avoids leaving zombie threads, preventing the problem. There is a new procedure for recovering from Raft group 0 majority loss. The procedure is safe for use with tablets. Repair of one tablet will no longer prevent another tablet from being migrated. The S3 driver has been made more robust. Audit syslog output was improved to make it machine parseable. The system can now constrain the replication factor to be equal to the number of racks. This greatly simplifies load balancing and will become the defaults. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/system-batchlog-table-tombstones-compactions-and-effect-on-performance/4645 Title: System batchlog table, tombstones compactions and effect on performance - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/system-batchlog-table-tombstones-compactions-and-effect-on-performance/4645 ## Headings Structure: H1: System batchlog table, tombstones compactions and effect on performance H3: Related topics ## Main Content: H1: System batchlog table, tombstones compactions and effect on performance H3: Related topics Originally from the User Slack @Chaitanya_Tondlekar: What is the significance of system.batchlog table ? i can see one of my cluster write latencies are huge compared to OPS and in logs can see heavy compactions are happening on same table. @avi @Botond_Dénes @Guy Can you help here with official documentation ? @Botond_Dénes: All write statements that are in: BEGIN BATCH ... APPLY BATCH are written into this table, and deleted once all statements were applied to the database successfully. So this table normally has a lot of tombstones and does need compaction to get rid of these. @Chaitanya_Tondlekar: Okay thanks. Any official documentation link for this table ? @Botond_Dénes: We have high level documentation on the batchlog feature, this table is part of the implementation. @Chaitanya_Tondlekar: For some reason , this table is not getting empty creating more and more compactions and resulting into high write latencies. What would be the reason for table not getting empty ? I can see that writes get executed successfully on db. If it’s getting executed successfully then batchlog table should be getting empty on it own. Is my understanding correct ? @Botond_Dénes: if you have batchlog workload this table never gets empty, because every batchlog insert will add entries to it, then delete those entries the tombstones stay around until they expire and need compaction to remove them @Chaitanya_Tondlekar: Ok. --- ### Page: https://forum.scylladb.com/t/high-payload-size-issue/4648 Title: High Payload Size issue - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 6.1.2-0.20240915.b60f9ef4c223 #Cluster size: 3 nodes os (RHEL/CentOS/Ubuntu/AWS AMI): Ubuntu Our payload size is around 3 - 5 kb. Is there anything we could do to lower the… Language: en Canonical URL: https://forum.scylladb.com/t/high-payload-size-issue/4648 ## Headings Structure: H1: High Payload Size issue H3: Related topics ## Main Content: H1: High Payload Size issue H3: Related topics Installation details #ScyllaDB version: 6.1.2-0.20240915.b60f9ef4c223 #Cluster size: 3 nodes os (RHEL/CentOS/Ubuntu/AWS AMI): Ubuntu Our payload size is around 3 - 5 kb. Is there anything we could do to lower the payload size? What are some potential issues it could cause? Could that cause memory issue because we saw in our older cluster it kept getting this error: The memory allocation failures are unrelated to the payload size. Rows or even cells which are 3-5KiB are considered to be of perfectly reasonable size. The memory errors are caused by something else, I cannot tell just based on these logs. --- ### Page: https://forum.scylladb.com/t/big-rows-do-not-compact-multi-sstable-to-one/4649 Title: Big rows do not compact multi sstable to one? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 5.4 #Cluster size: 6 nodes os (RHEL/CentOS/Ubuntu/AWS AMI): CentOS I update an item by adding attrs to make it a big row, but it do not compact to one. Could you please help m… Language: en Canonical URL: https://forum.scylladb.com/t/big-rows-do-not-compact-multi-sstable-to-one/4649 ## Headings Structure: H1: Big rows do not compact multi sstable to one? H3: Related topics ## Main Content: H1: Big rows do not compact multi sstable to one? H3: Related topics Installation details #ScyllaDB version: 5.4 #Cluster size: 6 nodes os (RHEL/CentOS/Ubuntu/AWS AMI): CentOS I update an item by adding attrs to make it a big row, but it do not compact to one. Could you please help me figure it out why it does not merge to one ? Large rows information see here Bash for test like this The system.large_* tables reflect the SSTables as they are on the disk. If ScyllaDB is not compacting the SSTables together, the entries in system.large_* also won’t be merged. You can force the SSTables to be all compacted into a single one with nodetool compact (major compaction), but this is unnecessary. In fact, I am trying to test a failure that result a node restart, which is due to compact big partition. Then, if there is big partition / row , and activate a compaction, will it cause bad_alloc ? Compacting a large partition should not cause any bad alloc, ScyllaDB doesn’t read entire partitions into memory. Very big rows on the other hand can cause problems because ScyllaDB reads the entire row into memory. If the row is big enough to take too much memory, it will cause problems for sure. Then large rows warn info is just to warn users, but does nothing to stop it, right? Maybe here is supposed to set some limitation to avoid such bad_alloc exception caused by big rows read at once? It is hard to implement such limitations. Such large rows can build up over time, with small individual writes adding up to a large row. If the database refuses to read the row when it becomes large, this will block compaction. Proper solution would be to not read all the row into memory, but this is a lot of complex work, for this edge case. So for now, we have the large_rows table and the user is expected to keep an eye on large rows and take action before/after they become a problem. Thanks for your nice answer. Then is there any metrics that may help us get large rows info more conveniently ? I think the large_* tables are also exported to monitoring, although they are an optional (opt-in) feature. @Amnon_Heiman can you point us to the documentation on this? --- ### Page: https://forum.scylladb.com/t/how-to-troubleshoot-an-issue-with-high-latency-happening-every-once-in-a-while/4650 Title: How to troubleshoot an issue with high latency happening every once in a while? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/how-to-troubleshoot-an-issue-with-high-latency-happening-every-once-in-a-while/4650 ## Headings Structure: H1: How to troubleshoot an issue with high latency happening every once in a while? H3: Related topics ## Main Content: H1: How to troubleshoot an issue with high latency happening every once in a while? H3: Related topics Originally from the User Slack @Roy: Hi, We see sometimes our Cluster has high latency (99% Read and Writes goes to 300+ ms) and it is gone after sometime. The amount of Reads and Writes from KS level generally DO NOT differs much during that time. So we are unable to clearly detect the culprit. We could also see many : reader_concurrency_semaphore in syslog. Beside many occurrence of prepared_cache_eviction in Scylla monitoring advisor. So to get a clear idea what is the culprit we are thinking of enabling more detail logging as per: https://opensource.docs.scylladb.com/stable/operating-scylla/nodetool-commands/setlogginglevel.html. Could someone suggest the logger we should enable ? Is it safe to enable debug level on all component in a High to Moderately loaded PROD cluster? Nodetool setlogginglevel | ScyllaDB Docs @Botond_Dénes: Blanked-enabling all loggers to debug is a really bad idea. You should start debugging this in monitoring, look for events correlating with the elevated timeouts. Switch the dashboards to “by shard” view, so aggregation doesn’t hide outlier shards. @Roy: Well at least can see spike on Compactions. But they are there all the time, when and when not the issue is there. We are using default compaction settings without any tuning in scylla.yaml. I will have a look at the Shard level in Monitoring. @Botond_Dénes: Check the compaction scheduling group shares – do they climb to 1000? @Roy: Hi, Compaction CPU Runtimes spikes to 150-220% Node level and 25-30% in Shard Level but the Compaction Shares stays close to 50 @Botond_Dénes: In that case, I don’t think compaction is to blame here. ScyllaDB has schedulers to isolate scheduling groups from each other, with 50 shares, compaction will not be able to impact the statement scheduling group (which has 1000 shares). @Roy: Also at node level “details” dashboard i can see high number of tombtone write and cell tombstone writes. @avi: In Advanced view, look for scheduling groups that have non-zero Task Quota Violations @Roy: could only see this for sl: default @Botond_Dénes: A common source of task quota violations are stalls. Do you see any in the logs? Stalls are known to cause high 99% latencies. @Roy: Hi morning. Yes see Reactor stalls and read-concurrency_semaphores in syslog @Botond_Dénes: There is a good chance the stalls are the direct cause of the elevated latencies. Please open a github issue with the stalls (please decode them with http://backtrace.scylladb.com/index.html) The Build ID can be obtained from the logs, it is printed right at the beginning of startup. The Build ID can also be obtained with scylla --build-id. --- ### Page: https://forum.scylladb.com/t/does-repair-only-read-rows-from-disk/4651 Title: Does repair only read rows from disk? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: all #Cluster size: os (RHEL/CentOS/Ubuntu/AWS AMI): Recently I read repair code about row_level_epair, and I see it has a function called read_rows_from_disk, then, I am wonde… Language: en Canonical URL: https://forum.scylladb.com/t/does-repair-only-read-rows-from-disk/4651 ## Headings Structure: H1: Does repair only read rows from disk? H3: Related topics ## Main Content: H1: Does repair only read rows from disk? H3: Related topics Installation details #ScyllaDB version: all #Cluster size: os (RHEL/CentOS/Ubuntu/AWS AMI): Recently I read repair code about row_level_epair, and I see it has a function called read_rows_from_disk, then, I am wondering will it read data from memtable? Or only sstable ? Here is an example, when there is a five-node cluster with 3 replicas, consistency is quorum, and it is doing repair, considering at the same time a write is settled to disk(sstable) on node-1, this write is still on memtable of the other two nodes. So, in this situation, how repair reads data ? If repair only reads data from sstable, will repair thinks that write is missed on the other two nodes? Correct cluster config: six-node cluster with 5 replicas Repair reads both from disk and memtables. Most data is expected to come from the disk, so repair code refers to data as coming from disk. Thanks a lot for your reply. And I see data is read to a cache with a limit value, then do a hash for it. So, if an item is stored in several sstable, how it compare the diff? Data from different SSTables and Memtables are merged into a single unified stream, this is what is stored in the repair buffers. This buffer is then used to compare against the unified stream on other replicas. --- ### Page: https://forum.scylladb.com/t/expected-p50-p99-latency-with-scylladb-aerospike-and-sla/4654 Title: Expected P50/P99 latency with ScyllaDB Aerospike and SLA - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/expected-p50-p99-latency-with-scylladb-aerospike-and-sla/4654 ## Headings Structure: H1: Expected P50/P99 latency with ScyllaDB Aerospike and SLA H3: Related topics ## Main Content: H1: Expected P50/P99 latency with ScyllaDB Aerospike and SLA H3: Related topics Originally from the User Slack @Antony_Mithun: Hi everyone, we’re evaluating databases for a low-latency use case and considering ScyllaDB vs. Aerospike. We want to understand ScyllaDB’s expected read/write latency (P50/P99) without having to conduct benchmarks ourselves. The docs don’t specify an SLA for latency, and I couldn’t find clear numbers in this ScyllaDB vs. Aerospike doc. Meanwhile, Aerospike claims different results in their benchmark. Can someone clarify what kind of latency we can expect under high-throughput workloads? Does ScyllaDB have an official SLA for latency? @dor: Hi Antony, latency depends on a ton of factors - your hardware, workload, schema, ram, .. In general, you can expect very low single digit p99 latency from Scylla. P50 below 1ms easy. Our latency is general on par with Aerospike, their might even be better for reads. However our throughput is better and so is the data model/featureset and elasticity and scalability @Robert: hi, the latency more depends on Your infrastructure, invest money, use case and it requirements etc., than on DB engine technology. For example network latency between nodes and ųService-nodes - which also depends on level of consistency. If Your application can enable any flag like SEP? Infra part - if You contain RAM to handle all data stored on disk in scylla case You can avoid IOPS - all data will be cached. But it cost a lot, maybe You want to have some ratio 1:5 (ram/disk)? Imo it’s a long story to just simple answer, but for my perspective on bare metal which has 1:2 ratio ram/disk I’m having <1ms p99 with 100k req/sec. @Antony_Mithun: Thank you for your quick and very helpful feedback. We currently use Redis, which provides sub-millisecond latency for most requests (I can get back to you with the exact p99 and p50 values and request counts). However, with our new features, we expect to store significantly more data, which could impact the current Redis performance (especially with AOF). Even 5ms of latency would be considered too high for us. Therefore, we were exploring options that could offer low latency with a hybrid disk-backed setup. Initially, we considered Bigtable but were concerned that it might not perform well in terms of latency. As you guys pointed out, it ultimately depends on the infrastructure setup, so it seems we’ll need to experiment and determine the best solution ourselves. Thanks again for your input! @dor: AOF is indeed awful. There shouldn’t be a problem to go below 5ms p99 --- ### Page: https://forum.scylladb.com/t/best-practice-for-multi-row-insertions/4656 Title: Best practice for multi-row insertions - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: What is considered “best practice” for inserting, say, 100.000 rows into a ScyllaDB table in general, or specifically for the Rust SDK? The options I know about are: Sequential insertion - One query at a time. I find … Language: en Canonical URL: https://forum.scylladb.com/t/best-practice-for-multi-row-insertions/4656 ## Headings Structure: H1: Best practice for multi-row insertions H3: Related topics ## Main Content: H1: Best practice for multi-row insertions H3: Related topics What is considered “best practice” for inserting, say, 100.000 rows into a ScyllaDB table in general, or specifically for the Rust SDK? The options I know about are: Please share your thoughts and opinions on what the best method of inserting many rows as fast as possible is. Parallel Inserttion is the best approach for bulk inserts in ScyllaDB: ScyllaDB is designed to handle parallelized inserts. It scales well with multiple concurrent connections and async tasks. You avoid the overhead of batches, and you get maximum throughput by leveraging the full cluster. --- ### Page: https://forum.scylladb.com/t/dial-tcp-127-0-0-1-connect-connection-refused-what-is-the-default-api-link-and-port/4658 Title: Dial tcp 127.0.0.1:5080: connect: connection refused. What is the default api link and port? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: os (RHEL/CentOS/Ubuntu/AWS AMI): Language: en Canonical URL: https://forum.scylladb.com/t/dial-tcp-127-0-0-1-connect-connection-refused-what-is-the-default-api-link-and-port/4658 ## Headings Structure: H1: Dial tcp 127.0.0.1:5080: connect: connection refused. What is the default api link and port? H3: Related topics ## Main Content: H1: Dial tcp 127.0.0.1:5080: connect: connection refused. What is the default api link and port? H3: Related topics Installation details #ScyllaDB version: os (RHEL/CentOS/Ubuntu/AWS AMI): I have a Scylla Node A and a Scylla Ops Server. Scylla is running on port 9042 on a different server. What do you mean by “Scylla Ops”? Please provide more details. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-88-2025-03-28/4659 Title: Last week in scylla-cluster-tests.git master (issue #88; 2025-03-28) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the bdc78f39…3787115a range are covered. There were 14 non-merge commits from 8 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-88-2025-03-28/4659 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #88; 2025-03-28) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #88; 2025-03-28) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the bdc78f39…3787115a range are covered. There were 14 non-merge commits from 8 authors in that period. Some notable commits: Fixed Hydra running on Mac by adding support for POSIX sed. Due to issues with Scylla-Bench 0.2.0, it was downgraded back to 0.1.15. Because Scylla now notifies aliveness to other nodes during boot, the rolling restart method no longer waits for all nodes to be in a normal state. latte tool was updated to version 0.28.4-scylladb, adding support for Counter, Date, Time, Duration, Varint, and Tuple CQL data types, along with a detailed version output including the underlying Scylla driver version. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/scylladb-university-live-online-training-event-happening-next-week/4662 Title: ScyllaDB University Live: online training event happening next week - University and Training - ScyllaDB Community NoSQL Forum Meta Description: ScyllaDB University Live is happening on Wednesday, April 9, 2025, 8 AM PDT – 10 AM PDT. It’s a free, online, instructor-led event. You’ll have a chance to ask questions, learn by running hands-on labs, and enhance you… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-university-live-online-training-event-happening-next-week/4662 ## Headings Structure: H1: ScyllaDB University Live: online training event happening next week H3: Related topics ## Main Content: H1: ScyllaDB University Live: online training event happening next week H3: Related topics ScyllaDB University Live is happening on Wednesday, April 9, 2025, 8 AM PDT – 10 AM PDT. It’s a free, online, instructor-led event. You’ll have a chance to ask questions, learn by running hands-on labs, and enhance your skills. You can read more about it in this blog post I wrote. I hope to see you there! --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-274-2025-03-30/4663 Title: Last week in scylladb.git master (issue #274; 2025-03-30) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 7646e1448a..0ee0696959 range are covered. There were 41 non-merge commits from 15 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-274-2025-03-30/4663 ## Headings Structure: H1: Last week in scylladb.git master (issue #274; 2025-03-30) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #274; 2025-03-30) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 7646e1448a..0ee0696959 range are covered. There were 41 non-merge commits from 15 authors in that period. Some notable commits: S3 driver streaming is now much faster. Instead of sending individual 128k requests (similar to how we read blocks from the filesystem), we now use a single HTTP GET request to stream the entire object. Seastar’s native TCP stack using DPDK is no longer built. It is not used in production. A bug that prevented tablets of materialized views from being split was fixed. A bug which prevented column renames from being propagated to materialized views was fixed. A type mismatch in some nodetool commands was fixed. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/how-is-null-handled-from-a-cql-and-a-storage-point-of-view/4664 Title: How is Null handled from a CQL and a storage point of view? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/how-is-null-handled-from-a-cql-and-a-storage-point-of-view/4664 ## Headings Structure: H1: How is Null handled from a CQL and a storage point of view? H3: Related topics ## Main Content: H1: How is Null handled from a CQL and a storage point of view? H3: Related topics Originally from the User Slack @Aleksandr_Udovenko: Hello i have table like this: I will have non NULL values for props1, props2 very rare, only when offset==0. Does scylla optimize store props1,props2 if it will be NULL in 99% ? (in my real DB i have many props) @Felipe_Cardeneti_Mendes: Storage wise, yes - null is… well null If you are wondering how NULL handling works see https://github.com/scylladb/scylladb/blob/6d7cb68aabea41024fac925aebc014139f4e5568/docs/cql/cql-extensions.md#null GitHub: scylladb/docs/cql/cql-extensions.md at 6d7cb68aabea41024fac925aebc014139f4e5568 · scylladb/scylladb @Aleksandr_Udovenko: NULL from point CQL it is intresting, but i need to undertand it from storage side) --- ### Page: https://forum.scylladb.com/t/release-scylladb-cdc-rust-0-4-0/4666 Title: [RELEASE] ScyllaDB CDC Rust 0.4.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce scylla-cdc-rust 0.4.0, a library that makes it easy to develop Rust applications that consume data from a Scylla Change Data Capture Log. This version bumps up the dependency on … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-cdc-rust-0-4-0/4666 ## Headings Structure: H1: [RELEASE] ScyllaDB CDC Rust 0.4.0 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB CDC Rust 0.4.0 H3: Related topics The ScyllaDB team is pleased to announce scylla-cdc-rust 0.4.0, a library that makes it easy to develop Rust applications that consume data from a Scylla Change Data Capture Log. This version bumps up the dependency on the scylla-rust-driver from 0.13.1 (in released 0.2) or 0.15.1 (in unreleased 0.3) to 1.0.0 (#127). --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-2-7/4668 Title: [RELEASE] ScyllaDB Enterprise 2024.2.7 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2024.2.7, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2024.2 Feature (Short Term Support) Release. Related Links Get ScyllaDB En… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-2-7/4668 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2024.2.7 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2024.2.7 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2024.2.7, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2024.2 Feature (Short Term Support) Release. The following issues are fixed in this release (with an open-source reference, if available): --- ### Page: https://forum.scylladb.com/t/migrating-from-sql-to-nosql-data-modeling-considerations-postgresql-joins-denormalization-and-materialized-views/4669 Title: Migrating from SQL to NoSQL, data modeling considerations, PostgresQL, joins, denormalization and Materialized Views - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/migrating-from-sql-to-nosql-data-modeling-considerations-postgresql-joins-denormalization-and-materialized-views/4669 ## Headings Structure: H1: Migrating from SQL to NoSQL, data modeling considerations, PostgresQL, joins, denormalization and Materialized Views H3: Related topics ## Main Content: H1: Migrating from SQL to NoSQL, data modeling considerations, PostgresQL, joins, denormalization and Materialized Views H3: Related topics Originally from the User Slack @Tushar_Vaswani: Hey guys, I am new to the nosql world and scylladb. I was looking out for migrating our codebase from sql to nosql. But I have two questions: @Robert: Long story to tell, but about point 2.: You can use some unique identifier which can’t be changed and that solves a problem of user name change. When username = A he wrote 1-100 messages, when was = B he wrote 101-200 messages. Other hand there is a static field in Scylla which You can be use to update and it of curse finally end like update in the all the rows (for whole partition) - so don’t need to update 100k records to change just 1 field. But also it’s not a good idea to have just 1 partition to handle 100k message, should be split by some date etc. so like 12 update to update whole year of messages per user. Btw. if it’s a common situation that users changes a name - also worth to think about that… Another way is to do a client-side join, so have 2 tables and if one is a small load it every 15min (invalidate) to memory and handle the mapping and do a “SQL join” on the application side. Of course can be a lot of different way to solve issue like that, but it mostly depends on Your business needs, what You want to achieve etc. If no-one in Your company is if fight (production) familiar with Scylla I would suggest to looking for some consultant to help with design, modeling, coding, best practice etc. @Tushar_Vaswani: thanks for explaining in detail @Robert. Now I am getting some idea about how it really is done. Looks like NoSQL world is not ideal and straightforward like SQL world. You just have to accept tradeoffs and move on for finding solutions. @Robert: SQL could solve all bad modeling by multiple indexes - even query like multiple Cartesian cross will be finally executed by rbdms @Tushar_Vaswani: yes although that makes it easier @Robert: easier means in that case sh.tty @Krasimir_Popov: I think you can’t just switch it is never that straight forward, it depends on many moving parts, I would suggest you just add ScyllaDB as a second database and move data there that your application access frequently and needs speed and performance. Like that your team and you will gather expertise over the time and you can make informed decision if you can go without the SQL database or you keep both or you go back to SQL @Tushar_Vaswani: Got it also any recommendation for updates issue? @Krasimir_Popov: uuid for example @avi: It depends on you data and performance scale. If it works with PostgresQL and will continue to work for the foreseeable future, don’t change it. If you need to manage many terabytes and/or have high performance needs, then the migration will be worthwhile. @Tushar_Vaswani: makes sense yeah postgres works fine dont have large scale rn. Like just a million rows max in the biggest table we have. but still just for knowledge purpose what would have been a possible solution for the updates? @avi: Materialized views @Guy: I’d also look into this lesson about denormalizing data in ScyllaDB University: Materialized Views, Secondary Indexes, and Filtering --- ### Page: https://forum.scylladb.com/t/release-scylladb-manager-3-5-0/4677 Title: [RELEASE] ScyllaDB Manager 3.5.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces the release of ScyllaDB Manager 3.5.0, a production-ready minor release of the stable 3.5 branch. ScyllaDB Manager is a centralized cluster administration and recurrent tasks automation tool. … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-manager-3-5-0/4677 ## Headings Structure: H1: [RELEASE] ScyllaDB Manager 3.5.0 H3: License Change H3: OpenShift H3: Leveraging ScyllaDB New Tablets Repair API H3: Improve Restore H3: Bug Fixes H3: Upgrade to the new release H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Manager 3.5.0 H3: License Change H3: OpenShift H3: Leveraging ScyllaDB New Tablets Repair API H3: Improve Restore H3: Bug Fixes H3: Upgrade to the new release H3: Related topics The ScyllaDB team announces the release of ScyllaDB Manager 3.5.0, a production-ready minor release of the stable 3.5 branch. ScyllaDB Manager is a centralized cluster administration and recurrent tasks automation tool. This release focuses on supporting the ScyllaDB 2025.1 release, but it also contains a few important changes to the ScyllaDB Manager license, container image, and new features. Below are the major changes in this release. ScyllaDB Manager license changed to AGPL (#4229). As ScyllaDB is moving to source available license, ScyllaDB Manager is changing its license to AGPL. See our blog post for a broader explanation and effective scope of this change. ScyllaDB Manager container image OpenShift certification (#4242). As a part of the effort of getting ScyllaDB Operator OpenShift certification, ScyllaDB Manager container images are now passing all the criteria in sections 2.1 (image content requirements) and 2.2 (image policy requirements) of the certification policy guide. Among other changes, ScyllaDB Manager image is now based on the Red Hat ubi9-minimal image. ScyllaDB Tablet repair API support (#4188, #4292, #4273). ScyllaDB 2025.1.0 introduced a new storage_service/tablets/repair API for repairing tablet replicated tables. The old /storage_service/repair_async/{keyspace} API proved to be unsafe when used with tablet tables. Moreover, it required stopping tablet load balancing for the time of repair, which could cause various problems for long running repairs. As it turned out, previous ScyllaDB Manager releases contained a bug which resulted in not stopping tablet load balancing when using the old repair API. This in combination with using non default tombstone gc mode ‘repair’ could lead to data resurrection. Because of that, all users are encouraged to upgrade to ScyllaDB Manager 3.5 even if they haven’t started using the ScyllaDB 2025.1 yet. New --dc-mapping restore task flag (#3829). This new flag allows for specifying which data center in the restored cluster will be responsible for downloading and streaming the data from which backed up data center. Using this flag in a multi data center scenario allows to save on cross region traffic costs and improve restore speed. This flag works only when the number of data centers in the restored cluster is the same as in the backup. For more information see the sctool documentation. Don’t allow for other tasks to run during restore (#4045). Forgetting to unschedule other tasks for the time of running restore was a common source of issues. Now, it is handled by ScyllaDB Manager automatically. For more information see restore task documentation. The ScyllaDB Manager 3.5.0 release also contains a few bug fixes: ScyllaDB customers are encouraged to upgrade to ScyllaDB Manager 3.5.0 in coordination with the ScyllaDB support team. The new release includes upgrades of both ScyllaDB Manager Server and Agent. Scylla Manager 3.5.0 supports the following ScyllaDB releases: You can install and run Scylla Manager on Kubernetes using Scylla Operator. More here. --- ### Page: https://forum.scylladb.com/t/release-scylladb-rust-driver-1-1-0/4678 Title: [RELEASE] ScyllaDB Rust Driver 1.1.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce ScyllaDB Rust Driver 1.1.0, an asynchronous CQL driver for Rust, optimized for Scylla, but also compatible with Apache Cassandra! Some interesting statistics: over 3.430k down… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-rust-driver-1-1-0/4678 ## Headings Structure: H1: [RELEASE] ScyllaDB Rust Driver 1.1.0 H2: Changes H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Rust Driver 1.1.0 H2: Changes H3: Related topics The ScyllaDB team is pleased to announce ScyllaDB Rust Driver 1.1.0, an asynchronous CQL driver for Rust, optimized for Scylla, but also compatible with Apache Cassandra! Some interesting statistics: New features / enhancements: API cleanups / better types: Internal API cleanups/refactors: CI / developer tool improvements: Congrats to all contributors and thanks everyone for using our driver! The source code of the driver can be found here: The official crates.io registry entry is here: Thank you for your attention, please do not hesitate to contact us if you have any questions, issues, feature requests, or are simply interested in our driver! Contributors since the last release: --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-89-2025-04-04/4680 Title: Last week in scylla-cluster-tests.git master (issue #89; 2025-04-04) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the eefedf5e…8859a386 range are covered. There were 24 non-merge commits from 7 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-89-2025-04-04/4680 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #89; 2025-04-04) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #89; 2025-04-04) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the eefedf5e…8859a386 range are covered. There were 24 non-merge commits from 7 authors in that period. Some notable commits: Elasticity performance test now uses the latest c-s version 3.17.3, switching from the prepared AMI to a Docker-based. Latte now supports ‘counter_write’ and ‘counter_read’ operations, mimicking c-s behavior, and Grafana dashboards reflect these separately from standard ‘write’ and ‘read’. Tier 1 tests now run with 3 virtual racks. Both cql-stress and Latte tools now report their versions, along with Rust driver details, to Argus. To reduce misses in config and doc updates, a new workflow was added that comments on PRs when these are likely needed. With HDR Histogram support in cql-stress, SCT was aligned to make use of it. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-275-2025-04-06/4683 Title: Last week in scylladb.git master (issue #275; 2025-04-06) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 0ee0696959..431de48df9 range are covered. There were 124 non-merge commits from 24 authors in that per… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-275-2025-04-06/4683 ## Headings Structure: H1: Last week in scylladb.git master (issue #275; 2025-04-06) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #275; 2025-04-06) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 0ee0696959..431de48df9 range are covered. There were 124 non-merge commits from 24 authors in that period. Some notable commits: There are now new compressor implementations which use dictionaries to improve the compression ratio. The dictionaries are shared across sstables and across all nodes in the cluster. The system will automatically generate new dictionaries when it sees a gain in compression ratio. A bug in audit of batch statements caused it not to record the query string (regression from enterprise). This is now fixed. The tablet load balancer will now prioritize table merge finalization over tablet migrations, so that if tablet count reduction is necessary it will execute promptly. Tracing now records the client-side timestamp for EXECUTE and BATCH statements. Encryption-at-rest now uses the Seastar HTTP client rather than a hand-rolled client to interact with key management APIs. Tablet repair will now watch for topology updates and allow migration of unrelated tablets to proceed. There is now configuration to enforce tablets mode for new keyspaces. The tablet load balancer will now strive to equalize tablet count across shards rather than nodes. This is important for heterogeneous clusters with different node sizes. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/how-do-i-run-repair-for-every-node-without-scylladb-manager-within-the-gc-grade-seconds-time-period/4684 Title: How do I run repair for every node without ScyllaDB Manager within the gc_grade_seconds time period? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/how-do-i-run-repair-for-every-node-without-scylladb-manager-within-the-gc-grade-seconds-time-period/4684 ## Headings Structure: H1: How do I run repair for every node without ScyllaDB Manager within the gc_grade_seconds time period? H3: Related topics ## Main Content: H1: How do I run repair for every node without ScyllaDB Manager within the gc_grade_seconds time period? H3: Related topics Originally from the User Slack @Hans: Hi, Without Scylla Manager, how do you guys do repair for every nodes in every gc_grace_seconds..? Please share your nice insights! @Roy: When we were without SM in early days, i used cronjob scheduler (or could be any scheduler) that runs at least twice before gc_grace_seconds @Hans: Thanks for your response. That’s most straightforward approach --- ### Page: https://forum.scylladb.com/t/revised-creating-role-with-options-is-not-supported/4688 Title: Revised: Creating ROLE with OPTIONS is not supported - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: Scylla Enterprise 2024.2.6 #Cluster size: 3 nodes (i4i.large) os (RHEL/CentOS/Ubuntu/AWS AMI): When I execute this query to create a role, it returns an error stating OPTIONS … Language: en Canonical URL: https://forum.scylladb.com/t/revised-creating-role-with-options-is-not-supported/4688 ## Headings Structure: H1: Revised: Creating ROLE with OPTIONS is not supported H3: Related topics ## Main Content: H1: Revised: Creating ROLE with OPTIONS is not supported H3: Related topics Installation details #ScyllaDB version: Scylla Enterprise 2024.2.6 #Cluster size: 3 nodes (i4i.large) os (RHEL/CentOS/Ubuntu/AWS AMI): When I execute this query to create a role, it returns an error stating OPTIONS option is not supported Am I doing it wrong? I was just referring to the official documentation The OPTIONS option is indeed not supported. based on your example, omitting the OPTIONS portion: Which may lead to conclusion that OPTIONS is indeed supported; however the below shows there is no column to store those options Hello, we support OPTIONS as a syntax but not all authenticators support them. The only one which does is SaslauthdAuthenticator for the rest there is simply nothing to set via OPTIONS. --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-16/4691 Title: [RELEASE] ScyllaDB Enterprise 2024.1.16 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2024.1.16, a bug-fix production-ready patch release for ScyllaDB Enterprise 2024.1 Long Term Support (LTS) Release. Note there is a later Feature (Short term Support) Rele… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-16/4691 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2024.1.16 H3: Related Links H2: Fixed Issue with an open-source reference: H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2024.1.16 H3: Related Links H2: Fixed Issue with an open-source reference: H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2024.1.16, a bug-fix production-ready patch release for ScyllaDB Enterprise 2024.1 Long Term Support (LTS) Release. Note there is a later Feature (Short term Support) Release 2024.2. --- ### Page: https://forum.scylladb.com/t/release-scylladb-2025-1/4692 Title: [RELEASE] ScyllaDB 2025.1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of ScyllaDB 2025.1.0 LTS, a production-ready ScyllaDB Enterprise Long Term Support Major Release. With the 2025.1 LTS release, ScyllaDB Enterprise 2023.1 become End … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-2025-1/4692 ## Headings Structure: H1: [RELEASE] ScyllaDB 2025.1 H3: 2025.1 is the first Source Available License release, replacing both ScyllaDB Open Source and Enterprise releases. The release accumulates all the improvements from Enterprise releases (up to 2024.2) and Open Source Releases (up to 6.2). H2: Relevant links H2: Tablets H3: vNodes will continue to be supported for existing Keyspaces as well as new Keyspaces using the tablets = { 'enabled': false } option. H3: Initial tablet number H3: Tablet Merge H3: Tablets Limitations H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 2025.1 H3: 2025.1 is the first Source Available License release, replacing both ScyllaDB Open Source and Enterprise releases. The release accumulates all the improvements from Enterprise releases (up to 2024.2) and Open Source Releases (up to 6.2). H2: Relevant links H2: Tablets H3: vNodes will continue to be supported for existing Keyspaces as well as new Keyspaces using the tablets = { 'enabled': false } option. H3: Initial tablet number H3: Tablet Merge H3: Tablets Limitations H3: Related topics The ScyllaDB team is pleased to announce the release of ScyllaDB 2025.1.0 LTS, a production-ready ScyllaDB Enterprise Long Term Support Major Release. With the 2025.1 LTS release, ScyllaDB Enterprise 2023.1 become End Of Life (EOL). More information on ScyllaDB Long Term Support (LTS) policy is available here. ScyllaDB 2025.1 focuses on making Tablets, first introduced in 2024.2, production-ready and enabled by default. Among others ScyllaDB 2025.1 improves performance, scaling speed, support for mix clusters (using different instance types), bug fixes and much more. ScyllaDB, with Tablets, and Mix instance Size support is the base for the upcoming ScyllaDB X Cloud, a new and improved Scylla Cloud offering which supports fast boot, fast scaling (out and in) and an upper limit of 90% storage utilization, compared to 70% today. ScyllaDB Enterprise customers are encouraged to upgrade to ScyllaDB 2025.1, and are welcome to contact our Support Team with questions. For best use of ScyllaDB 2025.1, use ScyllaDB Manager 3.5 and later, and ScyllaDB Monitoring Stack 4.9 and later. In this release, ScyllaDB makes tablets the default for new Keyspaces. Tablets are a new data distribution algorithm as a better alternative to the legacy vNodes approach inherited from Apache Cassandra. While the vNodes approach statically distributes all tables across all nodes and shards based on the token ring, the Tablets approach dynamically distributes each table to a subset of nodes and shards based on its size. In the future, distribution will use CPU, OPS, and other information to further optimize the distribution. In particular, Tablets provide the following: Note that you can run a cluster with some of the Keyspaces with Tablets disabled, and some with tablets enabled for as long as you wish. The scaling improvement will be partial, and limited to Keyspaces with Tablets enabled. By default, the number of tablets is dynamically set by the database, and there is no need to set it. However, if you know in advance the expected size of the data, you can speed up initial data ingestion, by setting the initial tablet number to: expected size (before replication) divided by 5GB. tablets = { 'initial': 2048}; Follow up patch release will deprecate this parameter for more flexible per table tablet options . The goal of Tablet Merge is to reduce the tablet count for a shrinking table. Similar to how split increases the count while the table is growing. The load balancer decision to merge is implemented today (came with infrastructure introduced for split), but it wasn’t handled until now. The topology coordinator will now detect tables that have shrunk, and merge adjacent tablets in order to meet the average tablet replica size goal. #18181 Tablets Keyspaces are not yet enabled for the following features: Alternator support Tablets with the following: Read more about Tablets here. Note: you can not ALTER an existing Keyspace to switch between Tablets and vNode based table and back. We will remove these restrictions in upcoming releases. More improvments here: [RELEASE] ScyllaDB 2025.1 - part 2 --- ### Page: https://forum.scylladb.com/t/release-scylladb-2025-1-part-2/4693 Title: [RELEASE] ScyllaDB 2025.1 - part 2 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: Additional Improvements Procedures With Tablets, the Replication Factor (RF) cannot be updated to a value higher than the number of nodes per Data Center (DC). This feature protects the Admin from setting an impossible-t… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-2025-1-part-2/4693 ## Headings Structure: H1: [RELEASE] ScyllaDB 2025.1 - part 2 H2: Additional Improvements H3: File-based streaming for Tablets H3: Arbiter and Zero-token Node H3: Strongly Consistent Topology Updates H3: Strongly Consistent Auth Updates H3: Strongly Consistent Service Levels H3: Improved network compression for intra-node RPC H3: Describe Schema with Internals H3: Native Nodetool H3: Removing the JMX Server H3: Maintenance Mode H3: Maintenance Socket H3: Deployment H3: Service Level Per Query H3: DESCRIBE SCHEMA enhance H3: Alternator RBAC H3: CQL H3: Deployments and packaging H3: Stability H3: Performance H3: Alternator H3: Tablets H3: API H3: Tooling H3: Monitoring H3: Security H3: Configuration H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 2025.1 - part 2 H2: Additional Improvements H4: Procedures H4: Monitor Tablets H4: Driver Support H3: File-based streaming for Tablets H3: Arbiter and Zero-token Node H3: Strongly Consistent Topology Updates H3: Strongly Consistent Auth Updates H3: Strongly Consistent Service Levels H3: Improved network compression for intra-node RPC H3: Describe Schema with Internals H3: Native Nodetool H3: Removing the JMX Server H3: Maintenance Mode H3: Maintenance Socket H3: Deployment H3: Service Level Per Query H3: DESCRIBE SCHEMA enhance H3: Alternator RBAC H3: CQL H3: Deployments and packaging H3: Stability H3: Performance H3: Alternator H3: Tablets H3: API H3: Tooling H3: Monitoring H3: Security H3: Configuration H3: Related topics With Tablets, the Replication Factor (RF) cannot be updated to a value higher than the number of nodes per Data Center (DC). This feature protects the Admin from setting an impossible-to-support RF. This affects the following operations: Node Decommission / Remove Starting from 2024.2, you cannot decommission or remove a node if the resulting number of nodes would be smaller than the largest non-zero replication factor (for any keyspace) in this DC 1 DC, 5 nodes, a KS with RF=5 The decommission request will fail The Replication Factor (RF) of Keyspaces must be less than or equal to the number of available nodes per Data Center (DC) Once a tablets-enabled Keyspace has tables, you can not ALTER its Replication Factor to be greater than the number of available nodes per DC. If you create such a Keyspace, you won’t be able to create Tables until you fix the RF or add more nodes. To Monitor Tablets in real time, upgrade ScyllaDB Monitoring Stack to release 4.7, and use the new dynamic Tablet panels, below. The following drivers versions and newer support Tablets Legacy ScyllaDB and Apache Cassandra drivers will continue to work with ScyllaDB but will be less efficient when working with tablet-based Keyspaces. File-based streaming is an optimization of tablet migration performance. In ScyllaDB Open Source, migrating tablets is performed by streaming mutation fragments, which involves deserializing SSTable files into mutation fragments and re-serializing them back into SSTables on the other node. In ScyllaDB 2025.1, migrating tablets is performed by streaming entire SStables, which does not require (de)serializing or processing mutation fragments. As a result, less data is streamed over the network, and less CPU is consumed, especially for data models that contain small cells. File-based streaming is used for tablet migration in all keyspaces created with tablets enabled. There is now support for zero-token nodes. Such nodes do not replicate any data, but can participate in query coordination, and in Raft quorum voting. One can use this to create an Arbiter: a tiebreaker node, with no data, that can help maintain quorum in the case of a symmetrical two-datacenter cluster. If one of the data centers fails, the Arbiter, deployed on a 3rd datacenter, keeps quorum on the node alive. Since the Arbiter has zero token, it does not replicate user data, and does not come with network and storage costs. #15360 You can use nodetool status for a list of zero token nodes. With Raft-managed topology enabled, all topology operations are internally sequenced consistently. A centralized coordination process ensures that topology metadata is synchronized across the nodes on each step of a topology change procedure. This makes topology updates fast and safe, as the cluster administrator can trigger many topology operations concurrently, and the coordination process will safely drive all of them to completion. For example, multiple nodes can be bootstrapped concurrently, which couldn’t be done with the previous gossip-based topology. Strongly Consistent Topology Updates is now the default for new clusters, and should be enabled after upgrade for existing clusters. System-auth-2 is a reimplementation of the Authentication and Authorization systems in a strongly consistent way on top of the Raft sub-system. This means that Role-Based Access Control (RBAC) commands like create role or grant permission are safe to run in parallel without a risk of getting out of sync with themselves and other metadata operations, like schema changes. As a result, there is no need to update system_auth RF or run repair when adding a DataCenter. Service Levels allow you to define attributes like timeout per workload. Service levels are now strongly consistent using Raft, like Schema, Topology and Auth. This release adds new RPC compression improvements for node to node communication: Below is a comparison of compressions algorithms on different types of data. Note that dictionary based compression can be used with either lz4 or zstd. Actual compression is very much workload-dependent and can vary between use cases. Until this release, CQL DESCRIBE SCHEMA was not sufficient to do a full schema restore from backup. For example, it lacks information about dropped columns. In 6.0, the DESC SCHEMA WITH INTERNALS command provides more information, streamlining the restore process. The nodetool utility provides simple command-line interface operations and attributes. ScyllaDB inherited the Java based nodetool from Apache Cassandra. In this release, the Java implementation was replaced with a backward-compatible native nodetool. The native nodetool works much faster. Unlike the Java version ,the native nodetool is part of the ScyllaDB repo, and allows easier and faster updates. With the Native Nodetool (above), the JMX server has become redundant and will no longer be part of the default ScyllaDB Installation or image. If you are using the JMX server directly, not via nodetool, you can either: Related issues: #15588 #18566 #18472 #18566 As part of moving to native tooling and away from Java tools, we deprecated SSTableloader. You can use the Load and Stream to upload SSTables directly to Scylla, either from Apache Cassandra or other ScyllaDB clusters. We are also deprecating the Java version of nodetool, which was replaced by a compatible native version (see above). Maintenance mode is a new mode in which the node does not communicate with clients or other nodes and only listens to the local maintenance socket and the REST API. It can be used to fix damaged nodes – for example, by using nodetool compact or nodetool scrub. In maintenance mode, ScyllaDB skips loading tablet metadata if it is corrupted to allow an administrator to fix it. The Maintenance Socket provides a new way to interact with ScyllaDB from within the node it runs on. It is mainly for debugging. You can use CQLSh with the Maintenance Socket as described in the Maintenance Socket docs. #16172 Enhancement: override service level per query using USING SERVICE LEVEL = name It is now possible to reroute an individual statement to a different service level; previously a different login session was required. This is useful for drivers to reduce the strain from the login queries. In OSS, this only affects the statement’s timeout. The DESCRIBE SCHEMA statement is now extended with statements to re-create roles and grants. This can be used to re-create not only the schema, but also the user and permission structure when restoring from backup. Authorization: Alternator now supports Role-Based Access Control (RBAC) via CQL commands. #5047 CQL3: implement NOT IN #21992 select * from TBL where v NOT IN (5,7) ALLOW FILTERING CQL3: Allow selecting map values and set elements, compatible to Cassandra 4.0 SELECT map['key'] FROM table SELECT map['key1']['key2'] FROM table The CREATE MATERIALIZED VIEW statement now supports the undocumented WITH ID clause, improving compatibility with Cassandra. #20616 The memtable_flush_period_in_ms option is now implemented. #20270 The CREATE ROLE USING SALTED HASH statement was renamed to CREATE ROLE USING HASHED PASSWORD for improved compatibility with Apache Cassandra. #21350. See Grant Authorization CQL Reference | ScyllaDB Docs CQL DESCRIBE statements for Change Data Capture (CDC) log tables have been improved. #21235 The DESC TABLE statement will now reject materialized views. #21026 The CQL PER PARTITION LIMIT clause is now respected for aggregating queries, fixing the an issue when combining PER PARTITION LIMIT and GROUP BY #5363 New COMPACT STORAGE tables can no longer be created. They have been deprecated for a long while. #16403 Failing to create table with sstable_compression=ZstdCompressor #22444 The index page cache will now generate fewer disk IOPS if an index read is partially cached. #20935 Raft-managed tables used for system metadata now have more eager garbage collection of tombstones, reducing performance problems with many schema or topology changes. #15607 The system.peers table was continuously updated even if no change was happening, stressing the disk. #20991 The sstable reader will now consult data in memtable before purging tombstones. This prevents data resurrection in scenarios involving very low write activity, which can lead to data staying in memtables for longer than a repair cycle. #20916 The efficiency of sstable reads rows within medium or large partitions, when column_index_size_in_kb has been reduced, is now improved. Such reads will generate less I/O. #10030 ScyllaDB tracks whether read requests are waiting for CPU or I/O. In one case, a disk read from the primary index was considered to be waiting on CPU, which reduced concurrency. This is now fixed. #21325 Materialized View building (initiated by CREATE MATERIALIZED VIEW or CREATE INDEX) is now performance-isolated from normal reads and writes. #21232 Repair flushes hints and batchlog in order to reduce the amount of work it has to do, but such flushes also generate work, so these flushes are now batched. #20259 Some performance bugs leading to extra I/O when reading the primary index for a large partition are fixed. #20897 Repair performance in mixed-shard configurations (where different nodes have different shard counts) has been improved. #21113 The sstable reader now frees memory more quickly, reducing memory requirements. #21160 During ordinary sstable compaction, we do not purge tombstones if they potentially delete data in commitlog, to avoid data resurrection on restart. However, this is unnecessary for the row cache, so row cache now ignores commitlog when purging tombstones. #16781 The materialized view update process updates views when the base tables are updated by an UPDATE or INSERT query. It is now able to avoid unnecessary updates when a view’s PK has a regular column #21652 Bootstrap and decommission now enable the small-table repair optimization. This speeds up bootstrap in large clusters when small or empty system tables have to be migrated to other nodes. #19131 ScyllaDB no longer takes snapshots of materialized views, since they regenerated from the base table at restore time. #21339 #20760 types: large allocations and stalls while comparing Decimal values #21716 The data plane coordination code (“storage_proxy”) now uses host UUIDs to track hosts rather than network addresses. This simplifies the code and brings a nice performance improvement. Part of #6403 Node rebuilds that use repair-based node operations now apply the small-table optimization when beneficial. #21951 We now create XFS filesystems with reduced metadata overhead. #22028 Materialized views pair each view replica with a base replica. This pairing is now rack-aware - the database will prefer to pair a base replica and a view replica on the same rack. This reduces rack crossings which can be expensive on public clouds, and generally have lower bandwidth and higher latency. #17147 ScyllaDB breaks long query results into pages to reduce transient memory consumption and latency. When it does so, it caches the query running on the replica and resumes it on the next page. This resuming broke when a paging decision was made due to a large number of tombstones, requiring the query to be restarted on the next page instead of resumed. This is now fixed. #22620 The reader concurrency semaphore tries to restrict the number of reads competing for CPU, since the competition delays all those reads. We now allow up to two reads to compete for the CPU in the default configuration. This allows common fast reads to compete and bypass rare long reads, reducing head-of-line blocking. #22450 The reader concurrency semaphore could accidentally use the query timeout for evicting cached queries, resulting in reduced performance. This is now fixed. #22629 Scylla Monitoring Stack released 4.9 and later supports ScyllaDB 2025.1 New config parameters: Until this release, the materialized view flow-control algorithm used a constant delay_limit_us hard-coded to one second, which means that when the size of view-update backlog reached the maximum (10% of memory), we delay every request by an additional second - while smaller amounts of backlog will result in smaller delays. So this patch replaces the hard-coded default with a live-updateable configuration parameter, view_flow_control_delay_limit_in_ms, which defaults to 1000ms as before. #18187 View_flow_control_delay_limit_in_ms: maximal amount of time that materialized-view update flow control may delay responses to try to slow down the client and prevent buildup of unfinished view updates. To be effective, this maximal delay should be larger than the typical latencies. Setting view_flow_control_delay_limit_in_ms to 0 disables view-update flow control. The small-table optimization for repair-based node operations is now enabled by default. This speeds up bootstrap and decommission operations for clusters with small amounts of data. #21861 --- ### Page: https://forum.scylladb.com/t/document-review-and-update/4716 Title: Document review and update - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi Team, I have analysed ur document and got this information. Can you please review and update ur document with below changes. Document link: Adding a New Node Into an Existing ScyllaDB Cluster (Out Scale) | ScyllaDB … Language: en Canonical URL: https://forum.scylladb.com/t/document-review-and-update/4716 ## Headings Structure: H1: Document review and update H3: Related topics ## Main Content: H1: Document review and update H3: Related topics I have analysed ur document and got this information. Can you please review and update ur document with below changes. Document link: Adding a New Node Into an Existing ScyllaDB Cluster (Out Scale) | ScyllaDB Docs Changes required: Under procedure section: after 2nd point ( On each node, edit the scylla.yaml file /etc/scylla/ to configure the parameters listed below.) 3rd point should not be the start scylla-server on new node. Before starting scylla-server, the 3rd point should be "we need to be update the /etc/scylla/cassandra-rackdc.properties file on new node with collecting info by the existing node /etc/scylla/cassandra-rackdc.properties file . After updating the DC, rack information on new node we can start the scylla-server Scylla Team, Please review this and update ur documentation accordingly with the required process. etc/scylla/cassandra-rackdc.properties Dear @CHARAN_CHINTHA , Thank you for your feedback, I have submitted an issue with your suggestion for our team to review, Best regards, Gabriel --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-90-2025-04-11/4719 Title: Last week in scylla-cluster-tests.git master (issue #90; 2025-04-11) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 09f87bb7…1208c230 range are covered. There were 14 non-merge commits from 9 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-90-2025-04-11/4719 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #90; 2025-04-11) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #90; 2025-04-11) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 09f87bb7…1208c230 range are covered. There were 14 non-merge commits from 9 authors in that period. Some notable commits: As the new minor release of Manager 3.5.0 is out, it’s now the default version used in tests, and we’re also covering the upgrade path from 3.4.*. A fix for automatic backporting to perf test branches was made, addressing previous regexp limitations. ScyllaYaml was updated to support tablets_mode_for_new_keyspaces, now preferred in ScyllaDB 2025.1+ over enable_tablets, which remains for backward compatibility. hydra commands like update-conf-docs, nemesis-list, and create-nemesis-pipelines no longer require Okta credentials. To avoid failing tests due to transient issues, we’re now ignoring raft topology connection close errors globally, which could appear due to race conditions between raft and gossip. This will be dropped once gossip is removed. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/reader-concurrency-semaphore-and-timeout-errors/4729 Title: Reader_concurrency_semaphore and timeout errors - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello guys, I have two ScyllaDB clusters one for production and one for test(specs are below). We are using only the alternator features but time to time on both we are facing the reader_concurrency_semaphore errors and… Language: en Canonical URL: https://forum.scylladb.com/t/reader-concurrency-semaphore-and-timeout-errors/4729 ## Headings Structure: H1: Reader_concurrency_semaphore and timeout errors H3: Related topics ## Main Content: H1: Reader_concurrency_semaphore and timeout errors H3: Related topics Hello guys, I have two ScyllaDB clusters one for production and one for test(specs are below). We are using only the alternator features but time to time on both we are facing the reader_concurrency_semaphore errors and sometimes the “Operation timed out” errors. (Added them below) The funny thing is on both clusters load is almost zero and total data size less than 100 MB. (Literally we got almost no load) I checked the logs and cluster health but everything looks good. Do you have any idea why we are getting these errors? One thing I notice is from the installation configs. The file /etc/scylla.d/cpuset.conf was empty all the nodes on the clusters. (The one other interesting this is I am able to run scylla when that file is empty) I am not sure that is the case but just for the test environment I changed it to “–cpuset 0-3” now. Example Logs From Test Environment: On one of the other nodes: #ScyllaDB version: 6.1.2 compaction strategy : SizeTieredCompactionStrategy #Cluster size: 3 Nodes - (TEST: 4 Core / 8 GB Ram) (PROD: 16 Core / 32 GB Ram) #OS: Rocky Linux 9.4 I want to add one more thing operations that uses LWT taking way too long to complete. here the screenshot from our monitoring stack. You can see there is not a lot of operations but especially Delete and Put Item operations taking too long. Also the CPU load is near zero. Hello again, After long long research. I changed HAProxy load balancing algorithm from roundrobin to source and now max latency around 500 - 700ms. We have a HAProxy before ScyllaDB nodes. I’m still not sure what the root cause of the problem is. I think these timeouts are related to LWT, which Alternator uses behind the scenes. I recall seeing such timeout reports with virtually no loads, due to problems in the LWT implementation. I recommend upgrading to a later version, we probably fixed this but didn’t backport to 6.1 due to EOL. --- ### Page: https://forum.scylladb.com/t/scylladb-community-version/4731 Title: ScyllaDB community version - Database Community - ScyllaDB Community NoSQL Forum Meta Description: Hello ScyllaDB community, I’m excited to learn about ScyllaDB. I’m looking for a community version of ScyllaDB for students that does not expire. ScyllaDB website offers a free 30 day trial of Enterprise version. Does … Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-community-version/4731 ## Headings Structure: H1: ScyllaDB community version H3: Related topics ## Main Content: H1: ScyllaDB community version H3: Related topics Hello ScyllaDB community, I’m excited to learn about ScyllaDB. I’m looking for a community version of ScyllaDB for students that does not expire. ScyllaDB website offers a free 30 day trial of Enterprise version. Does anyone know if the software completely stops working after 30 days? I want to know that I will be able to do hands-on practice before investing time and effort going through the ScyllaDB university courses. Just follow the instructions from Install ScyllaDB | ScyllaDB Docs You should be good to go, without any time/expiration limits. Best regards, Gabriel. --- ### Page: https://forum.scylladb.com/t/scylladb-manager-repair-taking-too-long-and-crashing/4732 Title: ScyllaDB Manager - repair taking too long and crashing - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-manager-repair-taking-too-long-and-crashing/4732 ## Headings Structure: H1: ScyllaDB Manager - repair taking too long and crashing H3: Related topics ## Main Content: H1: ScyllaDB Manager - repair taking too long and crashing H3: Related topics Originally from the User Slack @Matheus_Salvia: I’ve been having a lot of trouble running repairs on my cluster (scylla oss 6.2) repairs constantly go over 100% until it eventually crashes. also these tables take weeks to run. see attached Any ideas what can be going wrong here and how to debug it? running repair with sctool repair -c default/scylla -s now --intensity 2 @avi: Try running repair on the slowest table, with nodetool repair, to start decoupling components from the problem @Guy: Hey @Matheus_Salvia did you figure this out? @Matheus_Salvia: not yet, still investigating. looks like scylla manager wasn’t touching some nodes, as evidenced by some logs when I started a manual repair in a small table in one node repair session was 178 and in another it was 1, so I assume this was never ran before. What’s more funny, this table just showed as 100% in the scylla manager logs, I wouldn’t think there would be a problem here i.e. this isn’t one of the big problematic tables that takes a million hours, it’s a pretty small one that scylla manager showed as 100% done --- ### Page: https://forum.scylladb.com/t/gossip-and-raft-consensus-algorithm-usage-in-scylladb/4733 Title: Gossip and Raft Consensus Algorithm Usage in ScyllaDB - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/gossip-and-raft-consensus-algorithm-usage-in-scylladb/4733 ## Headings Structure: H1: Gossip and Raft Consensus Algorithm Usage in ScyllaDB H3: Related topics ## Main Content: H1: Gossip and Raft Consensus Algorithm Usage in ScyllaDB H3: Related topics Originally from the User Slack @Cong_Guo: Hi, I have a question that since Scylla has switched to Raft, why does the gossip module still exist? Are there any tasks that can only be handled by gossip and not by Raft? @avi: Yes, things like failure detection and disseminating some statistics --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-276-2025-04-14/4735 Title: Last week in scylladb.git master (issue #276; 2025-04-14) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 431de48df9..10589e966f range are covered. There were 104 non-merge commits from 23 authors in that per… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-276-2025-04-14/4735 ## Headings Structure: H1: Last week in scylladb.git master (issue #276; 2025-04-14) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #276; 2025-04-14) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 431de48df9..10589e966f range are covered. There were 104 non-merge commits from 23 authors in that period. Some notable commits: The S3 driver now avoids copying data during multipart uploads. A crash when TRUNCATE or DROP TABLE happened after tablet migration was fixed. The row cache garbage-collects tombstones in cache to improve read performance. It now checks for overlap with data in memtables, as this could cause data resurrection in some cases. There is now a dedicated nodetool cluster repair command to repair tablet keyspaces. Unlike most nodetool commands, it applies to the entire cluster. Backup will prioritize sstables that were deleted (usually as the result of compaction), as deletes sstables occupy space in snapshots. The CQL binary protocol server now throttles new connection processing in order to prevent connection storms from overwhelming the server. When rebuilding a tablet (due to the loss of a node), we will now stream data from just one replica, and use repair to fill in data from the rest. This saves bandwidth and reduces space amplification. A bug which could cause SELECT … PER PARTITION LIMIT or SELECT DISTINCT to terminate prematurely was fixed. There is now a new virtual table that describes load per node, and the tablet monitoring script was updated to make use of it. This is useful for heterogeneous clusters where different nodes have different storage capacity. The raft group 0 implementation now limit the number of raft voters in order to reduce the amount of work needed to reach consensus. Nodes are promoted to voters or demoted as needed. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-cpp-over-rust-driver-0-4-0/4738 Title: [RELEASE] ScyllaDB CPP-over-Rust Driver 0.4.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce ScyllaDB CPP Rust Driver 0.4.0, an API-compatible rewrite of GitHub - scylladb/cpp-driver: Scylla C/C++ Driver as a wrapper for the Rust driver. It will fully replace the CPP driv… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-cpp-over-rust-driver-0-4-0/4738 ## Headings Structure: H1: [RELEASE] ScyllaDB CPP-over-Rust Driver 0.4.0 H2: Changes H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB CPP-over-Rust Driver 0.4.0 H2: Changes H3: Related topics The ScyllaDB team is pleased to announce ScyllaDB CPP Rust Driver 0.4.0, an API-compatible rewrite of GitHub - scylladb/cpp-driver: Scylla C/C++ Driver as a wrapper for the Rust driver. It will fully replace the CPP driver, which is already reaching its End of Life. We are also happy to announce, that starting from this release, the driver version can be considered Beta: Some minor features still need to be included. See Limitations section in README.md. The underlying Rust driver used version: 1.1.0. Implemented API functions: New features / enhancements CI / developer tool improvements: Congrats to all contributors and thanks everyone for using our driver! The source code of the driver can be found here: Thank you for your attention, please do not hesitate to contact us if you have any questions, issues, feature requests, or are simply interested in our driver! Contributors since the last release: --- ### Page: https://forum.scylladb.com/t/tombstone-warnings-for-system-schema-tables-using-kubernetes-frequently-creating-and-deleting-tables/4740 Title: Tombstone warnings for system_schema.tables, using Kubernetes, frequently creating and deleting tables - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/tombstone-warnings-for-system-schema-tables-using-kubernetes-frequently-creating-and-deleting-tables/4740 ## Headings Structure: H1: Tombstone warnings for system_schema.tables, using Kubernetes, frequently creating and deleting tables H3: Related topics ## Main Content: H1: Tombstone warnings for system_schema.tables, using Kubernetes, frequently creating and deleting tables H3: Related topics Originally from the User Slack @Nick_Vladiceanu: Hey there! We are seeing warnings regarding the Tombstones for system_schema.tables: with that, we see some degradation in performance, mostly reads become slower. Could that be the cause? We do create and delete tables quite often. Thanks @avi: What version are you running? Are you constantly creating and dropping tables? It shouldn’t affect data-path performance. @Nick_Vladiceanu: we are currently using version 5.2.18 in Kubernetes. Every 2-3 hours a bunch of new tables get created and old ones deleted (~15 tables). Any suggestion where to look into? We start observing performance degradation, like slower responses and clients report errors connecting to Scylla @avi: 5.2 is long out of support. Newer versions don’t have that tombstone in schema problem esp. once you enable raft for schema management @Nick_Vladiceanu: could that tombstone problem be the reason of read performance degradation? @avi: It’s unlikely. It slows down driver connection, not not regular reads Maybe frequent driver connections are slowing down the cluster @Nick_Vladiceanu: ok, so the tombstone problem might be the reason of and as a consequence, this might trigger a reconnect on the driver(s) which might affect the performance of the cluster Thanks a lot, I’m looking into what is the next version I could upgrade to. Is there any page with the list of supported versions of OS Scylla? @Guy: AFAIR, the last two major versions are supported, so that would now be 6.2. You can read more in this FAQ. --- ### Page: https://forum.scylladb.com/t/benchmarking-with-one-node-how-to-properly-configure-i-o/4745 Title: Benchmarking with one node, how to properly configure I/O? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/benchmarking-with-one-node-how-to-properly-configure-i-o/4745 ## Headings Structure: H1: Benchmarking with one node, how to properly configure I/O? H3: Related topics ## Main Content: H1: Benchmarking with one node, how to properly configure I/O? H3: Related topics Originally from the User Slack @Chenhao_Ye: A very naive starter question: I want to make a simple benchmark with one node. I am able to run ./tools/toolchain/dbuild ./build/release/scylla --developer-mode 1 , but if removed --developer-mode 1, it will require me to run scylla_io_setup. Now if I directly run ./tools/toolchain/dbuild ./dist/common/scripts/scylla_io_setup I will get an error ModuleNotFoundError: No module named 'scylla_product'. I assume scylla_product.py is created by some other install scripts? My question is: Is there any easy way for me to run such a quick benchmark with I/O properly configured? @avi: Build a package (ninja dist-dev or dist-release) and install the package --- ### Page: https://forum.scylladb.com/t/release-scylladb-2025-1-1/4746 Title: [RELEASE] ScyllaDB 2025.1.1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB 2025.1.1, a bug-fix production-ready patch release for ScyllaDB 2025.1 LTS Release. capacity-aware load balancing of Tablets. Until this release, Tablets load-balancing was based so… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-2025-1-1/4746 ## Headings Structure: H1: [RELEASE] ScyllaDB 2025.1.1 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 2025.1.1 H3: Related topics The ScyllaDB team announces ScyllaDB 2025.1.1, a bug-fix production-ready patch release for ScyllaDB 2025.1 LTS Release. capacity-aware load balancing of Tablets. Until this release, Tablets load-balancing was based solely on the number of Tablets. This is suboptimal since, in a heterogeneous cluster, each node can have a different storage size. #23443 The following example is of a cluster with two nodes in each rack: i4i.8xlarge (7.5TB storage) and i4i.large (0.468TB). Before, disk usage differed between nodes, even though the bytes stored were similar. After - disk usage is now similar (around the 90% goal), while the bytes storaged differed 2025.1.1 also includes multiple bug fixes (below). The following issues are fixed in this release: --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-91-2025-04-18/4747 Title: Last week in scylla-cluster-tests.git master (issue #91; 2025-04-18) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the c4b2558f…d16d1243 range are covered. There were 10 non-merge commits from 7 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-91-2025-04-18/4747 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #91; 2025-04-18) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #91; 2025-04-18) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the c4b2558f…d16d1243 range are covered. There were 10 non-merge commits from 7 authors in that period. Some notable commits: We removed nemesis classes used for filtering and replaced them with a nemesis_selector option in test configs for SisyphusNemesis, simplifying logic and reducing duplication. Performance ‘gradual steps’ test now runs with c-s 3.17.5. disrupt_refuse_connection_with_* nemeses were extended to support IPv6, and improved to skip critical loader errors when connecting to banned nodes. Because in recent Scylla not all nodes are raft voters, SCT healthchecks were adjusted for that. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/mutliple-datacenter-cluster-diagnosing-high-latency-spike-and-performance-issues/4749 Title: Mutliple Datacenter cluster, diagnosing high latency spike and performance issues - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/mutliple-datacenter-cluster-diagnosing-high-latency-spike-and-performance-issues/4749 ## Headings Structure: H1: Mutliple Datacenter cluster, diagnosing high latency spike and performance issues H3: Related topics ## Main Content: H1: Mutliple Datacenter cluster, diagnosing high latency spike and performance issues H3: Related topics Originally from the User Slack @Igor_Q: hey, guys i have a cluster of 3 datacenters, using scylla v6.0.4 it worked fine for several months, until yesterday read latency went through the roof up to 5 seconds in two datacenters i’ve contemplated performing rolling restart, and after rebooting a single node everything seemed fine (read latency dropped to appropriate values) but today this happened again, now read latency maxed out in all 3 datacenters how do i diagnose the problem? i see a bunch of metrics in monitoring stack, but i can’t make out, which are the cause and which are the consequence also, i’ve noticed that upon rebooting compaction queue becomes really large (40+ compactions) – aren’t compactions supposed to run in the background? @avi: You can look at the Advanced dashboard (in per-shard mode) to see whether CPU or I/O are the bottleneck, and for which scheduling group. Compactions aren’t supposed to be affected by restarts. @Igor_Q: > Compactions aren’t supposed to be affected by restarts. you mean they shouldn’t spike when a node is restarted, right? @avi: Right Check if the compaction type is RESHAPE, we had bugs in this area where unnecessary reshapes were generated btw 40 compactions across the cluster isn’t a lot @Igor_Q: > Check if the compaction type is RESHAPE, we had bugs in this area where unnecessary reshapes were generated is there a way to see this in nodetool? i don’t see the type in nodetool compactionhistory also, is it possible that background compactions in my case are somehow broken? after i removed all load from the cluster and manually ran compaction on each node, the problem seems to be fixed if so, then how can i monitor for such cases? check compaction history once in a while? @avi: The logs show the compaction type. Hard to say what to do, I don’t have a clear image of what’s going on. @Igor_Q: i can provide the necessary info if you’re willing to look into it @avi: Post snapshots of the Advanced dashboard in per-shard mode when the event happens @Igor_Q: this proved to be a non-trivial task, but you should be able to see the snapshot here: https://grafs.sonc.top/dashboard/snapshot/c9S7NR1Alroyni1c0FAtpEcdved0Xx2E?orgId=0 this is the first occurrence: read latency started growing up to 5 seconds at 19:13 (correlates with C++ exceptions) this is the second occurrence: https://grafs.sonc.top/dashboard/snapshot/P7yJOe24LjU8A6iSKM0nvLz5I1rid5Vm?orgId=1 probably worth mentioning that we retry queries at most 1 time – this is due to high retransmit timeouts in our network but this never posed a problem until the incidents @avi: Looks like you’re out of CPU on shard 48. Probably have an imbalanced workload (use nodetool toppartitions) @Igor_Q: Thank you. But how is it possible that the same shard across all nodes in two datacenters (10.200., 10.144.) is getting overloaded? First of all, aren’t these all different shards? We have replication factor of 1, it seems pretty weird that the same shard 48 both on 10.144.65.37 and 10.144.65.14 is out of CPU. These are two different shards, they shouldn’t share any data, no? What am I missing here? I also see the same picture on “Detailed” dashboard. Annotation on “Reads per Shard – Coordinator” states “Amount of requests served as the coordinator. Imbalances here represent dispersion at the connection level, not your data model”. Should this be understood as “the driver decided to flood shard 48 with requests on multiple nodes due to its internal logic”? @avi Would you mind taking a look here, please? @avi: Reactor load per shard doesn’t give enough information, use the CPU panels in the advanced dashboard to see which scheduling group uses the CPU @Igor_Q: Didn’t you already link the corresponding panel on the Advanced tab here? I see no other significant load on these shards in this time interval. Does “statement” scheduling group include coordination? [April 9th, 2025 4:30 AM] avi: Looks like you’re out of CPU on shard 48. Probably have an imbalanced workload (use nodetool toppartitions) @avi: Statement includes coordination. With the shard-aware driver, the driver sends requests directly to the shards that own the data. So make sure you’re using a shard-aware driver. These problems are especially noticeable with large shard counts --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-277-2025-04-20/4750 Title: Last week in scylladb.git master (issue #277; 2025-04-20) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 10589e966f..2314feeae2 range are covered. There were 103 non-merge commits from 19 authors in that per… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-277-2025-04-20/4750 ## Headings Structure: H1: Last week in scylladb.git master (issue #277; 2025-04-20) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #277; 2025-04-20) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 10589e966f..2314feeae2 range are covered. There were 103 non-merge commits from 19 authors in that period. Some notable commits: The scylla-tools-java submodule (and associated scylla-tools rpm and deb subpackages) was removed from the source tree. It was previously used to supply nodetool (replaced by native nodetool) and cassandra-stress (here replaced by a pre-packaged cassandra-stress artifact) for profile guided optimization training. The managed_bytes type is used to represent blobs in the log-structured allocator, allowing the system to break up large blobs into small fragments. A bug in determining the boundaries where large blobs are broken up caused in turn a very old bug in the log-structured allocator to trigger, causing the system to think it ran out of memory where in fact it had plenty free. This resulted in crashes and in very small memtable flushes, causing performance problems. We now determine the split boundary correctly. Native backup now uploads from all shards simultaneously, rather than just from shard 0. This results in a substantial bandwidth improvement. The S3 driver now supports the CopyObject API, allowing copies to be performed by the object storage system. This is not yet used. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-17/4751 Title: [RELEASE] ScyllaDB Enterprise 2024.1.17 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2024.1.17, a bug-fix production-ready patch release for ScyllaDB Enterprise 2024.1 Long Term Support (LTS) Release. Note there is a later, better, LTS (Long term Support) R… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-17/4751 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2024.1.17 H3: Related Links H2: Fixed Issue with an open-source reference: H2: Stability H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2024.1.17 H3: Related Links H2: Fixed Issue with an open-source reference: H2: Stability H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2024.1.17, a bug-fix production-ready patch release for ScyllaDB Enterprise 2024.1 Long Term Support (LTS) Release. Note there is a later, better, LTS (Long term Support) Release 2025.1, and you are encouraged to upgrade to it. --- ### Page: https://forum.scylladb.com/t/multi-dc-cluster-query-performance-with-batches-gocql-driver/4752 Title: Multi DC cluster, query performance with batches, GOCQL driver - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/multi-dc-cluster-query-performance-with-batches-gocql-driver/4752 ## Headings Structure: H1: Multi DC cluster, query performance with batches, GOCQL driver H3: Related topics ## Main Content: H1: Multi DC cluster, query performance with batches, GOCQL driver H3: Related topics Originally from the User Slack @Igor_Q: Hi again, everyone. I have a cluster of 3 DC x 3 nodes of ScyllaDB 6.0.4. Table has the following schema: Keyspace has a replication factor of 1. At the moment we’re using it to read data by ID. Each request contains 1500 IDs, we split them into batches of 75 IDs and request them from Scylla in parallel (WHERE id IN ?). We’re using the fork of gocql. Now, I’ve read in multiple places, that it’s generally advisable to request data by a single key in parallel (WHERE id = ?) – this would allow the driver to select the correct shard by token-aware policy and remove the coordination load from Scylla. I’m exploring the possibility of moving to such implementation, but I see weird results in the benchmark. First of all, selecting the keys by one does not come near reading them in batches. Query traces nevertheless show that the queries are executed by the owning replica. Is network solely responsible for such difference in latency? Second, selecting keys by one in 1500 goroutines shows consistently slower results than creating two sessions and using only 75 goroutines in each (!). Is there some contention in *gocql.Session? What am I doing wrong? GitHub: GitHub - scylladb/gocql: Package gocql implements a fast and robust ScyllaDB client for the Go programming language. @avi: You have one node per DC, so the batches always land on the right node. You save same CPU by using a batch, and don’t lose anything due to queries that have to be rerouted, so batches win. In normal clusters that have more nodes, sending individual queries is better. About gocql, I don’t know enough to comment. @Igor_Q: > You have one node per DC 3 DC x 3 nodes = 3 nodes in each datacenter Total of 9 nodes across all datacenters. @avi: Then I don’t have an explanation, maybe it’s gocql. @Dmitry_Kropachev, any idea? See the discussion on GitHub for more info: Some kind of contention in driver? · Issue #432 · scylladb/gocql · GitHub --- ### Page: https://forum.scylladb.com/t/could-scylla-identify-az-level-failures/4754 Title: Could scylla identify AZ-level failures - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: #Cluster size: os (RHEL/CentOS/Ubuntu/AWS AMI): If there is failure in a datacenter, for example, a cluster with 6 replicas, which means every datacenter owns 2 replicas, when… Language: en Canonical URL: https://forum.scylladb.com/t/could-scylla-identify-az-level-failures/4754 ## Headings Structure: H1: Could scylla identify AZ-level failures H3: Best Practices H3: To summarize - H3: Related topics ## Main Content: H1: Could scylla identify AZ-level failures H4: Example: H4: Scenario: H3: Best Practices H3: To summarize - H3: Related topics Installation details #ScyllaDB version: #Cluster size: os (RHEL/CentOS/Ubuntu/AWS AMI): If there is failure in a datacenter, for example, a cluster with 6 replicas, which means every datacenter owns 2 replicas, when the 2 replica suddenly down without data stream, could scylla aware an az-level failure ? If there is failure in a datacenter, for example, a cluster with 6 replicas, which means every datacenter owns 2 replicas, when the 2 replica suddenly down without data stream, could scylla aware an az-level failure ? A best practice is to consider an Availability Zone (AZ) as a Rack ScyllaDB uses the NetworkTopologyStrategy and supports rack awareness via Snitch configuration (like GossipingPropertyFileSnitch). When you define racks within a datacenter, Scylla tries to spread replicas across racks to improve resilience — assuming you have more replicas than racks. Then yes, ScyllaDB will: Consider a failure like this, cluster as I describe above, a put req send to 6 nodes and two nodes in each az will get the req. Suddenly, two nodes in az-1 is down without response, another node in az-2 is down without response. Then in this case, 6 replicas with quorum write consistency level is bound to get failed. Would it retry, I means send to other nodes ?? And consider another case, one node suddenly down in each az ?? TLDR; you should use LOCAL_QUORUM instead of QUORUM - which also consumes quite significant bandwidth and will have perfomance impact (due to the need to read/write across DCs) Then, for 3 DCs and 3 Racks, create the keyspaces as below - For the above setup, LOCAL_QUORUM = 2, which is much tollerable for failures Normally I would not recommend a replication factor higher than 3. Thanks for your nice answer. And I have another problem bothering me, Scylla is rack-aware, not AZ-aware, but if each AZ is mapped to a rack correctly via snitch config, it achieves similar behavior. It doesn’t explicitly “detect AZ failure”, but it reacts to node-level failures via Gossip, which can reflect partial AZ failure. Node selection for read/write depends on replica placement, consistency level, and node liveness. When you issue a read/write at a given consistency level, Scylla: No, not directly. Scylla is rack-aware, and you can map AZs to racks via snitch. You use GossipingPropertyFileSnitch and set: for each node to reflect its AZ. Scylla will then try to spread replicas across racks (AZs), and avoid reading/writing to a single rack only when possible. Scylla handles this at the node level, not AZ-level: Scylla reacts like this: Q: If a node is down, how does Scylla select read/write nodes? A: Scylla selects replicas responsible for the token range of the partition. If a replica is down, it skips it and fails the request if not enough live replicas are available to satisfy the consistency level. Q: If two nodes are down in AZ-1 and one in AZ-2, would Scylla be AZ-aware? A: Scylla isn’t AZ-aware per se, but if each AZ is mapped to a separate rack via the snitch config, it behaves AZ-aware. It doesn’t explicitly “know” an AZ is down, but gossip tracks node liveness. If the remaining live replicas aren’t enough to satisfy the consistency level, the request fails. It doesn’t explicitly “know” an AZ is down, but gossip tracks node liveness. If the remaining live replicas aren’t enough to satisfy the consistency level, the request fails. Scylla would not send requests to shutdown nodes, right? if nodes in az-1 is all down, then requests would be sent to nodes that alive in other az , then new request would succeed ? Scylla would not send requests to shutdown nodes, right? if nodes in az-1 is all down, then requests would be sent to nodes that alive in other az , then new request would succeed ? Yes, Scylla will not send requests to nodes that are marked as down via gossip. Referencing the above example, If all nodes in AZ-1 are down, the coordinator will route requests only to live replicas in AZ-2 and AZ-3. The request will succeed if enough live replicas are available to meet the consistency level (e.g. 4 for QUORUM). If not enough live replicas exist, the request will fail immediately. So success depends on both replica distribution and consistency level. --- ### Page: https://forum.scylladb.com/t/will-tombstones-be-frequently-repaired-after-being-compressed/4755 Title: Will tombstones be frequently repaired after being compressed? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 5.2.18 #Cluster size: 35 os (RHEL/CentOS/Ubuntu/AWS AMI): centos I have a 35 node 5 replica cluster and the traffic pattern involves a lot of writes and deletes through altern… Language: en Canonical URL: https://forum.scylladb.com/t/will-tombstones-be-frequently-repaired-after-being-compressed/4755 ## Headings Structure: H1: Will tombstones be frequently repaired after being compressed? H3: Related topics ## Main Content: H1: Will tombstones be frequently repaired after being compressed? H3: Related topics Installation details #ScyllaDB version: 5.2.18 #Cluster size: 35 os (RHEL/CentOS/Ubuntu/AWS AMI): centos I have a 35 node 5 replica cluster and the traffic pattern involves a lot of writes and deletes through alternator. When there are a large number of tombstones in the cluster, cluster regular_compaction cannot clean up the tombstones in time. So I executed nodetool compact on each node one by one to clean up the tombstones. During this time, repair is still being executed. I found that the capacity of the node dropped to 15% after major_compact was executed, but after 5 days the capacity recovered to about 45%. After dumping some sstbale, I found that a large number of tombstones that were originally cleared existed again. I suspect that after my A node executed major_compaction, other nodes have not yet executed it. At this time, repair repaired the tombstone. This cycle continues and the tombstone cannot be cleared away. Do you have any suggestions? Should I stop repair first and then perform compaction on each node again? hi @denesb , do you have any suggestions? thanks. Hello, You can suppress repair from re-distributing expired tombstones by changing enable_tombstone_gc_for_streaming_and_repair. I should add, if the parameter is not (yet) available to you, then you would need to finish major on all nodes before running repair to ensure the tombstones are cleared on all nodes and do not get redistributed. --- ### Page: https://forum.scylladb.com/t/primary-replica-meaning-and-partitionaer-hash-function/4756 Title: Primary Replica meaning, and Partitionaer Hash Function - University and Training - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/primary-replica-meaning-and-partitionaer-hash-function/4756 ## Headings Structure: H1: Primary Replica meaning, and Partitionaer Hash Function H3: Related topics ## Main Content: H1: Primary Replica meaning, and Partitionaer Hash Function H3: Related topics Originally from the User Slack @Daria_Fedorova: Hi! I need some help (maybe scylla university lesson link) I noticed that documentation on nodetool rapair recommends option -pr (–partitioner-range) but I do not understand what it means. Web doc describes it like this --partitioner-range executes a repair only on the primary replica returned by the partitioner. But at my level of understanding partitioner is basicaly a hash function partitoner key to token. I do not know what is a primary replica in this context @Botond_Dénes: With the typical RF=3 replication factor, for each partition-range you have a primary replica and 2 additional replicas. The primary replica is simply the first in the list returned by get_replicas_for_token(t) . The ordering in this list is stable. The primary replica is not special in any way, that said there are some cases like repair -pr where it has a role to play. If you have a 3-node cluster and RF=3, calling nodetool repair on each node, will repair all data 3 times, because nodetool repair repairs all ranges that the node is a replica of. With nodetool repair -pr only ranges for which this node is a primary replica are repaired. Since there is just one primary replica for each range, using -pr guarantees that each range is repaired exactly once, if you call nodetool repair -pr on each node in the cluster. @Daria_Fedorova: thank you --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-2-8/4763 Title: [RELEASE] ScyllaDB Enterprise 2024.2.8 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2024.2.8, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2024.2 Feature (Short Term Support) Release. Note there is a later LTS (Long… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-2-8/4763 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2024.2.8 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2024.2.8 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2024.2.8, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2024.2 Feature (Short Term Support) Release. Note there is a later LTS (Long term Support) Release 2025.1, and you are encouraged to upgrade to it. The following issues are fixed in this release (with an open-source reference, if available): Since Scylla 6.0, cache and memtable cells between 13 kiB and 128 kiB are getting allocated in the standard allocator rather than inside LSA segments. This can result in out of memory issue #22941 #22389 #23781 compaction_manager::drain should wait if stop_ongoing_compactions is already in progress #20197 CDC: Race condition in handling of new generation during raft upgrade. new CDC generation from raft topology upgrade procedure not handled by every node; instead, legacy CDC code tries to obtain it (and fails) #21227 segfault when dumping semaphore diagnostics on SIGQUIT #22756 --- ### Page: https://forum.scylladb.com/t/migrating-from-scylladb-6-1-to-2025-1-snapshots-and-manager/4764 Title: Migrating from ScyllaDB 6.1 to 2025.1, snapshots, and Manager - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/migrating-from-scylladb-6-1-to-2025-1-snapshots-and-manager/4764 ## Headings Structure: H1: Migrating from ScyllaDB 6.1 to 2025.1, snapshots, and Manager H3: Related topics ## Main Content: H1: Migrating from ScyllaDB 6.1 to 2025.1, snapshots, and Manager H3: Related topics Originally from the User Slack @Anurag_Vishwakarma: Hello, I need help. We are currently migrating from ScyllaDB 6.1 to Scylla Enterprise, but the migration process seems very complicated. We need to migrate the entire dataset from one cluster to another, and it’s proving to be difficult. We tried using snapshots, but we’re encountering many errors. Setting up the ScyllaDB Agent and Scylla Manager also seems complicated. Can anyone help me understand the best way to back up and migrate data from one cluster to another? @Patrick_Bossman: https://docs.scylladb.com/manual/stable/upgrade/upgrade-guides/upgrade-guide-from-6.2-to-2025.1/upgrade-guide-from-6.2-to-2025.1.html Upgrade from ScyllaDB Open Source 6.2 to ScyllaDB 2025.1 | ScyllaDB Docs From 6.2, in-place migration steps to 2025.1 --- ### Page: https://forum.scylladb.com/t/iotune-tests-fails-because-it-saturates-evaluation-disks/4766 Title: Iotune tests fails because it saturates evaluation disks - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello. I’m using ScyllaDB Open Source 6.2.3 on physical machines. I’m running scylla_setup command : scylla_setup --no-raid-setup --no-selinux-setup --no-ntp-setup --no-coredump-setup --no-sysconfig-setup --io-setup … Language: en Canonical URL: https://forum.scylladb.com/t/iotune-tests-fails-because-it-saturates-evaluation-disks/4766 ## Headings Structure: H1: Iotune tests fails because it saturates evaluation disks H3: Related topics ## Main Content: H1: Iotune tests fails because it saturates evaluation disks H3: Related topics I’m using ScyllaDB Open Source 6.2.3 on physical machines. I’m running scylla_setup command : scylla_setup --no-raid-setup --no-selinux-setup --no-ntp-setup --no-coredump-setup --no-sysconfig-setup --io-setup 1 --no-version-check --no-fstrim-setup --no-rsyslog-setup I have NVME disks on my ScyllaDB servers organized like this : nvme0n1 259:0 0 1.7T 0 disk ├─data_commit_vg-6la_commitlog_lv 253:13 0 1000G 0 lvm /var/lib/scylla/commitlog ├─data_commit_vg-6la_hints_lv 253:16 0 250G 0 lvm /var/lib/scylla/hints └─data_commit_vg-6la_view_hints_lv 253:18 0 100G 0 lvm /var/lib/scylla/view_hint nvme2n1 259:1 0 1.7T 0 disk └─data_vg-6la_data_lv 253:15 0 1.7T 0 lvm /var/lib/scylla/data Unfortunately, the scylla_setup command fails because iotune saturates the disks it tests. Here is the iotune command performed : /usr/bin/iotune --format envfile --options-file /etc/scylla.d/io.conf --properties-file /etc/scylla.d/io_properties.yaml --evaluation-directory /var/lib/scylla/data --evaluation-directory /var/lib/scylla/commitlog --evaluation-directory /var/lib/scylla/hints --evaluation-directory /var/lib/scylla/view_hints /opt/scylladb/scripts/libexec/scylla_io_setup:41: SyntaxWarning: invalid escape sequence ‘\s’ pattern = re.compile(_nocomment + r"CPUSET=\s*"" + _reopt(_cpuset) + _reopt(_smp) + “\s*"”) ` WARN 2025-04-24 10:54:19,580 seastar - Kernel memory reservation (/proc/sys/vm/min_free_kbytes) unexpectedly high (5497558016), check your configuration INFO 2025-04-24 10:54:19,674 seastar - Reactor backend: linux-aio INFO 2025-04-24 10:54:20,415 [shard 0:main] iotune - /var/lib/scylla/view_hints passed sanity checks INFO 2025-04-24 10:54:20,416 [shard 0:main] iotune - Disk parameters: max_iodepth=1023 disks_per_array=1 minimum_io_size=4096 INFO 2025-04-24 10:54:20,420 [shard 0:main] iotune - Filesystem parameters: read alignment 512, write alignment 4096 ERROR 2025-04-24 10:54:35,453 [shard 0:main] seastar - Exiting on unhandled exception: std::system_error (error system:28, No space left on device) ERROR:root:Command ‘[’/usr/bin/iotune’, ‘–format’, ‘envfile’, ‘–options-file’, ‘/etc/scylla.d/io.conf’, ‘–properties-file’, ‘/etc/scylla.d/io_properties.yaml’, ‘–evaluation-directory’, ‘/var/lib/scylla/data’, ‘–evaluation-directory’, ‘/var/lib/scylla/commitlog’, ‘–evaluation-directory’, ‘/var/lib/scylla/hints’, ‘–evaluation-directory’, ‘/var/lib/scylla/view_hints’]’ returned non-zero exit status 1. ERROR:root:[‘/var/lib/scylla/data’, ‘/var/lib/scylla/commitlog’, ‘/var/lib/scylla/hints’, ‘/var/lib/scylla/view_hints’] did not pass validation tests, it may not be on XFS and/or has limited disk space. This is a non-supported setup, and performance is expected to be very bad. For better performance, placing your data on XFS-formatted directories is required. To override this error, enable developer mode as follow: sudo /opt/scylladb/scripts/scylla_dev_mode_setup --developer-mode 1 Disks are correctly formatted in XFS. I tried to decrease duration because by default it is 120s, but even with 20 seconds, iotune saturates view_hints disk and fails. Why does iotune saturate the disk pls ? Shoudn’t it stop its tests when it arrives at 80-90% of the disk ? What can I do to solve this problem and be able to execute scylla_setup ? Hi Gwenael, welcome to the ScyllaDB Community Forum! We don’t recommend separating ScyllaDB data components into separate partitions. You’ll likely end up wasting IO bandwidth by doing such splits as most of them will be idle most of the time. Our recommended setup is to RAID0 all available disks (assuming they have the same capacity and performance characteristics) into a single partition and mount it as /var/lib/scylla. This can be done via the scylla_setup process, or separately via the scylla_create_raid script. This way, all subdirectories are created within that device and ScyllaDB can leverage the RAID IO bandwidth to its maximum. You can explore more on our disks recommendations on this blog. Yes it is like this. The data is on multiple disks. But there is one disk used for commitlog, hints and view hints. And the problem of disk saturation occures on view hints because I set the FS to 100 GB. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-92-2025-04-25/4779 Title: Last week in scylla-cluster-tests.git master (issue #92; 2025-04-25) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 26a7a73c…0e303bb3 range are covered. There were 22 non-merge commits from 9 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-92-2025-04-25/4779 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #92; 2025-04-25) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #92; 2025-04-25) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 26a7a73c…0e303bb3 range are covered. There were 22 non-merge commits from 9 authors in that period. Some notable commits: Every time Scylla reports an oversized allocation, we now raise an error event. A new short performance test, based on the Java driver throughput test. It’s scheduled daily to catch regressions early. We now collect the contents of system.truncated in every test run. Nemesis discovery logic was extracted into a NemesisRegistry class, simplifying the code and improving discovery speed. Also unified the way disrupt methods are called across SisyphusMonkey and Complex monkeys. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/migrating-data-from-one-cluster-to-another-scylla-manager-migrator-and-nodetool-snapshot/4786 Title: Migrating data from one cluster to another, Scylla Manager, Migrator and nodetool snapshot - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/migrating-data-from-one-cluster-to-another-scylla-manager-migrator-and-nodetool-snapshot/4786 ## Headings Structure: H1: Migrating data from one cluster to another, Scylla Manager, Migrator and nodetool snapshot H3: Related topics ## Main Content: H1: Migrating data from one cluster to another, Scylla Manager, Migrator and nodetool snapshot H3: Related topics Originally from the User Slack @YGwan: Hello, I would like to ask you a question about SyllaDB Data Migration. I would like to migrate all the data in the table in a particular keyspace to a different syllaDB server. When I use the nodepool snapshot command, not all data is migrated, only certain parts of data are migrated. Each server is using the same version of syllaDB. Is there a best way to migrate these data? @Patrick_Bossman: You can use Scylla manager to backup on one scylla and restore a keyspace in another. You could also use Scylla Migrator, which would read using CQL from one Scylla to another. It operates one table at a time. @avi: nodetool snapshot should contain all the data up to the point in time where you took the snapshot @YGwan: @Patrick_Bossman > You can use Scylla manager to backup on one scylla and restore a keyspace in another. > You could also use Scylla migrator, which would read using CQL from one Scylla to another. It operates one table at a time. Do you have any reference materials for this part? I also checked the manager and the migrator, but I couldn’t find any reference materials related to data migration. @Patrick_Bossman: For Scylla manager, a restore takes the backup and … Restores it to the target. So the migration when the source of the backup is restored to a different target cluster, that’s the data movement. nodetool snapshot should contain all the data up to the point in time where you took the snapshot what’s mean “all the data up to the point in time” If I use the nodetool snapshot command to create a snapshot of a particular keyspace and table, isn’t it creating a snapshot of the entire data? Of course, I did the nodetool refresh command. @Patrick_Bossman: Scylla migrator is built on spark and uses CQL to move data from a source cluster to a target cluster. https://migrator.docs.scylladb.com/stable/ ScyllaDB Migrator Documentation | ScyllaDB Docs https://manager.docs.scylladb.com/stable/restore/ Restore | ScyllaDB Docs @YGwan: > For Scylla manager, a restore takes the backup and … Restores it to the target. So the migration when the source of the backup is restored to a different target cluster, that’s the data movement. Let’s check out the data backup using the scylla manager. Thank you. > Scylla migrator is built on spark and uses CQL to move data from a source cluster to a target cluster. thanks!!. i try to check this link --- ### Page: https://forum.scylladb.com/t/key-rotation-process-for-a-compromised-kmip-managed-encryption-key-in-scylladb/4787 Title: Key rotation process for a compromised KMIP-managed encryption key in ScyllaDB - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: We have set up a remote KMIP server and encrypted several tables, creating a new key for each table on the KMIP server. We would like to understand the process in case of a potential compromise. According to the document… Language: en Canonical URL: https://forum.scylladb.com/t/key-rotation-process-for-a-compromised-kmip-managed-encryption-key-in-scylladb/4787 ## Headings Structure: H1: Key rotation process for a compromised KMIP-managed encryption key in ScyllaDB H3: Related topics ## Main Content: H1: Key rotation process for a compromised KMIP-managed encryption key in ScyllaDB H3: Related topics We have set up a remote KMIP server and encrypted several tables, creating a new key for each table on the KMIP server. We would like to understand the process in case of a potential compromise. According to the documentation (Encryption at Rest | ScyllaDB Docs), we can use the ALTER command on a table to generate a new key on the KMIP server, while keeping the old key intact on the KMIP server. However, based on the key alias/UUID (available to us on the KMIP server), how can we determine which table is associated with a specific key UUID? Is there an internal table in Scylla that links tables to their respective key UUIDs? Or, if there is a better approach to key rotation, we would appreciate any guidance on that. To rotate a key using KMIP: First Modify your KMIP config so that queries for a key based on properties (such as length, cryptographic usage etc) does no longer return the key to be retired. Then either: --- ### Page: https://forum.scylladb.com/t/throttling-of-compaction-using-configuration-parameters/4789 Title: Throttling of compaction, using configuration parameters - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/throttling-of-compaction-using-configuration-parameters/4789 ## Headings Structure: H1: Throttling of compaction, using configuration parameters H1: Throttles compaction to the given total throughput across the entire H1: system. The faster you insert data, the faster you need to compact in H1: order to keep the sstable count down, but in general, setting this to H1: 16 to 32 times the rate you are inserting data is more than sufficient. H1: Setting… H3: Related topics ## Main Content: H1: Throttling of compaction, using configuration parameters H1: Throttles compaction to the given total throughput across the entire H1: system. The faster you insert data, the faster you need to compact in H1: order to keep the sstable count down, but in general, setting this to H1: 16 to 32 times the rate you are inserting data is more than sufficient. H1: Setting… H3: Related topics Originally from the User Slack @Active_questions_tagged_scylla_-_Stack_Overflow: ScyllaDB, throttling for compaction I would like to define throttling for compaction in ScyllaDB, but I did not see this setting in scylla.yaml file. I only saw this option for Apache Cassandra in cassandra.yaml file, see the setting: Stack Overflow: ScyllaDB, throttling for compaction @JIST: Do you see the way for throttling of compaction e.g. via different configuration parameters in scylla.yaml file? NOTE: I would like to migrate cassandra cluster to scylla cluster. @avi: You can set compaction_static_shares or compaction_throughput_mb_per_sec, both can be updated without restarting the server (just send SIGHUP) --- ### Page: https://forum.scylladb.com/t/release-scylladb-2025-1-2/4790 Title: [RELEASE] ScyllaDB 2025.1.2 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB 2025.1.2, a bug-fix production-ready patch release for ScyllaDB 2025.1 LTS Release. Cluster Level Repair for Tablets This version includes a new nodetool command: nodetool cluster, … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-2025-1-2/4790 ## Headings Structure: H1: [RELEASE] ScyllaDB 2025.1.2 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 2025.1.2 H3: Related topics The ScyllaDB team announces ScyllaDB 2025.1.2, a bug-fix production-ready patch release for ScyllaDB 2025.1 LTS Release. Cluster Level Repair for Tablets This version includes a new nodetool command: nodetool cluster, for cluster wide operations. The first cluster-level command is nodetool cluster repair for Tablets. The command uses a new admin REST API /storage_service/tablets/repair . Unlike “nodetool repair” which runs at the node level, running cluster-level repair synchronizes all data on all nodes in the cluster, for Tablets only. If you are using both vNode and Tablets Keyspaces, make sure to use both commands. Scylla Manager version 3.5 and later automatically use both commands when required. 2025.1.2 also includes multiple bug fixes (below). The following issues are fixed in this release: This enforcement will be used in the upcoming ScyllaDB Cloud X Cloud clusters. A rare race condition may cause assertion `_u.st == state::future’ failed. (reader concurrency semaphore related) #22919 Read-repair issue may lead to use-after-move and potential exit, for example when running a test with read repair and trace log #21907 #21714 #23512 #23513 Exception during node shutdown, for example, replace node with the same IP #23325 #23305 #21815 Tablets: Truncate or drop table after tablet migration might cause assert and unexpected exit #18059 Tablets: Failed to complete splitting of table after removing an MV during tablet splitting, causing infinite split retry loop #21859 Since Scylla 6.0, cache and memtable cells between 13 kiB and 128 kiB are getting allocated in the standard allocator rather than inside LSA segments. This can result in out of memory issue. #22941 #22389 #23781 Access to disengaged optional when rewriting bloom filter of a new sstable #23484 Coredump after scylla starting after audit log was enabled #22973 Coredump during RefuseConnectionWithBlockScyllaPortsOnBannedNode nemesis #23348 Possible race condition in the task manager may cause an error, for example during repair #22316 Tablets: Finalize tablet splits earlier. If there is a large load balancing backlog, split finalization may be delayed arbitrarily long and we end up with large tablets. #21762 Tablets: handle_tablet_migration: do not continue if a global metadata barrier is executed #22792 segfault when dumping semaphore diagnostics on SIGQUIT #22756 Tablets: Tablet allocation on table creation overloads nodes with fewer shards #23378 Tablets: rare race condition when aborting when adding a node to the cluster #23222 Rare integer overflow in abstract_read_executor::execute when running cross DC repair #23314 Stop node while restarting failed with seastar::gate_closed_exception (gate closed) #23153 Tablet split may fail with Assertion `_promise’ failed #22715 Possible out-of-space issue when adding a new DC. Currently, when we rebuild a tablet, we stream data from all replicas. This creates a lot of redundancy, wastes bandwidth and CPU resources. This is now fixed by splitting the streaming stage of tablet rebuild into two phases: first we stream tablet’s data from only one replica and then repair the tablet. scylla-nodetool: rapidjson::GenericValue::GetInt() can trigger assert if integer value overflows 32 bit int. The result nodetool netstats return a bad formatted JSON output. #23394 Update node_exporter to release to 1.9.0 #22884 --- ### Page: https://forum.scylladb.com/t/need-help-to-identify-the-scylla-down-issue/4801 Title: Need help to identify the scylla down issue - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi , I’m running a 3-node ScyllaDB cluster, with each node provisioned with 8 vCPUs and 32 GB of memory. Scylla service is down with below logs- ay 2 17:57:44 vm-wynk-prd-01 scylla[1684]: Reactor stalled for 33 ms on … Language: en Canonical URL: https://forum.scylladb.com/t/need-help-to-identify-the-scylla-down-issue/4801 ## Headings Structure: H1: Need help to identify the scylla down issue H3: Related topics ## Main Content: H1: Need help to identify the scylla down issue H3: Related topics Hi , I’m running a 3-node ScyllaDB cluster, with each node provisioned with 8 vCPUs and 32 GB of memory. Scylla service is down with below logs- ay 2 17:57:44 vm-wynk-prd-01 scylla[1684]: Reactor stalled for 33 ms on shard 3. Backtrace: 0x4d768f2 0x4d75550 0x4d76800 0x7f3fea8c0a1f 0x1dccc73 0x2129437 0x21414ad 0x272d0fa 0x276b54d 0x276543d 0x272f5f2 0x1dfe7de 0x4d860e4 0x4d874c7 0x4da6515 0x4d59ffa 0x92a4 0x100322 May 2 17:57:44 vm-wynk-prd-01 scylla[1684]: Reactor stalled for 33 ms on shard 5. Backtrace: 0x4d768f2 0x4d75550 0x4d76800 0x7f3fea8c0a1f 0x2129480 0x21414ad 0x272d0fa 0x276b54d 0x276543d 0x272f5f2 0x1dfe7de 0x4d860e4 0x4d874c7 0x4da6515 0x4d59ffa 0x92a4 0x100322 May 2 17:57:44 vm-wynk-prd-01 scylla[1684]: Reactor stalled for 33 ms on shard 4. Backtrace: 0x4d768f2 0x4d75550 0x4d76800 0x7f3fea8c0a1f 0x4d47e50 0x4d4bc8d 0x2202af0 0x21294d0 0x21414ad 0x272d0fa 0x276b54d 0x276543d 0x272f5f2 0x1dfe7de 0x4d860e4 0x4d874c7 0x4da6515 0x4d59ffa 0x92a4 0x100322 May 2 17:57:44 vm-wynk-prd-01 scylla[1684]: Reactor stalled for 33 ms on shard 2. Backtrace: 0x4d768f2 0x4d75550 0x4d76800 0x7f3fea8c0a1f 0x21294c8 0x21414ad 0x272d0fa 0x276b54d 0x276543d 0x272f5f2 0x1dfe7de 0x4d860e4 0x4d874c7 0x4da6515 0x4d59ffa 0x92a4 0x100322 May 2 17:58:30 vm-wynk-prd-01 systemd[1]: Starting GCE Workload Certificate refresh… May 2 17:58:30 vm-wynk-prd-01 gce_workload_cert_refresh[2192]: 2025/05/02 17:58:30: Done May 2 17:58:30 vm-wynk-prd-01 systemd[1]: gce-workload-cert-refresh.service: Succeeded. May 2 17:58:30 vm-wynk-prd-01 systemd[1]: Started GCE Workload Certificate refresh. May 2 17:58:44 vm-wynk-prd-01 scylla[1684]: [shard 1] BatchStatement - Batch modifying 206 partitions in iptv.iptv_users_error_events is of size 715358 bytes, exceeding specified WARN threshold of 204800 by 510558. May 2 18:00:42 vm-wynk-prd-iptv-scylla-01 sshd[2221]: rexec line 124: Deprecated option UsePrivilegeSeparation May 2 18:00:52 vm-wynk-prd-iptv-scylla-01 scylla[1684]: [shard 4] BatchStatement - Batch modifying 249 partitions in iptv.iptv_users_error_events is of size 711383 bytes, exceeding specified WARN threshold of 204800 by 506583. May 2 18:02:24 vm-wynk-prd-01 scylla[1684]: [shard 1] BatchStatement - Batch modifying 226 partitions in iptv.iptv_users_error_events is of size 714514 bytes, exceeding specified WARN threshold of 204800 by 509714. 5.1.19-0.20231127.0f2269afbf9c You are running an unsupported version of ScyllaDB, please upgrade to the latest version and see if the issue still persists. --- ### Page: https://forum.scylladb.com/t/error-while-repairing-system-keyspace-which-keyspaces-should-be-repaired/4803 Title: Error while repairing system keyspace, which keyspaces should be repaired? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/error-while-repairing-system-keyspace-which-keyspaces-should-be-repaired/4803 ## Headings Structure: H1: Error while repairing system keyspace, which keyspaces should be repaired? H3: Related topics ## Main Content: H1: Error while repairing system keyspace, which keyspaces should be repaired? H3: Related topics Originally from the User Slack @Daria_Fedorova: Hi, I got an unfamiliar error while repairing keyspace=system on staging scylla cluster. (version is 5.1.19 Т_Т) I did not know the number of shards could be different for tables. How could this happen? Is was misconfigured ? Someone forgot to clean up tables when retrying setup ? And how do i fix it ? @avi: You’re not supposed to repair the system keyspace @Daria_Fedorova: ok ) forgot it system_distributed system_schema Are also skipped when repairing all spaces But system_traces system_auth system_distributed_everywhere should be repaired ? @avi: system_traces can be repaired, but it’s not important so better to skip it. The others must be repaired regularly --- ### Page: https://forum.scylladb.com/t/node-crashing-when-server-encryption-options-is-set/4804 Title: Node crashing when server_encryption_options is set - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello Team, Node crashing when server_encryption_options is set. Details : Scylla version : scylla:2025.1.2 in docker If not set server_encryption_options, it works fine, but with server_encryption_options configured… Language: en Canonical URL: https://forum.scylladb.com/t/node-crashing-when-server-encryption-options-is-set/4804 ## Headings Structure: H1: Node crashing when server_encryption_options is set H3: Related topics ## Main Content: H1: Node crashing when server_encryption_options is set H3: Related topics Hello Team, Node crashing when server_encryption_options is set. Details : Scylla version : scylla:2025.1.2 in docker If not set server_encryption_options, it works fine, but with server_encryption_options configured, it keeps crashing and restarting after startup, and the same configuration works fine in version 6.2.3 如果没有设置server_encryption_options,工作正常,但是配置了server_encryption_options,启动后就会不断崩溃重启,相同的配置在6.2.3版本中运行的很好 Let us know if anything is needed for debugging ? @Yaniv_Kaul you may want to take a look I think it’s the same as Node crashing when server_encryption_options is set · Issue #23994 · scylladb/scylladb · GitHub ? --- ### Page: https://forum.scylladb.com/t/last-fortnight-in-scylladb-git-master-issue-278-2025-05-04/4805 Title: Last fortnight in scylladb.git master (issue #278; 2025-05-04) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 2314feeae2..8ffe4b0308 range are covered. There were 59 non-merge commits from 17 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-fortnight-in-scylladb-git-master-issue-278-2025-05-04/4805 ## Headings Structure: H1: Last fortnight in scylladb.git master (issue #278; 2025-05-04) H3: Related topics ## Main Content: H1: Last fortnight in scylladb.git master (issue #278; 2025-05-04) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 2314feeae2..8ffe4b0308 range are covered. There were 59 non-merge commits from 17 authors in that period. Some notable commits: The snitch reports the datacenter and rack of a node. We will now refuse to start the service if the snitch reports a different datacenter or rack, preventing data loss. When creating a new table, we will prefer to distribute tablets for the new table equally among all nodes, rather than distributing all-table tablet count equally. This improves performance for the new table. Alternator, ScyllaDB’s implementation of the DynamoDB API, now limits attribute name length in the same name as DynamoDB. Materialized views are now more robust during schema changes. The change removes the possibility of accessing an outdated schema that no longer exists or is incompatible with the view schema. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/c-driver-exceptions-logs-metrics-and-timeouts/4806 Title: C++ driver, exceptions, logs, metrics and timeouts - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/c-driver-exceptions-logs-metrics-and-timeouts/4806 ## Headings Structure: H1: C++ driver, exceptions, logs, metrics and timeouts H3: Related topics ## Main Content: H1: C++ driver, exceptions, logs, metrics and timeouts H3: Related topics Originally from the User Slack @Nick_Vladiceanu: Hey folks we started to observe C++ exceptions in our cluster (Scylla version 5.2), resulting in errors on the clients. Where should we look first, and how can we find the C++ exceptions? Any help appreciated The only unusual logs I find on some of the nodes is @avi: 5.2 is long out of support, please upgrade to a supported version @Nick_Vladiceanu: I guess this is a more generic question: where the C++ exceptions can be found? among regular logs? @avi: Some C++ exceptions (e.g. timeouts) are swallowed by the code and converted into metrics (e.g. timeout metrics). The rest are in the logs. @Nick_Vladiceanu: would they appear in the logs for the log level WARN, or require a lower log level? It was a bit difficult to understand what those C++ exceptions were about, no exceptions in the logs. @avi: WARN or ERROR. They’re likely timeouts. Check the dedicated timeout metrics. --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-2-9/4808 Title: [RELEASE] ScyllaDB Enterprise 2024.2.9 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2024.2.9, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2024.2 Feature (Short Term Support) Release. Note there is a later LTS (Long… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-2-9/4808 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2024.2.9 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2024.2.9 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2024.2.9, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2024.2 Feature (Short Term Support) Release. Note there is a later LTS (Long term Support) Release 2025.1, and you are encouraged to upgrade to it. The following stability issue is fixed in this release (with an open-source reference): --- ### Page: https://forum.scylladb.com/t/long-connection-times-connection-timeout-using-rust-driver-connection-ports-with-docker/4809 Title: Long connection times, Connection timeout using Rust driver, connection ports with Docker - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/long-connection-times-connection-timeout-using-rust-driver-connection-ports-with-docker/4809 ## Headings Structure: H1: Long connection times, Connection timeout using Rust driver, connection ports with Docker H3: Related topics ## Main Content: H1: Long connection times, Connection timeout using Rust driver, connection ports with Docker H3: Related topics Originally from the User Slack @Mattia: Hello everyone I am trying ScyllaDB using the official rust driver. I have spawned 3 nodes in a cluster using docker compose and I build the Session from rust code. The problem I’m having is that the function that builds the connection blocks the executing until the timeout specified with .connection_timeout(Duration::from_secs(..)) runs out, even if the connection is fine. For example if I set the timeout to 30 seconds I have to wait exactly 30 seconds before having a successful connection to the db. Is this the expected behavior? Thanks @Karol_Baryła: Definitely not an expected behavior. Could you share what is your OS, how exactly are you setting up the cluster, and how are you connecting to it? @Mattia: Sure, thanks for the reply. I’m using docker-compose up -d with podman as the backend. With tracing I see that it waits for x seconds before exiting the function, with x determined by the connection_timeout func Also, it is a debug build @Karol_Baryła: The problem is most likely that scylla2 and scylla3 are listening on localhost (in their own network namespaces), and broadcasting this address, so the driver can’t connect to them. For such docker setup you need to set up rpc address and listen address (and remove the “ports” section from scylla1). See for example the docker compose from tests of Rust Driver: https://github.com/scylladb/scylla-rust-driver/blob/main/test/cluster/docker-compose.yml GitHub: scylla-rust-driver/test/cluster/docker-compose.yml at main · scylladb/scylla-rust-driver @Mattia: Thank you I will look into it. I got the docker compose from the docker’s hub page but what you are saying makes sense Just as a note: I’m only supplying “127.0.0.1:9042” as an address to the SessionBuilder so I thought it was sufficient to expose the port of one of the containers and everything would work. @Karol_Baryła: It would indeed allow the driver to connect to this one node. However the driver tries to connect to each node (each shard in fact). Connection process is roughly as follows: • One of provided contact points is chosen as control connection • Driver tries to connect to it • If successfull, driver then fetches information about cluster from system tables using this connection. • This allows the driver to learn about other nodes, to which it then tries to connect (this is the step that is hanging in your case) • After connecting, the Session is returned. If you want the driver to really connect only to this single node (which is usually a very bad idea, and only makes sense for tools like cqlsh) you could use https://docs.rs/scylla/latest/scylla/policies/host_filter/trait.HostFilter.html ( https://docs.rs/scylla/latest/scylla/client/session_builder/type.SessionBuilder.html#method.host_filter ) to prevent the driver from connecting to other nodes. HostFilter in scylla::policies::host_filter - Rust SessionBuilder in scylla::client::session_builder - Rust @Mattia: Thank you so much for the explanation and your time. I will later try with the compose you have linked and see how it goes. For some reasons I get that the scylla1 node is unhealthy when trying to bring up the cluster. Any idea? @Karol_Baryła: Did you remove the ports section from scylla1? @Mattia: Yes, I have completely removed the first file @Karol_Baryła: Can you show me the full docker compose? @Mattia: Hello, I am trying to make a contribution to ScyllaDB rust driver but I am unable to start the cluster to test the changes I have made. I am using the Makefile and running make test or make ci but both docker and podman hangs indefinitely (I think on the healthcheck of the first container). Can anyone help me with this? (this is the docker compose used) GitHub: scylla-rust-driver/test/cluster/docker-compose.yml at 00856129603b39676c9a31c26626d0ad5ec19c68 · scylladb/scylla-rust-driver @Karol_Baryła: Sorry for leaving you without response on a previous thread! Tbh this cluster sometimes (but rarely) also fails to start for me. In that case it is usually enough to try again (make down && make up ). I never really investigated why that happens. You can check logs of this node with docker logs cluster-scylla1-1 and link them here (use pastebin / gist, or send a file), maybe we’ll find something there. @Mattia: Unfortunately I have tried restarting it many times, I will try to share the logs now Here are the file for the logs from the container: stdout.txt (2.9 KB) stderr.txt (109.1 KB) Even tough after 355 seconds the job stops due to unhealthy condition in scylla1 if I run nodetools I get the following status: @Karol_Baryła: I have no idea why it doesn’t start. Scylla logs look fine I think. What happens if you execute the healthcheck command inside the container manually? @Karol_Baryła: Definitely without root. docker inspect --format "{{json .State.Health }}" cluster-scylla-1 | jq should show the output of healthcheck commands executed by docker - maybe it will give us some clue. @Mattia: Thank you, this is rather telling: @Karol_Baryła: I don’t see how it could possibly interpret * as port number in our compose file Did you modify the compose file? I see that your container name is different than mine, that may have something to do with it. This is the one I am using, straight from the github repo @Karol_Baryła: What is the output of docker --version? @Mattia: Docker version 28.1.1, build 4eba377327 I also tried with podman with the same problem (podman version 5.4.2) @Karol_Baryła: In that case I’m unfortunately out of ideas, sorry @Mattia: Thank you for the help anyway! I think I will try to remove the docker-compose/podman layer. That might create some problems I guess Seems like I was still using podman, for some reasons it took precedence over docker when running the commands. With docker it works. (at least the first container is healthy, now the second seems to be hanging) @Mattia: Thank you for the help anyway! I think I will try to remove the docker-compose/podman layer. That might create some problems I guess @Mattia: Seems like I was still using podman, for some reasons it took precedence over docker when running the commands. With docker it works. (at least the first container is healthy, now the second seems to be hanging) --- ### Page: https://forum.scylladb.com/t/release-added-support-for-social-logins-google-linkedin-github-6-may-2025/4810 Title: [RELEASE] Added Support for Social Logins (Google, LinkedIn, GitHub) - 6 May 2025 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: Introducing Social Login Support: You can now log in to ScyllaDB Cloud using your existing Google, LinkedIn, or GitHub account — no need to manage a separate password. If you currently log in with an email/password, yo… Language: en Canonical URL: https://forum.scylladb.com/t/release-added-support-for-social-logins-google-linkedin-github-6-may-2025/4810 ## Headings Structure: H1: [RELEASE] Added Support for Social Logins (Google, LinkedIn, GitHub) - 6 May 2025 H3: Related topics ## Main Content: H1: [RELEASE] Added Support for Social Logins (Google, LinkedIn, GitHub) - 6 May 2025 H3: Related topics Introducing Social Login Support: You can now log in to ScyllaDB Cloud using your existing Google, LinkedIn, or GitHub account — no need to manage a separate password. If you currently log in with an email/password, you can switch to social login at any time — as long as the email on your social account matches your ScyllaDB Cloud account email. This update makes it easier and faster to access your clusters while keeping your login experience secure and seamless. To try it out, go to the ScyllaDB Cloud login page and click the provider you’d like to use. --- ### Page: https://forum.scylladb.com/t/kubernetes-doet-not-create-correct-directory/4815 Title: Kubernetes doet not create correct directory - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details kubernetes version: microk8s #ScyllaDB version: 2025.1 #Cluster size: 3 nodes os (RHEL/CentOS/Ubuntu/AWS AMI): Ubuntu When I try to follow the gitops documentation for the scylla operator with … Language: en Canonical URL: https://forum.scylladb.com/t/kubernetes-doet-not-create-correct-directory/4815 ## Headings Structure: H1: Kubernetes doet not create correct directory H3: Related topics ## Main Content: H1: Kubernetes doet not create correct directory H3: Related topics Installation details kubernetes version: microk8s #ScyllaDB version: 2025.1 #Cluster size: 3 nodes os (RHEL/CentOS/Ubuntu/AWS AMI): Ubuntu When I try to follow the gitops documentation for the scylla operator with a modified setup for the nodeconfig: apiVersion: scylla{DOT}scylladb{DOT}com/v1alpha1 kind: NodeConfig metadata: name: scylladb-pool-1 spec: localDiskSetup: raids: - name: nvmes type: RAID0 RAID0: devices: nameRegex: ^/dev/sdb$ filesystems: - device: /dev/md/nvmes type: xfs mounts: - device: /dev/md/nvmes mountPoint: /var/lib/persistent-volumes unsupportedOptions: - prjquota placement: nodeSelector: scylla{DOT}scylladb{DOT}com/node-type: scylla tolerations: - effect: NoSchedule key: scylla-operator{DOT}scylladb{DOT}com/dedicated operator: Equal value: scyllaclusters When a pod tries to use the local csi driver (installed like the instructions in the gitops documentation). I get the following error in pod describe while the pod stays on init indefinitely: MountVolume.SetUp failed for volume “pvc-9effa2df-4f8b-44b6-94a2-0697c2980a17” : rpc error: code = Internal desc = Failed to publish volume: can’t create target path at “/var/snap/microk8s/common/var/lib/kubelet/pods/ad3827af-2ee6-4913-9853-55205cd6f8c2/volumes/kubernetes{DOT}io~csi/pvc-9effa2df-4f8b-44b6-94a2-0697c2980a17/mount”: mkdir /var/snap/microk8s/common/var/lib/kubelet/pods/ad3827af-2ee6-4913-9853-55205cd6f8c2/volumes/kubernetes{DOT}io~csi/pvc-9effa2df-4f8b-44b6-94a2-0697c2980a17/mount: no such file or directory All folders exist apart form the pvc-9effa2df-4f8b-44b6-94a2-0697c2980a17 one When manually creating this folder it gets deleted when the pod tries to mount the volume again. @mflendrich / @zimnx can you take a look ? On MicroK8s the kubelet root lives under /var/snap/microk8s/common/var/lib/kubelet, not /var/lib/kubelet. In the upstream manifest the three hostPath volumes that the pods rely on still point to /var/lib/kubelet (see local-csi-driver/deploy/kubernetes/local-csi-driver/50_daemonset.yaml at master · scylladb/local-csi-driver · GitHub). Because that directory is empty on a MicroK8s node, the driver container cannot see (or create) /var/snap/…/kubelet/pods//volumes/kubernetes.io~csi//mount and fails exactly as you observed. There are two options to fix it, quick and dirty, on every node you could create a symlink sudo ln -s /var/snap/microk8s/common/var/lib/kubelet /var/lib/kubelet, or patch the DaemonSet (or create your own overlay) so that all host paths reference the MicroK8s kubelet root. --- ### Page: https://forum.scylladb.com/t/last-2-weeks-in-scylla-cluster-tests-git-master-issue-93-2025-05-09/4818 Title: Last 2 weeks in scylla-cluster-tests.git master (issue #93; 2025-05-09) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last 2 weeks. Commits in the ba8d5d6f…5644f077 range are covered. There were 39 non-merge commits from 12 authors i… Language: en Canonical URL: https://forum.scylladb.com/t/last-2-weeks-in-scylla-cluster-tests-git-master-issue-93-2025-05-09/4818 ## Headings Structure: H1: Last 2 weeks in scylla-cluster-tests.git master (issue #93; 2025-05-09) H3: Related topics ## Main Content: H1: Last 2 weeks in scylla-cluster-tests.git master (issue #93; 2025-05-09) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last 2 weeks. Commits in the ba8d5d6f…5644f077 range are covered. There were 39 non-merge commits from 12 authors in that period. Some notable commits: When a test ends, SCT now saves the full schema including internals for better debugging and post-analysis. The monitoring stack has been upgraded to 4.9.2, bringing in various improvements and fixes. To catch crashes in SCT itself, we now collect core dumps and report them via Argus. podman support was fixed by correcting tmpfs mount parameters. Timeouts for decommission and add node operations were reduced when using tablets; if any such operation takes over 3 hours, a critical event is now published to abort the test. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/operator-error-message-cant-apply-statefulset-update-cant-update/4819 Title: Operator error message: can't apply statefulset update: can't update - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/operator-error-message-cant-apply-statefulset-update-cant-update/4819 ## Headings Structure: H1: Operator error message: can't apply statefulset update: can't update H3: Related topics ## Main Content: H1: Operator error message: can't apply statefulset update: can't update H3: Related topics Originally from the User Slack @Ishaan_Raghav: hey, So I am using scylla-cluster deployed using scylla-operator. I want to vertically scale my cluster. To achieve this, I updated the resources in kind: scyllacluster which should make changes in my sts and hence the cluster but I get the following: What’s the way to upscale using the operator? @Maciej_Zimnoch: which operator version are you using? what have you changed exactly? @Ishaan_Raghav: Image versions: Scylla: 5.2.7 Scylla-manager: 3.2.5 scylla-operator: 1.11.0 Scylla-manager-agent: 3.1.2 Helm chart: Scylla: v1.11.5 Scylla-Operator: v1.11.0 Scylla-manager: v1.11.5 Changes I made: kubectl edit scyllacluster myscylladb and then edit memory in this section: And then I got the above error. @Maciej_Zimnoch: You’re using quite old versions of scylla products, all of them ended their support long time ago. This specific issue - updating immutable sts field - was fixed in 1.11.1. @Ishaan_Raghav: Noted @Maciej_Zimnoch I’ll upgrade and check more. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-279-2025-05-11/4820 Title: Last week in scylladb.git master (issue #279; 2025-05-11) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 8ffe4b0308..092a88c9b9 range are covered. There were 83 non-merge commits from 19 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-279-2025-05-11/4820 ## Headings Structure: H1: Last week in scylladb.git master (issue #279; 2025-05-11) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #279; 2025-05-11) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 8ffe4b0308..092a88c9b9 range are covered. There were 83 non-merge commits from 19 authors in that period. Some notable commits: Alternator, ScyllaDB’s implementation of the DynamoDB API, now automatically retries schema table changes that can fail due to contention. ScyllaDB uses a data structure called partition_sstable_set to rapidly find relevant sstables for run-based compaction strategies (Leveled Compaction Strategy and Incremental Compaction Strategy). This data structure is now careful to avoid quadratic space complexity when pushed to extreme situations, which could cause excessive memory use in the past. This is achieved by noting if sstables are actually organized in runs, and if not, using a different data structure. ScyllaDB warns on large allocations in excess of 1MB as they can cause high latency and thrash the cache. The warning threshold is now reduced to 128kB to flush out smaller violations. ScyllaDB build normally runs in a container that provides all the build dependencies. The container now supports running nested containers, so some of the build dependencies can be container images themselves. Raft group 0 manages metadata (topology and the schema) it restricts the number of voters to improve throughput. It is now careful not to change the voter quorum too much and so reduces leader re-election, which causes minor disruptions. A crash related to referencing a freed statistics data structure for memtables was fixed. The version number in the master branch was changed to 2025.3, indicating the start of the 2025.2 release stabilization cycle. A possible use-after-free during schema changes related to the sstable_set type was fixed. Sstable compression can use dictionaries to improve the compression ratios. We now distribute the dictionaries across all shards in a node, but we are also careful to have each NUMA node own a copy of the dictionary, to minimize performance loss due to cross-NUMA-node memory accesses. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/bad-alloc-when-trying-to-remove-a-node-large-collections-topology-changes-errors/4822 Title: Bad_alloc when trying to remove a node, large collections, topology changes errors - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/bad-alloc-when-trying-to-remove-a-node-large-collections-topology-changes-errors/4822 ## Headings Structure: H1: Bad_alloc when trying to remove a node, large collections, topology changes errors H3: Related topics ## Main Content: H1: Bad_alloc when trying to remove a node, large collections, topology changes errors H3: Related topics Originally from the User Slack @Cong_Guo: Hi there, I am trying to remove a node via curl -X POST “http://127.0.0.1:10000/storage_service/remove_node/xxx” running in scylla 6.2.2, but see below exception, do you know why? @avi: Probably the node is sick, check its logs @Cong_Guo: I once tried to decommission one node, but it triggered other nodes to restart suddenly, before that I see some warning logs like and I do see bad collection design warning like They are some nodes in the DN status, but not majority. No clue if it caused the bad_alloc when running decommission. something like (running on 6.2.2) My shallow thought is the bad collection design (non frozen list) consumes too much memory which impact the memory claim needed for decommission. @avi: Check if you have large collections or large rows in the system.large_* tables @Cong_Guo: Yes, it does. Some bad designed schema like unfrozen list. But I am a little worried that should not trigger so many nodes to restart, especially the majority nodes restarts. Actually see below log when triggered decommission, that should be the cause of nodes restart. Curious why gossip negotiation needs so much memory? N7seastar12continuationINS_8internal22promise_base_with_typeIvEENS_6futureIvE12finally_bodyIZNS_3smp9submit_toIZNS_7shardedIN3gms8gossiperEE9invoke_onINS_20noncopyable_functionIFS5_RSB_EEEJES5_Qsr3stdE9invocableITL0__RT_DpOTL0_0_EEET1_jNS_21smp_submit_to_optionsEOSJ_DpOT0_EUlvE_EENS_8futurizeINSt13invoke_resultISJ_JEE4typeEE4typeEjSP_SQ_EUlvE_Lb0EEEZNS5_17then_wrapped_nrvoIS5_S12_EENSV_ISJ_E4typeEOT0_EUlOS3_RS12_ONS_12future_stateINS1_9monostateEEEE_vEE @avi: It’s likely https://github.com/scylladb/scylladb/issues/23781 GitHub: managed_bytes violates preferred contiguous allocation size · Issue #23781 · scylladb/scylladb @Cong_Guo: Guess below problem is caused by the same issue? https://scylladb-users.slack.com/archives/C2NLNBXLN/p1745586119428079 I am afraid if I trigger any topology change will bring the cluster to down for raft module seems doesn’t work caused by bad_alloc. [April 25th, 2025 6:01 AM] cong.guo: Hi, do you know that does this error mean when trying to change the replica number of system_auth? raft::state_machine_error NoHostAvailable: ('Unable to complete the operation against any hosts', {: (State machine error at raft/server.cc:1360): std::bad_alloc (std::bad_alloc)"">}) @Cong_Guo: Just curious, if this error means raft instance is stopped, why the process is still running? Will raft instance still recover to running later? If stopping the read/write to make more room in the Non-LSA, will it be a chance for the alloc to be success for a while? @avi: Don’t know enough about it --- ### Page: https://forum.scylladb.com/t/release-added-support-for-n2d-highmem-and-n2d-standard-instance-types-in-scylladb-cloud/4823 Title: [RELEASE] Added Support for n2d-highmem and n2d-standard instance types in ScyllaDB Cloud - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: We are happy to announce support for two new instance types in ScyllaDB Cloud: n2d-highmem and n2d-standard. These additions give users more flexibility when deploying workloads on Google Cloud Platform. Currently Avai… Language: en Canonical URL: https://forum.scylladb.com/t/release-added-support-for-n2d-highmem-and-n2d-standard-instance-types-in-scylladb-cloud/4823 ## Headings Structure: H1: [RELEASE] Added Support for n2d-highmem and n2d-standard instance types in ScyllaDB Cloud H3: Related topics ## Main Content: H1: [RELEASE] Added Support for n2d-highmem and n2d-standard instance types in ScyllaDB Cloud H3: Related topics We are happy to announce support for two new instance types in ScyllaDB Cloud: n2d-highmem and n2d-standard. These additions give users more flexibility when deploying workloads on Google Cloud Platform. Currently Available Regions and Instance Types: n2d-highmem • Region: asia-south1 (Mumbai) • Instance Types: n2d-standard • Regions: asia-south1 (Mumbai) and asia-southeast1 (Singapore) • Instance Types: To try them out, simply create a new cluster in ScyllaDB Cloud and select your preferred region and instance type. We’d love your feedback — let us know how these new options are performing for your workloads! --- ### Page: https://forum.scylladb.com/t/scylladb-cloud-bring-your-own-account-byoa-aws-setup-auth-token-backups-and-repairs/4825 Title: ScyllaDB Cloud. Bring your own account (BYOA), AWS setup, auth_token, backups and repairs - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-cloud-bring-your-own-account-byoa-aws-setup-auth-token-backups-and-repairs/4825 ## Headings Structure: H1: ScyllaDB Cloud. Bring your own account (BYOA), AWS setup, auth_token, backups and repairs H3: Related topics ## Main Content: H1: ScyllaDB Cloud. Bring your own account (BYOA), AWS setup, auth_token, backups and repairs H3: Related topics Originally from the User Slack @Mayank_Joshi: Hello. I have a BYOA (aws) setup. How do I get auth_token to run sctool cluster add @Patrick_Bossman: When you install scylla-manager-agent on all the nodes, you create the token on one of the nodes and copy it to all the others. The that is also what you use to add cluster to Scylla manager. So take a look at the scylla-manager-agent install section. @Mayank_Joshi: Thanks for this. But I think we cannot ssh into the node. Can we? @Patrick_Bossman: Are you in Scylla cloud? @Mayank_Joshi: Yes, I am in scylla cloud. using BYOA offering. @Patrick_Bossman: Backups and repairs are included in the managed service. To change the schedule, you’d engage with support. @Mayank_Joshi: Oh alright. I’ll raise a ticket. Thanks a lot @Patrick_Bossman: You’re welcome! --- ### Page: https://forum.scylladb.com/t/nodetool-refresh-claims-toc-is-empty-but-its-not/4827 Title: Nodetool refresh claims TOC is empty, but it's not - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 2024.2.9-0.20250423.9aa1f3a492bc #Cluster size: 1 os (RHEL/CentOS/Ubuntu/AWS AMI): Ubuntu Hello all Trying to load some SSTables using nodetool refresh gibing me the followin… Language: en Canonical URL: https://forum.scylladb.com/t/nodetool-refresh-claims-toc-is-empty-but-its-not/4827 ## Headings Structure: H1: Nodetool refresh claims TOC is empty, but it's not H3: Related topics ## Main Content: H1: Nodetool refresh claims TOC is empty, but it's not H3: Related topics Installation details #ScyllaDB version: 2024.2.9-0.20250423.9aa1f3a492bc #Cluster size: 1 os (RHEL/CentOS/Ubuntu/AWS AMI): Ubuntu Trying to load some SSTables using nodetool refresh gibing me the following error: The content of the file mc-855-big-TOC.txt is: Relevant code is in sstables.cc line from 815: The TOC file looks alright. Did anyone edit this file manually? The only potential problem I see now is that the file is parsed with “\n” as line terminator and maybe that’s not the case with your file. Also, do you see any “Unrecognized TOC component was found:” info logs? You were exactly right. The backup was comming from a Windows System. So the line endings are CR+LF. After eliminating the CRs from the file, the import could be processed without a problem. @daprodigy nice, I’m glad you fixed the issue! --- ### Page: https://forum.scylladb.com/t/disabling-or-modifying-the-cpuset-configuration-number-of-cores-cpu-core-mismatch/4829 Title: Disabling or modifying the cpuset configuration, number of cores, CPU core mismatch - Kubernetes Operator - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/disabling-or-modifying-the-cpuset-configuration-number-of-cores-cpu-core-mismatch/4829 ## Headings Structure: H1: Disabling or modifying the cpuset configuration, number of cores, CPU core mismatch H3: Related topics ## Main Content: H1: Disabling or modifying the cpuset configuration, number of cores, CPU core mismatch H3: Related topics Originally from the User Slack @Vitaly_Ivanov: Hi! I have a question about cpuset configuration. How can I disable or modify it? ScyllaDB fails to start because the scylla-operator assigned cpuset 0-31, but my system requires it to be set to 0-11 @Maciej_Zimnoch: Make sure to follow this guide: https://operator.docs.scylladb.com/stable/architecture/tuning.html Without staticcpuManagerPolicy and Guaranteed QoS class kubelet cannot assign full exclusive CPUs, so you end up with shares in entire cpu pool. Tuning | ScyllaDB Docs @Vitaly_Ivanov: Thank you! Right now, I’m not interested in tuning performance. I just want ScyllaDB to start normally. My processors don’t support cpuset 0-31 It supports cpuset 0-11 only. I didn’t change any special settings, just using the default ones. I really wish I could disable cpuset completely so ScyllaDB would just started to work I found the getCPUsAllowedList() function call at: https://github.com/scylladb/scylla-operator/blob/6c26a631e1a6409dbf59f4cb2a72a56d6cd22882/pkg/sidecar/config/config.go#L217 This function works similarly to this Bash command: In my case, inside the ScyllaDB pod, it returns 0-31, but my system has 6 cores only (so the maximum should be 0-11). It turns out that the default cpuset detection is incorrect @Maciej_Zimnoch: well kernel see 32 vcpus, I don’t think it’s wrong. Double check if your Pod lands where you expect @Vitaly_Ivanov: After reviewing the scylla-operator code, I found that downgrading it to v1.14 is the only viable option as it allows adding cpuset: false configuration to the Kubernetes manifest https://github.com/scylladb/scylla-operator/blob/ddcc1582499b346e01d9750ef30b7f3c75114960/pkg/sidecar/config/config.go#L237C5-L237C24 @Patrick_Bossman: @Vitaly_Ivanov Maciej is the SME in this space. I recommend listening to him @Vitaly_Ivanov: Nice to meet you @Patrick_Bossman Thanks! > Double check if your Pod lands where you expect Done. I have a 4-node K3s cluster and would like to deploy ScyllaDB on all 4 nodes. I got an error message in the Scylla container logs Manually setting cpuset=0-11 in /etc/scylla.d/cpuset.conf resolved the startup issue. So I decided to downgrade to v1.14. Now it starts successfully with cpuset: false setting. My problem is solved, but I think: • proper cpuset option is important for starting /usr/bin/scylla • here is some unknown issue with cpuset I can revert operator to 1.16.2 for debuging Container started with command and it started python script @Maciej_Zimnoch: looks like k3s is doing some nasty stuff, kernel allows process to run on all vCPUs while k3s restricts it to only few. To overcome this, I suggest to follow guide I linked at the beginning, this way Scylla Pods will only use resources they exclusively receive (subset of vcpus) and you shouldn’t hit this issue. Note that k3s is not supported platform. @Vitaly_Ivanov: Thank you for you help! Outside of K3s and in the systemd I see the same CPU core mismatch: cat /proc/1/status | grep Cpus_allowed cat /proc/cpuinfo | grep 'core id' 12 cores (with Hyper-threading) 12 and 32 are mismatch I’ve checked a bunch of our servers and it happened only on AMD Ryzen 5 3600 6-Core Processor servers. No ideas what’s wrong with this cpu model @Patrick_Bossman: @vladzcloudius @vladzcloudius: Hi. The output above it absolutely normal AFAICT. The content of /proc/cpuinfo has the actual information about the CPU HW. The value of the Cpus_allowed from /proc/[pid]/status on the other hand tells on which CPUs the corresponding process (init in your case) is allowed to run: https://man7.org/linux/man-pages/man5/proc_pid_status.5.html And I couldn’t find anywhere any requirement for the latter mask to include only actually present CPUs. This means that the user should always use a bitwise AND between the masks from cpuinfo and the Cpus_allowed in order to get the mask of the actual physically present CPU on which the corresponding process is allowed to run. @Patrick_Bossman: @Maciej_Zimnoch ^^ @Vitaly_Ivanov: What does this mean for scylla-operator? Should it use a bitwise AND between the cpuinfo and Cpus_allowed masks to correctly set --cpuset? @Maciej_Zimnoch: I would need to check how those constraints look like in container environment, maybe it’s standardized - i’m not sure. So far all major cloud providers I checked fills cpus_allowed into cpus where container is actually able to run, all available when using CFS shares or subset when running on exclusively allocated ones. --- ### Page: https://forum.scylladb.com/t/release-scylladb-operator-1-17-0/4830 Title: [RELEASE] ScyllaDB Operator 1.17.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of ScyllaDB Operator 1.17.0. ScyllaDB Operator is an open-source project that helps users run ScyllaDB on Kubernetes. The ScyllaDB Operator manages ScyllaDB cluste… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-operator-1-17-0/4830 ## Headings Structure: H1: [RELEASE] ScyllaDB Operator 1.17.0 H1: Managed Multi-DC H1: Other notable changes H1: Upgrade instructions H1: Supported versions H1: Getting started with ScyllaDB Operator H1: Related Links H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Operator 1.17.0 H1: Managed Multi-DC H1: Other notable changes H1: Upgrade instructions H1: Supported versions H1: Getting started with ScyllaDB Operator H1: Related Links H3: Related topics The ScyllaDB team is pleased to announce the release of ScyllaDB Operator 1.17.0. ScyllaDB Operator is an open-source project that helps users run ScyllaDB on Kubernetes. The ScyllaDB Operator manages ScyllaDB clusters deployed to Kubernetes and automates tasks related to operating a ScyllaDB cluster, like installation, vertical and horizontal scaling, as well as rolling upgrades. As with each minor version, ScyllaDB Operator 1.17 adds new features and improves stability. What makes this release special is the technical preview of managed multi-datacenter ScyllaDB clusters. ScyllaDB Operator 1.17 introduces a new deployment pattern: a ScyllaDB cluster spanning across multiple datacenters (regions) running in distinct Kubernetes clusters: Image credit: Maciej Zimnoch @zimnx This functionality allows you to run a “Multi-DC control plane” in one place, and specify your ScyllaDB installation spanning across many (e.g. 3) datacenters (for example, cloud regions), each running its own Kubernetes cluster. ScyllaDB Operator will then manage a consistent, highly-available deployment, utilizing ScyllaDB’s multi-datacenter replication technology. Multi-DC installations are governed by the newly introduced ScyllaDBCluster CRD. Try it out by following the AWS EKS and Google Cloud GKE quickstart guides. See documentation for more. Managed Multi-DC in ScyllaDB Operator 1.17 comes as a technical preview, enabling configuration and core lifecycle management of the Multi-DC ScyllaDB cluster. Monitoring and backup/repair (using ScyllaDB Manager) are available, but require manual setup (automation to be incrementally included in subsequent minor releases - please stay tuned to future announcements). Please refer to the Multi-DC operator enhancements tracking issue on GitHub. This release brings support for new ScyllaDB 2025.1 and ScyllaDB Manager 3.5.0, and includes a major Prometheus version upgrade (v2.54.1 → v3.1.0) as well as updated Grafana (11.3.0 → 11.4.3) and our newest ScyllaDB monitoring dashboards (#2623). The default storage capacity in the Helm chart and the deploy manifests has been adjusted from 10 GB to 120 GB which is more useful for typical workloads (#2565). Also, ScyllaDB Operator 1.17 fixes an important bug that could lead to storage exhaustion upon upgrade on the first node of the rack due to leftover keyspace snapshots. The fix is available in 1.16.2 as well. (#2528) For more changes and details, check out the GitHub release notes. Upgrading from v1.16.x with kubectl apply doesn’t require any extra action, just take the manifest from v1.17.0 tag and substitute the released image. Using helm requires a mandatory manual step for every release because helm can’t handle CRDs updates. For details, see our upgrade documentation. We’ll welcome your feedback! Feel free to open an issue or reach out on the #scylla-operator channel in ScyllaDB User Slack. The ScyllaDB Operator Team --- ### Page: https://forum.scylladb.com/t/need-to-move-a-large-cluster-from-1-rack-to-3-racks/4833 Title: Need to move a large cluster from 1 rack to 3 racks - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 5.4 #Cluster size: 18 nodes w/ ~4 TiB used. os (RHEL/CentOS/Ubuntu/AWS AMI): Ubuntu We are currently using scylla operator with kube and have built out a large cluster in a si… Language: en Canonical URL: https://forum.scylladb.com/t/need-to-move-a-large-cluster-from-1-rack-to-3-racks/4833 ## Headings Structure: H1: Need to move a large cluster from 1 rack to 3 racks H3: Related topics ## Main Content: H1: Need to move a large cluster from 1 rack to 3 racks H3: Related topics Installation details #ScyllaDB version: 5.4 #Cluster size: 18 nodes w/ ~4 TiB used. os (RHEL/CentOS/Ubuntu/AWS AMI): Ubuntu We are currently using scylla operator with kube and have built out a large cluster in a single availability zone. However, we want to move to having it in 3 AZs and wish to use racks to provide rack awareness to the zones to help mitigate any interzonal network costs. However, we’ve hit issues with operator crashing or putting the entire data set on a single node when attempting this before. Are there any recommended ways of doing this so that we don’t create a major hot spot? You’re using an unsupported version. You might be hitting an issue that was already fixed. Please upgrade to the latest version and see if this still happens. Do you happen top know which changes may have fixed it so that we can validate? --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-94-2025-05-16/4834 Title: Last week in scylla-cluster-tests.git master (issue #94; 2025-05-16) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 743d66c9…5c55c7c0 range are covered. There were 25 non-merge commits from 10 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-94-2025-05-16/4834 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #94; 2025-05-16) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #94; 2025-05-16) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 743d66c9…5c55c7c0 range are covered. There were 25 non-merge commits from 10 authors in that period. Some notable commits: We fixed one-line backtrace decoding so it properly handles single-line traces, such as from oversized allocations. cql-stress was updated to v0.2.2, bringing scylla-rust-driver 1.1.0 and a -transport flag for encrypted client workloads. Gemini was upgraded to 1.9.1 with improvements and is now used in Tier 1 tests. A new rust-driver based performance test was added for custom workload d1/w1, targeting latency in DB cluster steady state. It leverages multiple latte/rune functions, requiring improved HDR tag support and Argus integration. Rolling upgrade tests with vnodes (without tablets) were added to the test tree. Introduced a new ‘actions’ log summarizing SCT’s operations on the cluster in a concise format. So far used in artifact and upgrade tests, with plans to expand across more test components (e.g. Nemeses). pylint disables comments were removed, as pylint has not been used for a while. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/resharindg-hyperthreading-and-scylladb-performance/4837 Title: Resharindg, Hyperthreading and ScyllaDB performance - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/resharindg-hyperthreading-and-scylladb-performance/4837 ## Headings Structure: H1: Resharindg, Hyperthreading and ScyllaDB performance H3: Related topics ## Main Content: H1: Resharindg, Hyperthreading and ScyllaDB performance H3: Related topics Originally from the User Slack @Mahdi_Kamali: Is hyper threading good for ScyllaDB performance? Does it have too much effect? @avi: We think it has a small positive effect, but long time since we last tested it @Mahdi_Kamali: Thanks for your previous answer @avi. I am having trouble resharding production nodes. When a node is resharded, the cluster performance drops drastically. I think the communication between nodes becomes slower. What is the solution? @avi: When a node is resharded it’s offline, so it doesn’t affect cluster performance (apart from the fact that it doesn’t serve requests) --- ### Page: https://forum.scylladb.com/t/comparing-cassandra-to-scylla-what-can-i-do-to-show-scylla-in-its-best-possible-light/4841 Title: Comparing Cassandra To Scylla - what can I do to show scylla in its best possible light? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 6.2.3 #Cluster size: 3 m6g.large-1000 EC2 instances os (RHEL/CentOS/Ubuntu/AWS AMI): AWS AMI I have been some load testing of a small Cassandra system (3 m6g.large-600 nodes -… Language: en Canonical URL: https://forum.scylladb.com/t/comparing-cassandra-to-scylla-what-can-i-do-to-show-scylla-in-its-best-possible-light/4841 ## Headings Structure: H1: Comparing Cassandra To Scylla - what can I do to show scylla in its best possible light? H3: Related topics ## Main Content: H1: Comparing Cassandra To Scylla - what can I do to show scylla in its best possible light? H3: Related topics Installation details #ScyllaDB version: 6.2.3 #Cluster size: 3 m6g.large-1000 EC2 instances os (RHEL/CentOS/Ubuntu/AWS AMI): AWS AMI I have been some load testing of a small Cassandra system (3 m6g.large-600 nodes - Cassandra 4.0.9) for a new data design I am working on. While I have been doing this, we are also testing the waters with scylla so I pointed my load test tooling at it instead. The test tool has two parts: (1) A python tool to insert an average of x number of records for y customers. On Cassandra I did this with batching records in the same partition → this seemed to be the fastest method from what we have learnt from Cassandra use (2) A second Java tool which tries to read all records for each customer at a predetermined rate (e.g. 50 records per second or perhaps 500 records per second) I thought ScyllaDB would smoke it…but at least from what I have seen so far it is working slower. Often ScyllaDB would even fail on the bulk inserts (a couple of billion records inserted over several hours) which were fine for Cassandra. I may not be using Scylla right of course. I have done some research and I think I have found some things I should change. Anything else? Would these make a significant difference to performance?: (1) Switch to massive concurrent write instead of batching. (2) Switch to the ScyllaDB specific drivers (3) Switch to using the shard port (19042) Is the system potentially too small to a fair comparison? Any responses are greatly appreciated. thanks in advance, Gareth Collins Thanks for your question, First, I see you are executing your tests on m6g.large instances, I strongly suggest you stick to one of the instance types ScyllaDB recommends, either i3en, i4i, i7i or i7ie, for AWS. (Cloud Instance Recommendations for AWS, GCP, and Azure | ScyllaDB Docs) As per your questions and assumptions, using our drivers, and performing concurrent inserts would definelty improve performance. Last, consider testing with larger instances (more vCPUs), that will also make the differene as ScyllaDB performance scales better with more cores and memory bandwidth. On small nodes, the shard-per-core design gets bottlenecked easily. I am referencing a couple of my collagues who could give you additional extra tips: cc @avikivity @GuyCarmin Hi @gcollins - a quick question from my end - did you meant 3 nodes of m6g.large? If so - what does the “1000” and “600” number refer to? As @Gabriel mentioned - ScyllaDB is a real-time low-latency DB’s and as such we realy on fast and consist disk performance hence ScyllaDB would work much better on the i4i/i7i machines. on top of that - could you pleas share more about what and how you tested, schema, query pattern etc? And yes - concurrent - or connection per shard is a better way to utilise ScyllaSB. a hint- in order to see where is the bottle neck it is highly recommended to deploy and connect the Scylla Monitoring stack which will give you a full visibility of the cluster. feel free to ping me if you have more info / more questions Thanks very much for the responses! The “-600” and “-1000” indicate the disk sizes. So the schema we were using looks like this: PRIMARY KEY (id, timestamp, unique_id) ) WITH CLUSTERING ORDER BY (timestamp DESC, unique_id ASC) AND compaction = {‘class’: ‘org.apache.cassandra.db.compaction.LeveledCompactionStrategy’} We are using leveled compaction as we are optimizing to minimize latency on reading. The load testing involves reading all the records for random ids on a configurable number of threads. And we are testing varying the number of ids, the number of records per id and the size of the data blob. Interestingly, at least for Cassandra, fewer records with more data per record appears to do much better for read speed than more records with much less data (but the same total data overall). A question - @GuyCarmin We have systems in production which are 27-36 nodes for example (m6g.4xlarge). What sort of size cut should I be able to expect from a Scylla DB system (very roughly)? We tend to currently get limited by the CPU rather than I/O. For Scylla, is it that much more CPU efficient that it is way more likely to be limited by I/O instead of CPU thus the recommendation for instances that use SSDs? --- ### Page: https://forum.scylladb.com/t/row-cache-for-different-fields-in-the-same-record/4842 Title: Row cache for different fields in the same record - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: SELECT email, username, age FROM users WHERE id = ? LIMIT 1 SELECT id, email, created_at, username FROM users WHERE id = ? LIMIT 1 SELECT city FROM users WHERE id = ? LIMIT 1 I have a few more similar queries in scyll… Language: en Canonical URL: https://forum.scylladb.com/t/row-cache-for-different-fields-in-the-same-record/4842 ## Headings Structure: H1: Row cache for different fields in the same record H3: Related topics ## Main Content: H1: Row cache for different fields in the same record H3: Related topics SELECT email, username, age FROM users WHERE id = ? LIMIT 1 SELECT id, email, created_at, username FROM users WHERE id = ? LIMIT 1 SELECT city FROM users WHERE id = ? LIMIT 1 I have a few more similar queries in scylladb, different cases require different fields and I think that I do more performant by selecting only the necessary ones, instead of simply selecting the entire record with all fields for all cases. But is it really more efficient to do this? And at the scylladb level too, because it caches at the row level and doesn’t this break the cache by essentially caching the same record several times, only different fields? Like If the row is not in the row cache, it is read from the disk. If I have three different queries that read by id, but different fields, they will all create separate row cache entries, which inflates the cache and can reduce the hit ratio. I have a few more similar queries in scylladb, different cases require different fields and I think that I do more performant by selecting only the necessary ones, instead of simply selecting the entire record with all fields for all cases. But is it really more efficient to do this? Yes! Providing the required columns to be fetched for a given SELECT is more efficient because it reduces the processing overhead for columns you won’t use, as well as the server to client response payload (be wary of network costs!). And at the scylladb level too, because it caches at the row level and doesn’t this break the cache by essentially caching the same record several times, only different fields? Like If the row is not in the row cache, it is read from the disk. If I have three different queries that read by id, but different fields, they will all create separate row cache entries, which inflates the cache and can reduce the hit ratio. This is where it gets interesting. Yes, ScyllaDB will always cache the whole row, despite only a limited subset of columns were provided in the SELECT query. As we discussed above, providing the columns on SELECT is great to reduce serialization and deserialization as well as network overheads. But it doesn’t affect cache population and efficiency. You are on the right track to make more efficient use of your cache. The next question you should look into is how often different groups of columns are used, especially in relation to each other. If you can break most-used columns into their own table, and keep a separate table for columns that are less frequently used, you may improve your cache utilization. Especially if you have larger cells that are rarely used. --- ### Page: https://forum.scylladb.com/t/tombstones-warning-threshold-creating-and-removing-tables-every-few-hours-performance-impacts/4843 Title: Tombstones warning, threshold, creating and removing tables every few hours, performance impacts - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/tombstones-warning-threshold-creating-and-removing-tables-every-few-hours-performance-impacts/4843 ## Headings Structure: H1: Tombstones warning, threshold, creating and removing tables every few hours, performance impacts H3: Related topics ## Main Content: H1: Tombstones warning, threshold, creating and removing tables every few hours, performance impacts H3: Related topics Originally from the User Slack @Nick_Vladiceanu: Hi folks, we are seeing quite a lot of warnings regarding Tombstones in system_schema keyspace, Scylla 5.4.9: As far as I understand, this can hurt the cluster performance. Can we do something to clean up those tombstones faster? Since the gc_grace_seconds = 604800 and it’s immutable for this keyspace, what are the options? Thanks @Botond_Dénes: Major compaction on the keyspace can help. These tables are not read all the time, only on startup and on schema changes. @Nick_Vladiceanu: I have tried to execute a few times Although it runs without issues, the warnings about tombstones are still present We create and remove tables every 2-3 hours @Botond_Dénes: In a recent version (don’t remember which exactly) we added a different method to purge tombstones, this would allow getting rid of most of them. > We create and remove tables every 2-3 hours With this use-case you will have many tombstones, no way to avoid it. The above mentioned improvement will help though. You can look around in the release notes of later releases (6.0+) to see which one has this improvement. In any case I suggest upgrading, 5.4 is EOL for a while now. @Nick_Vladiceanu: Thanks a lot Botond. We have recently upgraded to 5.4 as preparation for the 6.x upgrade, which will follow soon. I will post here updates after the upgrade to 6.x. --- ### Page: https://forum.scylladb.com/t/using-wasm-binaries-instead-of-wat/4846 Title: Using WASM Binaries instead of WAT - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I’m attempting to embed the Parquet crate as a UDF so that I can query the metadata for a blob without taking the network hit of sending the whole thing over the wire. It’s probably not a common scenario, but I wanted to… Language: en Canonical URL: https://forum.scylladb.com/t/using-wasm-binaries-instead-of-wat/4846 ## Headings Structure: H1: Using WASM Binaries instead of WAT H3: Wasm support for user-defined functions | ScyllaDB Docs H3: Related topics ## Main Content: H1: Using WASM Binaries instead of WAT H3: Wasm support for user-defined functions | ScyllaDB Docs H3: Related topics I’m attempting to embed the Parquet crate as a UDF so that I can query the metadata for a blob without taking the network hit of sending the whole thing over the wire. It’s probably not a common scenario, but I wanted to push the limits of UDFs to see how far I got. The current implementation of UDFs is pretty neat for using Wasmtime. However, I’m running into some sizing issues after optimizing my binary size. My WASM binary is 644KB and my WAT file is 19MB. The WAT loader does not like such a big bundle and it sits at 100% CPU for hours until I eventually kill it. Since the current issue is with loading, would it be possible to allow loading the WASM binary instead of the WAT file? As far as I can tell, the only way to load a UDF is through a CQL query embedding the WAT directly as a string. ScyllaDB requires WebAssembly UDFs to be provided in the WebAssembly Text (WAT) format, which must be embedded directly as a string in the CQL CREATE FUNCTION statement. While you can write the WAT by hand, typically you compile your code to a Wasm binary and then convert it to WAT using a tool like wasm2wat. Currently, loading Wasm binaries directly into ScyllaDB UDFs is not supported, so you need to provide the full WAT text even if it’s large. This is the recommended and only supported method as of the latest ScyllaDB versions. Here’s a link for the WASM relevant official documentation - ScyllaDB is an Apache Cassandra-compatible NoSQL data store that can handle 1 million transactions per second on a single server. --- ### Page: https://forum.scylladb.com/t/failed-to-restart-scylla-server-service-unit-scylla-image-setup-service-not-found/4849 Title: Failed to restart scylla-server.service: Unit scylla-image-setup.service not found - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: #Cluster size: os (RHEL/CentOS/Ubuntu/AWS AMI): We are using AWS Scylla DB AMI 5.2.8 version. We wanted it to be added to the cluster which has a node of 5.4.9. So we were tryi… Language: en Canonical URL: https://forum.scylladb.com/t/failed-to-restart-scylla-server-service-unit-scylla-image-setup-service-not-found/4849 ## Headings Structure: H1: Failed to restart scylla-server.service: Unit scylla-image-setup.service not found H2: Root Cause H2: Fix: Reinstall Scylla AMI Tools H3: Step 1: Reinstall Scylla AMI-specific scripts H3: Step 2: Re-run the image setup H3: Related topics ## Main Content: H1: Failed to restart scylla-server.service: Unit scylla-image-setup.service not found H2: Root Cause H2: Fix: Reinstall Scylla AMI Tools H3: Step 1: Reinstall Scylla AMI-specific scripts H3: Step 2: Re-run the image setup H3: Related topics Installation details #ScyllaDB version: #Cluster size: os (RHEL/CentOS/Ubuntu/AWS AMI): We are using AWS Scylla DB AMI 5.2.8 version. We wanted it to be added to the cluster which has a node of 5.4.9. So we were trying to upgrade it to 5.4.9. Somehow it got corrupted and we are getting “Failed to restart scylla-server.service: Unit scylla-image-setup.service not found.” while running sudo systemctl restart scylla-server. nstallation details #ScyllaDB version: #Cluster size: os (RHEL/CentOS/Ubuntu/AWS AMI): We are using AWS Scylla DB AMI 5.2.8 version. We wanted it to be added to the cluster which has a node of 5.4.9. So we were trying to upgrade it to 5.4.9. Somehow it got corrupted and we are getting “Failed to restart scylla-server.service: Unit scylla-image-setup.service not found.” while running sudo systemctl restart scylla-server. You’re dealing with a broken ScyllaDB upgrade on an AWS AMI, and the error: means that the ScyllaDB image setup service that prepares the AMI environment is either missing or wasn’t properly migrated during the upgrade. This is specific to ScyllaDB AMIs, which rely on custom units and setup scripts (not found in vanilla packages). You likely did a package-level upgrade (e.g. via yum or apt) on an AWS Scylla AMI, which is not recommended unless you’re following the official Scylla upgrade path for AMIs. AWS Scylla AMIs come pre-configured with systemd units like scylla-image-setup.service. If those are deleted, replaced, or not installed during a manual upgrade, Scylla won’t start correctly. You need to restore the missing image setup scripts and systemd services. Try this: This package contains scylla-image-setup.service and related systemd configuration required for the AMI. Thanks @Gabriel for the response. I followed your instruction. However, I am unable to find that package. Please check the below error that I received sudo apt install -y scylla-ami-setup Reading package lists… Done Building dependency tree… Done Reading state information… Done E: Unable to locate package scylla-ami-setup I followed your instruction. However, I am unable to find that package. Please check the below error that I received sudo apt install -y scylla-ami-setup Reading package lists… Done Building dependency tree… Done Reading state information… Done E: Unable to locate package scylla-ami-setup I’d therefore recommend to decommission the corrupted and launch a new one using the official Scylla AMIs for the desired version. @Yaron_Kaikov - would you have an alternate suggestion ? @Gabriel We had a single node set up. We have to set it up in cluster. If we can’t upgrade the node, we can’t add node to it. In this case, we are not sure how to migrate the data. We have around 800GB of data. Could you please suggest any suitable strategy for this? @Amit_Sahoo I think you are trying to install the wrong package, try sudo apt install -y scylla-machine-image Thanks @Yaron_Kaikov. That worked. --- ### Page: https://forum.scylladb.com/t/new-instance-of-aws-ami-scylladb-5-2-8-failed-to-start-scylla-server/4850 Title: New instance of AWS AMI ScyllaDB 5.2.8 failed to start Scylla Server - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 5.2.8 #Cluster size: single node aws AMI: Getting Below error: [shard 0] init - I/O Scheduler is not properly configured! This is a non-supported setup, and performance is ex… Language: en Canonical URL: https://forum.scylladb.com/t/new-instance-of-aws-ami-scylladb-5-2-8-failed-to-start-scylla-server/4850 ## Headings Structure: H1: New instance of AWS AMI ScyllaDB 5.2.8 failed to start Scylla Server H3: Related topics ## Main Content: H1: New instance of AWS AMI ScyllaDB 5.2.8 failed to start Scylla Server H3: Related topics Installation details #ScyllaDB version: 5.2.8 #Cluster size: single node aws AMI: Getting Below error: [shard 0] init - I/O Scheduler is not properly configured! This is a non-supported setup, and performance is expected to be unpredictably bad. Reason found: none of --max-io-requests, --io-properties and --io-properties-file are set. To properly configure the I/O Scheduler, run the scylla_io_setup utility shipped with Scylla. Is it supposed to work like this? You are running an old version that is no longer supported; please upgrade. What happens when you run the utility as the message suggests? I had to add developer_mode: true in the scylla.yaml to up the server . --- ### Page: https://forum.scylladb.com/t/cassandra-scylladb-driver-which-protocol-version-has-support/4854 Title: Cassandra ScyllaDB Driver, which protocol version has support? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/cassandra-scylladb-driver-which-protocol-version-has-support/4854 ## Headings Structure: H1: Cassandra ScyllaDB Driver, which protocol version has support? H3: Related topics ## Main Content: H1: Cassandra ScyllaDB Driver, which protocol version has support? H3: Related topics Originally from the User Slack @Joey_de_Waal: Hi, for educational purposes I am writing a CassandraDB/ScyllaDB database driver and started implementing native protocol V5 while testing against cassandra db. Now I am trying to connect to scyllaDB (in Docker) it seems that V5 is not supported (and V4 is used in cqlsh). Will there be support for V5 or am I missing something? @avi: v5 will come one day, but not being implemented now @Joey_de_Waal: Thanks for the quick reply, I’ll implement V4 as well then! --- ### Page: https://forum.scylladb.com/t/adding-a-node-in-another-data-center-with-a-different-port-runtime-error-message/4855 Title: Adding a node in another data center with a different port, runtime error message - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/adding-a-node-in-another-data-center-with-a-different-port-runtime-error-message/4855 ## Headings Structure: H1: Adding a node in another data center with a different port, runtime error message H3: Related topics ## Main Content: H1: Adding a node in another data center with a different port, runtime error message H3: Related topics Originally from the User Slack @novan.bram: Hi folks, I’m running a 3 node scyllaDB cluster v4.3 in GCP dc, and i want to add new node in another dc which is cannot using default scylla port like (e.g. 7000, 9042, etc). when i want to join the new node into existing scylla cluster in GCP, i got this error is there anyone can help for this? Thankyou @Felipe_Cardeneti_Mendes: 4.3 is an ancient release — IIRC it had to do with not being able to gossip with your seed node but rather than figuring it out upgrade first @novan.bram: @Felipe_Cardeneti_Mendes got it, may i know what is the stable version to support the multi DC cluster and support for ipv6 and custom internode port? @Felipe_Cardeneti_Mendes: 2025.1 is our latest major release @novan.bram: is it for enterprise only no? what if the open source version? @Felipe_Cardeneti_Mendes: 2025.1 a Source Available License 6.2 is the latest GPL release --- ### Page: https://forum.scylladb.com/t/stuck-topology-requests/4856 Title: Stuck topology requests - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 6.2 #Cluster size: 12 nodes x 3 dc #OS: Ubuntu 24.10 Hi. I had issues with couple of nodes which led to them being down, unable to be removed and replaced. SELECT * FROM sys… Language: en Canonical URL: https://forum.scylladb.com/t/stuck-topology-requests/4856 ## Headings Structure: H1: Stuck topology requests H3: Related topics ## Main Content: H1: Stuck topology requests H3: Related topics Installation details #ScyllaDB version: 6.2 #Cluster size: 12 nodes x 3 dc #OS: Ubuntu 24.10 I had issues with couple of nodes which led to them being down, unable to be removed and replaced. I don’t have exact steps that led to this, one of them ran out of disk space and I’ve tried to replace it with the other. As I understand they are blocking other cluster ops, and removing records from topology_requests table will not help. How can i remove/cancel stuck topology requests? Also I have ks with tablets enabled, so no manual recovery. It would help if you attached logs. I have two guesses from your description: Before logs, the contents of system.topology and system.cluster_status would also help. Hi. Thank You for the reply. two nodes are down, yes I don’t think removenode can take more than one host id removenode was executed with --ignore-dead-nodes for both subsequents attempts result in “Concurrent request for removal already in progress” for both no, quorum is not lost, i have 12 nodes per each of 3 dc does restoring group0 implies “manual recovery”? because per documentation it is incompatible with tablets being enabled I also have a ghost node, which might be the cause of this: Ghost node in none state I have ungodly amount of logs. What should I look for and on which node? I have yet to find anything of practical value. Maybe we don’t need logs. First, let’s check the status of topology. On one of the nodes: select * from system.cluster_status; select * from system.topology; removenode was executed with --ignore-dead-nodes for both subsequents attempts result in “Concurrent request for removal already in progress” for both That’s fine, the nodes should still be marked as excluded (can be observed in system.topology in the “ignore_nodes” column). no, quorum is not lost, i have 12 nodes per each of 3 dc does restoring group0 implies “manual recovery”? because per documentation it is incompatible with tablets being enabled Yes, but it’s not needed here. If you have a ghost node, then indeed it could block progress. You can manually fix it like this: Observe if things start moving. If they are, the node on which topology coordinator runs should start logging progress. Leader runs on the node which was the last to log raft_group0 - gaining leadership. On topology updates, so also when tablet migration is moving forward, it will logs “raft_topology - updating topology state:” --- ### Page: https://forum.scylladb.com/t/performance-differences-with-kubernetes-and-without-it/4858 Title: Performance differences with Kubernetes and without it - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/performance-differences-with-kubernetes-and-without-it/4858 ## Headings Structure: H1: Performance differences with Kubernetes and without it H3: Related topics ## Main Content: H1: Performance differences with Kubernetes and without it H3: Related topics Originally from the User Slack @JIST: Hi, do you see performance differences in case of run ScyllaDB in K8s vs VM (with the “same” HW configuration)? NOTE: K8s => ingress, autoscaler, … VM => higher HW configuration, typically better I/O, … @avi: Yes, in general VMs give better I/O since our kernel tuning works on VMs and not on Kubernetes without special magic --- ### Page: https://forum.scylladb.com/t/fix-installation-failure-on-debian-testing-trixie-possibly-others-due-to-wrong-gpg-output-key-format/4861 Title: [FIX] installation failure on Debian testing/trixie, possibly others, due to wrong GPG output key format - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details Installing from apt repo on Debian testing, following the getting started guide. I’m just regurgitating here, but AFAIU, there’s been a bug for awhile in gpg where it’s creating apt-compatible GPG … Language: en Canonical URL: https://forum.scylladb.com/t/fix-installation-failure-on-debian-testing-trixie-possibly-others-due-to-wrong-gpg-output-key-format/4861 ## Headings Structure: H1: [FIX] installation failure on Debian testing/trixie, possibly others, due to wrong GPG output key format H3: Related topics ## Main Content: H1: [FIX] installation failure on Debian testing/trixie, possibly others, due to wrong GPG output key format H3: Related topics Installation details Installing from apt repo on Debian testing, following the getting started guide. I’m just regurgitating here, but AFAIU, there’s been a bug for awhile in gpg where it’s creating apt-compatible GPG keys though it shouldn’t, so that the instruction in the guide, sudo gpg --homedir /tmp --no-default-keyring --keyring /etc/apt/keyrings/scylladb.gpg --keyserver hkp://keyserver.ubuntu.com:80 --recv-keys a43e06657bac99e3 works. Recent updates have fixed this, so i now creates the proper format, which is not apt compatible (source on github: /freedomofpress/dangerzone/issues/1052#issuecomment-2594065503). This means apt-update will fail to include the Scylla repo, and installation will fail. One solution, inspired by a random fix elsewhere (also on github: /freedomofpress/dangerzone/commit/1398c4d354822b2cdd67cdea857f73ce6117cf6c): sudo apt install qv sudo sq network keyserver --server hkp://keyserver.ubuntu.com:80 search a43e06657bac99e3 --overwrite --output /etc/apt/keyrings/scylladb.gpg This now leaves a valid key for apt, and I can continue the installation. Just FYI, gotta go install and play around now #ScyllaDB version: Latest one, not sure, haven’t installed it just yet #Cluster size: os (RHEL/CentOS/Ubuntu/AWS AMI): Debian Trixe (testing) --- ### Page: https://forum.scylladb.com/t/release-scylladb-rust-driver-1-2-0/4864 Title: [RELEASE] ScyllaDB Rust Driver 1.2.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce ScyllaDB Rust Driver 1.2.0, an asynchronous CQL driver for Rust, optimized for Scylla, but also compatible with Apache Cassandra! Some interesting statistics: over 3.877k down… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-rust-driver-1-2-0/4864 ## Headings Structure: H1: [RELEASE] ScyllaDB Rust Driver 1.2.0 H2: Changes H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Rust Driver 1.2.0 H2: Changes H3: Related topics The ScyllaDB team is pleased to announce ScyllaDB Rust Driver 1.2.0, an asynchronous CQL driver for Rust, optimized for Scylla, but also compatible with Apache Cassandra! Some interesting statistics: New features / enhancements: Internal API cleanups/refactors: CI / developer tool improvements: Congrats to all contributors and thanks everyone for using our driver! The source code of the driver can be found here: The official crates.io registry entry is here: Thank you for your attention, please do not hesitate to contact us if you have any questions, issues, feature requests, or are simply interested in our driver! Contributors since the last release: --- ### Page: https://forum.scylladb.com/t/procedure-for-removing-a-node-from-a-cluster-replication-factor/4865 Title: Procedure for removing a node from a cluster, replication factor - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/procedure-for-removing-a-node-from-a-cluster-replication-factor/4865 ## Headings Structure: H1: Procedure for removing a node from a cluster, replication factor H3: Related topics ## Main Content: H1: Procedure for removing a node from a cluster, replication factor H3: Related topics Originally from the User Slack @Shivaprasad_Bhat: Hi team! We had a Scylla cluster with 3 nodes containing a keyspace with RF=3 (NetworkTopolgyStrategy).. We removed one node due to some cost reasons (in a dev environment). The node status was showing as DN on sctool. What is the process to remove the node from this? Do we need to reduce RF first? @Piotr_Smaroń: down scale procedure is summarized here https://opensource.docs.scylladb.com/stable/operating-scylla/procedures/cluster-management/remove-node.html Remove a Node from a ScyllaDB Cluster (Down Scale) | ScyllaDB Docs --- ### Page: https://forum.scylladb.com/t/last-two-weeks-in-scylla-cluster-tests-git-master-issue-95-2025-05-30/4867 Title: Last two weeks in scylla-cluster-tests.git master (issue #95; 2025-05-30) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the aa5f1007…642370c5 range are covered. There were 16 non-merge commits from 8 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-two-weeks-in-scylla-cluster-tests-git-master-issue-95-2025-05-30/4867 ## Headings Structure: H1: Last two weeks in scylla-cluster-tests.git master (issue #95; 2025-05-30) H3: Related topics ## Main Content: H1: Last two weeks in scylla-cluster-tests.git master (issue #95; 2025-05-30) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the aa5f1007…642370c5 range are covered. There were 16 non-merge commits from 8 authors in that period. Some notable commits: A new longevity test was added to ensure no coordination traffic reaches AZs without loaders, with rack-aware policy validation added to support it. Latte was updated to 0.30.0-scylladb, bringing in scylla-rust-driver 1.1, with support for validating read query row counts and fixed DB version detection and output. Some CI pipelines now include the separator plugin for clearer job configuration visibility. cql-stress updated to 0.2.3, fixing coordinated omission latency reporting for sub-millisecond ranges. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/release-scylla-doctor-v1-6/4868 Title: [RELEASE]: Scylla Doctor v1.6 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: Scylla Doctor v1.6 is released. Fixes: scylla-doctor: CPUSetCollector: skip the collector if ‘–cpuset’ is not set in CPUSET OSSupportAnalyzer: Add support for newer editions scylla-doctor: Makefile: add support for Po… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-doctor-v1-6/4868 ## Headings Structure: H1: [RELEASE]: Scylla Doctor v1.6 H3: Related topics ## Main Content: H1: [RELEASE]: Scylla Doctor v1.6 H3: Related topics Scylla Doctor v1.6 is released. --- ### Page: https://forum.scylladb.com/t/rust-driver-serialize-enum-into-text/4869 Title: Rust Driver, serialize enum into text - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/rust-driver-serialize-enum-into-text/4869 ## Headings Structure: H1: Rust Driver, serialize enum into text H3: Related topics ## Main Content: H1: Rust Driver, serialize enum into text H3: Related topics Originally from the User Slack @Mattia: Is there a straightforward way to serialize a Rust enum into a text inside scylla? Since we can not just derive the macro Does this makes sense? When MyCustomEnum also implements Display @Karol_Baryła: This impl looks correct. You could possibly avoid the allocation caused by creating a String , but its hard to tell without seeing more of your code. @Mattia: I have standard enums like so: Do you have any pointer on how I could avoid the allocation? @Karol_Baryła: Can’t you implement as_str method, returning a string literal different for each case? I don’t see why you need a String. @Mattia: You’re right, thanks for the heads up --- ### Page: https://forum.scylladb.com/t/last-three-weeks-in-scylladb-git-master-issue-280-2025-06-01/4870 Title: Last three weeks in scylladb.git master (issue #280; 2025-06-01) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last three weeks. Commits in the 092a88c9b9..7d562c24b1 range are covered. There were 198 non-merge commits from 31 authors in t… Language: en Canonical URL: https://forum.scylladb.com/t/last-three-weeks-in-scylladb-git-master-issue-280-2025-06-01/4870 ## Headings Structure: H1: Last three weeks in scylladb.git master (issue #280; 2025-06-01) H3: Related topics ## Main Content: H1: Last three weeks in scylladb.git master (issue #280; 2025-06-01) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last three weeks. Commits in the 092a88c9b9..7d562c24b1 range are covered. There were 198 non-merge commits from 31 authors in that period. Some notable commits: When loading sstables into the database with nodetool, there is now a –skip-cleanup option when the user wishes to defer cleanup to a later time. This allows a single cleanup operation to be run for many sstable loads. Materialized view updates perform a read-modify-write operation on the base table. To prevent overload, we queue some of the reads. Previously, this queue had a limit of some number of entries beyond which reads would be rejected. This could cause base/view inconsistencies. The queue limit is now removed and we rely on throttling the base table writes to control its length. Some alternator tests were migrated from dtest to test.py. This marks the beginning of an effort to migrate all dtests into scylladb.git. We now manage temporary memory for zstd inter-node compression ourselves, in order to reduce allocation stalls. The compaction history table now has additional columns for statistics. There is now a custom index class (useful in CREATE INDEX) for vector indexes. Note it is not yet possible to query using the new index. The native CQL transport now supports the metadata ID extension (from protocol version 5) that allows updating row metadata for prepared statements for SELECT * queries (when columns were added or removed) or SELECT udt queries (when the user define type definition changed). Note a driver that supports the extension is required to make use of this. The nodetool refresh command now supports the –scope option, allowing data to be streamed only to the local node, local rack or local datacenter. This is useful when restoring from a backup that contains individual sstable sets for each rack. A bug which could cause data resurrection in materialized views (but not in the base tables) due to mixing up purge times for regular tombstones and shadowable tombstones was fixed. The process for merging tablets (which happens when a table’s size decreases) did not update the row cache about the merged sstables, causing some data to be missed during full scans. This is now fixed. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/p99-and-p95-spikes-hot-partitions-performance-and-data-modeling/4873 Title: P99 and p95 spikes, hot partitions, performance and data modeling - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/p99-and-p95-spikes-hot-partitions-performance-and-data-modeling/4873 ## Headings Structure: H1: P99 and p95 spikes, hot partitions, performance and data modeling H3: Related topics ## Main Content: H1: P99 and p95 spikes, hot partitions, performance and data modeling H3: Related topics Originally from the User Slack must-gather.zip (9.2 MB) @Sh4d1: Hello all! I’m hitting some performance issue on my scylla cluster (5 nodes on GKE with scylla operator, with 7 cores and 56GB reserved for each scylla node). You can see the p95/99 on the attached image. During peak load we are a bit below 60k write/s and 35k read/s. From what I read scylla should not be that slow with that number of requests. In the logs, I see quite a lot of which iiuc, is that the cpu is overloaded, which makes sense with the view per shard (second image). My guess is that we’re using scylla in a bad way for some table: • I have around 60 large cells (no more than 5M) • 305 large partitions (some going over 10M), and CDC enabled on one of the table with the most large partitions • a lot of different tables/keyspaces, more or less used. Is that enough to show the slowness I’m seeing ? Or could it be something else too? Feel free to ask more questions if needed, and thanks in advance! The spike on the read latency on the left was due to a huge spike in writes at that time* Also, could this be linked I guess? @avi: What kind of storage are you using? @Maciej_Zimnoch: Have you followed https://operator.docs.scylladb.com/stable/architecture/tuning.html ? Tuning | ScyllaDB Docs btw Avi is disabling write back cache still a recommended thing to do on GCP Local SSDs? @Sh4d1: Yep sorry, I’m using local nvme with raid0 (with write cache disabled on xfs), and the tuning should have been done by the operator (the jobs ran correctly) And of course cpu sets are used (hence 7 and not 8, cause of k8s system component using some cpu) @Maciej_Zimnoch: would you mind sharing a must-gather? I can double check if tuning was executed correctly Gathering data with must-gather | ScyllaDB Docs @Sh4d1: Yes, doing that, do you need the --all-resources ? @Maciej_Zimnoch: no, default collects what we usually need @Sh4d1: There you go! @Maciej_Zimnoch: tuning looks ok. You’re using quite ancient Scylla version (5.4.x), I would recommend trying out the latest one @Sh4d1: Yes the upgrade is planned somewhere on my backlog I guess it can always help indeed! Better to upgrade and then try to see if the issues are still there according to you? @Maciej_Zimnoch: yes, newer versions contains lots of improvements @Sh4d1: Makes sense will try to do it end of this week, will keep you updated here! Sorry for the delay, I’m currently upgrading my cluster to 6.2.3. Should I see improvements directly? Looks like the latency decreased quite a bit, still have some p99 above 50ms, with spikes up to 1-2s, but I guess it’s linked to our hot partitions? --- ### Page: https://forum.scylladb.com/t/is-high-cardinality-partition-key-a-problem-in-scylladb/4874 Title: Is High-Cardinality Partition Key a Problem in ScyllaDB? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 5.2 os (RHEL/CentOS/Ubuntu/AWS AMI): CentOS I’m designing a ScyllaDB table where each partition is tied to a unique job_id, which is a UUID. Each job typically inserts a small … Language: en Canonical URL: https://forum.scylladb.com/t/is-high-cardinality-partition-key-a-problem-in-scylladb/4874 ## Headings Structure: H1: Is High-Cardinality Partition Key a Problem in ScyllaDB? H3: Related topics ## Main Content: H1: Is High-Cardinality Partition Key a Problem in ScyllaDB? H3: Related topics Installation details #ScyllaDB version: 5.2 os (RHEL/CentOS/Ubuntu/AWS AMI): CentOS I’m designing a ScyllaDB table where each partition is tied to a unique job_id, which is a UUID. Each job typically inserts a small number of rows — sometimes just one or two. Over time, the number of jobs may grow to several million, resulting in a large number of small partitions. My current schema looks like this: Is it acceptable in ScyllaDB to have millions of small partitions with high-cardinality UUIDs? Or Am I misunderstanding it? Since I read it’s very adviced to focus on high cardinality for partition keys or secondary indexes In ScyllaDB, it’s absolutely acceptable (and often desirable) to have millions of small, high-cardinality partitions, especially when: --- ### Page: https://forum.scylladb.com/t/ghost-node-in-none-state/4875 Title: Ghost node in none state - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 6.2 with tablets enabled #Cluster size: 12 x 3dc os (RHEL/CentOS/Ubuntu/AWS AMI): Ubuntu 24.10 Hi. I have a ghost node select * from system.cluster_status; peer | dc … Language: en Canonical URL: https://forum.scylladb.com/t/ghost-node-in-none-state/4875 ## Headings Structure: H1: Ghost node in none state H3: Related topics ## Main Content: H1: Ghost node in none state H3: Related topics Installation details #ScyllaDB version: 6.2 with tablets enabled #Cluster size: 12 x 3dc os (RHEL/CentOS/Ubuntu/AWS AMI): Ubuntu 24.10 It does not appear in system.raft_state, but I got its host_id from nodetool gossipinfo. Attempts to removenode it lead to I can not assasinate it, because raft. Rest api’s /gossiper/force_remove_endpoint/10.0.0.1 would happily report that it “Finished to force remove node” and do nothing (it would do so for 8.8.8.8 also). I think it prevents getting my cluster in a working state by blocking topology ops. As with my other problem I don’t have exact steps that led to this, one of them ran out of disk space and I’ve tried to replace it with the others. A rolling restart will frequently clear ghost nodes. --- ### Page: https://forum.scylladb.com/t/debian-dependency-conflict-with-scylla-node-exporter/4876 Title: Debian dependency conflict with scylla.node-exporter - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 6.2.3 #Cluster size:16 os (debian bookworm): We want to install some prometheus plugins. These do depend on prometheus node exporter. The scylla node exporter has a “conflict… Language: en Canonical URL: https://forum.scylladb.com/t/debian-dependency-conflict-with-scylla-node-exporter/4876 ## Headings Structure: H1: Debian dependency conflict with scylla.node-exporter H3: Related topics ## Main Content: H1: Debian dependency conflict with scylla.node-exporter H3: Related topics Installation details #ScyllaDB version: 6.2.3 #Cluster size:16 os (debian bookworm): We want to install some prometheus plugins. These do depend on prometheus node exporter. The scylla node exporter has a “conflicts” on prometheus node exporter which prevents installing those plugins on debian. I think the debian package for scylla node exporter should contain a “provides prometheus node exporter” declaration, which would allow to install more prometheus plugins. Is this a valid solution or has somebody a better advice? Thanks for helping, Michael scylla-node-exporter is not a dependency for scylla-server just replace it with whatever version you want Thanks for the reply. I was confused because of the message: “The following packages will be removed: scylla scylla-node-exporter” But yes, it’s only the meta package “scylla” which gets removed and not the database. --- ### Page: https://forum.scylladb.com/t/nodetool-cleanup-per-node-cluster-table-or-keyspace/4878 Title: Nodetool cleanup per node cluster table or keyspace? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/nodetool-cleanup-per-node-cluster-table-or-keyspace/4878 ## Headings Structure: H1: Nodetool cleanup per node cluster table or keyspace? H3: Related topics ## Main Content: H1: Nodetool cleanup per node cluster table or keyspace? H3: Related topics Originally from the User Slack @Renato_Nascimento: Hi team, since I have upgraded my clusters to 6.1 (with the official AMI), I noticed that the nodetool cleanup operation is cluster wide, even with tables disabled. Is this the expected behavior? Is there a way to run cleanups just on a selected node? Thank you! @Thomas_Braly: Hey @Renato_Nascimento, nice to see you again! I’ll talk with my team and get you an answer asap. In ScyllaDB 6.1 nodetool cleanup only runs on the node you point it at, but if you don’t give it a keyspace (and optionally a table), it will scan every keyspace on that host—even tables you’ve otherwise disabled. To limit it: More details here: https://opensource.docs.scylladb.com/branch-6.1/operating-scylla/nodetool.html Let me know if you have any additional questions! Nodetool | ScyllaDB Docs --- ### Page: https://forum.scylladb.com/t/scylladb-upgrade-path-repositories-download-location-and-will-there-be-downtime/4879 Title: ScyllaDB upgrade path, repositories download location and will there be downtime? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-upgrade-path-repositories-download-location-and-will-there-be-downtime/4879 ## Headings Structure: H1: ScyllaDB upgrade path, repositories download location and will there be downtime? H3: Related topics ## Main Content: H1: ScyllaDB upgrade path, repositories download location and will there be downtime? H3: Related topics Originally from the User Slack @Some_Random_Guy: hey all. I’m reading documentation in preparation to bring a 5.2.2 cluster up to date. According to what i’m reading, there seems to be an upgrade path from oss to 2023.1 enterprise. Since the newer enterprise editions replace the oss version, can I upgrade 5.2.2 to 2023.1 > 2024.1 > 2024.2 > 2025.1 > 2025.2 instead of going through the entire point release process for everything in the oss tree from 5.2.2 to 6.2.x? @Felipe_Cardeneti_Mendes: yes @Some_Random_Guy: sweet! thanks as always Felipe! actually one more question. Will I incur downtime due to licensing issues before I’m on the latest version? @Felipe_Cardeneti_Mendes: nope, the upgrade process works the same way all the way, rolling restart @Some_Random_Guy: where can i get the enterprise package? the method listed in the documentation tells me the package can’t be found https://opensource.docs.scylladb.com/branch-5.2/upgrade/upgrade-to-enterprise/upgrade-guide[…]o-2023.1/upgrade-guide-from-5.2-to-2023.1-generic.html Upgrade Guide - ScyllaDB 5.2 to 2023.1 | ScyllaDB Docs @Felipe_Cardeneti_Mendes @Felipe_Cardeneti_Mendes: You may find the repositories here https://downloads.scylladb.com/deb/ubuntu/ @Some_Random_Guy: oh! thanks! is enterprise the same where I have to do all the point releases ? or can i go from 2023.1 > 2024.1 i say that under the assumption that i need to go from 5.2.2 to 2023.1.0 is that the case or can i go straight to 2023.1.11? nm, i guess the upgrade answered that question for me --- ### Page: https://forum.scylladb.com/t/release-scylladb-monitoring-4-10-0/4880 Title: [RELEASE] ScyllaDB Monitoring 4.10.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of ScyllaDB Monitoring Stack 4.10.0 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB based on Prometheus and Grafana. ScyllaDB Monitoring Sta… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-monitoring-4-10-0/4880 ## Headings Structure: H1: [RELEASE] ScyllaDB Monitoring 4.10.0 H1: New Information in ScyllaDB Dashboards H2: Overview Dashboard Change H2: Detailed Dashboard Change H2: Alternator Dashboard Change H2: OS Dashboard Change H2: Advanced Dashboard Change H1: Bug Fixes H1: Operational Changes H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Monitoring 4.10.0 H1: New Information in ScyllaDB Dashboards H2: Overview Dashboard Change H2: Detailed Dashboard Change H2: Alternator Dashboard Change H2: OS Dashboard Change H2: Advanced Dashboard Change H1: Bug Fixes H1: Operational Changes H3: Related topics The ScyllaDB team is pleased to announce the release of ScyllaDB Monitoring Stack 4.10.0 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB based on Prometheus and Grafana. ScyllaDB Monitoring Stack 4.10.0 supports: This release includes multiple updates to the overview, detailed, alternator, and advanced dashboards, including in CPU Utilization, Disk utilization, and others. Version updates for ScyllaDB Monitoring Stack 4.10.0 ScyllaDB’s advanced use of Service Level for query processing and background operations can cause confusion when observing CPU consumption. When monitoring ScyllaDB, high CPU usage is not necessarily an indication of system overload. To clarify this, the overview dashboard now displays only the query-related priority group consumption, split by priority group. In non-heterogeneous clusters, mixing different instance size nodes, each with a different number of cores, storage volume and disk usage is harder to interpret when viewed in bytes. Instead, the graph now shows percentile usage, displaying the average percentage of disk space used. The descriptions for tombstones in SSTables graphs have been clarified. The relative numbers represent the number of tombstones found in an SSTable and are updated after flush, compaction, or streaming operations. This clarification helps users understand when to expect updates to these values. When viewing the compressed bytes sent graph, it’s useful to track the aggregated total number of bytes. A new total graph within the panel makes this easier to follow. When examining tablet balance across the system, the total number of tablets per node can be misleading in non-heterogeneous clusters. Instead, the tablet load balancer’s load metric should be used. This metric defines load in proportion to each node’s capacity. In a balanced non-heterogeneous cluster, the load balancer load metric will be equal across nodes, even if the number of tablets is not. The new RPC delay graph in the RPC section shows the total round-trip time of an RPC message between the verb caller and the server. The querier cache stores queries paused due to paging and resumes them later, reducing query startup costs. If it misbehaves due to overload or bugs, performance can degrade. The new querier cache section displays population, lookup rate, and miss rate. Alternator relies on the HTTP protocol. There is now an HTTP section in the Alternator dashboard, currently showing open connections and new connections. This helps identify situations where there are too few or too many connections. Non-token-aware queries result in performance loss. When viewing the non-token-aware graph, it’s helpful to distinguish whether the source is reads or writes. The panel now displays two separate graphs: one for reads and one for writes. Large partitions can lead to performance degradation. To address this, ScyllaDB collects information about large partitions. The updated panel now includes additional columns: dead rows and range tombstones. TCP retransmission segments may indicate a network problem. A new graph now shows the rate of TCP retransmissions. The commit log information was previously always aggregated using averages. This caused confusion, and in some cases, it was necessary to aggregate using other methods, such as sum. It now uses the same aggregation functions available from the drop-down menu as the rest of the graphs. --- ### Page: https://forum.scylladb.com/t/release-scylladb-2025-1-3/4881 Title: [RELEASE] ScyllaDB 2025.1.3 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB 2025.1.3, a bug-fix production-ready patch release for ScyllaDB 2025.1 LTS Release. Related Links Get ScyllaDB 2025.1 Upgrade from ScyllaDB Enterprise 2024.x to ScyllaDB 2025.1 Upg… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-2025-1-3/4881 ## Headings Structure: H1: [RELEASE] ScyllaDB 2025.1.3 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 2025.1.3 H3: Related topics The ScyllaDB team announces ScyllaDB 2025.1.3, a bug-fix production-ready patch release for ScyllaDB 2025.1 LTS Release. The following issues are fixed in this release: --- ### Page: https://forum.scylladb.com/t/abnormal-clusters-node-s-behaviour-high-cpu-usage-on-4-5-nodes/4885 Title: Abnormal cluster's node(s) behaviour. High CPU usage on 4/5 nodes - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details ScyllaDB version: 6.2.2 Cluster size: 5 x 8vCPU/32GB RAM OS (RHEL/CentOS/Ubuntu/AWS AMI): ubuntu-jammy-22.04-amd64-server-20250327 Hi, we running ScyllaDB in k8s on dedicated EC2 instances with 5 n… Language: en Canonical URL: https://forum.scylladb.com/t/abnormal-clusters-node-s-behaviour-high-cpu-usage-on-4-5-nodes/4885 ## Headings Structure: H1: Abnormal cluster's node(s) behaviour. High CPU usage on 4/5 nodes H3: ScyllaDB version: 6.2.2 H3: Cluster size: 5 x 8vCPU/32GB RAM H3: OS (RHEL/CentOS/Ubuntu/AWS AMI): ubuntu-jammy-22.04-amd64-server-20250327 H3: Related topics ## Main Content: H1: Abnormal cluster's node(s) behaviour. High CPU usage on 4/5 nodes H3: ScyllaDB version: 6.2.2 H3: Cluster size: 5 x 8vCPU/32GB RAM H3: OS (RHEL/CentOS/Ubuntu/AWS AMI): ubuntu-jammy-22.04-amd64-server-20250327 H3: Related topics Hi, we running ScyllaDB in k8s on dedicated EC2 instances with 5 nodes. On each node exists only scyllaDB-related pod and k8s service pods. After one month of usage increased CPU usage for only 4 nodes of 5 were detected, when 5th node use CPU for ~20-40%; while all others are in ~70% Is there a way to see what tasks ScyllaDB perform in the background? - I investigated via ‘nodetool tasks’ but did not found anything suspicious Datacenter: eu-central-1 ======================== Status=Up/Down |/ State=Normal/Leaving/Joining/Moving – Address Load Tokens Owns Host ID Rack UN 100.65.105.82 500.14 GB 256 ? 99c64a19-f0eb-46c3-80c4-e290f7a5fd3e eu-central-1 UN 100.65.61.129 544.89 GB 256 ? 51af2d90-587c-4951-819d-309c4ed1268e eu-central-1 UN 100.66.85.110 531.18 GB 256 ? efe0d085-9a42-4fd0-abda-fa7bc3532752 eu-central-1 UN 100.70.9.99 558.64 GB 256 ? 9cc55415-d641-450d-8b3e-d1c59c9eb047 eu-central-1 UN 100.71.73.244 544.99 GB 256 ? c6a865b9-7030-4012-8c30-fb0addb22a0e eu-central-1 node 100.115.164.196 - is a node with small CPU usage ScyllaDB monitoring stack dashboards Load - comparation node where CPU ~20%(100.) vs all others. On last image graph appear not for all period because we did restart the node: k8s Node exporter metrics (random node vs node with small CPU usage): I want to believe that one node with small CPU consumption is NORMAL one and all others are stuck somewhere in the high intensive CPU task. Will provide any extra details if needed. k8s Node exporter metrics - ScyllaDB node with small CPU usage: ~20%(100.) = ~20%(100.115.164.196) its a note to first screenshot on post We are facing similar sort of issue, can someone guide how to fix this @mflendrich take a look --- ### Page: https://forum.scylladb.com/t/how-can-i-change-a-tables-name-in-cql/4886 Title: How can I change a table's name in CQL? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/how-can-i-change-a-tables-name-in-cql/4886 ## Headings Structure: H1: How can I change a table's name in CQL? H3: Related topics ## Main Content: H1: How can I change a table's name in CQL? H3: Related topics Originally from the User Slack @Mayank_Joshi: Hi. How do I rename a table in scylladb? @Botond_Dénes: CQL doesn’t support table renaming. You have to create a new table and migrate the data over. @Mayank_Joshi: Oh. Thanks a lot --- ### Page: https://forum.scylladb.com/t/support-for-having-multiple-queries-that-scan-the-same-data-share-a-scan-scan-sharing-or-synchronized-scans/4887 Title: Support for having multiple queries that scan the same data share a scan (Scan sharing or Synchronized scans) - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/support-for-having-multiple-queries-that-scan-the-same-data-share-a-scan-scan-sharing-or-synchronized-scans/4887 ## Headings Structure: H1: Support for having multiple queries that scan the same data share a scan (Scan sharing or Synchronized scans) H3: Related topics ## Main Content: H1: Support for having multiple queries that scan the same data share a scan (Scan sharing or Synchronized scans) H3: Related topics Originally from the User Slack @Em: Does Scylla support ‘scan sharing’ or ‘synchronized scans’, where multiple queries that scan over the same data can share a scan? @Botond_Dénes: No. You can achieve this indirectly, to some degree, by letting range scans use the cache. But there are caveats here: if your dataset is larger than cache, this won’t work very well. Also data populated into the cache by range scans can evict data populated by single partition queries, which are usually more latency sensitive. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-281-2025-06-09/4888 Title: Last week in scylladb.git master (issue #281; 2025-06-09) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 7d562c24b1..2d716f3ffe range are covered. There were 54 non-merge commits from 16 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-281-2025-06-09/4888 ## Headings Structure: H1: Last week in scylladb.git master (issue #281; 2025-06-09) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #281; 2025-06-09) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 7d562c24b1..2d716f3ffe range are covered. There were 54 non-merge commits from 16 authors in that period. Some notable commits: Alternator, ScyllaDB’s implementation of the DynamoDB API, now supports time-to-live (TTL) attributes on tables running with tablets. Alternator, ScyllaDB’s implementation of the DynamoDB API, now hides tags intended for internal use from the public API. Alternator now supports a set of per-table metrics. The mutation_fragment data structure, used to represent transient rows flowing through memory, had its memory footprint optimized. This reduces the probability of stalls related to memory allocation. A race condition with auto-parallel aggregation queries that involve user-defined functions for reduction was fixed. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/is-it-good-to-have-a-single-separate-physical-ssd-disk-not-raid-0-for-commitlog/4889 Title: Is it good to have a single separate physical SSD disk (not raid-0) for commitlog? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/is-it-good-to-have-a-single-separate-physical-ssd-disk-not-raid-0-for-commitlog/4889 ## Headings Structure: H1: Is it good to have a single separate physical SSD disk (not raid-0) for commitlog? H3: Related topics ## Main Content: H1: Is it good to have a single separate physical SSD disk (not raid-0) for commitlog? H3: Related topics Originally from the User Slack @Mahdi_Kamali: Is it good to have a single separate physical SSD disk (not raid-0) for commitlog? Does it give better performance? (Assume we have 6 same SSD disks with raid-0 for /var/lib/scylla directory) @avi: It’s better to share the data and commitlog disks so that the available bandwidth can go where it’s most needed. --- ### Page: https://forum.scylladb.com/t/issues-with-repair-not-functioning-properly-in-scylladb-version-5-2-6/4891 Title: Issues with repair not functioning properly in ScyllaDB version 5.2.6 - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 5.2.6 #Cluster size: 1 DC/ 9 Nodes os (RHEL/CentOS/Ubuntu/AWS AMI):AWS AMI Environment Details Issue: nodetool repair -pr appears to be stuck during system_traces keyspace re… Language: en Canonical URL: https://forum.scylladb.com/t/issues-with-repair-not-functioning-properly-in-scylladb-version-5-2-6/4891 ## Headings Structure: H1: Issues with repair not functioning properly in ScyllaDB version 5.2.6 H2: Environment Details H2: Current Situation H2: Log Analysis H2: Questions H2: Additional Context H2: Request H3: Related topics ## Main Content: H1: Issues with repair not functioning properly in ScyllaDB version 5.2.6 H2: Environment Details H2: Current Situation H2: Log Analysis H2: Questions H2: Additional Context H2: Request H3: Related topics Installation details #ScyllaDB version: 5.2.6 #Cluster size: 1 DC/ 9 Nodes os (RHEL/CentOS/Ubuntu/AWS AMI):AWS AMI I executed nodetool repair -pr and the repair process was progressing normally through the system_traces keyspace. However, the repair seems to have gotten stuck and is no longer generating any log entries. From the repair logs, I can see: After this timestamp, no further repair progress logs are being generated. How can I identify what’s happening with shard 9? What are the safe approaches to resolve this stuck repair? I’m looking for guidance on: Any insights or similar experiences would be greatly appreciated! Has anyone encountered similar shard-specific repair hanging issues? What was your resolution approach? ScyllaDB version: 5.2.6 It looks like you are running a very old ScyllaDB version, I’d suggest upgrading to a more recent version and see if the issues is resolved. In addition, I’d suggest to you use Scylla Manager for repairs and backups. I solved this problem. I could do repair after restarting node each other. I’m not sure how they solved it. --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-2-10/4894 Title: [RELEASE] ScyllaDB Enterprise 2024.2.10 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2024.2.10, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2024.2 Feature (Short Term Support) Release. Note there is a later LTS (Lon… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-2-10/4894 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2024.2.10 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2024.2.10 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2024.2.10, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2024.2 Feature (Short Term Support) Release. Note there is a later LTS (Long term Support) Release 2025.1, and you are encouraged to upgrade to it. Moving an existing cluster to TLS without downtime This release introduces a new Transitional state for node-node encryption. The transitional state encrypts all outgoing traffic but allows non-encrypted incoming traffic, allowing an upgrade from non-encrypted to encrypted node-to-node traffic without downtime. The following issues are fixed in this release (with an open-source reference, if available): Share IO queues between mountpoints seastar#2733 scylladb#23820 Currently, Seastar assigns IO queues based on mountpoints listed in io-properties.yaml, mapping each mountpoint’s device number to an IO queue. However, this approach fails when a single physical disk is split into multiple virtual block devices with separate mountpoints, each gets its own IO queue, which is incorrect since they share the same underlying disk. This PR enhances the configuration format to allow a single disk entry to specify multiple mountpoints. Seastar will then create one IO queue for all specified mountpoints, correctly reflecting the shared physical disk. --- ### Page: https://forum.scylladb.com/t/adding-nodes-to-a-cluster-getting-errors-tags-and-versions/4896 Title: Adding nodes to a cluster, getting errors, tags and versions - Kubernetes Operator - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/adding-nodes-to-a-cluster-getting-errors-tags-and-versions/4896 ## Headings Structure: H1: Adding nodes to a cluster, getting errors, tags and versions H3: Related topics ## Main Content: H1: Adding nodes to a cluster, getting errors, tags and versions H3: Related topics Originally from the User Slack @Terence_Liu: How to expand the cluster from 3 members to, say 4 with the operator? I tried to update members: and restarted, but the new node could not join the existing cluster due to missing certain files I’m assuming upon initial cluster setup, proper number of hosts are prepopulated with the right files to form a cluster. I was expecting adjusting that number would automatically set up the new joiners. But apparently it didn’t happen. I tried to refer to this page https://operator.docs.scylladb.com/stable/resources/scyllaclusters/nodeoperations/replace-node.html but the tagging & restarting didn’t do anything. Replacing a Scylla node | ScyllaDB Docs @Maciej_Zimnoch: I guess you’re using latest as your Operator version, and between when you initially installed it, and when new node was added new version was released. Is it true? @Terence_Liu: it was indeed latest as of 11/14/2024 I upgrade the existing nodes from 6.2.1 to 6.2.3 before introducing the new node @Maciej_Zimnoch: That’s why you should never use rolling tags. Because they resolve to different versions at different time. Make sure to use stable one and use digest format. > I upgrade the existing nodes from 6.2.1 to 6.2.3 before introducing the new node I’m talking about Operator version, not Scylla version. @Terence_Liu: In my case, I don’t think scylla operator was updated or new resources under scylla-operator was created though? So I’m essentially still looking at whatever latest version on Nov 14 @Maciej_Zimnoch: Operator runs sidecar alongside Scylla and image used by it is the same as Operator image. Because new node was added, latest resolved to different version that what Operator deployment is using. These two must be in sync. @Terence_Liu: oooh So… if I were to make this work, I should pin scylla operator latest explicitly to that Nov 14 version, and recreate the new node? That way, new node will take the same operator sidecar version @Maciej_Zimnoch: yes make sure to specify proper version in two places: https://github.com/scylladb/scylla-operator/blob/61ddd3c444712fc74a2e29286ff4564e2cce0dcd/deploy/operator/50_operator.deployment.yaml#L25-L30 GitHub: scylla-operator/deploy/operator/50_operator.deployment.yaml at 61ddd3c444712fc74a2e29286ff4564e2cce0dcd · scylladb/scylla-operator both container image and env var as the env var controls sidecar image @Terence_Liu: Got it. Thank you for the tip - will report back. How do people update scylla operator though? @Maciej_Zimnoch: https://operator.docs.scylladb.com/stable/installation/overview.html#installation-modes Overview | ScyllaDB Docs @Terence_Liu: oh, awesome - N+1, very helpful! @Bradley_Stock: Hi, so I’ve updated the operator to pin the version (to get the operator deployed, I had to use helm template and copy the manifests to git in order to make manual modifications to support istio, and those were set to latest by default) from latest to 1.14. The operator itself updated and seems to have run fine, but my ScyllaCluster object seems to be a bit de-sync’d from the operator. I’m seeing errors such as: Looks like the operator I tried to go to was lower, we found that the ones with latest were running 1.16-alpha.0, so I bumped to 1.16 and the operator is communicating properly with the scyllacluster object --- ### Page: https://forum.scylladb.com/t/backup-failing-up-giving-indexing-file-error/4898 Title: Backup failing up giving indexing file error - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details Running a 3 node scylla cluster in GKE. #ScyllaDB version: 6.2.0 #Cluster size: 3 os (RHEL/CentOS/Ubuntu/AWS AMI): Ubuntu Hello guys, Using scylla manager for scheduled backup. After some succ… Language: en Canonical URL: https://forum.scylladb.com/t/backup-failing-up-giving-indexing-file-error/4898 ## Headings Structure: H1: Backup failing up giving indexing file error H3: Related topics ## Main Content: H1: Backup failing up giving indexing file error H3: Related topics Installation details Running a 3 node scylla cluster in GKE. #ScyllaDB version: 6.2.0 #Cluster size: 3 os (RHEL/CentOS/Ubuntu/AWS AMI): Ubuntu Hello guys, Using scylla manager for scheduled backup. After some successful backup iterations, it gets failed showing indexing file error Looking at your error, the backup is failing because Scylla Manager can’t find any files to index on node 10.240.5.5. Since this worked before and now suddenly fails, here’s what I’d check: First, verify if there’s actually data to backup: Check what happened to your data: Verify backup configuration: This is the logs of scylla-manger {“L”:“ERROR”, “M”:“Indexing snapshot files failed on host”, “N”:“backup.index”, “S”:“github/scylladb/go-log.Logger.log github/scylladb/go-log@v0.0.7/logger.go:101 github/scylladb/go-log.Logger.Error github/scylladb/go-log@v0.0.7/logger.go:84 github/scylladb/scylla-manager/v3/pkg/service/backup.(*worker).Index.func2 githubcom/scylladb/scylla-manager/v3/pkg/service/backup/worker_index.go:33 githubcom/scylladb/scylla-manager/v3/pkg/service/backup.hostsInParallel.func2 githubcom/scylladb/scylla-manager/v3/pkg/service/backup/parallel.go:89 githubcom/scylladb/scylla-manager/v3/pkg/util/parallel.Run.func1 githubcom/scylladb/scylla-manager/v3/pkg/util@v0.0.0-20241104134613-aba35605c28b/parallel/parallel.go:79”, “T”:“2025-06-22T06:00:08.273Z”, “_trace_id”:“UORGpSjeQWKBAeP43xVHKw”, “error”:“10.244.11.91: could not find any files”, “errorStack”:"githubcom/scylladb/scylla-manager/v3/pkg/service/backup.hostsInParallel.func1 githubcom/scylladb/scylla-manager/v3/pkg/service/backup/parallel.go:86 githubcom/scylladb/scylla-manager/v3/pkg/util/parallel.Run.func1 githubcom/scylladb/scylla-manager/v3/pkg/util@v0.0.0-20241104134613-aba35605c28b/parallel/parallel.go:72 runtime.goexit runtime/asm_amd64.s:1700 ", “host”:“10.244.11.91”} Backup’s coming properly after reschedule until this disaster comes. --- ### Page: https://forum.scylladb.com/t/how-to-use-scylla-datasource-in-scylladb-monitoring/4899 Title: How to use scylla-datasource in scyllaDB Monitoring - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 6.2.3 #Cluster size: 4 os (RHEL/CentOS/Ubuntu/AWS AMI): RHEL9.2 Hello. I’m installing scylladb monitoring with docker on a server and I modified the file grafana/datasource.s… Language: en Canonical URL: https://forum.scylladb.com/t/how-to-use-scylla-datasource-in-scylladb-monitoring/4899 ## Headings Structure: H1: How to use scylla-datasource in scyllaDB Monitoring H3: Related topics ## Main Content: H1: How to use scylla-datasource in scyllaDB Monitoring H3: Related topics Installation details #ScyllaDB version: 6.2.3 #Cluster size: 4 os (RHEL/CentOS/Ubuntu/AWS AMI): RHEL9.2 I’m installing scylladb monitoring with docker on a server and I modified the file grafana/datasource.scylla.yml with a user/password to benefit from the scylla-datasource in grafana. But now I have done that, how can I use it, please ? I don’t find any documentation about that. I tried to use the datasource scylla-datasource in grafana web ui, I tried to explore the datasource but all the queries I performed returns No Data. For example : SELECT keyspace_name, table_name FROM system_schema.tables; The user I use has SELECT permissions on everything in ScyllaDB. The plugin scylla-datasource appears correctly in the list of plugins. Thank you in advance for your help. @Amnon_Heiman, any idea? I suggest that you’ll start with checking ScyllaDB’s own table. In the CQL dashboard open the CQL System tables section: This is how it should look like: If your configuration is correct, you should see values. What can happen, is that if the plugin was never connected to a cluster and no cluster ip was giving to it (not a must) it doesn’t know the cluster ips. From the explorer you can add an ip: Once the plugin find one node, it should find the rest. Don’t be surprise is that once the cql dashboard show you the values, the rest would work, that’s because one it learn about the cluster it will be able to connect to it --- ### Page: https://forum.scylladb.com/t/operator-sidecar-issue-getting-error-syncing-key-failed-cant-sync-the-hostid-annotation/4900 Title: Operator sidecar issue, getting error: syncing key failed: can't sync the HostID annotation - Kubernetes Operator - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/operator-sidecar-issue-getting-error-syncing-key-failed-cant-sync-the-hostid-annotation/4900 ## Headings Structure: H1: Operator sidecar issue, getting error: syncing key failed: can't sync the HostID annotation H3: Related topics ## Main Content: H1: Operator sidecar issue, getting error: syncing key failed: can't sync the HostID annotation H3: Related topics Originally from the User Slack @Cong_Guo: Hi community, does anyone know what below error mean? also see below log in scylla-operator @Maciej_Zimnoch: it means operator sidecar cannot get host id from Scylla due to bad_alloc exception raised inside Scylla. Looks like you’re having memory issues. status.conditions[15].reason: Too long: may not be longer than 1024, these kind of errors will be fixed in next Operator release. It cannot update ScyllaCluster status due to this validation error. @Cong_Guo: > It cannot update ScyllaCluster status due to this validation error. Does that mean it won’t impact the functionality of scylla-operator, just status couldn’t be updated? May I know if there is a issue opened for it? @Maciej_Zimnoch: yes, just status isn’t updated, main resource is reconciled. there’s no issue, just PR fixing other issue with Condition.reason, but it reduces length of it as well: https://github.com/scylladb/scylla-operator/pull/2581 GitHub: Remove Type prefix for aggregated Condition Reasons by zimnx · Pull Request #2581 · scylladb/scylla-operator --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-18/4903 Title: [RELEASE] ScyllaDB Enterprise 2024.1.18 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2024.1.18, a bug-fix production-ready patch release for ScyllaDB Enterprise 2024.1 Long Term Support (LTS) Release. Note there is a later, better, LTS (Long term Support) R… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-18/4903 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2024.1.18 H3: Related Links H2: Fixed Issue with an open-source reference: H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2024.1.18 H3: Related Links H2: Fixed Issue with an open-source reference: H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2024.1.18, a bug-fix production-ready patch release for ScyllaDB Enterprise 2024.1 Long Term Support (LTS) Release. Note there is a later, better, LTS (Long term Support) Release 2025.1, and you are encouraged to upgrade to it. Moving an existing cluster to TLS without downtime This release introduces a new Transitional state for node-node encryption. The transitional state encrypts all outgoing traffic but allows non-encrypted incoming traffic, allowing an upgrade from non-encrypted to encrypted node-to-node traffic without downtime. Share IO queues between mountpoints seastar#2733 scylladb#23820 Currently, Seastar assigns IO queues based on mountpoints listed in io-properties.yaml, mapping each mountpoint’s device number to an IO queue. However, this approach fails when a single physical disk is split into multiple virtual block devices with separate mountpoints, each gets its own IO queue, which is incorrect since they share the same underlying disk. This PR enhances the configuration format to allow a single disk entry to specify multiple mountpoints. Seastar will then create one IO queue for all specified mountpoints, correctly reflecting the shared physical disk. compaction_manager::perform_task_on_all_files always executes a task, even if there are no sstables to compact #16803 chunked_managed_vector violates the preferred max contiguous allocation size #23854 clone semantics is incorrect for sstable runs in partitioned_sstable_set #17878 ScyllaDB enables integration with external Key Management Interoperability Protocol (KMIP) servers for managing encryption keys, supporting Encryption at Rest. Encryption at Rest: KMIP LOCATE operation request has incorrect attribute names #23970 As a result the KMIP server rejected the request, leading ScyllaDB to assume that a key with these specifications doesn’t exist, and creates a new key in the KMIP server. --- ### Page: https://forum.scylladb.com/t/getting-address-already-in-use-error-number-of-client-side-connections/4905 Title: Getting address already in use error, number of client side connections - Kubernetes Operator - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/getting-address-already-in-use-error-number-of-client-side-connections/4905 ## Headings Structure: H1: Getting address already in use error, number of client side connections H3: Related topics ## Main Content: H1: Getting address already in use error, number of client side connections H3: Related topics Originally from the User Slack @Terence_Liu: I’m having a lot of trouble getting our istio service mesh to live together with the scylla operator and scylla cluster. Is there a guide that help set this up right? @Terence_Liu: We’ve excluded a few ports on the operator and the ScyllaCluster, notably 7000 (node-to-node communication) and 8080 (from the scylladb-api-status-probe), the service is up, and client can read from it. But when writing becomes anywhere around 500~1000 reqs/s, the scylladb-api-status-probe starts to report errors, and the istio sidecar starts to throw out many messages like This forces the scylla livenessProbe to fail after 12 tries, and restarts scylla. My 3-node cluster goes into a rotating restart wave as a result. When no writes happen, some reads seem totally fine. This is the pod level manifest. You can see from ISTIO_KUBE_APP_PROBERS our istio remaps these health endpoints (notably on 8080) to We figured it out - we opened too many CQL connections from the client side, ~3000 to a 3-node Scylla cluster. After packing and trimming there, the cluster became stable to write to. The excessive connections were not only a strain on the cluster, but more importantly our istio service mesh. --- ### Page: https://forum.scylladb.com/t/scylladb-sink-connecter-error-in-debezium/4908 Title: Scylladb-sink-connecter error in debezium - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: it is possible to use scylldb-sink-connecter in debezium or it will only work in kafka-connect ? at java.base/java.lang.ClassLoader.defineClass(ClassLoader.java:1027) at java.base/java.security.SecureClassLoader.defi… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-sink-connecter-error-in-debezium/4908 ## Headings Structure: H1: Scylladb-sink-connecter error in debezium H3: Related topics ## Main Content: H1: Scylladb-sink-connecter error in debezium H3: Related topics it is possible to use scylldb-sink-connecter in debezium or it will only work in kafka-connect ? The ScyllaDB sink connector is not designed to run inside Debezium; it is meant to run as a Kafka Connect sink connector. --- ### Page: https://forum.scylladb.com/t/query-hangs-advice-on-timeout-retry-strategy/4910 Title: Query Hangs – Advice on Timeout/Retry Strategy - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi everyone, I’m running into an issue with my Rust-based system using ScyllaDB and would appreciate some guidance from more experienced folks in this area. System Overview: I’ve built a ScyllaClient abstraction in Rus… Language: en Canonical URL: https://forum.scylladb.com/t/query-hangs-advice-on-timeout-retry-strategy/4910 ## Headings Structure: H1: Query Hangs – Advice on Timeout/Retry Strategy H3: System Overview: H3: Problem: H3: Questions: H3: Questions: H3: Related topics ## Main Content: H1: Query Hangs – Advice on Timeout/Retry Strategy H3: System Overview: H3: Problem: H3: Questions: H3: Questions: H3: Related topics I’m running into an issue with my Rust-based system using ScyllaDB and would appreciate some guidance from more experienced folks in this area. I’ve built a ScyllaClient abstraction in Rust that: The system includes two main binaries: Both programs use the same shared Scylla client implementation. The data_manager process occasionally hangs during a query execution. There’s no panic or error — it just stalls silently. Logging shows the last event before hanging is a Scylla query attempt. I suspect this might be due to resource constraints on my local development setup (Dockerized Scylla cluster with 2 nodes, each using --smp 1, --memory 750M, and --developer-mode 1). However, I’m surprised that a query could hang indefinitely with no error. I expected either: I realize this might be a side-effect of my limited dev environment, but I’d like to understand best practices around timeouts and retry strategies when using Scylla in high-throughput systems. It absolutely does not do anything near the ~25 million simulations. It already hangs far below ~1000 simulations. Definitely not. There are two kinds of timeouts involved in statement execution: server-side timeout and client-side timeout. While the driver may have no client-side timeout set (though the default is 30 seconds), the server-side timeout is always set (5 seconds by default) and should make a request that runs for too long return a {Read,Write}Timeout error. This is thus unexpected that your application hangs indefinitely. Both retry and timeout mechanisms are already there (backoff is not yet there, but planned for the future): The answer is simple. Every Session has a significant overhead: For this reason, it’s recommended to use a single Session. For customising various workloads, use multiple ExecutionProfiles in one Session. Please let me know if the above helps, and don’t hesitate to share your findings on what caused the hanging. --- ### Page: https://forum.scylladb.com/t/last-two-weeks-in-scylla-cluster-tests-git-master-issue-96-2025-06-13/4913 Title: Last two weeks in scylla-cluster-tests.git master (issue #96; 2025-06-13) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 8e238298…e88b7dfc range are covered. There were 81 non-merge commits from 18 authors in t… Language: en Canonical URL: https://forum.scylladb.com/t/last-two-weeks-in-scylla-cluster-tests-git-master-issue-96-2025-06-13/4913 ## Headings Structure: H1: Last two weeks in scylla-cluster-tests.git master (issue #96; 2025-06-13) H3: Related topics ## Main Content: H1: Last two weeks in scylla-cluster-tests.git master (issue #96; 2025-06-13) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 8e238298…e88b7dfc range are covered. There were 81 non-merge commits from 18 authors in that period. Some notable commits: New perf test based on PerformanceRegressionPredefinedStepsTest verifies that lz4 and zstd dict compression don’t cause unwanted latency regressions. Monitoring was updated to 4.10. ycsb stress tool was updated to the latest version, fixing multithreading issues with HDR output and using a newer Java in Dockerfile. scylla-bench was updated to 0.2.4, bringing package and driver updates, stability fixes, and version info collection. “Simulated racks” are now globally enabled, except for multidc scenarios. Schema info is archived separately for easier and faster access. gemini version is now reported to Argus, along with gocql version. Client encryption support was added to cql-stress. This tool is now used in several sanity, tier1, and upgrade tests. Disruption list shuffling for SisyphusMonkey was fixed, so expect changes in disruption order across some tests. With latte 0.31.0, rack awareness became possible and is now supported in SCT for both single- and multi-dc deployments. To align with Scylla Cloud and improve logging, support for vector.dev was added and enabled for tier1 tests for evaluation. In next steps we will work on optimizing it, compression, formatting and sending to centralized logs service that can be searched. Dependencies are now managed via pyproject.toml, improving consistency with other tools. See updated dev setup docs. SCT now supports AWS dedicated hosts, which can be used to reduce interference from noisy neighbors. hydra list-images supports multiple regions and uses cloud-specific defaults. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/scaling-while-using-cdc-to-capture-events-from-scylladb-to-kafka-redpanda-rust-driver/4914 Title: Scaling while using CDC to capture events from ScyllaDB to Kafka (Redpanda), Rust driver - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/scaling-while-using-cdc-to-capture-events-from-scylladb-to-kafka-redpanda-rust-driver/4914 ## Headings Structure: H1: Scaling while using CDC to capture events from ScyllaDB to Kafka (Redpanda), Rust driver H3: Related topics ## Main Content: H1: Scaling while using CDC to capture events from ScyllaDB to Kafka (Redpanda), Rust driver H3: Related topics Originally from the User Slack @Sh4d1: Hello! I was wondering if there is a way to horizontally scale the CDC driver (rust one in my case) ? @Krasimir_Popov: What exactly is your use case? Are you trying to capture events from scylla to do actions in external system or something else? @Sh4d1: Yes we’re currently capturing all inserts to a redpanda (Kafka) cluster And I’m wondering if it’s possible to horizontally scale the CDC driver @Krasimir_Popov: Theoretically you can, but I never tried such a thing. You need to take into consideration something like Token Range Sharding. ScyllaDB’s CDC data is partitioned by token ranges. Distribute these ranges across driver instances to parallelize ingestion. Each instance of your event consumer (rust microservice pod)will follow a different range. Maybe start small and try to split the range by 2 statically hardcoded and if it works you can keep hacking to make it capable to scale up and down exchanging who(which instance of the rust pod) is following what range via etcd, redis, consul or special scylladb table @Sh4d1: Yeah that was kind of what I had in mind! Thanks @Krasimir_Popov: Let me know if the hardcoded approach works, this is actually a good idea for an open source project to easy grab some github stars I am not very good with rust but I can definitely make something on top of the go driver @Sh4d1: Won’t have a lot of time to in the near future, but I’ll definitely need this at some point I’ll try to find some time to hack something up on top of the rust driver (if I manage ) --- ### Page: https://forum.scylladb.com/t/rate-limiting-dropped-498-similar-messages-waiting-for-2-live-nodes-to-show-up-in-gossip-currently-1-present/4917 Title: (rate limiting dropped 498 similar messages) Waiting for 2 live nodes to show up in gossip, currently 1 present - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 5.4.9 #Cluster size: 3 os (RHEL/CentOS/Ubuntu/AWS AMI): AWS AMI I have a cluster having 3 nodes. Out of them 2 are seed node. I am trying to add 1 more node. Going forward thi… Language: en Canonical URL: https://forum.scylladb.com/t/rate-limiting-dropped-498-similar-messages-waiting-for-2-live-nodes-to-show-up-in-gossip-currently-1-present/4917 ## Headings Structure: H1: (rate limiting dropped 498 similar messages) Waiting for 2 live nodes to show up in gossip, currently 1 present H3: Related topics ## Main Content: H1: (rate limiting dropped 498 similar messages) Waiting for 2 live nodes to show up in gossip, currently 1 present H3: Related topics Installation details #ScyllaDB version: 5.4.9 #Cluster size: 3 os (RHEL/CentOS/Ubuntu/AWS AMI): AWS AMI I have a cluster having 3 nodes. Out of them 2 are seed node. I am trying to add 1 more node. Going forward this node is going to be the seed node. We would obsolete existing 3 nodes. What would be the correct strategy for this. Currently I am trying to add 1 new node to existing 2 seed node by giving seed as xx.5.2.1,xx.5.1.1 I am getting an error shard 0:main] init - Startup failed: std::runtime_error (Timed out waiting for 2 live nodes to show up in gossip) I have checked 7000,9042 ports. Its connecting from the new node. Can anybody help in this? I would revalidate the configuration on the new node - on an existing node - check the nodes health, using nodetool status; all nodes should show “UN” As for the replacement strategy itself - --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-2-11/4919 Title: [RELEASE] ScyllaDB Enterprise 2024.2.11 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2024.2.11, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2024.2 Feature (Short Term Support) Release. Note there is a later LTS (Lon… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-2-11/4919 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2024.2.11 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2024.2.11 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2024.2.11, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2024.2 Feature (Short Term Support) Release. Note there is a later LTS (Long term Support) Release 2025.1, and you are encouraged to upgrade to it. The following issues are fixed in this release (with an open-source reference, if available): ScyllaDB enables integration with external Key Management Interoperability Protocol (KMIP) servers for managing encryption keys, supporting Encryption at Rest. --- ### Page: https://forum.scylladb.com/t/release-scylladb-cpp-over-rust-driver-0-5-0/4920 Title: [RELEASE] ScyllaDB CPP-over-Rust Driver 0.5.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce ScyllaDB CPP Rust Driver 0.5.0, an API-compatible rewrite of GitHub - scylladb/cpp-driver: Scylla C/C++ Driver as a wrapper for the Rust driver. It will fully replace the CPP driv… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-cpp-over-rust-driver-0-5-0/4920 ## Headings Structure: H1: [RELEASE] ScyllaDB CPP-over-Rust Driver 0.5.0 H2: Changes H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB CPP-over-Rust Driver 0.5.0 H2: Changes H3: Related topics The ScyllaDB team is pleased to announce ScyllaDB CPP Rust Driver 0.5.0, an API-compatible rewrite of GitHub - scylladb/cpp-driver: Scylla C/C++ Driver as a wrapper for the Rust driver. It will fully replace the CPP driver, which is already reaching its End of Life. The driver version shall be considered Beta. Some minor features still need to be included. See Limitations and Unimplemented functions from cassandra.h sections in README.md. The underlying Rust driver used version: 1.2.0. Implemented API functions: New features / enhancements CI / developer tool improvements: Congrats to all contributors and thanks everyone for using our driver! The source code of the driver can be found here: Thank you for your attention, please do not hesitate to contact us if you have any questions, issues, feature requests, or are simply interested in our driver! Contributors since the last release: --- ### Page: https://forum.scylladb.com/t/scylladb-university-live-july-30th/4921 Title: ScyllaDB University LIVE - July 30th - Announcements - ScyllaDB Community NoSQL Forum Meta Description: Our next training event, ScyllaDB University Live, will take place on July 30th. We’re going to be hosting two tracks, Essentials and Advanced, covering different topics to help you improve your ScyllaDB skills. The ev… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-university-live-july-30th/4921 ## Headings Structure: H1: ScyllaDB University LIVE - July 30th H3: Related topics ## Main Content: H1: ScyllaDB University LIVE - July 30th H3: Related topics Our next training event, ScyllaDB University Live, will take place on July 30th. We’re going to be hosting two tracks, Essentials and Advanced, covering different topics to help you improve your ScyllaDB skills. The event is online, live, and free. Save your spot here, hope to see you there! The event is happening in 10 days! Save your spot here . --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-97-2025-06-20/4922 Title: Last week in scylla-cluster-tests.git master (issue #97; 2025-06-20) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 35a2b964…5dd5e9a2 range are covered. There were 24 non-merge commits from 7 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-97-2025-06-20/4922 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #97; 2025-06-20) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #97; 2025-06-20) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 35a2b964…5dd5e9a2 range are covered. There were 24 non-merge commits from 7 authors in that period. Some notable commits: We reworked the way we generate the nemesis list file and individual nemesis jobs by introducing the create-nemesis-yaml Hydra command. AddDropColumnNemesis runs faster now by using DESC SCHEMA WITH INTERNALS to get table info instead of describing each table one by one. The hydra list-resources command now scans all regions by default, making it easier to spot forgotten or hanging resources in AWS, Azure, or GCP. perf-simple-query result validation can now be tuned per test configuration to be stricter or more relaxed, depending on test needs. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/last-fortnight-in-scylladb-git-master-issue-282-2025-06-22/4925 Title: Last fortnight in scylladb.git master (issue #282; 2025-06-22) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last two weeks. Commits in the 2d716f3ffe..a9a53d9178 range are covered. There were 105 non-merge commits from 22 authors in tha… Language: en Canonical URL: https://forum.scylladb.com/t/last-fortnight-in-scylladb-git-master-issue-282-2025-06-22/4925 ## Headings Structure: H1: Last fortnight in scylladb.git master (issue #282; 2025-06-22) H3: Related topics ## Main Content: H1: Last fortnight in scylladb.git master (issue #282; 2025-06-22) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last two weeks. Commits in the 2d716f3ffe..a9a53d9178 range are covered. There were 105 non-merge commits from 22 authors in that period. Some notable commits: HTTP TLS support now shares credentials across connections, reducing expensive TLS handshakes. The nodetool refresh command gained a –skip-reshape switch. Reshaping can be an expensive operation and one may want to defer it. The S3 driver support is now better at avoiding resource exhaustion due to too many open connections. C++ unit test execution now runs under pytest. There is now a framework for testing upgrades in the main repository. Previously, upgrades were tested in external tests. The first use is to test sstable dictionary compression during upgrades. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/io-tune-check-ignoring-smp-restrictions/4928 Title: Io tune check ignoring smp restrictions - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 2025.1.2-0.20250422.502c62d91d48 #Cluster size: 128 core os (RHEL/CentOS/Ubuntu/AWS AMI): Ubuntu 22.04.5 LTS Scylla doesn’t start because the iotune checks are failing. Deeper… Language: en Canonical URL: https://forum.scylladb.com/t/io-tune-check-ignoring-smp-restrictions/4928 ## Headings Structure: H1: Io tune check ignoring smp restrictions H3: Related topics ## Main Content: H1: Io tune check ignoring smp restrictions H3: Related topics Installation details #ScyllaDB version: 2025.1.2-0.20250422.502c62d91d48 #Cluster size: 128 core os (RHEL/CentOS/Ubuntu/AWS AMI): Ubuntu 22.04.5 LTS Scylla doesn’t start because the iotune checks are failing. Deeper inspection showed that they were running the check assuming all Cpu cores on the node was available for it. I have too many cores on the node and can’t dedicate all of them to scylla. I exec-ed into the scylla container and manually ran the iotune test but with the smp flag /usr/bin/iotune --format envfile --options-file /etc/scylla.d/io.conf --properties-file /etc/scylla.d/io_properties.yaml --evaluation-directory /var/lib/scylla/data --default-log-level debug --smp=1 This causes the scylla pod to start working. I have to repeat the above process for all pods in my statefulset. I have been able to replicate this many times in different clusters. Forgot to add I am installing this using helm v1.17 and passing smp restriction through ScyllaArgs I’d suspect you need to set --developer-mode 1 I think that is related to the validation tests not passing. I can’t be setting developer mode for production workloads ScyllaDB has firm requirements to bootstrap, the developer-mode will make them soft-requirements, meaning the service will load. then it’s ‘on your own risk’ running on a sub-optimal environment for production. --- ### Page: https://forum.scylladb.com/t/read-latency-issues-on-al2023/4929 Title: Read Latency issues on Al2023 - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 4.6.11 #Cluster size: 18 os (RHEL/CentOS/Ubuntu/AWS AMI): Al2023 We recently migrated the cluster from Al2 to Al2023. Post that, We started seeing the latency spikes for read… Language: en Canonical URL: https://forum.scylladb.com/t/read-latency-issues-on-al2023/4929 ## Headings Structure: H1: Read Latency issues on Al2023 H3: Related topics ## Main Content: H1: Read Latency issues on Al2023 H3: Related topics Installation details #ScyllaDB version: 4.6.11 #Cluster size: 18 os (RHEL/CentOS/Ubuntu/AWS AMI): Al2023 We recently migrated the cluster from Al2 to Al2023. Post that, We started seeing the latency spikes for read operations. Hi team, Did anyone notice this issue ? any suggestions in OS level tuning? You’re running a very old ScyllaDB version, that is no longer supported, please update to the latest version, and update here if your problem is solved. --- ### Page: https://forum.scylladb.com/t/release-scylla-operator-1-16-3/4932 Title: [RELEASE] Scylla Operator 1.16.3 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the patch release of Scylla Operator 1.16.3. This patch release include two bug fixes along with minor documentation updates. Fixed a bug that prevented setting custom ScyllaDB… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-operator-1-16-3/4932 ## Headings Structure: H1: [RELEASE] Scylla Operator 1.16.3 H3: Related topics ## Main Content: H1: [RELEASE] Scylla Operator 1.16.3 H3: Related topics The ScyllaDB team is pleased to announce the patch release of Scylla Operator 1.16.3. This patch release include two bug fixes along with minor documentation updates. Check full 1.16.3 changelog. --- ### Page: https://forum.scylladb.com/t/release-scylla-operator-1-17-1/4933 Title: [RELEASE] Scylla Operator 1.17.1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the patch release of Scylla Operator 1.17.1. This patch release include two bug fixes along dependency update. Fixed a bug that prevented setting custom ScyllaDB arguments that… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-operator-1-17-1/4933 ## Headings Structure: H1: [RELEASE] Scylla Operator 1.17.1 H3: Related topics ## Main Content: H1: [RELEASE] Scylla Operator 1.17.1 H3: Related topics The ScyllaDB team is pleased to announce the patch release of Scylla Operator 1.17.1. This patch release include two bug fixes along dependency update. Check full 1.17.1 changelog. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-98-2025-06-27/4935 Title: Last week in scylla-cluster-tests.git master (issue #98; 2025-06-27) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 5a597e2d…6485769c range are covered. There were 35 non-merge commits from 11 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-98-2025-06-27/4935 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #98; 2025-06-27) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #98; 2025-06-27) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 5a597e2d…6485769c range are covered. There were 35 non-merge commits from 11 authors in that period. Some notable commits: Pre-commit hook now auto-generates nemesis YAMLs and pipelines, reducing manual actions. actions.log will include steps from multiple nemeses, improving traceability during test investigations. Instead of raw S3 URLs, SCT now shows Argus-proxied links, making test artifacts private by default. An example rune script was added for “row count validation” in latte . It ensures the expected number of rows remain post-test, including verification of deletions. Introduced a test for Manager 1-1 restore. By default, OSS versions are no longer selected as base versions for upgrade tests. Manual branch selection remains available. New test pipelines support RHEL10 artifacts on AWS, for both ARM and x86. Durations of Scylla operations and test timeouts are now stored in Argus (visible in Results tab), instead of ElasticSearch. This can be disabled via adaptive_timeout_store_metrics: false in test config. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-283-2025-06-29/4937 Title: Last week in scylladb.git master (issue #283; 2025-06-29) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the a9a53d9178..9d70e7a067 range are covered. There were 81 non-merge commits from 27 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-283-2025-06-29/4937 ## Headings Structure: H1: Last week in scylladb.git master (issue #283; 2025-06-29) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #283; 2025-06-29) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the a9a53d9178..9d70e7a067 range are covered. There were 81 non-merge commits from 27 authors in that period. Some notable commits: ScyllaDB tracks the amount of memory in memtables that was spooled to an sstable and tries to ensure new writes don’t consume memory faster that it is written to disk. A crash due to a rare edge case when tracking this memory was fixed. The maximum length of keyspace, table, and view names was extended from 48 characters to 192 characters. The nodetool repair command now rejects tablet keyspaces; those are repaired via a different command. There is now a queue for topology requests. Requests that cannot be processed in parallel will be queued one after the other. Lightweight transactions now work correctly for tablets when a tablet is migrated to another shard on the same node. Note LWT for tablets isn’t enabled yet. UUID SSTable generations (introduced in ScyllaDB 5.4 / 2024.1) are now mandatory. ScyllaDB can automatically parallelize some aggregation queries via an internal map/reduce service. This automatic parallelization is now optimized for tablets. An edge case with DECIMAL type parsing was fixed. A regression caused the casasndra role to be recreated even if it was explicitly dropped. This is now fixed. A regression that could lead to possibly stale LWT reads was fixed. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/datastax-enterprise-to-scylladb-migration/4939 Title: Datastax enterprise to scylladb migration - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 3.08 os (RHEL/CentOS/Ubuntu/AWS AMI): RHEL can i migrate datastax enterprise (version 3.11) sstable to scylladb (3.08) using scylla sstable loader? Language: en Canonical URL: https://forum.scylladb.com/t/datastax-enterprise-to-scylladb-migration/4939 ## Headings Structure: H1: Datastax enterprise to scylladb migration H3: Related topics ## Main Content: H1: Datastax enterprise to scylladb migration H3: Related topics Installation details #ScyllaDB version: 3.08 os (RHEL/CentOS/Ubuntu/AWS AMI): RHEL can i migrate datastax enterprise (version 3.11) sstable to scylladb (3.08) using scylla sstable loader? There are a few ways to migrate to ScyllaDB. I’d start by reading the Migrator documentation. ScyllaDB 3.08 is years old, I suggest you use the latest version. --- ### Page: https://forum.scylladb.com/t/release-scylladb-2025-2-0/4943 Title: [RELEASE] ScyllaDB 2025.2.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of ScyllaDB 2025.2, a production-ready ScyllaDB Short Term Support (STS) Minor Feature Release. Upon the release of ScyllaDB 2025.2 STS, support for ScyllaDB 2024.2 h… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-2025-2-0/4943 ## Headings Structure: H1: [RELEASE] ScyllaDB 2025.2.0 H2: Relevant links H2: New features H3: Capacity Aware Tablets Load balancing H3: Storage ZSTD + dictionary compression H3: New Tablets Guardrail: enforce tablets mode for new keyspaces H3: New Topology guardrail: prevent Snitch update and DC and/or RACK name change H3: Cluster Level Repair for Tablets H3: Raft majority loss Recovery H3: Raft Voters H3: Security H3: Vector Type H3: Cassandra-stress no longer part of the default package H3: Native Backup Enhancement - Experimental H2: Additional Updates H3: CQL H3: Alternator H3: Correctness H3: Tablets H3: Deployment H3: Stability H3: Performance H3: Tooling H3: Security H3: Tracing H3: Monitoring H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 2025.2.0 H2: Relevant links H2: New features H3: Capacity Aware Tablets Load balancing H3: Storage ZSTD + dictionary compression H3: New Tablets Guardrail: enforce tablets mode for new keyspaces H3: New Topology guardrail: prevent Snitch update and DC and/or RACK name change H3: Cluster Level Repair for Tablets H3: Raft majority loss Recovery H3: Raft Voters H3: Security H3: Vector Type H3: Cassandra-stress no longer part of the default package H3: Native Backup Enhancement - Experimental H2: Additional Updates H3: CQL H3: Alternator H3: Correctness H3: Tablets H3: Deployment H3: Stability H3: Performance H3: Tooling H3: Security H3: Tracing H3: Monitoring H3: Related topics The ScyllaDB team is pleased to announce the release of ScyllaDB 2025.2, a production-ready ScyllaDB Short Term Support (STS) Minor Feature Release. Upon the release of ScyllaDB 2025.2 STS, support for ScyllaDB 2024.2 has officially ended. ScyllaDB Cloud users currently on the 2024.x release will be contacted to schedule an upgrade. ScyllaDB Enterprise users utilizing the 2024.x release should reach out to the support team for upgrade assistance. More information on ScyllaDB’s Long Term Support (LTS) policy is available here. The 2025.2 release adds new features like improved storage compression, Tablets rebalancing, and multiple additional improvements. ScyllaDB Enterprise customers are encouraged to upgrade to ScyllaDB 2025.2, and are welcome to contact our Support Team with questions. To get the most from ScyllaDB 2025.1, use ScyllaDB Manager 3.5 and later, and ScyllaDB Monitoring Stack 4.9 and later. Tablets load-balancing is now aware of each node’s capacity. Different nodes can have different ratios between storage size and shard count. This change prevents some nodes from reaching 100% utilization while others have free space. #23079 New compressor implementations use dictionaries to improve the compression ratio. The dictionaries are shared across SSTables and across all nodes in the cluster. The system automatically generates new dictionaries when it sees a gain in compression ratio. The new compression is NUMA aware. It distributes the dictionaries across all shards in a node, with one copy per NUMA node. Performance loss is minimized due to cross-NUMA-node memory accesses. Below is a comparison of the storage used for a dataset from Tutorials and Example Datasets | ClickHouse Docs You can either CREATE or ALTER table to use the new ‘sstable_compression’ option: ALTER TABLE keyspace.table WITH compression = {‘sstable_compression’: ‘ZstdWithDictsCompressor’}; Source: Shared-Dictionary Compression for SSTables Docs A new Guardrail allows the ScyllaDB admin to force Tablets only Keyspaces on the cluster. This feature is used in the latest ScyllaDB X Cloud feature. Changing DC or rack on a node which was already bootstrapped is unsafe for Vnodes , and not supported for Tablets. With the new guardrail changing snitch is possible only if it uses the same DCs and Racks. This version includes a new nodetool command: nodetool cluster, for cluster wide operations. The first cluster-level command is nodetool cluster repair for Tablets. The command uses a new admin REST API /storage_service/tablets/repair . Unlike “nodetool repair” which runs at the node level, running cluster-level repair synchronizes all data on all nodes in the cluster, for Tablets only. ScyllaDB Manager automatically uses the proper API. A new procedure for recovering from Raft group 0 majority loss. The procedure is safe for use with tablets. #20657 Use this procedure only as a last resort, when there is no other way to recover a failed node and get a quorum of the nodes running. The full procedure is here. The raft group 0 implementation now limits the number of raft voters in order to reduce the amount of work needed to reach consensus. Nodes are promoted to voters or demoted as needed. ScyllaDB automatically selects a maximum of 5 voters per cluster, and replaces them dynamically as needed. No configuration is required. #18793 #23786 #23950 #23588 Audit syslog output was improved to make it machine parseable. #23099 There is now support for vector types in CQL. Vectors are fixed-size arrays of another data type, commonly used for AI. Note nearest-neighbor vector search is not yet supported. To minimize image size and remove unnecessary dependencies, ScyllaDB’s default package no longer includes cassandra-stress. The recommended way to run cassandra-stress is via Docker: docker run --rm scylladb/cassandra-stress:3.17.0 For more details, refer to the official documentation here. This release includes experimental support for native backup. Backups to S3 relied on the Scylla Manager Agent (using rclone), managed by the Scylla Manager server. In this release adds the infrastructure for direct connection between ScyllaDB and S3 for backup and restore. Native Backup proves to complete the backup file upload much faster, but in the same cases, it hurts online request latency. We are working toward making it production ready in an upcoming release. You can already experiment with direct backup, by setting up the S3 connectivity. compacting_reader: decorated-key passed by reference to the compactor is moved. The compactor might use this moved-from key later to obtain tombstone GC information, which will result in incorrect tombstone GC decisions and possibly data resurrection. #23291 Partitions are (temporarily) missing when combining range scans with SELECT DISTINCT or PER PARTITION LIMIT, and can result in omitting records from the table if they do read-repair. #20084 Materialized view updates perform a read-modify-write operation on the base table. To prevent overload, we queue some of the reads. Previously, this queue had a limit of some number of entries beyond which reads would be rejected. This could cause base/view inconsistencies. The queue limit is now removed and we rely on throttling the base table writes to control its length. #23319 A bug which could cause data resurrection in materialized views (but not in the base tables) due to mixing up purge times for regular tombstones and shadowable tombstones was fixed. #23272 Tablet merges happen when the load balancer wants to reduce the number of tablets in a table. To merge tablets, the load balancer performs a “colocation migration” to move one tablet of the pair to the same node and shard as the other. It will now prefer migrating within the same rack. #22994 Add tablets enforcing option #22273 There are now per-table tablet configuration options used to control how many tablets are created for a table, allowing planning ahead for performance. #22090 Full docs: Data Definition | ScyllaDB Docs Tables created with tablets now have topology that is better prepared for immediate ingestion. changes default number of tablet replicas per shard to be 10 in order to reduce load imbalance between shards introduces a global goal for tablets replica count per shard and adds logic to tablet scheduler to respect it by controlling per-table tablet count #21967 Tablet repair can now filter by host or datacenter. #22417 Repair of one tablet will no longer prevent another tablet from being migrated. #22408 storage_service: fix tablet split of materialized views #23335 Finalize tablet splits earlier. If there is a large load balancing backlog, split finalization may be delayed arbitrarily long and we end up with large tablets. #21762 repair: Topology operations such as tablet migration can now run concurrently with repair of other tablets. #23453 Tablet allocation on table creation overloads nodes with fewer shards #23378 Truncate or drop table after tablet migration might cause assert and unexpected exit #18059 When rebuilding a tablet (due to the loss of a node), we will now stream data from just one replica, and use repair to fill in data from the rest. This saves bandwidth and reduces space amplification. #17174 There is now a new virtual table that describes load per node, and the tablet monitoring script was updated to make use of it. This is useful for heterogeneous clusters where different nodes have different storage capacity. #23584 Improved tablet load distribution is implemented to address situations where a new table is created on an already unevenly balanced cluster. #23631 The process for merging tablets (which happens when a table’s size decreases) did not update the row cache about the merged sstables, causing some data to be missed during full scans. This is now fixed. #23313 ScyllaDB no longer crashes when creating a table while there is a rack that has no nodes in the NORMAL state. #22625 Some configuration parameters can be live-updated on a running server by sending SIGHUP. We now prevent parameters that are not designed to be live updated from being updated in the same manner, as it can cause unpredictable behavior. Updateable Config: SIGHUP will make unsupported parameters effective in runtime · Issue #5382 · scylladb/scylladb · GitHub Error handling while streaming mutations was improved. #20227 ScyllaDB will automatically parallelize some aggregation queries, such as SELECT count(*) FROM table. Such parallelized queries are now cancelled if a node is being shut down. #22337 The Raft implementation now limits consumption of memory for replication. #14411 Handling of TRUNCATE statements while previous TRUNCATE statements are still processing was improved. #22166 A race condition between splitting tablets of a table, and a DROP of the same table, was fixed. #21859 A case where Raft initialization loads incorrect values from disk was fixed. #21114 A race condition between the cleanup operation and snapshot operation was fixed. #23049 s3_client: Add retries to Security Token Service/EC2 instance metadata credentials providers #21933 (see Native Backup above) Node shutdown now cancels draining hints. This reduces problems shutting down a node if the rest of the cluster is not healthy. #21949 A bug which prevented column renames from being propagated to materialized views was fixed. #22194 The CQL binary protocol server now throttles new connection processing in order to prevent connection storms from overwhelming the server. #22844 Since Scylla 6.0, cache and memtable cells between 13 kiB and 128 kiB are getting allocated in the standard allocator rather than inside LSA segments. This can result in an out of memory issue Materialized views are now more robust during schema changes. The change removes the possibility of accessing an outdated schema that no longer exists or is incompatible with the view schema. #9059 #21292 #22194 #22410 Fix potential use after free in replica/database: memtable_list: save ref to memtable_table_shared_data #23762 A possible use-after-free during schema changes related to the sstable_set type was fixed #22040 The memtable_flush_period_in_ms option now works for system tables. #21223 The scylla sstable command can now access sstables stored on S3 rather than local disk. (see S3 backend above) #20535 scylla-nodetool: rapidjson::GenericValue::GetInt() can trigger assertions if integer value overflows 32 bit int. #23394 When loading sstables into the database with nodetool, there is now a –skip-cleanup option when the user wishes to defer cleanup to a later time. This allows a single cleanup operation to be run for many sstable loads. #24136 The nodetool refresh command now supports the –scope option, allowing data to be streamed only to the local node, local rack or local datacenter. This is useful when restoring from a backup that contains individual sstable sets for each logical scope and its main purpose is to speed up restoring of 1:1 scenarios by reducing the number of copies streamed. #23861 Scylla Monitoring Stack 4.10 and later support ScyllaDB 2025.2 --- ### Page: https://forum.scylladb.com/t/getting-error-when-i-type-scylla-on-terminal/4945 Title: Getting error when I type "scylla" on terminal - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details: curl -sSf get.scylladb.com/server | sudo bash ( run this script to install) #ScyllaDB version: 2025.2.0-0.20250625.33e947e75342 #Cluster size: os (RHEL/CentOS/Ubuntu/AWS AMI): Ubuntu (Oracle VP… Language: en Canonical URL: https://forum.scylladb.com/t/getting-error-when-i-type-scylla-on-terminal/4945 ## Headings Structure: H1: Getting error when I type "scylla" on terminal H3: Related topics ## Main Content: H1: Getting error when I type "scylla" on terminal H3: Related topics Installation details: curl -sSf get.scylladb.com/server | sudo bash ( run this script to install) #ScyllaDB version: 2025.2.0-0.20250625.33e947e75342 #Cluster size: os (RHEL/CentOS/Ubuntu/AWS AMI): Ubuntu (Oracle VPS) I dont know in which location it is asking for conf folder. how can I fix this issue? You should start ScyllaDB with systemctl start scylla-server. If you just want to play around with ScyllaDB on a development machine, you can start it with something like scylla -c2 -m2G --developer-mode=1 --options-file=/path/to/scylla.yaml. You can just create an empty yaml file with touch scylla.yaml. --- ### Page: https://forum.scylladb.com/t/scylla-db-read-performance-impact-in-amazon-al2023/4948 Title: Scylla db read performance impact in Amazon Al2023 - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 5.1.6 #Cluster size:12 os (RHEL/CentOS/Ubuntu/AWS AMI):Al2023 We migrated the cluster from Al2 to Al2023.Post that we are seeing latency spike for read operations.Any suggesti… Language: en Canonical URL: https://forum.scylladb.com/t/scylla-db-read-performance-impact-in-amazon-al2023/4948 ## Headings Structure: H1: Scylla db read performance impact in Amazon Al2023 H3: OS Support by Linux Distributions and Version | ScyllaDB Docs H3: Related topics ## Main Content: H1: Scylla db read performance impact in Amazon Al2023 H3: OS Support by Linux Distributions and Version | ScyllaDB Docs H3: Related topics Installation details #ScyllaDB version: 5.1.6 #Cluster size:12 os (RHEL/CentOS/Ubuntu/AWS AMI):Al2023 We migrated the cluster from Al2 to Al2023.Post that we are seeing latency spike for read operations.Any suggestion in OS level tuning. As a first step, I’d recommend updating to the latest ScyllaDB version. You’re running an old version, and you might encounter issues already solved. Only the latest version of scylla supports Amazon Linux, none of the OSS versions supported Amazon Linux ScyllaDB is an Apache Cassandra-compatible NoSQL data store that can handle 1 million transactions per second on a single server. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-99-2025-07-04/4954 Title: Last week in scylla-cluster-tests.git master (issue #99; 2025-07-04) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the a6571cc3…1936d3c4 range are covered. There were 12 non-merge commits from 5 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-99-2025-07-04/4954 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #99; 2025-07-04) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #99; 2025-07-04) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the a6571cc3…1936d3c4 range are covered. There were 12 non-merge commits from 5 authors in that period. Some notable commits: Due to reaching the AWS Security Groups limit, the cloud cleanup process will now periodically remove any unused security groups that aren’t tagged with keep:alive. We’ve refactored how nodes are selected and assigned to nemeses, introducing the new NemesisNodeAllocator entity, which centralizes node selection and tracking to avoid the complexity and race conditions from locking in multiple places. To minimize capacity issues and prevent tests from impacting each other, we’ve enabled Jenkins throttle-concurrents plugin for performance tests, so only one test will run at a time in the same region. Finally, rsync will now be installed by default since it better fits our synchronization needs and is generally preferred for file transfers. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/handling-large-partition-and-full-scan-query-when-operating-scylladb/4956 Title: Handling large partition and full scan query when operating ScyllaDB - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 5.4 #Cluster size: 3 each DC1,2 os (RHEL/CentOS/Ubuntu/AWS AMI): CentOS Hi all — I’m using ScyllaDB to store transactional data and had a few questions around query patterns an… Language: en Canonical URL: https://forum.scylladb.com/t/handling-large-partition-and-full-scan-query-when-operating-scylladb/4956 ## Headings Structure: H1: Handling large partition and full scan query when operating ScyllaDB H3: Related topics ## Main Content: H1: Handling large partition and full scan query when operating ScyllaDB H3: Related topics Installation details #ScyllaDB version: 5.4 #Cluster size: 3 each DC1,2 os (RHEL/CentOS/Ubuntu/AWS AMI): CentOS Hi all — I’m using ScyllaDB to store transactional data and had a few questions around query patterns and cluster impact. I’d love to hear your experience on: Has anyone tackled similar patterns or have suggestions from an app-level or cluster tuning perspective? Thanks a lot! Look into Workload Priotization, it was designed exactly for this use-case: isolate heavy but non-latency sensitive scans from other queries. Thank you for the reply! Just to confirm, isn’t Workload Prioritization an enterprise-only feature? With the move to source-available, we no longer have enterprise-only features. All features are available in the single source-available release stream. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-284-2025-07-06/4959 Title: Last week in scylladb.git master (issue #284; 2025-07-06) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 9d70e7a067..33225b730d range are covered. There were 141 non-merge commits from 24 authors in that per… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-284-2025-07-06/4959 ## Headings Structure: H1: Last week in scylladb.git master (issue #284; 2025-07-06) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #284; 2025-07-06) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 9d70e7a067..33225b730d range are covered. There were 141 non-merge commits from 24 authors in that period. Some notable commits: The nodetool backup command learned the --move-files option which moves files instead of copying them. The data structure used for building mutations now has additional sanity checks for row clustering keys. The sstable writer will detect invalid row clustering keys and redirect those rows into a new system table for corrupted data. Support for Ubuntu 20.04, which has reached end-of-life, was removed. Lightweight transactions (LWT) now synchronize with tablet migrations. This is a step towards enabling LWT with tablets. The toolchain used to build ScyllaDB is now based on Fedora 42 with clang 20.1 and libstdc++ 15. Repair will now be delayed if there is an ongoing tablet merge finalization, to avoid triggering an internal error. There is now support for co-locating tablets of different tables. Co-located tablets are migrated, split, and merged together. They will be used for lightweight transactions, change data capture, and local materialized views. We now avoid large contiguous allocations during DESCRIBE statements with tables that have many (possibly deleted) columns. There is now support for converting CQL3 data type representations to a new byte-comparable representation. This is a step towards implementation of Trie sstable indexes. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/detecting-hot-partitions/4963 Title: Detecting Hot Partitions - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/detecting-hot-partitions/4963 ## Headings Structure: H1: Detecting Hot Partitions H3: Related topics ## Main Content: H1: Detecting Hot Partitions H3: Related topics Originally from the User Slack @Daria_Fedorova: Hi, does scylla have any method of detecting hot partitions ? We had a bug in one of our clients that caused a bunch of async writes in one partition leading to high latencies for some calls I am still thinking how we could have detected the problem faster because my first thoughts were about problems with network or hardware @avi: nodetool toppartitions --- ### Page: https://forum.scylladb.com/t/repair-by-token-ranges/4964 Title: Repair by token ranges - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello. I’m trying to use repair API with token ranges and have a number of questions. Empty [start | end]Token: Consider the following result: curl 10.0.32.70:10000/storage_service/range_to_endpoint_map/markb | jq -… Language: en Canonical URL: https://forum.scylladb.com/t/repair-by-token-ranges/4964 ## Headings Structure: H1: Repair by token ranges H3: Related topics ## Main Content: H1: Repair by token ranges H3: Related topics Hello. I’m trying to use repair API with token ranges and have a number of questions. The following command returns: {"message": "boost::wrapexcept (bad lexical cast: source type value could not be interpreted as target)", "code": 500}. Q1.1: Is it possible to repair this token range using the ranges parameter? Q1.2: If A1.1 is “no”, then what’s the proper command to repair such a token range (if I need to repair it at all)? Notice, that there is no 10.0.32.71 host in the endpoint list of this token range. But the command above “successfully” repairs this token range according to the result of the following command (with X returned by the above command). The result doesn’t change, if I don’t use the primaryRange parameter. Q2: How is it possible to successfully run a repair command of a token range on a host which doesn’t present in the endpoint list of this token range? Does this command really repairs a token range on other hosts, when a repair master host doesn’t have any replica of this data? Please share more info about your deployment (ScyllaDB Version, OS, etc.) Why don’t you use Scylla Manager to manage cluster repair and backup? It’s a test system. Scylladb 5.4.4 Community Debian 10 3 DC x 3 nodes each The reason I cat’t use Scylla Manager is its 5-node limit for open-source editions. But I believe, that results of these API calls are the same on all Scylla versions. BTW, the nodetool repair [-pr] -st X -et Y keyspace commands return SUCCESSFUL as well on whatever X, Y values disregarding the host where I run it. For example, the following command finishes “successfully” as well. So, if all such commands return “success”, it would be good to know if they really do some useful work. And if not, that I need to know the correct algorithm of my cluster reparation by token ranges. I’m not able to construct such an algorithm based on available documentation at the moment, so I need some help here. Lets presume, that I need to repair a keyspace markb in my cluster. My algorithm with not clear enough info is below. I believe, that a (start_token, end_token) pair must be provided to the nodetool utility as is from the storage_service/range_to_endpoint_map API call (both return/accept such an interval as exclusive-inclusive, so I don’t need any ±1 math here). It’s not clear here what to do with token ranges returned by storage_service/range_to_endpoint_map without start_token or end_token. Should I ignore these ranges? Should I run some special repair command on them? It’s not clear if I need to use the -pr parameter here. Seems, that it should be no difference, because 10.0.32.70 holds a primary range and I run the command namely on this host. I must (or could?) omit -pr running this command on, say, 10.0.32.74, because it holds replica of this token range. Please, correct me if it’s not a correct algorithm to repair a keyspace by token ranges. Seems, that the ScyllaDB’s implementation of nodetool repair with -st / -et is not compatible with the Cassandra’s one. ScyllaDB doesn’t check parameters at all. Below is a couple of examples which must finish with an error as on, say, Cassandra 4.1.9. ScyllaDB returns “SUCCESSFUL” in both cases, which is incorrect in my opinion. P.S.: The same behavior is on ScyllaDB 6.2.3 Community… As explained under a different venue the SUCCESSFUL message indicates the non-replica node works only as a repair coordinator. In fact, if you shutdown one of the actual replicas involved in the repair, then its status should change to FAILED. With regards to tokens missing a or , this indicates the range wrap arounds on /range_to_endpoint_map. ScyllaDB returns “SUCCESSFUL” in both cases, which is incorrect in my opinion. That said, it is arguable whether the behavior your observe is incorrect or not, as repair is taking place and thus fulfilling its goal, though not in most efficient way. Forcing it to only accept a replica coordinator could be an option, warning the user about it another, or simply adding a boolean flag yet another. Note that both /range_to_endpoint_map/{keyspace} as well as (my favorite) /describe_ring/{keyspace} include the relevant endpoints owning a particular range. That said, incorrectly triggering repair to a non-replica is unlikely to manifest in practice unless manually forced by the user as your example demonstrates. Please file a GitHub issue and explain your use case should our APIs be missing on anything. the SUCCESSFUL message indicates the non-replica node works only as a repair coordinator It becomes more interesting Did I get it right, that Scylla nodetool repair implementation can run on non-replica node to repair a full token range? This means, that it’s possible to run it on a single node to repair the whole cluster. Is that correct? Well, it’s cool enough, but it should be clearly documented somewhere. We all know, that it’s impossible to do such things in Cassandra, Datastax. Even ScyllaDB docs say, that we must run nodetool repair on all cluster hosts! I’ll definitely open an issue on this, if I have time. Either with a request to clarify this incompatible behavior (for example, to clearly state, that a successful run on a non-replica node really repairs the corresponding range and not silently swallows the call doing nothing useful as one could expect), or to forbid such calls… Did I get it right, that Scylla nodetool repair implementation can run on non-replica node to repair a full token range? This means, that it’s possible to run it on a single node to repair the whole cluster. Is that correct? When you specify a token range to repair, yes. This alone, is a very particular use case, which on its own already require retrieving the token-to-replica mapping on its own. So the subsequent repair will often get executed on a replica node anyway. You may check the scylla-server journalctl logs as you invoke it on a non-replica node and compare it versus a replica node. In the former case, all replicas plus the non-replica node will be involved during the repair. In the latter, only the actual replicas, hence why the latter is more efficient. I don’t remember the logic around nodetool repair alone (-pr isn’t subject to this as it already assumed the primary ranges of the particular invoked replica), but if mind serves me well ScyllaDB would distribute the tasks accordingly across the natural endpoints. But best to check logs to confirm for a relatively sparse table. And here is the Slack thread for reference: Originally from the User Slack @Mark_Barinstein: Hello. I’m trying to use repair API with token ranges and have a number of questions. The following command returns {"message": "boost::wrapexcept (bad lexical cast: source type value could not be interpreted as target)", "code": 500} Q1.1: Is it possible to repair this token range using the ranges parameter? Q1.2: If A1.1 is “no”, then what’s the proper command to repair such a token range (if I need to repair it at all)? Notice, that there is no 10.0.32.71 in the endpoints list of this token range. But the command above “successfully” repairs this token range according to the result of the following command (with X returned by the above command). The result doesn’t change, if I don’t use the primaryRange parameter. Q2: How is it possible to successfully run a repair command of a token range on a host which doesn’t present in the endpoints list of this token range? Is it some “feature”? Thanks in advance. FYI: The command above behaves the same: it always returns success disregarding the host where I run it. Moreover, I even get success if I specify whatever start & end tokens like: How could one trust such a result? @Felipe_Cardeneti_Mendes: well, your first repair_async command doesn’t contain an ending token, so clearly it fails. The former range_to_endpoint_map output means it wraps around, so you should apply the relevant values instead. Personally, I find /storage_service/describe_ring/{keyspace} more convenient. Consider: And the following list of operations: 172.31.4.30 isn’t a replica and received the request to repair the range owned by 2 other replicas. In that sense, it is a repair coordinator, and you’ll observe the result of this coordination under its logs. That said, you may call it a “feature”, but as you can imagine it is a bit pointless to start a repair task using a non-replica coordinator, as you’ll transfer data around unnecessarily. Overall, Scylla Manager streamlines all this logic, including when tablets are used which have their own APIs for repairing. @Mark_Barinstein: Indeed, the storage_service/describe_ring API result doesn’t have such a “feature” with empty key boundary unlike storage_service/range_to_endpoint_map: Will use describe_ring instead. I started to use range_to_endpoint_map because of its smaller result set - w/o useless fields in my case. Thanks! BTW, more problems with the Scylla nodetool repair utility (and similar API). Its result is not compatible with Cassandra’s one sometimes. https://forum.scylladb.com/t/repair-by-token-ranges/4964 ScyllaDB Community NoSQL Forum: Repair by token ranges When you specify a token range to repair, yes. This alone, is a very particular use case, which on its own already require retrieving the token-to-replica mapping on its own. So the subsequent repair will often get executed on a replica node anyway. Q1: Do I get it right, that such a non-replica node becomes a “replay master” node anyway, and this node can’t participate in any other parallel repairs until the end of this one, even this non-replica node doesn’t hold any data being repaired? Q2: Do I get it right that it’s some “fake” token range above (with start_token > end_token) which can’t hold any data, and we can omit its reparation? --- ### Page: https://forum.scylladb.com/t/scylladb-jmx-listening-on-ports/4965 Title: ScyllaDB JMX listening on ports - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-jmx-listening-on-ports/4965 ## Headings Structure: H1: ScyllaDB JMX listening on ports H3: Related topics ## Main Content: H1: ScyllaDB JMX listening on ports H3: Related topics Originally from the User Slack @Mahdi_Kamali: Why scylla-jmx is opening a random port and listening on 0.0.0.0:xxxx?! What this port is used for? tcp 0 0 0.0.0.0:xxxxx 0.0.0.0:* LISTEN 87644/java @Some_Random_Guy: it isn’t. my guess is you’re running netstat as a non-root user @avi: also, modern versions no longer have jmx --- ### Page: https://forum.scylladb.com/t/number-of-client-connections-per-scylladb-node-shard-core/4970 Title: Number of client connections per ScyllaDB node/shard/core? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/number-of-client-connections-per-scylladb-node-shard-core/4970 ## Headings Structure: H1: Number of client connections per ScyllaDB node/shard/core? H3: Related topics ## Main Content: H1: Number of client connections per ScyllaDB node/shard/core? H3: Related topics Originally from the User Slack @Mahdi_Kamali: How many connections is too much for a scylla node or core? We have a cluster with multiple client applications (about 100 clients). Each client opens 1 connection per core (default for golang driver). Is this too much? More than 100 connections per core and about 10K connections per node @avi: Several thousand connections per shard aren’t a problem --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-2-12/4971 Title: [RELEASE] ScyllaDB Enterprise 2024.2.12 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2024.2.12, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2024.2 Feature (Short Term Support) Release. Note that 2024.2 has become en… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-2-12/4971 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2024.2.12 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2024.2.12 H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2024.2.12, a bug-fix production-ready ScyllaDB Enterprise patch release for ScyllaDB Enterprise 2024.2 Feature (Short Term Support) Release. Note that 2024.2 has become end of life, and this will be the last patch release for it. You are encouraged to upgrade to Long Term Support Release 2025.1, or Feature Release 2025.2 in coordination with the ScyllaDB Support team. The following issues are fixed in this release (with an open-source reference, if available): --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-100-2025-07-11/4972 Title: Last week in scylla-cluster-tests.git master (issue #100; 2025-07-11) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 89af4b3a…6fa21ca6 range are covered. There were 15 non-merge commits from 8 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-100-2025-07-11/4972 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #100; 2025-07-11) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #100; 2025-07-11) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 89af4b3a…6fa21ca6 range are covered. There were 15 non-merge commits from 8 authors in that period. Some notable commits: Added a bit of documentation for loaders configuration parameters. Introduced an enhanced script for copying an AMI to multiple regions, applying tags, and waiting for the image to become available. It detects if the image already exists before copying, supports dry-run mode, and can fetch the default region list from sct_config.py. Continued the refactor of nemesis machinery: run_nemesis was moved from the Nemesis class to the node allocator abstraction. This change allows marking target nodes without requiring a Nemesis object and still preserves backward compatibility. A performance test using latte for steady-state latency under a custom schema is now ready for regular triggering. It supports two flavors: one for vnodes and one for tablets. Switched SCT test execution from unittest.main() to pytest, gaining better maintainability and access to pytest features like fixtures, subtests, and hooks. Docker backend now uses the official ScyllaDB Docker image instead of rebuilding scylla-sct. A new DockerCmdRunner abstraction was added to run commands via Docker API instead of SSH. scylla-bench was updated from version 0.2.4 to 0.2.5, fixing an issue where the -help flag was incorrectly printed to stderr instead of stdout. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/why-does-a-token-unaware-query-involving-a-local-secondary-index-require-a-round-trip/4976 Title: Why does a token-unaware query involving a local secondary index require a round trip? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: According to the documentation on local secondary indexes, the read path for token-unaware queries involving a local secondary index (LSI) are not completely satisfied locally. Instead, the keys retrieved from the LSI on… Language: en Canonical URL: https://forum.scylladb.com/t/why-does-a-token-unaware-query-involving-a-local-secondary-index-require-a-round-trip/4976 ## Headings Structure: H1: Why does a token-unaware query involving a local secondary index require a round trip? H3: Related topics ## Main Content: H1: Why does a token-unaware query involving a local secondary index require a round trip? H3: Related topics According to the documentation on local secondary indexes, the read path for token-unaware queries involving a local secondary index (LSI) are not completely satisfied locally. Instead, the keys retrieved from the LSI on a given node are sent back to the coordinator node, which sends a request back to the node to retrieve the data associated with those keys. If the data which an LSI on a given node references also exists on the same node, why does the coordinator node need to be involved at all in the retrieval of such data on that node? A token-unaware query using a local secondary index (LSI) in ScyllaDB requires a round trip—even when the data referenced by the LSI is colocated on a single node—due to the way query coordination and result gathering is structured in ScyllaDB. Even though LSI guarantees that both the index and the corresponding base table data reside on the same node (because the index shares the base table’s partition key), the step-wise query process still involves a coordinator node. For a token-unaware driver (i.e., when the client does not direct the query to the correct node), here is what happens: Why can’t the initial node simply serve the data immediately? Optimization is possible with token-aware queries: In short: A token-unaware query needs a round trip because ScyllaDB’s query flow separates index lookups and base data retrieval, always routing through the coordinator node, even if the queried node (which executes the index lookup) could have served the data itself. This path is necessary for consistency and feature uniformity, though it can be avoided with token-aware clients. Check also this great blog post for more details, --- ### Page: https://forum.scylladb.com/t/how-to-run-nodetool-repair-all-nodes-or-only-one-datacenter/4978 Title: How to run nodetool repair, all nodes or only one datacenter? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/how-to-run-nodetool-repair-all-nodes-or-only-one-datacenter/4978 ## Headings Structure: H1: How to run nodetool repair, all nodes or only one datacenter? H3: Related topics ## Main Content: H1: How to run nodetool repair, all nodes or only one datacenter? H3: Related topics Originally from the User Slack @Mahdi_Kamali: Should I run nodetool repair -pr in all nodes of cluster for full cluster repair? Or running on only one datacenter is enough? (Assuming data is replicated even in all datacenters) @avi: All nodes of the cluster --- ### Page: https://forum.scylladb.com/t/release-scylladb-2025-1-4/4979 Title: [RELEASE] ScyllaDB 2025.1.4 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB 2025.1.4, a bug-fix production-ready patch release for ScyllaDB 2025.1 LTS Release. Note there is a new Short Term Support (STS) Feature release 2025.2. You are welcome to upgrade to… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-2025-1-4/4979 ## Headings Structure: H1: [RELEASE] ScyllaDB 2025.1.4 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 2025.1.4 H3: Related topics The ScyllaDB team announces ScyllaDB 2025.1.4, a bug-fix production-ready patch release for ScyllaDB 2025.1 LTS Release. Note there is a new Short Term Support (STS) Feature release 2025.2. You are welcome to upgrade to it for the latest and greatest features, or stay on the 2025.1 track for long term support. The following issues are fixed in this release: Encryption at Rest: KMIP LOCATE operation request has incorrect attribute names #23970 As a result the KMIP server rejected the request, leading ScyllaDB to assume that a key with these specifications doesn’t exist, and creates a new key in the KMIP server. A race condition between the cleanup operation and snapshot operation was fixed. #23049 non-full row keys cause mis-parsing of sstables #24489 Avoid killing a node when reading from sstables #20845 Assertion `_mt._flushed_memory <= _mt.occupancy().total_space()’ failed during bulk ingestion #21413 Backport seastar fix for nested stack unwind crash #24464 seasstar fix: stall_detector: no backtrace if exception #2714 Core dump during node bootstrap accessing the mapping before it is initialized #24479 Potential data race in utils::alien_worker #24751 Don’t start maintenance auth service if not enabled #24528 Empty clustering keys generated and spread freely in the system #24506 A possible use-after-free during schema changes related to the sstable_set type was fixed #22040 Materialized views are now more robust during schema changes. The change removes the possibility of accessing an outdated schema that no longer exists or is incompatible with the view schema. #9059 #21292 #22194 #22410 Joining node enters synchronize state after joining group0, which races with streaming and causes join to fail #23536 paxos_response_handler timeout logging #24591 A bug which prevented column renames from being propagated to materialized views was fixed. #22194 repair: to_repair_rows_on_wire stalls destroying input list #24725 SSTable compaction fails due to Chunk count mismatch between CRC and Data.db #23728 user_functions_test.TestUserFunctions.test_restart AssertionError: Expected [[‘aaaaaaaaa’]] from … but got [[‘aaaa’]] #20662 root cause is race condition in the mapreduce_service --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-19/4980 Title: [RELEASE] ScyllaDB Enterprise 2024.1.19 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2024.1.19, a bug-fix production-ready patch release for ScyllaDB Enterprise 2024.1 Long Term Support (LTS) Release. Note there is a later, better, LTS (Long term Support) R… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-19/4980 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2024.1.19 H3: Related Links H2: Fixed Issue with an source available reference: H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2024.1.19 H3: Related Links H2: Fixed Issue with an source available reference: H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2024.1.19, a bug-fix production-ready patch release for ScyllaDB Enterprise 2024.1 Long Term Support (LTS) Release. Note there is a later, better, LTS (Long term Support) Release 2025.1, and a feature release 2025.2. You are encouraged to upgrade in coordination with the ScyllaDB Support team. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-285-2025-07-14/4981 Title: Last week in scylladb.git master (issue #285; 2025-07-14) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 33225b730d5..fdcaa9a7e79 range are covered. There were 76 non-merge commits from 12 authors in that pe… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-285-2025-07-14/4981 ## Headings Structure: H1: Last week in scylladb.git master (issue #285; 2025-07-14) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #285; 2025-07-14) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 33225b730d5..fdcaa9a7e79 range are covered. There were 76 non-merge commits from 12 authors in that period. Some notable commits: Tablet metadata can grow very large in large clusters. Its creation already avoided reactor stalls, and now its destruction as well. The repair small table optimization is used to speed up repairs on tables that are empty or almost empty (like many tables in system_distributed). On large clusters, the optimization could lead to an out-of-memory condition. This is now fixed. The vector search system implements the client that accesses the external index. Some stalls when managing schemas with thousands of tables were eliminated. A new schema application framework was merged, which will allow to improve atomicity when managing cluster metadata. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-cluster-level-email-notifications-i3-no-longer-available-in-scylladb-cloud-portal-15-july-2025/4984 Title: [RELEASE] Cluster-Level Email Notifications + i3 No Longer Available in ScyllaDB Cloud Portal - 15 July 2025 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: Here’s what’s new (and what’s changed) in ScyllaDB Cloud: :open_mailbox_with_raised_flag: Cluster-Level Notification Emails You can now define notification emails per cluster. Until now, email notifications were manage… Language: en Canonical URL: https://forum.scylladb.com/t/release-cluster-level-email-notifications-i3-no-longer-available-in-scylladb-cloud-portal-15-july-2025/4984 ## Headings Structure: H1: [RELEASE] Cluster-Level Email Notifications + i3 No Longer Available in ScyllaDB Cloud Portal - 15 July 2025 H3: Cluster-Level Notification Emails H3: i3 Instance Types No Longer Available in ScyllaDB Cloud Portal H3: Related topics ## Main Content: H1: [RELEASE] Cluster-Level Email Notifications + i3 No Longer Available in ScyllaDB Cloud Portal - 15 July 2025 H3: Cluster-Level Notification Emails H3: i3 Instance Types No Longer Available in ScyllaDB Cloud Portal H3: Related topics Here’s what’s new (and what’s changed) in ScyllaDB Cloud: You can now define notification emails per cluster. Until now, email notifications were managed at the account level using a shared list of default contact emails. With this new feature, you can customize who gets notified for each cluster by specifying a unique set of email addresses per cluster. Head to Profile > Notification Settings in your ScyllaDB Cloud portal to set up per-cluster overrides. We’ve removed AWS i3 instance types from the ScyllaDB Cloud Portal. While i3 instances (Nitro-based) are still supported for existing deployments, new clusters should be launched using i4i instance types, which offer improved performance and efficiency. If you have any questions or would like help evaluating i4i options, your Account Manager and our Support team are here to assist. --- ### Page: https://forum.scylladb.com/t/duplicate-user-define-data-type-list-value-while-restoring-data-using-sstable-loader/4985 Title: Duplicate User Define data type list value while restoring data using sstable loader - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 5.2 #Cluster size:12 os (RHEL/CentOS/Ubuntu/AWS AMI):Amazon Linux 2 We could see data corruption (Duplicate list Value) for user define data type while restoring data using ss… Language: en Canonical URL: https://forum.scylladb.com/t/duplicate-user-define-data-type-list-value-while-restoring-data-using-sstable-loader/4985 ## Headings Structure: H1: Duplicate User Define data type list value while restoring data using sstable loader H3: Related topics ## Main Content: H1: Duplicate User Define data type list value while restoring data using sstable loader H3: Related topics Installation details #ScyllaDB version: 5.2 #Cluster size:12 os (RHEL/CentOS/Ubuntu/AWS AMI):Amazon Linux 2 We could see data corruption (Duplicate list Value) for user define data type while restoring data using sstable loader. rstats list> It’s hard to know what’s happening without more details. In any case, you’re using an old version that’s no longer supported. Please update to the latest version; perhaps this issue has already been resolved. --- ### Page: https://forum.scylladb.com/t/release-scylladb-cpp-over-rust-driver-0-5-1/4990 Title: [RELEASE] ScyllaDB CPP-over-Rust Driver 0.5.1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce ScyllaDB CPP Rust Driver 0.5.1, an API-compatible rewrite of GitHub - scylladb/cpp-driver: Scylla C/C++ Driver as a wrapper for the Rust driver. It will fully replace the CPP driv… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-cpp-over-rust-driver-0-5-1/4990 ## Headings Structure: H1: [RELEASE] ScyllaDB CPP-over-Rust Driver 0.5.1 H2: Changes H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB CPP-over-Rust Driver 0.5.1 H2: Changes H3: Related topics The ScyllaDB team is pleased to announce ScyllaDB CPP Rust Driver 0.5.1, an API-compatible rewrite of GitHub - scylladb/cpp-driver: Scylla C/C++ Driver as a wrapper for the Rust driver. It will fully replace the CPP driver, which is already reaching its End of Life. The driver version shall be considered Beta. Some minor features still need to be included. See Limitations and Unimplemented functions from cassandra.h sections in README.md. The underlying Rust driver used version: 1.3.0. New features / enhancements CI / developer tool improvements: Congrats to all contributors and thanks everyone for using our driver! The source code of the driver can be found here: Thank you for your attention, please do not hesitate to contact us if you have any questions, issues, feature requests, or are simply interested in our driver! Contributors since the last release: --- ### Page: https://forum.scylladb.com/t/release-scylladb-cloud-bring-your-own-account-byoa-now-supports-google-cloud-platform-17-july-2025/4993 Title: [RELEASE] ScyllaDB Cloud Bring Your Own Account (BYOA) now supports Google Cloud Platform - 17 July 2025 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: ScyllaDB Cloud Bring Your Own Account (BYOA) is available for Google Cloud Platform We’re pleased to announce that ScyllaDB Cloud Bring Your Own Account (BYOA) now supports Google Cloud Platform. With BYOA, ScyllaDB Cl… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-cloud-bring-your-own-account-byoa-now-supports-google-cloud-platform-17-july-2025/4993 ## Headings Structure: H1: [RELEASE] ScyllaDB Cloud Bring Your Own Account (BYOA) now supports Google Cloud Platform - 17 July 2025 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Cloud Bring Your Own Account (BYOA) now supports Google Cloud Platform - 17 July 2025 H3: Related topics ScyllaDB Cloud Bring Your Own Account (BYOA) is available for Google Cloud Platform We’re pleased to announce that ScyllaDB Cloud Bring Your Own Account (BYOA) now supports Google Cloud Platform. With BYOA, ScyllaDB Cloud runs directly inside your private Google Cloud account. Your data remains fully under your control and never leaves your cloud environment. It combines the convenience of having a managed product with the control of your own infrastructure. The BYOA feature for both AWS/GCP is available today on cloud.scylladb.com for the Professional plan and above. Key benefits include: You can find more information in our cloud documentation: Deploy ScyllaDB to Your Own Cloud Account - GCP | ScyllaDB Docs In addition to GCP, we greatly improved the process of linking your account for AWS. Now it is a fully automated procedure with the help of CloudFormation. Updated documentation can be found here: Deploy ScyllaDB to Your Own Cloud Account - AWS | ScyllaDB Docs --- ### Page: https://forum.scylladb.com/t/release-scylladb-manager-3-5-1/4994 Title: [RELEASE] ScyllaDB Manager 3.5.1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces the release of ScyllaDB Manager 3.5.1, a production-ready patch release of the stable 3.5 branch. ScyllaDB Manager is a centralized cluster administration and recurrent tasks automation tool. … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-manager-3-5-1/4994 ## Headings Structure: H1: [RELEASE] ScyllaDB Manager 3.5.1 H3: Bug Fixes H3: Upgrade to the new release H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Manager 3.5.1 H3: Bug Fixes H3: Upgrade to the new release H3: Related topics The ScyllaDB team announces the release of ScyllaDB Manager 3.5.1, a production-ready patch release of the stable 3.5 branch. ScyllaDB Manager is a centralized cluster administration and recurrent tasks automation tool. This release focuses on many important bugfixes. ScyllaDB customers are encouraged to upgrade to ScyllaDB Manager 3.5.1 in coordination with the ScyllaDB support team. The new release includes upgrades of both ScyllaDB Manager Server and Agent. Scylla Manager 3.5.1 supports the following ScyllaDB releases: --- ### Page: https://forum.scylladb.com/t/release-scylladb-2025-1-5/4995 Title: [RELEASE] ScyllaDB 2025.1.5 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB 2025.1.5, a bug-fix production-ready patch release for ScyllaDB 2025.1 LTS Release. Note there is a new Short Term Support (STS) Feature release 2025.2. You are welcome to upgrade to… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-2025-1-5/4995 ## Headings Structure: H1: [RELEASE] ScyllaDB 2025.1.5 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 2025.1.5 H3: Related topics The ScyllaDB team announces ScyllaDB 2025.1.5, a bug-fix production-ready patch release for ScyllaDB 2025.1 LTS Release. Note there is a new Short Term Support (STS) Feature release 2025.2. You are welcome to upgrade to it for the latest and greatest features, or stay on the 2025.1 track for long term support. The following issues are fixed in this release: --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-101-2025-07-18/4997 Title: Last week in scylla-cluster-tests.git master (issue #101; 2025-07-18) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 42cc466b…9a2e7756 range are covered. There were 19 non-merge commits from 9 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-101-2025-07-18/4997 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #101; 2025-07-18) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #101; 2025-07-18) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 42cc466b…9a2e7756 range are covered. There were 19 non-merge commits from 9 authors in that period. Some notable commits: The monitoring stack has been upgraded to version 4.11. As the loaders operating system is no longer supported, they now use Ubuntu 24.04 as the base. Loaders are expected to start faster due to Docker installation during the cloud-init process. The latte tool has been upgraded to version 0.31.1, incorporating the latest Scylla Rust driver version 1.3. The sanity performance tests have been temporarily disabled. Skipping nemesis due to capacity issues could result in a cluster configuration differing from the initial setup. To address this, SCT now stops the test if such issues occur. When uploading artifact files to Argus using the upload Hydra command, users can specify whether the upload should be public or private with the --public or --no-public flag. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/how-to-define-an-inner-class-as-a-udt-in-entity-and-use-the-udt/5001 Title: How to define an inner class as a UDT in @Entity and use the UDT? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: How to define an inner class as a UDT in @Entity and use the UDT? com.scylladb java-driver-core 4.19.0.1 <dependency> <groupId>com.scylladb</groupId> <… Language: en Canonical URL: https://forum.scylladb.com/t/how-to-define-an-inner-class-as-a-udt-in-entity-and-use-the-udt/5001 ## Headings Structure: H1: How to define an inner class as a UDT in @Entity and use the UDT? H3: Related topics ## Main Content: H1: How to define an inner class as a UDT in @Entity and use the UDT? H3: Related topics How to define an inner class as a UDT in @Entity and use the UDT? package com.xgy.entity; import com.datastax.oss.driver.api.mapper.annotations.*; import com.datastax.oss.driver.api.mapper.entity.naming.NamingConvention; import com.fasterxml.jackson.annotation.JsonInclude; import com.fasterxml.jackson.annotation.JsonProperty; import lombok.AllArgsConstructor; import lombok.Data; import lombok.NoArgsConstructor; import java.util.List; import java.util.UUID; import static com.datastax.oss.driver.api.mapper.annotations.SchemaHint.TargetElement.UDT; @Data @AllArgsConstructor @NoArgsConstructor @JsonInclude(JsonInclude.Include.NON_NULL) @Entity(defaultKeyspace = “novel”) @NamingStrategy(convention = NamingConvention.LOWER_CAMEL_CASE) @CqlName(“character_”) public class Character { //CREATE TYPE novel.story_node ( // from_chapter text, // to_chapter text, // action text, // context text //); // //CREATE TABLE novel.character_ ( // novel_id uuid, // name text, // summary text, // story_chain list>, // PRIMARY KEY (novel_id, name) //); Caused by: java.lang.IllegalArgumentException: The CQL ks.table: novel.character_ defined in the entity class: com.xgy.entity.Character declares type mappings that are not supported by the codec registry: Field: story_chain, Entity Type: com.xgy.entity.Character$StoryNode, CQL type: UDT(novel.story_node) The core of your issue is that the ScyllaDB Java driver mapper cannot properly map your inner class (StoryNode) as a UDT when used in your @Entity (Character), resulting in an IllegalArgumentException about unsupported type mappings. This is a common mapping challenge with Cassandra/Scylla Java driver, especially with the use of inner/static nested classes for UDTs. Key points and steps for a correct mapping setup: So, in short - Move your UDT Java class (StoryNode) to its own file and annotate it with @Udt . Update your Entity to use this top-level class in its list field. This aligns with ScyllaDB/Datastax Java driver requirements and will resolve the mapping exception you encountered --- ### Page: https://forum.scylladb.com/t/scylladb-x-cloud-bring-your-own-account-byoa-vpc-peering-and-accessing-a-cluster-from-outside-aws/5003 Title: ScyllaDB X Cloud, Bring your own account (BYOA), VPC Peering and accessing a cluster from outside AWS - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-x-cloud-bring-your-own-account-byoa-vpc-peering-and-accessing-a-cluster-from-outside-aws/5003 ## Headings Structure: H1: ScyllaDB X Cloud, Bring your own account (BYOA), VPC Peering and accessing a cluster from outside AWS H3: Related topics ## Main Content: H1: ScyllaDB X Cloud, Bring your own account (BYOA), VPC Peering and accessing a cluster from outside AWS H3: Related topics Originally from the User Slack @Alex_Ioannides: Hello, I have a question. I have setup a ScyllaDB X Cloud with VPC peering. However, I would like to access the cluster from outside AWS as well. I am trying to avoid setting Network Type to Public Internet so that whole communication goes through public. I see that the created EC2 instances in BYOA setup don’t have public ip addresses. Any suggestions? @Buff: Hi @Alex_Ioannides The clusters have public IP addresses only when you choose Network Type to public. However, to improve the security in this scenario, you can still limit access to specific IP addresses using the host allowed list. Another option is to use AWS Transit Gateway, The transit gateway allows you to route traffic between various AWS services, including AWS VPN, which can be connected to another network. ScyllaDB can be attached to your transit gateway and you can route the traffic from your services. The purpose of the VPC peering option is to limit traffic to the trusted internal AWS infrastructure. You cannot use this from the outside unless you configure your peered VPC network to pass through traffic from another network (VPN, GCP etc.) based on your specific network configuration. VPC peering is not transitive, which comes with many limitations. More about transitivity and VPC peering features can be found here Amazon Web Services, Inc.: Network Gateway - AWS Transit Gateway - AWS How VPC peering connections work - Amazon Virtual Private Cloud @Alex_Ioannides: Thanks for your answer @Buff. For our usecase, its mostly about cost-savings because some of the traffic will be outside AWS but the majority of it isn’t. It’s the same region, same and different AZ. I can handle the security of public Network Type. If I opt for public, all the traffic will go through public internet though? @Patrick_Bossman: AWS controls the answer to this question. It was asked and answered here, and there are a couple useful links. https://repost.aws/questions/QUHhUaRUfTS3mR0AX7ubhUkg/is-traffic-between-two-ec2-public-instance-over-the-internet-or-on-aws-backbone-network Amazon Web Services, Inc.: Is traffic between two EC2 public instance over the internet or on AWS Backbone network ? --- ### Page: https://forum.scylladb.com/t/release-scylladb-monitoring-stack-4-11-0/5004 Title: [RELEASE] ScyllaDB monitoring Stack 4.11.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of ScyllaDB Monitoring Stack 4.11.0 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB based on Prometheus and Grafana. ScyllaDB Monitoring Sta… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-monitoring-stack-4-11-0/5004 ## Headings Structure: H1: [RELEASE] ScyllaDB monitoring Stack 4.11.0 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB monitoring Stack 4.11.0 H3: Related topics The ScyllaDB team is pleased to announce the release of ScyllaDB Monitoring Stack 4.11.0 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB based on Prometheus and Grafana. ScyllaDB Monitoring Stack 4.11.0 supports: This release includes multiple updates to the overview CQL, keyspace, alternator, manager, and advanced dashboards, including CPU Utilization, Alternator per-table panels, and others. Version updates for ScyllaDB Monitoring Stack 4.11.0 Grafana security updates 4.11.0 New Information in ScyllaDB Dashboards Overview Dashboard Change The Load stats panel indicates how busy the CPU is. ScyllaDB uses a priority group mechanism to protect User-Facing activity from background activity. Combining the load from User-Facing and background activity is misleading, as ScyllaDB utilizes periods of lower user activity for background operations like compaction and backup. Instead, the load panel will now show CPU usage by User-Facing activity only. The ‘Live’ column in the node table was hard to follow. Instead, in the case of a split-brain situation, the Status column will now display Split-Brain instead of Normal. Using a multi-partition Logged BATCH is considered bad practice. It provides weak atomicity guarantees and reduces performance, as the request is sent to a random node rather than the correct replica or shard. A new gauge and graph have been added to the CQL Optimization section to indicate when a user is using a multi-partition Logged BATCH. Keyspace Dashboard Change Multiple enhancements were made to the tables table, including adding units to values where applicable. The table name now serves as a quick navigation link to that table. New columns show the maximum reads and writes over the last 24 hours. A new Keyspace Table shows a summary of all existing keyspaces. It includes the number of tables and the number of active tables, meaning tables that were used in the last hour. It also displays aggregated information about the tables in each keyspace, including disk size, reads, writes, and latencies. The table is especially helpful in cases where the user maintains many unused keyspaces. Alternator Dashboard Change Starting with ScyllaDB upcoming version 2025.3, Alternator includes per-table metrics. The Alternator dashboard was updated to allow users to select a table and view the relevant operations specific to that table. Manager Dashboard Change ScyllaDB Manager 3.6 adds the 1-1 Restore procedure, which restores an entire cluster that exactly matches the backup cluster. New panels show the progress and remaining bytes of the restore procedure. Advanced Dashboard Change Stalls occur when a specific task runs continuously on the CPU without allowing other operations to execute. The stalls graph is used by ScyllaDB support and engineers to identify potential issues. The Keyspace Dashboard allows users to select which table to view. In cases where there are many tables, the dropdown can become too long to manage. To address this, the dropdown now shows the top 20 most active tables based on reads and writes. The updated Tables Table now also serves as a quick selection tool, allowing users to select a table by clicking its name. --- ### Page: https://forum.scylladb.com/t/release-scylladb-java-driver-4-19-0-1/5006 Title: [RELEASE] ScyllaDB Java Driver 4.19.0.1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces release of ScyllaDB Java Driver 4.19.0.1. Previous release post: 4.18.0.2 This is a joint release note for all versions released since that time (4.18.1.0, 4.19.0.0, 4.19.0.1). It is always… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-java-driver-4-19-0-1/5006 ## Headings Structure: H1: [RELEASE] ScyllaDB Java Driver 4.19.0.1 H1: 4.19.0.1 H2: 4.19.0.0 H2: 4.18.1.0 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Java Driver 4.19.0.1 H1: 4.19.0.1 H2: 4.19.0.0 H2: 4.18.1.0 H3: Related topics The ScyllaDB team announces release of ScyllaDB Java Driver 4.19.0.1. Previous release post: 4.18.0.2 This is a joint release note for all versions released since that time (4.18.1.0, 4.19.0.0, 4.19.0.1). It is always recommended to use the latest version. More detailed changelog can be found here. More detailed changelog can be found here. More detailed changelog can be found here As always the source code and full history is available on GitHub https://github.com/scylladb/java-driver/tree/scylla-4.x For how to use the driver please refer to the “Getting the driver” section. --- ### Page: https://forum.scylladb.com/t/release-gocql-v1-15-1/5008 Title: [RELEASE]: GoCQL v1.15.1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: Driver Release Summary: v1.15.1 Release link: Release v1.15.1 · scylladb/gocql · GitHub What’s Changed Fixes Fix usage of cassandra-specific system.peers_v2 table (#477) Fix possible panic on session.Close (#492) Imp… Language: en Canonical URL: https://forum.scylladb.com/t/release-gocql-v1-15-1/5008 ## Headings Structure: H1: [RELEASE]: GoCQL v1.15.1 H2: What’s Changed H3: Fixes H3: Improvements H3: Tests H2: New Contributors H3: Related topics ## Main Content: H1: [RELEASE]: GoCQL v1.15.1 H2: What’s Changed H3: Fixes H3: Improvements H3: Tests H2: New Contributors H3: Related topics Driver Release Summary: v1.15.1 Release link: Release v1.15.1 · scylladb/gocql · GitHub Full Changelog: Comparing v1.15.0...v1.15.1 · scylladb/gocql · GitHub --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-102-2025-07-25/5011 Title: Last week in scylla-cluster-tests.git master (issue #102; 2025-07-25) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 4efd286d…c8c40572 range are covered. There were 29 non-merge commits from 9 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-102-2025-07-25/5011 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #102; 2025-07-25) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #102; 2025-07-25) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 4efd286d…c8c40572 range are covered. There were 29 non-merge commits from 9 authors in that period. Some notable commits: A proof-of-concept support for scylla-cloud backend was implemented, enabling SCT to provision clusters in Scylla Cloud without depending on the siren-tests framework. This marks the initial step toward full Scylla Cloud support in SCT. Support for Z3 small instances was added. cassandra-stress was upgraded to version 3.18.1, built with Java 21, reinstating the counter_write table creation feature. Gemini has been upgraded to version 2.1, bringing significant improvements in stability, performance, and usability. A notable addition is a new statement logger, which operates independently of loader disk constraints by utilizing oracle’s disk. In case of data validation errors, all mutations applied to inconsistent partitions (for both test and Oracle clusters) are reported, aiding in detailed issue analysis. Additionally, users can now configure the ratio of inserts, updates, and deletes for greater control. Gemini is still undergoing a stabilization period. To help uncover potential issues, the gemini_seed parameter is now set to None by default, ensuring some runs use different data distributions and schemas each time. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/k8s-operator-fixing-an-unbalanced-deployment-on-different-availability-zones/5015 Title: K8s Operator, fixing an unbalanced deployment on different availability zones - Kubernetes Operator - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/k8s-operator-fixing-an-unbalanced-deployment-on-different-availability-zones/5015 ## Headings Structure: H1: K8s Operator, fixing an unbalanced deployment on different availability zones H3: Related topics ## Main Content: H1: K8s Operator, fixing an unbalanced deployment on different availability zones H3: Related topics Originally from the User Slack @Igor_Domrev: Hello, I have a question. We are currently running scylla in production on k8s , on spot instances. Using scylla k8s operator. It works like a clock for three years already. Our problem is that we accidentally created an un-balanced deployment of scylla on different availability zones. 1 node in the first AZ 2 nodes in the second AZ We want to fix that and move to a balanced deployment. However, because each pod has a PVC in its zone. We first need to create a new PVC in the new zone. we are currently on scylla 5.2 community. cluster of three nodes. What is the best strategy todo proceed ? If we just delete the disk to one of scylla pods, when it re-spawns and provisions a new and empty disk in the new zone. Will it just re-sync all the data from the other cluster nodes ? We can also create PVC in the new ZONE from backup and make the new pod to mount it. Will appreciate any advise on the matter. Thanks in advance @Maciej_Zimnoch: Use replace dead node procedure: https://operator.docs.scylladb.com/stable/resources/scyllaclusters/nodeoperations/replace-node.html Replacing a Scylla node | ScyllaDB Docs @Igor_Domrev: TY for the reply. will this work with operator 1.8 ? can i upgrade the operator to 1.17 without upgrading the scylla cluster ? @Maciej_Zimnoch: I don’t know, 1.8 is EOL. 5.2 is also EOL and is not supported by 1.17 --- ### Page: https://forum.scylladb.com/t/last-fortnight-in-scylladb-git-master-issue-286-2025-07-27/5016 Title: Last fortnight in scylladb.git master (issue #286; 2025-07-27) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last two weeks. Commits in the fdcaa9a7e79..a1d7722c6d6 range are covered. There were 138 non-merge commits from 28 authors in t… Language: en Canonical URL: https://forum.scylladb.com/t/last-fortnight-in-scylladb-git-master-issue-286-2025-07-27/5016 ## Headings Structure: H1: Last fortnight in scylladb.git master (issue #286; 2025-07-27) H3: Related topics ## Main Content: H1: Last fortnight in scylladb.git master (issue #286; 2025-07-27) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last two weeks. Commits in the fdcaa9a7e79..a1d7722c6d6 range are covered. There were 138 non-merge commits from 28 authors in that period. Some notable commits: A gap between node bootstrap and becoming eligible to be a Raft voter could expose clusters to loss-of-quorum. Nodes now become eligible to be voters earlier, closing the gap. Encryption using AWS Key Management Service (KMS) can now use externally-provided credentials. This makes recovery tasks simpler. Cleanup of Key Management Interoperability Protocol (KMIP) connections, used by encryption, is more robust, avoiding TLS errors. A crash with unreasonably sized alternator table name was fixed. Password authentication now offloads the password hash computation to a separate thread, avoiding a reactor stall. The small table repair optimization is used when bootstrapping or repairing tables with little data, like system_tracing tables. It now optimizes token range calculations for larger clusters. A deadlock when removing a failed node in a cluster with materialized views was fixed. Alternator, ScyllaDB’s implementation of the DynamoDB API, is now more careful to avoid large contiguous allocations during query/scan operations, as these can cause stalls. Typically secondary indexes are backed by materialized views. This does not hold for vector indexes, so we now avoid creating a view for them. The CQL binary protocol server now avoids exceptions when rejecting protocol versions it does not support. Unsupported protocol versions are common when drivers probe for server capability, and preventing exceptions improves robustness during connection storms. Background writes are now cancelled when a node is shut down, so shutdown completes more quickly. A large remote procedure call (RPC) hash table is now sized in advance to prevent resizing from causing latency jumps. Encryption at Rest can use the Azure Key Provider to manage keys. Progress reporting in the internal task manager (exposed by nodetool tasks) now uses consistent units. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/scylladb-crashes-with-iostream-errors/5017 Title: ScyllaDB crashes with iostream errors - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello, I’d like to share the logs from one of my ScyllaDB nodes. I’m not sure what’s causing the issue. This has happened for the second time. The first time it occurred, I was trying to shut down the node or one of o… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-crashes-with-iostream-errors/5017 ## Headings Structure: H1: ScyllaDB crashes with iostream errors H3: Related topics ## Main Content: H1: ScyllaDB crashes with iostream errors H3: Related topics Hello, I’d like to share the logs from one of my ScyllaDB nodes. I’m not sure what’s causing the issue. This has happened for the second time. The first time it occurred, I was trying to shut down the node or one of other nodes and it crashed. (I don’t remember clearly) Do you have any idea why this might be happening? Jul 21 09:23:20 scylladb-node-3 scylla[26690]: scylla: seastar/include/seastar/core/iostream.hh:420: seastar::output_stream::~output_stream() [CharType = char]: Assertion `!_end && !_zc_bufs && “Was this stream properly closed?”’ failed. Jul 21 09:23:20 scylladb-node-3 scylla[26690]: Aborting on shard 0, in scheduling group streaming. Jul 21 09:23:20 scylladb-node-3 scylla[26690]: Backtrace: Jul 21 09:23:20 scylladb-node-3 scylla[26690]: 0x59d7144 Jul 21 09:23:20 scylladb-node-3 scylla[26690]: 0x59965bb Jul 21 09:23:20 scylladb-node-3 scylla[26690]: 0x59cbf16 Jul 21 09:23:20 scylladb-node-3 scylla[26690]: /opt/scylladb/libreloc/libc.so.6+0x40cff Jul 21 09:23:20 scylladb-node-3 scylla[26690]: /opt/scylladb/libreloc/libc.so.6+0x994a3 Jul 21 09:23:20 scylladb-node-3 scylla[26690]: /opt/scylladb/libreloc/libc.so.6+0x40c4d Jul 21 09:23:20 scylladb-node-3 scylla[26690]: /opt/scylladb/libreloc/libc.so.6+0x28901 Jul 21 09:23:20 scylladb-node-3 scylla[26690]: /opt/scylladb/libreloc/libc.so.6+0x2881d Jul 21 09:23:20 scylladb-node-3 scylla[26690]: /opt/scylladb/libreloc/libc.so.6+0x38d86 Jul 21 09:23:20 scylladb-node-3 scylla[26690]: 0x1b485d4 Jul 21 09:23:20 scylladb-node-3 scylla[26690]: 0x5d8b600 Jul 21 09:23:20 scylladb-node-3 scylla[26690]: 0x59a691f Jul 21 09:23:20 scylladb-node-3 scylla[26690]: 0x59a7e8a Jul 21 09:23:20 scylladb-node-3 scylla[26690]: 0x59a9077 Jul 21 09:23:20 scylladb-node-3 scylla[26690]: 0x59a8428 Jul 21 09:23:20 scylladb-node-3 scylla[26690]: 0x5938773 Jul 21 09:23:20 scylladb-node-3 scylla[26690]: 0x5937ad3 Jul 21 09:23:20 scylladb-node-3 scylla[26690]: 0x13842a5 Jul 21 09:23:20 scylladb-node-3 scylla[26690]: 0x1385c60 Jul 21 09:23:20 scylladb-node-3 scylla[26690]: 0x13826c3 Jul 21 09:23:20 scylladb-node-3 scylla[26690]: /opt/scylladb/libreloc/libc.so.6+0x2a087 Jul 21 09:23:20 scylladb-node-3 scylla[26690]: /opt/scylladb/libreloc/libc.so.6+0x2a14a Jul 21 09:23:20 scylladb-node-3 scylla[26690]: 0x137fd44 Jul 21 09:24:18 scylladb-node-3 scylla[200784]: Scylla version 6.1.2-0.20240915.b60f9ef4c223 with build-id c713ac9e819492d7560aa3ad461c43cf404c977b starting … Jul 21 09:24:18 scylladb-node-3 scylla[200784]: command used: “/usr/bin/scylla --log-to-syslog 1 --log-to-stdout 0 --default-log-level info --network-stack posix --io-properties-file=/etc/scylla.d/io_properties.yaml --cpuset 0-3 --lock-memory=1” Installation details #ScyllaDB version: 6.1.2 #Cluster size: 3 Node (4 Core - 8 GB Ram) os (RHEL/CentOS/Ubuntu/AWS AMI): Rocky Linux 9 6.2 is not supported anymore, please update. In general, when posting backtraces, please resolve them via Scylla relocatable packages For posterity, here is the resolved backtrace: I am not familiar with this crash, but it is very likely that we have fixed it since, 6.1 is quite an old release. Please upgrade and see if it reproduces with a supported version, if so we will investigate. --- ### Page: https://forum.scylladb.com/t/how-can-i-limit-a-scylladb-cluster-to-use-less-cpu-and-memory/5018 Title: How can I limit a ScyllaDB cluster to use less CPU and memory? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Where can I set this? Language: en Canonical URL: https://forum.scylladb.com/t/how-can-i-limit-a-scylladb-cluster-to-use-less-cpu-and-memory/5018 ## Headings Structure: H1: How can I limit a ScyllaDB cluster to use less CPU and memory? H3: Related topics ## Main Content: H1: How can I limit a ScyllaDB cluster to use less CPU and memory? H3: Related topics Where can I set this? The --smp option (for instance, --smp 2 ) will restrict Scylla to fewer CPUs. It will still use 100 % of those CPUs, but at least won’t take your system out completely. An analogous option exists for memory: -m . Use these flags with care, ScyllaDB is optimized to use all the resources of the machine and not to be hosted on a machine that it running a different workload. Also see the Best Practices for Running on Docker documentation for more information about configuring resource limits. --- ### Page: https://forum.scylladb.com/t/trouble-installing-scylla/5026 Title: Trouble installing Scylla - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version:2025.2 #Cluster size:3 os (RHEL/CentOS/Ubuntu/AWS AMI):Ubuntu 22.04.5 (but also tried 24.04.2) Hi, I am trying to install Scylla 2025.2 on Ubuntu 22.04 LTS via the Web Installer… Language: en Canonical URL: https://forum.scylladb.com/t/trouble-installing-scylla/5026 ## Headings Structure: H1: Trouble installing Scylla H3: Related topics ## Main Content: H1: Trouble installing Scylla H3: Related topics Installation details #ScyllaDB version:2025.2 #Cluster size:3 os (RHEL/CentOS/Ubuntu/AWS AMI):Ubuntu 22.04.5 (but also tried 24.04.2) Hi, I am trying to install Scylla 2025.2 on Ubuntu 22.04 LTS via the Web Installer but keep getting this error: Any ideas about this ? maybe this question is rather uncommon, but I was able to get around it by installing the system packeges. On Ubuntu this gave me 2025.2 and it worked out of the box. Now I have a different Question about latencies, but post in another topic. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-103-2025-08-01/5028 Title: Last week in scylla-cluster-tests.git master (issue #103; 2025-08-01) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 249298b2…43acc451 range are covered. There were 17 non-merge commits from 8 authors in tha… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-103-2025-08-01/5028 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #103; 2025-08-01) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #103; 2025-08-01) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 249298b2…43acc451 range are covered. There were 17 non-merge commits from 8 authors in that period. Some notable commits: RackAwarePolicy testing saw improvements: a new scylla-bench-based test was added and tier1 triggers now use scylla-bench with rackaware validation, replacing one test with java driver/cassandra-stress. If an error occurs during resource provisioning, Argus status is now set to “TEST ERROR”. The clean-resources command in sct.py now supports cleaning up SCT runners. To use it, add --clean-runners to clean-resources hydra command. Latte container configuration was optimized for performance by switching to host networking and disabling secure computing mode. Latte was also bumped to version 0.32.0, bringing support for ‘SERIAL’ and ‘LOCAL_SERIAL’ consistency types (for testing LWT) and updating to scylla-rust-driver 1.3.1. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-java-driver-3-11-5-7/5029 Title: [RELEASE] ScyllaDB Java Driver 3.11.5.7 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces release of ScyllaDB Java Driver 3.11.5.7. This is a late release post for both versions 3.11.5.6 and 3.11.5.7. If possible it is advised to use the newer major version of the driver, like 4.… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-java-driver-3-11-5-7/5029 ## Headings Structure: H1: [RELEASE] ScyllaDB Java Driver 3.11.5.7 H3: 3.11.5.7 H3: 3.11.5.6 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Java Driver 3.11.5.7 H3: 3.11.5.7 H3: 3.11.5.6 H3: Related topics The ScyllaDB team announces release of ScyllaDB Java Driver 3.11.5.7. This is a late release post for both versions 3.11.5.6 and 3.11.5.7. If possible it is advised to use the newer major version of the driver, like 4.19.0.1. Previous release post: 3.11.5.5 Detailed changelog is available here. Detailed changelog is available here. Latest code, tagged versions and full history can be checked out on GitHub: https://github.com/scylladb/java-driver/tree/scylla-3.x Check out “Getting the driver” section for how to incorporate the driver in your project: --- ### Page: https://forum.scylladb.com/t/real-time-machine-learning-with-scylladb-as-a-feature-store-app-examples-and-use-cases/5030 Title: Real time machine learning with ScyllaDB as a feature store, app examples and use cases - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/real-time-machine-learning-with-scylladb-as-a-feature-store-app-examples-and-use-cases/5030 ## Headings Structure: H1: Real time machine learning with ScyllaDB as a feature store, app examples and use cases H3: Related topics ## Main Content: H1: Real time machine learning with ScyllaDB as a feature store, app examples and use cases H3: Related topics Originally from the User Slack @Attila_Tóth: ScyllaDB + Feature Store Hi All, recently we published quite a few blog posts and tutorials to help you build feature stores with ScyllaDB, here are some quick links: • Real-Time Machine Learning with ScyllaDB as a Feature Store (blog post) • ScyllaDB as an online feature store (YouTube) • Used car price prediction app (GitHub) Let me know if you are interested in more examples or have questions! GitHub: scylladb-feature-store/used_cars at main · scylladb/scylladb-feature-store --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-287-2025-08-03/5031 Title: Last week in scylladb.git master (issue #287; 2025-08-03) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the a1d7722c6d..1c25aa891b range are covered. There were 111 non-merge commits from 23 authors in that per… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-287-2025-08-03/5031 ## Headings Structure: H1: Last week in scylladb.git master (issue #287; 2025-08-03) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #287; 2025-08-03) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the a1d7722c6d..1c25aa891b range are covered. There were 111 non-merge commits from 23 authors in that period. Some notable commits: The main data structure holding vnode tokens was changed to avoid large allocations, which can produce stalls on very large clusters. After replacing a node, the load balancer will immediately check available space, to avoid the up to 60 second delay from regularly scheduled checks that can delay load balancing. Tablets now support LightWeight Transactions. Paxos state is stored in a side table allocated on demand for each table using LWT, rather than system.paxos. Compaction start/end messages were demoted to DEBUG level, as they were deemed too noisy. Repair will send smaller messages for partition differences with many small rows (like tobmstone-only rows), reducing stalls. The CQL port listener will now first half-close the receiving side of the connection, drain all responses, then close the sending side. This prevents queries from being lost during graceful shutdown. The writing side of the sstable Trie index writer was merged. Trie indexes will still not be written. When the coordinator for a LWT transaction cannot find a co-located shard (because the coordinator was not also a replica), it will now allocate a random (but consistent) shard for processing, rather than always falling back to shard 0 and overloading it. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-2025-2-1/5032 Title: [RELEASE] ScyllaDB 2025.2.1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB 2025.2.1, a bug-fix production-ready patch release for ScyllaDB 2025.2 Feature Release. Related Links Get ScyllaDB 2025.2 Upgrade from ScyllaDB 2025.1 to ScyllaDB 2025.2 Sub… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-2025-2-1/5032 ## Headings Structure: H1: [RELEASE] ScyllaDB 2025.2.1 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 2025.2.1 H3: Related topics The ScyllaDB team announces ScyllaDB 2025.2.1, a bug-fix production-ready patch release for ScyllaDB 2025.2 Feature Release. Upgrade from ScyllaDB 2025.1 to ScyllaDB 2025.2 The following issues are fixed in this release: non-full row keys cause mis-parsing of sstables #24489 Fix stability issue with KMS caching testing #24574 Assertion `_mt._flushed_memory <= _mt.occupancy().total_space()’ failed during bulk ingestion #21413 Avoid killing a node when reading from sstables #20845 Coredump right after seed node decommission #23911 Potential data race in utils::alien_worker #24751 Don’t start maintenance auth service if not enabled #24528 Empty clustering keys generated and spread freely in the system #24506 Enabling a disabled compaction manager throws assertion #24504 Gossip: Failed to add server: “No ip address for … when one is expected” #23407 Joining node enters synchronize state after joining group0, which races with streaming and causes join to fail #23536 Node shutdown is stuck waiting for batchlog manager drain #24599 paxos_response_handler timeout logging #24591 RBNO[small_table_optimization=true]: memory amplification by a factor that is equals to number of peer nodes | test_add_many_nodes_under_load: node40 runs out of memory #22244 repair: to_repair_rows_on_wire stalls destroying input list #24725 Unregister raft_topology_get_cmd_status on shutdown #24910 big_decimal incorrectly parses decimals with a large negative exponent. #24581 LWT: possible stale data read #24630. The issue is limited to rare cases where Replication Factor >= 4 , the most recent commit is stored (“learned”) only on one replica, other replicas must have failed to receive an update last time LWT was executed for this partition key. tablets: stop storage group on deallocation #24857 #24828. This can cause errors when a node is restarted during tablet cleanup and other use cases. repair: tablet repair does not exclude with global topology operations #24195. The fix delays repair until topology updates are completed. 2025.2.0 release is missing some commits: nodetool refresh --skip-cleanup --skip-reshape doesn’t work. #24913. This fix enable the upcoming vNode 1-1 restore feature available in Scylla Manager 3.6 Add REST API to the storage service to check topology cmd status to improve raft topology debugability #24860: see /storage_service/raft_topology/cmd_rpc_status in API reference docs. generic_server: uninitialized_connections_semaphore_cpu_concurrency config is not live updatable #24557 nodetool: repair: skip tablet keyspaces #24040 nodetool backup: add support for the move_files options, used by native backup. An experimental feature in 2025.2 #24372 scylla_sysconfig_setup prints warnings in 2025.2.0 #24915 Docker: Container image lacks PS command used by Scylla scripts #24827 --- ### Page: https://forum.scylladb.com/t/read-errors-in-a-cluster-io-setup-failure-disk-performance-and-optimizations/5033 Title: Read errors in a cluster, IO setup failure, disk performance and optimizations - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/read-errors-in-a-cluster-io-setup-failure-disk-performance-and-optimizations/5033 ## Headings Structure: H1: Read errors in a cluster, IO setup failure, disk performance and optimizations H3: Related topics ## Main Content: H1: Read errors in a cluster, IO setup failure, disk performance and optimizations H3: Related topics Originally from the User Slack @Erik-Jan_van_de_WalErik-Jan_van_de_Wal**:** Hi all, I’m running into intermittent read errors under load on a 3-node Scylla 2025.1 cluster: Jul 31 16:01:02 learner01 taskset[121958]: Error read_moving_average_data: Failed to fetch the first page of the result: Database returned an error: Not enough nodes responded to the read request in time to satisfy required consistency level (consistency: Quorum, received: 1, required: 2, data_present: false) Cluster hardware: (servers are about 8 years old) Node 1: 40 cores / 164 GB RAM / 20 TB disk (DELL PowerEdge R720) Node 2: 20 cores / 164 GB RAM / 20 TB disk (DELL PowerEdge R720) Node 3: 56 cores / 64 GB RAM / 20 TB disk (Supermicro) The servers are in the same rack, and the machine that runs the application that uses Scylla as well. Keyspace: NetworkTopologyStrategy, RF = 3 Failing Table Schema: Query function (Rust): Load Profile: • App is a simulator running 8–10 jobs, each with max 15 sub-simulations in parallel • Average: ~~7k ops/sec, peaks at ~~10k ops/sec • Errors occur mostly under high read pressure Identified bottlenecks: CPU, disk Node 3 Diagnostics (Problem Node): total_successful_reads: 1,469,616 total_failed_reads: 3,234,429 reads_enqueued: 3,655,955 reads_queued_count: 3,610,289 permits: 100/100 Disk layout: /dev/sda3 (LVM): Note: scylla_io_setup failed on node 3 during install; I bybassed this by explicitly only test /var/lib/scylla/data and /var/lib/scylla/commitlog (because commitlog, hints, view_hints, saved_caches are the same disk) The problem mostly looks like node3, it looks like it can not keep up and node1 and node2 need to pick up the work for node3 where eventually node1 and node2 also crash. See the screenshot (the blue line is node3) Any help that push me in the right direction is helpful, I have exhausted everything I know so far. I do understand that node3 is the weakest link, and of course the first thing that needs to be solved is the scylla_io_setup, but I am stuck here This is the error from scylla_io_setup coconut@scylla-03:~$ sudo scylla_io_setup [sudo] password for coconut: tuning /sys/devices/pci0000:80/0000:80:02.0/0000:81:00.0/host0/target0:2:1/0:2:1:0/block/sdb/sdb1 tuning /sys/devices/pci0000:80/0000:80:02.0/0000:81:00.0/host0/target0:2:1/0:2:1:0/block/sdb already tuned: /sys/devices/pci0000:80/0000:80:02.0/0000:81:00.0/host0/target0:2:1/0:2:1:0/block/sdb/queue/nomerges tuning /sys/devices/virtual/block/dm-1 tuning: /sys/devices/virtual/block/dm-1/queue/nomerges 2 tuning /sys/devices/pci0000:80/0000:80:02.0/0000:81:00.0/host0/target0:2:0/0:2:0:0/block/sda/sda3 tuning /sys/devices/pci0000:80/0000:80:02.0/0000:81:00.0/host0/target0:2:0/0:2:0:0/block/sda already tuned: /sys/devices/pci0000:80/0000:80:02.0/0000:81:00.0/host0/target0:2:0/0:2:0:0/block/sda/queue/nomerges tuning /sys/devices/virtual/block/dm-2 tuning: /sys/devices/virtual/block/dm-2/queue/nomerges 2 tuning /sys/devices/pci0000:80/0000:80:02.0/0000:81:00.0/host0/target0:2:0/0:2:0:0/block/sda/sda3 tuning /sys/devices/virtual/block/dm-4 tuning: /sys/devices/virtual/block/dm-4/queue/nomerges 2 tuning /sys/devices/pci0000:80/0000:80:02.0/0000:81:00.0/host0/target0:2:0/0:2:0:0/block/sda/sda3 tuning /sys/devices/virtual/block/dm-5 tuning: /sys/devices/virtual/block/dm-5/queue/nomerges 2 tuning /sys/devices/pci0000:80/0000:80:02.0/0000:81:00.0/host0/target0:2:0/0:2:0:0/block/sda/sda3 INFO 2025-07-31 20:34:48,499 seastar - Reactor backend: io_uring INFO 2025-07-31 20:34:49,051 [shard 0:main] iotune - /var/lib/scylla/view_hints passed sanity checks INFO 2025-07-31 20:34:49,052 [shard 0:main] iotune - Disk parameters: max_iodepth=916 disks_per_array=1 minimum_io_size=4096 INFO 2025-07-31 20:34:49,056 [shard 0:main] iotune - Filesystem parameters: read alignment 512, write alignment 4096 Starting Evaluation. This may take a while… Measuring sequential write bandwidth: ERROR 2025-07-31 20:36:12,578 [shard 0:main] seastar - Exiting on unhandled exception: std::system_error (error system:28, No space left on device) ERROR:root:Command ‘[’/usr/bin/iotune’, ‘–format’, ‘envfile’, ‘–options-file’, ‘/etc/scylla.d/io.conf’, ‘–properties-file’, ‘/etc/scylla.d/io_properties.yaml’, ‘–evaluation-directory’, ‘/var/lib/scylla/data’, ‘–evaluation-directory’, ‘/var/lib/scylla/commitlog’, ‘–evaluation-directory’, ‘/var/lib/scylla/hints’, ‘–evaluation-directory’, ‘/var/lib/scylla/view_hints’, ‘–evaluation-directory’, ‘/var/lib/scylla/saved_caches’, ‘–cpuset’, ‘1,2,3,4,5,6,7,8,9,10,11,12,13,15,16,17,18,19,20,21,22,23,24,25,26,27,29,30,31,32,33,34,35,36,37,38,39,40,41,43,44,45,46,47,48,49,50,51,52,53,54,55’]’ returned non-zero exit status 1. ERROR:root:[‘/var/lib/scylla/data’, ‘/var/lib/scylla/commitlog’, ‘/var/lib/scylla/hints’, ‘/var/lib/scylla/view_hints’, ‘/var/lib/scylla/saved_caches’] did not pass validation tests, it may not be on XFS and/or has limited disk space. This is a non-supported setup, and performance is expected to be very bad. For better performance, placing your data on XFS-formatted directories is required. To override this error, enable developer mode as follow: sudo /opt/scylladb/scripts/scylla_dev_mode_setup --developer-mode 1 It is only on the Scylla 3 node, it has the same harddrive capacity. I have reformatted the disks many times already The partitions are formatted as XFS └─sda3 LVM2_member LVM2 001 Ct4FCa-TwgT-4DjE-7dBR-6cgn-0b0B-MomIEC ├─ubuntu–vg-ubuntu–lv ext4 1.0 60fe3131-3982-44ab-8698-b3895c152ca0 79.5G 14% / ├─ubuntu–vg-scylla_commitlog xfs addf9c73-d96a-43f3-80ab-cf7b40de31c2 20.8G 58% /var/lib/scylla/commitlog ├─ubuntu–vg-scylla_hints xfs 8249b669-f41e-4b2e-aabd-79fefe11a55f 24.4G 2% /var/lib/scylla/hints ├─ubuntu–vg-scylla_caches xfs 03a5ab30-3099-4b9a-b7b9-3bce7b6aca39 ├─ubuntu–vg-scylla_view_hints xfs c8581434-4c36-4fcd-b287-ba9238f7e101 9.7G 2% /var/lib/scylla/view_hints └─ubuntu–vg-scylla_saved_caches xfs 23c34a60-8c32-4f10-9478-89616cc79a96 24.4G 2% /var/lib/scylla/saved_caches sdb └─sdb1 xfs b00edbbb-08b8-4e7e-a77c-dfd81755f618 17.8T 2% /var/@Felipe_Cardeneti_Mendesib/scylla/data @Felipe_Cardeneti_Mendes**:** Hm… there are quite a few problems here. Node 3 has a smaller memory per vCPU than the rest of the nodes. Node1 ~~ 4GB/vCPU, Node2 ~~ 8GB/vCPU, Node 3 ~1GB/vCPU Your iotune (what scylla_io_setup calls) fails due to: Given the sanity check is on view_hints I suppose that’s the directory it failed on. But why do you have a VG for each directory, rather than simply having a single /var/lib/scylla ? If these are slow disks, then maybe keep just the commitlog. The 2% utilization under /var/lib/scylla/data also looks strange - is it all of your data, or is it simply not persisting any data? If the data is different than from the other nodes, then the read could be failing because it could be a large partiti@Erik-Jan_van_de_Waln trying to read-repair to that node? @Erik-Jan_van_de_Wal**:** I though it would be wise to keep my actual data on a seperate disk (the 20TB) and keep the other data on@Felipe_Cardeneti_Mendesthe primary (OS) disk. As recommended in the config file @Felipe_Cardeneti_Mendes**:** if at all possible,@Erik-Jan_van_de_WalI’d suggest you try to keep similar shard count and memory ratios per vCPU. Hmm @Erik-Jan_van_de_Wal**:** the 2% utlization is probably correct. I only have about 5 millions records I need to work with the hardware I have (wooh@Felipe_Cardeneti_Mendeso startups). Would it be better if I virtualize (docker) Scylla per server to make them identical? @Felipe_Cardeneti_Mendes**:** > I though it would be wise to keep my actual data on a seperate disk (the 20TB) and keep the other data on the primary (OS) disk. As recommended in the config file Yeah, if these disks are NVMes you should be good with just a single mount. The separate filesystem recommendation for the commitlog is mai@Erik-Jan_van_de_Wally when the disks are known to be slow, which is no@Felipe_Cardeneti_Mendes the case with NVMes (hints are rarely used, and we never use caches) @Erik-Jan_van_de_Wal**:** Right. The disks are magnetic disks. Not SSDs @Felipe_Cardeneti_Mendes**:** > Would it be better if I virtualize (docker) Scylla per server to make them identical? Or just reduce the shard count (/etc/scylla.d/cpuset.conf). You can also reduce memory in /etc/scylla.d/memory.conf. We often use ~8GB per vCPU. 4GB works, but I wouldn’t go lower than that. The more memor@Erik-Jan_van_de_Wal, the more caching space you have, plus headroom for metadata. oh ok, with magnetic, definitely keep the commitlog, b@Felipe_Cardeneti_Mendest drop the rest, you don’t need these LVs. @Erik-Jan_van_de_Wal**:** got it! so I take the weakest link, and make sure that the other servers are locked to that config too 8GB/vCPU @Felipe_Cardeneti_Mendes**:** also, increase the commitlog LV size to about the memory size you have, otherwise recycling their segments may be too frequent. yeah, 8 vCPUs should work fine for that amount of RAM one last @Erik-Jan_van_de_Walote before you move forward - since you’re on 2025.1 - IF you are using tablets, then you’d @Felipe_Cardeneti_Mendesee@avi to replace the node. Or j@Erik-Jan_van_de_Walst recreate the clu@avit@Felipe_Cardeneti_Mendesr with the right configs if you are ok with losing the data. @Erik-Jan_van_de_Wal**:** Thanks@Felipe_Cardeneti_Mendes I opted out for tablets so I should be good with reconfiguring them one by one @Felipe_Cardeneti_Mendes**:** @avi drops a tear silently @Erik-Jan_van_de_Wal**:** aaaaaw i’m sorry @avi @Felipe_Cardeneti_Mendes could I@Erik-Jan_van_de_Walalso pass @Felipe_Cardeneti_Mendeshe cpu and mem config via SCYLLA_ARGS in /etc/default/scylla-server? @Felipe_Cardeneti_Mendes**:** yes all these are sourced by the systemd scylla-server unit, so ensure you dont have duplicates so the cmdline isn’t messed up. But other than that, you have the flexibility to adjust it as you see better fit. @Erik-Jan_van_de_Wal**:** Thanks! @Felipe_Cardeneti_Mendes so I did what you suggested. • Removed the LVMS • Increased the commit partition • Configured each server to use 8CPU/64GB RAM And I still get the following error on only one (the 3rd) server. As you can s@Erik-Jan_van_de_Wale in the screenshot blue is completly flatlined. The problem so far is that when I have a low req/s everything seems to go well, but when I got to 5K+ req/s stuff starts to break. Any other advice? I tried to lower the num_tokens to 128, but that doesnt change anything. The load is also low @Felipe_Cardeneti_Mendes**:** what’s the io_properties.yaml values compared to other nodes? @Erik-Jan_van_de_Wal**:** server_1: coconut@scylla-01:/etc/scylla.d$ cat io_properties.yaml disks: coconut@scylla-02:/etc/scylla.d$ cat io_properties.yaml disks: server_3 (problem server): coconut@scylla-03:/etc/scylla.d$ cat io_properties.yaml disks: @Felipe_Cardeneti_Mendes**:** Yeah, it’s definitely disk. The semaphore dump earlier shows all read permits are consumed (100/100) with reads in progress. Once the queue gets full, other requests accumulate until they timeout or we kill it due to overload. Is it possible most reads from n1/n2 are served from cache? You can check the misses/hits in the Cache section within the Detailed panel. It is probably wise to see the IOPS metrics on the Advanced dashboard for your classes (memtable, compaction, sl:default) and compare these nodes. My wild guess is that your reads on N3 are saturating its IOPS, and the course of action (despite using faster disks) would be to tune the system to use a less IOPS as possible. This would involve chunk_length_kb (https://www.scylladb.com/2017/08/01/compression-chunk-sizes-scylla/) depending on your average read size, bloom filters, and (at the extreme) SSTable summary ratios (https://medium.com/agoda-engineering/exploring-scylla-disk-performance-and-optimizations-at-agoda-65a6dcdd6fe7). You can also play with different Seastar settings - like io-latency-goal-ms (https://www.scylladb.com/2023/07/17/top-mistakes-with-scylladb-storage/): Lastly, triple check the number of CPU/shards stand as you configured them - /usr/lib/scylla/seastar-cpu-map.sh -n scylla Medium: Exploring Scylla Disk Performance and Optimizations at Agoda @Erik-Jan_van_de_Wal**:** Thanks. I will research further and dive deeper in the documentation. I think this is a case of “never let a software engineer handle database engineering”. We can run our tests and potentially move to live. I’ve built a L1 cache between the application and data and it seems to hold (I also removed N3 from the cluster and put RF=2, not ideal, but for now we need to work with what we’ve got) Thanks again for the quick response and helpful insights @Felipe_Cardeneti_Mendes! --- ### Page: https://forum.scylladb.com/t/how-to-enable-https-for-all-components-of-scylladb-monitoring/5034 Title: How to enable https for all components of scylladb monitoring? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: ScyllaDB Monitoring 4.9.4 #Cluster size:8 os (RHEL/CentOS/Ubuntu/AWS AMI): RHEL Hello, I would like to enable https for components in ScyllaDB monitoring. I can see there are p… Language: en Canonical URL: https://forum.scylladb.com/t/how-to-enable-https-for-all-components-of-scylladb-monitoring/5034 ## Headings Structure: H1: How to enable https for all components of scylladb monitoring? H2: Grafana HTTPS Configuration H2: Prometheus and Alertmanager HTTPS H2: Configuration Files H3: Related topics ## Main Content: H1: How to enable https for all components of scylladb monitoring? H2: Grafana HTTPS Configuration H2: Prometheus and Alertmanager HTTPS H2: Configuration Files H3: Related topics Installation details #ScyllaDB version: ScyllaDB Monitoring 4.9.4 #Cluster size:8 os (RHEL/CentOS/Ubuntu/AWS AMI): RHEL Hello, I would like to enable https for components in ScyllaDB monitoring. I can see there are parameters to the start-all.sh script, but there is also a lot of “http” directly in the code. Just wanted to know if there was a cheatcode to enable https to all components of scylladb monitoring or if I need to change directly file content like this one grafana/datasource.yml: Asking @Amnon_Heiman to see if he can assist with an answer. To enable HTTPS for all components of ScyllaDB Monitoring (Grafana, Prometheus, Alertmanager, etc.), there is no universal “cheatcode” that automatically switches all URLs in configs from HTTP to HTTPS. You typically need to configure HTTPS explicitly for each component and update relevant configuration files accordingly. Grafana supports HTTPS natively via its grafana.ini config file. You need to provide Grafana with an SSL certificate and key (either self-signed or CA-signed). Set the following in grafana.ini: [server] protocol = https cert_file = /path/to/cert.pem cert_key = /path/to/key.pem Then restart Grafana for changes to take effect. This encrypts the web UI connections.​ Prometheus and Alertmanager do not have built-in HTTPS for their web UIs by default. A common approach is to put them behind a reverse proxy (e.g., Nginx, HAProxy, Traefik) configured with SSL termination. Update Alertmanager’s URL in Grafana’s datasource to use HTTPS via the proxy. Similarly, use HTTPS URLs for Prometheus targets accessed by Grafana and alertmanager clients.​ For instance, in grafana/datasource.yml, replace: url: http://DB_ADDRESS url: https://DB_ADDRESS Similarly update Alertmanager URLs to use HTTPS. Note that some URLs can be hardcoded in configs or scripts, so you may need to manually edit those files. --- ### Page: https://forum.scylladb.com/t/release-added-support-for-gcp-z3-instance-family-and-n2d-highmem-us-east4-in-scylladb-cloud/5035 Title: [RELEASE] Added Support for GCP z3 Instance Family and n2d-highmem (us-east4) in ScyllaDB Cloud - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: We’re excited to announce two enhancements to ScyllaDB Cloud: :white_check_mark: New z3 Instance Family Support ScyllaDB Cloud now supports the z3 instance family, giving you access to high-memory, high-performance SSD… Language: en Canonical URL: https://forum.scylladb.com/t/release-added-support-for-gcp-z3-instance-family-and-n2d-highmem-us-east4-in-scylladb-cloud/5035 ## Headings Structure: H1: [RELEASE] Added Support for GCP z3 Instance Family and n2d-highmem (us-east4) in ScyllaDB Cloud H3: Related topics ## Main Content: H1: [RELEASE] Added Support for GCP z3 Instance Family and n2d-highmem (us-east4) in ScyllaDB Cloud H3: Related topics We’re excited to announce two enhancements to ScyllaDB Cloud: New z3 Instance Family Support ScyllaDB Cloud now supports the z3 instance family, giving you access to high-memory, high-performance SSD-backed instances - ideal for I/O-intensive workloads. Available Instance Types: z3-highmem-8-highlssd z3-highmem-16-highlssd z3-highmem-22-highlssd z3-highmem-32-highlssd z3-highmem-44-highlssd us-east1 (South Carolina) asia-southeast1 (Singapore) Expanded Support for n2d-highmem Instances We’ve also expanded regional support for the n2d-highmem instance types — now available in us-east4 (Northern Virginia) in addition to existing regions. To try these out, create a new cluster in ScyllaDB Cloud and choose your preferred region and instance type. We’d love to hear from you — let us know how these new options perform for your workloads! --- ### Page: https://forum.scylladb.com/t/inter-dc-replication-latency-with-strong-consistency/5036 Title: Inter DC Replication Latency with strong consistency - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details Small cluster, 2 DCs , 3 Nodes in each DC, RF 2 #ScyllaDB version: 2025.2 #Cluster size: 2 x 3 os (RHEL/CentOS/Ubuntu/AWS AMI): Ubuntu 22.04 LTS Hi, assuming there are 2 DCs with 100ms latency … Language: en Canonical URL: https://forum.scylladb.com/t/inter-dc-replication-latency-with-strong-consistency/5036 ## Headings Structure: H1: Inter DC Replication Latency with strong consistency H3: Related topics ## Main Content: H1: Inter DC Replication Latency with strong consistency H3: Related topics Installation details Small cluster, 2 DCs , 3 Nodes in each DC, RF 2 #ScyllaDB version: 2025.2 #Cluster size: 2 x 3 os (RHEL/CentOS/Ubuntu/AWS AMI): Ubuntu 22.04 LTS assuming there are 2 DCs with 100ms latency amongst them and 3 nodes per DC. All nodes are created equal (same hardware, connectivity, load). The latency inside each DC is 0.5ms Lets say for simplicity the given link latencies are constant and all nodes are healthy, what would be the expected replication times in the whole cluster, especially between the DCs ? Cases of interest are EACH_QUORUM and ALL . Would be glad for any answer or hint or even wild speculation Well theres a reason for asking .. A while back on a similar project Mongo was giving the team serious headaches with casual replication backlogs going from a few seconds to sometimes even minutes in the same datacenter and even more on remote. (setup was very similar: 2 DCs with 3 nodes in each for a replica set ( 1 master, 2 replicas ), moderate load, healthy nodes) - Never again! Just want to take care not getting into similar hot water with Scylla When you say RF=2 do you mean one in each DC or 2 in each DC? Either way, if you are using quorum or each quorum, Scylla will write to both DC’s before acknowledging to the application. So it’s not any replication lag - the write itself will be slower. Also note - Scylla does not have a binary log. It attempts every write to all targets all replicas at time of execution. If some replicas can’t be written but aren’t required to meet consistency, if hints are on, and hint will be generated (to a point). ScyllaDB uses repair, read repair as anti-entropy to re-establish consistency, rather then playing a replication backlog as you’re imagining. I recommend taking a look at Scylla U for consistency levels and anti-entropy. thanks a lot for your response ! RF=2 was a typo on my side. We go with RF=3 so a quorum can be reached. Sorry for the confusion. So its 2 DCs, 3 nodes each with replication factor set to 3 in each datacenter so we can reach LOCAL_QUORUM quite fast (needed most of the times). So presuming that everything runs ok, a write will attempt to propagate immediately to all nodes with a replica (in our case all 6) and the write attempts on the remote DC will start immediately after arrival there as if the write command has been issued there ? RF=3 in both DCs, great! With LOCAL_QUORUM, 2 out of 3 replicas in the local DC need to ACK the write to be considered successful. Writes are sent to all replicas immediately. For remote DC, there is an optimization to send to a coordinator on the remote and it then replicates to the other two replicas. In the LOCAL DC, if replica 3 was down, and a read cones along at LOCAL_QUORUM, and replica 3 is one of the two returning results - the other replica - 1 and 2 have the current data. The difference would be discovered on the LOCAL_QUORUM read and read repair would be triggered to fix up replica 3. Or standard regular repair would catch it up. The latest replicas would be caught up if writes failed should be via regular repairs. Thanks for the clarification! I get the optimization with the coordinator in the remote DC. Data travels only once over the WAN. Only one small question: Is this coordinator a specific node or randomized ? What happens if the coordinator does not respond ? There can be write failures. If you have hints enabled, then it could generate hints (there are limits to this) and they can get replayed. The backstop to restore full consistency is regular repairs. --- ### Page: https://forum.scylladb.com/t/release-scylladb-c-driver-v3-22-0-1/5038 Title: [RELEASE] ScyllaDB C# Driver v3.22.0.1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce first release of ScyllaDB C# Driver v3.22.0.1, a fork of DataStax C# Driver v3.22.0, optimized for Scylla. It contains important Scylla-specific features and optimizations: Sh… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-c-driver-v3-22-0-1/5038 ## Headings Structure: H1: [RELEASE] ScyllaDB C# Driver v3.22.0.1 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB C# Driver v3.22.0.1 H3: Related topics The ScyllaDB team is pleased to announce first release of ScyllaDB C# Driver v3.22.0.1, a fork of DataStax C# Driver v3.22.0, optimized for Scylla. It contains important Scylla-specific features and optimizations: LWT prepared statements metadata mark --- ### Page: https://forum.scylladb.com/t/scylladb-node-spikes-at-00h00-utc/5039 Title: ScyllaDB node spikes at 00h00 UTC - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 5.0.5-0.20221009 #Cluster size: 6 nodes (3 - us-east-1; 3 - us-east-2) Replication factor 3 os (debian10-base-amd64-202408061511): Hello! We randomly have ScyllaDB load/latenc… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-node-spikes-at-00h00-utc/5039 ## Headings Structure: H1: ScyllaDB node spikes at 00h00 UTC H3: Related topics ## Main Content: H1: ScyllaDB node spikes at 00h00 UTC H3: Related topics Installation details #ScyllaDB version: 5.0.5-0.20221009 #Cluster size: 6 nodes (3 - us-east-1; 3 - us-east-2) Replication factor 3 os (debian10-base-amd64-202408061511): We randomly have ScyllaDB load/latency spikes in some nodes every day shortly after 00h00 UTC. This lasts for about 30 seconds, in each node, and during that period, all the queries to node time out, it looks like that node is not available. In the metrics of the EC2 instance, we see this load spike: The compactions are running during the day and nothing out of normal is running at this time. We also don’t have a spike in throughput, in fact is decreasing at this time of the day. We don’t have any scheduled jobs, backup, repair, etc, scheduled for this time frame. Do you have any idea why this might be occurring? We are running out of ideas… 5.0.5 has reached end-of-life ages ago. The problem you’re describing is likely due to fstrim running to discard unused disk space. It has been replaced (in 5.1 timeframe) by online discard (see dist: scylla_raid_setup: mount XFS with online discard · scylladb/scylladb@a19d00e · GitHub). Note that upgrading won’t transition the nodes to online discard, you either have to apply the changes manually, or bootstrap new nodes and decommission the old ones. Thank you for your reply. We’ve checked and fstrim is disabled. Nevertheless, upgrading the ScyllaDB version seems like a good idea. --- ### Page: https://forum.scylladb.com/t/does-seastar-use-one-thread-per-core/5041 Title: Does Seastar use one thread per core? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I thought that ScyllaDB’s underlying Seastar framework uses one thread per core, but I see more than two threads per core. Why is that? Language: en Canonical URL: https://forum.scylladb.com/t/does-seastar-use-one-thread-per-core/5041 ## Headings Structure: H1: Does Seastar use one thread per core? H3: Related topics ## Main Content: H1: Does Seastar use one thread per core? H3: Related topics I thought that ScyllaDB’s underlying Seastar framework uses one thread per core, but I see more than two threads per core. Why is that? Seastar creates an extra thread per core for blocking syscalls (like open()/ fsync() / close() ); this allows the Seastar reactor to continue executing while a blocking operation takes place. Those threads are usually idle, so they don’t contribute to significant context switching activity. You can learn more about this in the documentation. --- ### Page: https://forum.scylladb.com/t/twcs-with-many-unique-partitions/5044 Title: TWCS with many unique partitions - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi, I have a table for storing user refresh tokens. CREATE TABLE IF NOT EXISTS user_refresh_tokens ( refresh_token text, user_id timeuuid, PRIMARY KEY (refresh_token) ) WITH compaction = { 'class' : 'Tim… Language: en Canonical URL: https://forum.scylladb.com/t/twcs-with-many-unique-partitions/5044 ## Headings Structure: H1: TWCS with many unique partitions H2: Is TWCS Appropriate for a Table with Unique Partitions? H2: Would STCS Be Better for Unique Partitions with TTL? H2: Summary Table H2: Recommendation H3: Related topics ## Main Content: H1: TWCS with many unique partitions H2: Is TWCS Appropriate for a Table with Unique Partitions? H2: Would STCS Be Better for Unique Partitions with TTL? H2: Summary Table H2: Recommendation H3: Related topics Hi, I have a table for storing user refresh tokens. Since access has a lifetime of 5 minutes, each user of the application will read and create a new record in this table every 5 minutes. Since I only need to store them for 30 days, I thought that the window strategy would be a good fit here. But am I right? How correct is it to have only unique records in such a table? Maybe it is better to use STCS with the same default_time_to_live ? Also, here is screen of this table in prometheus from dev server, where low rate of requests and about 100 users that do refresh token every 5 minut. But is it normal that table has 122 sstables ? TWCS (Time Window Compaction Strategy) is primarily optimized for time-series data, where each partition receives new data over time and records may be expired after a predictable duration. In your schema, each partition key (the refresh_token) is unique and will only receive a single write, never updated or appended. Strengths of TWCS: It efficiently compacts SSTables within time windows and works best when partitions accumulate data and eventually expire all at once (typical in time-series workloads). Limitations for Unique Keys: When every write generates a new partition (unique token), your data is essentially an “insert-only” workload with no clustering within partitions and no benefit from windowed compaction; each partition will expire independently, and TWCS’s windowing gains are minimal. SSTable Management: You may end up with many very small SSTables if each token insert creates a new partition (which is probably what you already see, as mentioned in your second post), and TWCS will not compact across windows, so fragmentation and file count may increase, putting pressure on the system. Conclusion: TWCS is not an ideal fit for this workload, since your data is not time-series in the classic sense; compaction efficiency is likely to be low for many unique partitions with no clustering or updates per partition. STCS (Size-Tiered Compaction Strategy) is the default strategy for general-purpose, insert-heavy workloads in Cassandra/ScyllaDB. It compacts together similarly sized SSTables, reducing fragmentation and efficiently handling lots of unique partitions—especially when all rows have a reasonably aggressive TTL. Strengths for Unique Partitions: STCS will compact small files into larger ones and remove expired data from older SSTables according to TTL, avoiding the file explosion risk of TWCS in your scenario. Efficient for Expiring Data: With all rows using default_time_to_live (30 days), expired tokens are purged during compaction events without risk of indefinitely persisting small SSTables. Disk Utilization: Make sure your disk is sized appropriately, as STCS can require extra disk space during compactions. Conclusion: For a use case where every write is a unique partition and data expires after a fixed period, STCS is usually the more effective option. It handles unique-partition workloads and repeated inserts efficiently, especially if you’re not reading or deleting tokens before their TTL expires. For a table with many unique, short-lived records (one per refresh token), STCS with TTL is generally preferable to TWCS. Reserve TWCS for workloads where partitions accumulate over time and deletions occur in time-based batches. For your user_refresh_tokens example, switch to STCS to avoid numerous small SSTables and inefficient compaction behavior inherent to TWCS with highly unique partition keys. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-104-2025-08-08/5046 Title: Last week in scylla-cluster-tests.git master (issue #104; 2025-08-08) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the cb8e427b…d8d67bdb range are covered. There were 38 non-merge commits from 13 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-104-2025-08-08/5046 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #104; 2025-08-08) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #104; 2025-08-08) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the cb8e427b…d8d67bdb range are covered. There were 38 non-merge commits from 13 authors in that period. Some notable commits: The longevity-harry-2h test was added back to weekly runs, bringing it to regular tier1 triggers. New performance scale-out/scale-in tests were added for elastic-cloud, checking performance under load (~50%) while scaling with different instance types. These manually triggered tests validate that there’s no performance regression during scaling operations. Argus client was updated to version 0.15.5, improving logging and resource submission. Instance types are displayed in the Argus Resources tab. It’s helpful for tests with various DB instance types. There was great effort this week to improve our unit tests through multiple commits. E.g. the test_coredump was fixed to use events fixture, setting up the events registry and eliminating warnings about unhandled exceptions. Azure rate limiting was fixed with exponential backoff and error isolation, ensuring that temporary Azure API issues don’t prevent the entire cloud monitoring system from generating reports for other providers. This was a contribution by GitHub Copilot agent without much human intervention. The cql-stress Docker tag was updated to v0.2.4 bringing rust driver 1.3.0 and support for extra_definitions in user profiles. Rolling upgrade pipeline improvements were made: a missing provision stage was added and the pipeline was switched to use centralized runSctTest. Scylla Cloud backend gained the ability to run stress loads, building upon the existing PoC by adding loader node instantiation capabilities. The longevity pipeline was also updated to support running SCT tests on Scylla Cloud backend in Jenkins. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/tablets-support-for-lwt-mv-cdc-timeline-and-roadmap/5047 Title: Tablets support for LWT, MV, CDC, timeline and roadmap - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/tablets-support-for-lwt-mv-cdc-timeline-and-roadmap/5047 ## Headings Structure: H1: Tablets support for LWT, MV, CDC, timeline and roadmap H3: Related topics ## Main Content: H1: Tablets support for LWT, MV, CDC, timeline and roadmap H3: Related topics Originally from the User Slack @Mikael_HedbergMikael_Hedberg**:** I’m curious about the tablets roadmap. Are you aiming to support LWTs, counters and MVs in the future with tablets? If so, what’s the@dortimeline? @dor**:** Yes, all will support tablets. LWT in master has 99% of what tablets need The timeline is to have MV, CDC, LWT code complete by the end of Sep, to be released as 2025.4 Counters will@Mikael_Hedbergarrive later @Mikael_Hedberg**:** Tablets is pretty much the greatest innovation in this space for a long time. Really looking forward to it. Did you require to compromise a lot with the initial promises of@dortablets to get it working? @dor**:** I don’t understand the compromises you refer to. It’s quite a big change. They do impose a challenge since each tablet is a LSM tree but they work just fine @Mikael_Hedberg**:** Never mind, I was just curious. Seems like you’ve been working on it for a long time - which is understandable. Really great to hear that you guys been doing such progress on it. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-288-2025-08-10/5048 Title: Last week in scylladb.git master (issue #288; 2025-08-10) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 1c25aa891b..f3d9d0c1c7 range are covered. There were 76 non-merge commits from 20 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-288-2025-08-10/5048 ## Headings Structure: H1: Last week in scylladb.git master (issue #288; 2025-08-10) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #288; 2025-08-10) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 1c25aa891b..f3d9d0c1c7 range are covered. There were 76 non-merge commits from 20 authors in that period. Some notable commits: The Raft recovery procedure is run when a cluster suffers a majority loss and the quorum needs to be re-established. It is now simplified. Tests can now be labeled regarding their stability and criticallity. The system.clients virtual table now lists active Alternator requests in addition to CQL connections. An internal error in truncate was relaxed into a warning, after it was determined the condition can legitimately happen. The CQL SELECT statement gained the ability to use a vector secondary index. There is now a REST API endpoint to drop quarantined sstables and reclaim their storage. Sstables become quarantined when corruption is detected in them. It is now possible to write into system tables using Alternator rather than just read them. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-20/5049 Title: [RELEASE] ScyllaDB Enterprise 2024.1.20 - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2024.1.20, a bug-fix production-ready patch release for ScyllaDB Enterprise 2024.1 Long Term Support (LTS) Release. Note there is a later, better, LTS (Long term Support) R… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-20/5049 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2024.1.20 H3: Related Links H2: Fixed Issue with an source available reference: H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2024.1.20 H3: Related Links H2: Fixed Issue with an source available reference: H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2024.1.20, a bug-fix production-ready patch release for ScyllaDB Enterprise 2024.1 Long Term Support (LTS) Release. Note there is a later, better, LTS (Long term Support) Release 2025.1, and a feature release 2025.2. You are encouraged to upgrade in coordination with the ScyllaDB Support team. Get ScyllaDB Enterprise 2024.1 (customers only, or 30-day evaluation) Upgrade from ScyllaDB Enterprise 2023.1.x to 2024.1.y Upgrade from ScyllaDB Enterprise 2022.2.x to 2024.1.y Upgrade from ScyllaDB Open Source 5.4 to ScyllaDB Enterprise 2024.1.x A rare seastar issue causes segmentation fault in logger during early startup. Fix is now backported to 2024.1 seastar#1974 non-full row keys cause mis-parsing of sstables #24489 Avoid killing a node when reading from sstables #20845 --- ### Page: https://forum.scylladb.com/t/monitoring-how-to-publish-rust-driver-metrics-to-prometheus-histogram-total-count/5050 Title: Monitoring - how to publish Rust driver metrics to Prometheus, histogram total count - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/monitoring-how-to-publish-rust-driver-metrics-to-prometheus-histogram-total-count/5050 ## Headings Structure: H1: Monitoring - how to publish Rust driver metrics to Prometheus, histogram total count H3: Related topics ## Main Content: H1: Monitoring - how to publish Rust driver metrics to Prometheus, histogram total count H3: Related topics Originally from the User Slack Hi, if one wants to publish the metrics available from rust driver to Prometheus, how does one get total count from histogram? I only see the methods that export percentiles from internal histogram, but not the count. @amnon a histogram comes with a built in count and sum Михаил_Доронин: Can you point me on API one needs to use? @amnon if a histogram name is called scylla_alternator_op_latency for example. It will have additional two metrics: scylla_alternator_op_latency_count scylla_alternator_op_latency_sum --- ### Page: https://forum.scylladb.com/t/release-scylla-operator-1-18-0/5051 Title: [RELEASE] Scylla Operator 1.18.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of ScyllaDB Operator 1.18.0. ScyllaDB Operator is an open-source project that helps users run ScyllaDB on Kubernetes. The ScyllaDB Operator manages ScyllaDB cluste… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-operator-1-18-0/5051 ## Headings Structure: H1: [RELEASE] Scylla Operator 1.18.0 H1: Multi-DC backup & repair tasks scheduling H1: Other notable changes H1: Upgrade instructions H1: Supported versions H1: Getting started with ScyllaDB Operator H1: Related Links H3: Related topics ## Main Content: H1: [RELEASE] Scylla Operator 1.18.0 H1: Multi-DC backup & repair tasks scheduling H1: Other notable changes H1: Upgrade instructions H1: Supported versions H1: Getting started with ScyllaDB Operator H1: Related Links H3: Related topics The ScyllaDB team is pleased to announce the release of ScyllaDB Operator 1.18.0. ScyllaDB Operator is an open-source project that helps users run ScyllaDB on Kubernetes. The ScyllaDB Operator manages ScyllaDB clusters deployed to Kubernetes and automates tasks related to operating a ScyllaDB cluster, like installation, vertical and horizontal scaling, as well as rolling upgrades. As with each minor version, ScyllaDB Operator 1.18 adds new features and improves stability. The highlight of the release is the enhancement of the technical preview support for managed multi-DC clusters by adding support for ScyllaDB Manager backup and repair tasks scheduling. ScyllaDB Operator 1.18 enhances the technical preview support for managed multi-DC setups with the ability to schedule backup and repair ScyllaDB Manager tasks using a Kubernetes-native approach. Registering a multi-DC cluster with a global ScyllaDB Manager and scheduling tasks for it becomes a matter of configuring proper Kubernetes resources. From now on, you can run backups and repairs on your multi-DC ScyllaDB cluster by simply adding a label to your ScyllaDBCluster object (to inform ScyllaDB Manager about the multi-DC cluster’s existence). Then, you can associate ScyllaDBManagerTask objects that define backup or repair tasks with your ScyllaDBCluster to get them scheduled. ScyllaDBManagerTasks can be seen as a new generation of the ScyllaCluster spec’s backups and repairs APIs. Assuming you have an existing multi-DC ScyllaDBCluster, you can register it in the global ScyllaDB Manager by labeling it as follows: Then, you can create a task (in this example, it will be a repair task), using the following manifest: Please refer to the ScyllaDBManagerTask API reference to learn more about our ScyllaDB Manager integration with multi-DC clusters. Please note that we’ve added support for backup and repair tasks as feature parity with the existing stable single-DC ScyllaCluster. Restoring backups using CRDs is still not supported, but can be achieved manually. Please refer to ScyllaDB Manager integration documentation for more details. We’re still working towards promoting the managed multi-DC deployment to general availability. If you’re waiting for it, please refer to #2230 to track our progress and plans. Along with multi-DC backup & repair tasks scheduling, we simplified ScyllaDB Operator’s architecture by removing a confusing standalone deployment manager-controller that was deployed alongside Operator to save resources and give you one less component to operate. That means there will no longer be a separate standalone application orchestrating ScyllaDB Manager required in your cluster, it will now run as part of the ScyllaDB Operator process. (#2681) This release also includes a patch update of ScyllaDB (2025.1.2 → 2025.1.5; #2845), a major Grafana upgrade (11.4.3 → 12.0.2), as well as updates of Prometheus (v3.1.0 → v3.5.0) and our newest ScyllaDB monitoring dashboards (#2799). For more changes and details, check out the GitHub release notes. Upgrading from v1.17.x with kubectl apply requires extra actions due to the removal of the standalone ScyllaDB Manager controller. The controller becomes an integral part of the Operator. Because of that, users have to delete this deployment before installing v1.18.x manifests: Using Helm requires a mandatory manual step for every release because Helm can’t handle CRDs updates. Please refer to the Upgrade guide and its 1.17 to 1.18 section for more information. ScyllaDB >=2024.1 && <=2025.1 Container Runtime Interface API == v1 ScyllaDB Manager >=3.5.0 && <=3.5.1 ScyllaDB Operator Documentation Learn how to deploy ScyllaDB on Google Kubernetes Engine (GKE) here Learn how to deploy ScyllaDB on Amazon Elastic Kubernetes Engine (EKS) here Learn how to deploy ScyllaDB on a Kubernetes Cluster here ScyllaDB Operator source (on GitHub) ScyllaDB Operator image on DockerHub ScyllaDB Operator Helm Chart repository ScyllaDB Operator documentation ScyllaDB Operator for Kubernetes lesson in ScyllaDB University We’ll welcome your feedback! Feel free to open an issue or reach out on the #scylla-operator channel in ScyllaDB User Slack. Regards, The ScyllaDB Operator Team --- ### Page: https://forum.scylladb.com/t/release-scylladb-manager-3-6-0/5052 Title: [RELEASE] ScyllaDB Manager 3.6.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces the release of ScyllaDB Manager 3.6.0, a production-ready minor release of the stable 3.6 branch. ScyllaDB Manager is a centralized cluster administration and recurrent tasks automation tool. … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-manager-3-6-0/5052 ## Headings Structure: H1: [RELEASE] ScyllaDB Manager 3.6.0 H3: 1-1-restore H3: Native backup support H3: Stability improvements H3: Upgrade to the new release H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Manager 3.6.0 H3: 1-1-restore H3: Native backup support H3: Stability improvements H3: Upgrade to the new release H3: Related topics The ScyllaDB team announces the release of ScyllaDB Manager 3.6.0, a production-ready minor release of the stable 3.6 branch. ScyllaDB Manager is a centralized cluster administration and recurrent tasks automation tool. This release focuses on the new 1-1-restore procedure and support for the experimental native backup. Below are the changes in this release. Scylla Manager 3.6.0 introduces a new sctool restore 1-1-restore command that significantly speeds up the restore process at the cost of stricter requirements: It sets restored tables and views `tombstone_gc’ mode to ‘repair’ which is required to avoid running repair operation as part of the restore procedure. After restoration, the ‘tombstone_gc’ mode can only be changed once the tables have been repaired - otherwise, data resurrection may occur. Here is a benchmark of the regular (Load&Stream) restore and the 1-1-restore against a 3-node i4i.4xlarge cluster of ScyllaDB 2025.2.1 with a single table with 3TB of data after replication RF=3. The new ScyllaDB 2025.2.0 release includes experimental support for native backup, which allows uploading sstables directly from ScyllaDB to S3 without ScyllaDB Manager Agent (rclone) proxy. Native backup proves to be much faster, but because of that, it might also impact online request latency. Its performance and impact can be controlled by setting stream_io_throughput_mb_per_sec in ‘scylla.yaml’. The ScyllaDB Manager 3.6.0 release allows for experimenting with the native backup by setting the new sctool backup --method flag. It specifies the API used for uploading sstables: All methods still require configuring the ScyllaDB Manager Agent with ‘scylla-manager-agent.yaml’. The native backup also requires object_storage_endpoints to be configured in ‘scylla.yaml’. –rate-limit and –transfers flags do not take effect when using –method=native, as the upload performance is controlled by the ScyllaDB server (see stream_io_throughput_mb_per_sec in ‘scylla.yaml’). The ScyllaDB Manager 3.6.0 release also contains stability improvements: ScyllaDB customers are encouraged to upgrade to ScyllaDB Manager 3.6.0 in coordination with the ScyllaDB support team. The new release includes upgrades of both ScyllaDB Manager Server and Agent. Scylla Manager 3.6.0 supports the following ScyllaDB releases: You can install and run ScyllaDB Manager on Kubernetes using ScyllaDB Operator. More here. --- ### Page: https://forum.scylladb.com/t/scylla-manager-indexing-issue/5058 Title: Scylla manager indexing issue - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi guys, I’m getting indexing files error in scylla backup and because of this the backups got disturbed abruptly. We are currently addressing this issue manually by first deleting the old task and then recreating the s… Language: en Canonical URL: https://forum.scylladb.com/t/scylla-manager-indexing-issue/5058 ## Headings Structure: H1: Scylla manager indexing issue H3: Related topics ## Main Content: H1: Scylla manager indexing issue H3: Related topics Hi guys, I’m getting indexing files error in scylla backup and because of this the backups got disturbed abruptly. We are currently addressing this issue manually by first deleting the old task and then recreating the same task in Scylla Manager. This process helps us reset the task, ensuring that the new configuration or settings are applied properly. While this is a temporary workaround, we’re monitoring the situation closely to ensure it resolves the issue for now. This is an exact (same screenshot) dup of Backup fails with indexing files error which is waiting on user reply. @Guy is there a way of marking this issue as dup and closing it in our forum? Thanks @Michal_Leszczynski , closing this as a duplicate. --- ### Page: https://forum.scylladb.com/t/multi-data-centers-cluster-on-k8s-ipvs-loadbalancer-environment/5060 Title: Multi Data Centers Cluster on K8s + IPVS Loadbalancer environment - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 2025.2 #Cluster size: 2 k8s cluster with 5 nodes per each os (RHEL/CentOS/Ubuntu/AWS AMI): I want to configure SycllaDB multi-data center clustering between two k8s clusters. O… Language: en Canonical URL: https://forum.scylladb.com/t/multi-data-centers-cluster-on-k8s-ipvs-loadbalancer-environment/5060 ## Headings Structure: H1: Multi Data Centers Cluster on K8s + IPVS Loadbalancer environment H3: Exposing ScyllaDB cluster | ScyllaDB Docs H3: Related topics ## Main Content: H1: Multi Data Centers Cluster on K8s + IPVS Loadbalancer environment H3: Exposing ScyllaDB cluster | ScyllaDB Docs H3: Related topics Installation details #ScyllaDB version: 2025.2 #Cluster size: 2 k8s cluster with 5 nodes per each os (RHEL/CentOS/Ubuntu/AWS AMI): I want to configure SycllaDB multi-data center clustering between two k8s clusters. Our company doesn’t use a public cloud, but rather our own IDC. The k8s LoadBalancer is configured via IPVS, so from the pod’s perspective, the destination IP is the LoadBalancer IP. Therefore, if I bind the listen-address to the pod IP, the connection fails due to RST packets. However, if I bind the listen-address to the LoadBalancer IP, the source IP of outgoing packets from the pod is the listen-address (LoadBalancer IP), so the connection fails to establish properly. Since the two k8s clusters don’t share a common network, access is only possible through the LoadBalancer. How can I configure multi-data center clustering? Your IPVS LB appears to use DSR mode. Packets reach the pod with the destination IP = VIP, which the pod doesn’t own, so the kernel rejects the SYN and you see RST. Binding Scylla to the VIP “fixes” the inbound path but breaks outbound (the pod now sources packets from the VIP). That’s expected with DSR and won’t work for Scylla pods. You need the LB to perform NAT or L4 proxying so that packets arriving to the pod have destination = podIP (or NodePort DNAT). Then you can keep listen_address on the pod and advertise a different broadcast address for cross DC peers. This is exactly what the Operator broadcast options are for. ScyllaDB is an Apache Cassandra-compatible NoSQL data store that can handle 1 million transactions per second on a single server. Please note that exposeOptions are immutable, they cannot be changed after a ScyllaDB cluster is created. In general, what you need to do is the following: --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-105-2025-08-15/5062 Title: Last week in scylla-cluster-tests.git master (issue #105; 2025-08-15) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the bdab118b…959e222a range are covered. There were 27 non-merge commits from 8 authors in tha… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-105-2025-08-15/5062 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #105; 2025-08-15) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #105; 2025-08-15) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the bdab118b…959e222a range are covered. There were 27 non-merge commits from 8 authors in that period. Some notable commits: Added fallback to Azure SDK reboot when walinuxagent is absent ensuring Azure VMs based on non-official images can still be restarted reliably. Switched artifact I/O tuning to use scylla_io_setup instead of direct iotune to align with current Scylla logic and improve result correctness. Preserved original spacing in events sent to Argus fixing unintended trimming in some cases. Added rack-aware support for the Python driver tests. A new Rust performance test using cql-stress was added. New performance tests branch branch-perf-v17 was created. Key workloads migrated include rolling upgrade vnodes, elasticity, and microbenchmark jobs. Introduced vector-store service deployment for the Docker backend enabling VS testing in containerized environments. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/how-do-i-upgrade-from-scylladb-5-0-minor-and-major-versions-and-upgradesstables/5065 Title: How do I upgrade from ScyllaDB 5.0, minor and major versions, and upgradesstables - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/how-do-i-upgrade-from-scylladb-5-0-minor-and-major-versions-and-upgradesstables/5065 ## Headings Structure: H1: How do I upgrade from ScyllaDB 5.0, minor and major versions, and upgradesstables H3: Related topics ## Main Content: H1: How do I upgrade from ScyllaDB 5.0, minor and major versions, and upgradesstables H3: Related topics Originally from the User Slack @bryam_castillobryam_castillo**:** HI scylla team, is it ok to upgrade from 5.0 direct to 5.4 and then to 6.0? wonder if we can skip minor versions also, for which versions upgrade, should we execute “nodetool upgradesstables” command? @avi**:** No, only minor upgrades are tested. 5.0 → 5.1 → 5.2 → 5.4 → 6.0 (which is already unsupported) there’s no need to run upgradesstables --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-289-2025-08-18/5068 Title: Last week in scylladb.git master (issue #289; 2025-08-18) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the f3d9d0c1c7..f689d41747 range are covered. There were 106 non-merge commits from 25 authors in that per… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-289-2025-08-18/5068 ## Headings Structure: H1: Last week in scylladb.git master (issue #289; 2025-08-18) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #289; 2025-08-18) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the f3d9d0c1c7..f689d41747 range are covered. There were 106 non-merge commits from 25 authors in that period. Some notable commits: Raft group 0 will now enforce an odd number of voters to reduce risk of a network partition preventing consensus from being achieved. The row cache is able to purge expired tombstones in order to improve performance of reads that later touch the same key. To prevent data resurrection, it checks memtables for overlapping data. It now avoids checking memtables for which it can prove there is no overlapping order data, reducing false positives and increasing the number of tombstones purged. The CREATE KEYSPACE ‘class’ option now defaults to NetworkTopologyStrategy and can be omitted. It is the only sensible choice for user keyspaces. Internode communication errors are now reported to the user during LWT transaction failures. The system will no longer create sstables with numeric generation numbers (only UUIDs). It can still read such sstables. The vector store client can now be dynamically updated with URLs to the vector store server, allowing for hot reconfiguration. The nodetool getsstables and similar commands that accept a key now work if the key contains a colon (apart from composite keys). See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-python-driver-3-29-4/5069 Title: [RELEASE] Python Driver 3.29.4 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: Driver Release Summary: 3.29.4 Release Link: Release 3.29.4 · scylladb/python-driver · GitHub Driver Release Notes Summary Improvements & Fixes Added ApplicationInfo API to supply information about application to the… Language: en Canonical URL: https://forum.scylladb.com/t/release-python-driver-3-29-4/5069 ## Headings Structure: H1: [RELEASE] Python Driver 3.29.4 H3: Driver Release Notes Summary H3: Related topics ## Main Content: H1: [RELEASE] Python Driver 3.29.4 H3: Driver Release Notes Summary H3: Related topics Driver Release Summary: 3.29.4 Release Link: Release 3.29.4 · scylladb/python-driver · GitHub Environment & Compatibility Updates --- ### Page: https://forum.scylladb.com/t/release-scylladb-2025-1-6/5070 Title: [RELEASE] ScyllaDB 2025.1.6 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB 2025.1.6, a bug-fix production-ready patch release for ScyllaDB 2025.1 LTS Release. Note there is a new Short Term Support (STS) Feature release 2025.2. You are welcome to upgrade to… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-2025-1-6/5070 ## Headings Structure: H1: [RELEASE] ScyllaDB 2025.1.6 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 2025.1.6 H3: Related topics The ScyllaDB team announces ScyllaDB 2025.1.6, a bug-fix production-ready patch release for ScyllaDB 2025.1 LTS Release. Note there is a new Short Term Support (STS) Feature release 2025.2. You are welcome to upgrade to it for the latest and greatest features, or stay on the 2025.1 track for long term support. Upgrade from ScyllaDB Enterprise 2024.x to ScyllaDB 2025.1 Upgrade from ScyllaDB Open Source 6.2 to ScyllaDB Enterprise 2025.1.x The following issues are fixed in this release: auth: default user creation may cause a crash if done during upgrade to raft topology #24975 Coredump during truncate: Data written after truncation time was incorrectly truncated #25013 Coredump right after seed node decommission #23911 Node crashes at shutdown with “Assertion this->_con->get()->sink_closed() failed” while restore in progress #25165 non-RBNO streaming may hang on failure if IO permits are deadlocked #24925 partitioned_sstable_set go to quadratic space consumption mode in a case of a lot of unleveled sstables #23634 streaming,repair: db::view::check_needs_view_update_path() can deadlock with the repair/streaming permit #24807. This issue happens with removenode, when RBNO is disabled, so range test_mv_tablets_empty_ip: node got stuck during shutdown #23665 token_metadata_ptr may be destroyed without gentle cleaning #13381 Upgrade to raft topology may cause reads from system_distributed, leading to timeouts and defuncting group0 fiber #24963 tablets: stop storage group on deallocation #24857 #24828. This can cause errors when a node is restarted during tablet cleanup and other use cases. repair: tablet repair does not exclude with global topology operations #24195. The fix delays repair until topology updates are completed. Change Data Capture (CDC) node crashes when dropping a column while writing to it with CDC #24952 It’s possible to manually drop columns from CDC log #24643 Destroying tablet_metadata holding many tables or even clearing it gently queues a long burst of seastar tasks #24814 seastar_memory - oversized allocation (locator::describe_ring) #24158 service levels: cache is reloaded too often #25114 #23065 utils::http::dns_connection_factory creates new system trust credential on each instance #24447 --- ### Page: https://forum.scylladb.com/t/release-scylladb-2025-2-2/5072 Title: [RELEASE] ScyllaDB 2025.2.2 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB 2025.2.2, a bug-fix production-ready patch release for ScyllaDB 2025.2 Feature Release. Related Links Get ScyllaDB 2025.2 Upgrade from ScyllaDB 2025.1 to ScyllaDB 2025.2 Upg… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-2025-2-2/5072 ## Headings Structure: H1: [RELEASE] ScyllaDB 2025.2.2 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 2025.2.2 H3: Related topics The ScyllaDB team announces ScyllaDB 2025.2.2, a bug-fix production-ready patch release for ScyllaDB 2025.2 Feature Release. Upgrade from ScyllaDB 2025.1 to ScyllaDB 2025.2 Upgrade from ScyllaDB 2025.2.x to ScyllaDB 2025.2.y The following issues are fixed in this release: Remove throwing (some) protocol_exceptions when handling new connections #24567 #25272. avoid negative side effects of throwing these exceptions during a connection storm. auth: default user creation may cause a crash if done during upgrade to raft topology #24975 Coredump during truncate: Data written after truncation time was incorrectly truncated #25013 Node crashes at shutdown with “Assertion this->_con->get()->sink_closed() failed” while restore in progress #25165 non-RBNO streaming may hang on failure if IO permits are deadlocked #24925 seastar_memory - oversized allocation (locator::describe_ring) #24158 Segmentation fault in scheduling group streaming during tablets streaming after node was decommissioned #23162 streaming,repair: db::view::check_needs_view_update_path() can deadlock with the repair/streaming permit #24807. This issue happens with removenode, when RBNO is disabled, so range test_mv_tablets_empty_ip: node got stuck during shutdown #23665 token_metadata_ptr may be destroyed without gentle cleaning #13381 Upgrade to raft topology may cause reads from system_distributed, leading to timeouts and defuncting group0 fiber #24963 Change Data Capture (CDC) CDC is a feature that allows you to not only query the current state of a database’s table, but also query the history of all changes made to the table. It’s possible to manually drop columns from CDC log #24643 node crashes when dropping a column while writing to it with CDC #24952 The system.clients virtual table now lists active Alternator requests in addition to CQL connections. Change severity of Failed force_keyspace_cleanup - gate_closed_exception error to WARNING #16732 Destroying tablet_metadata holding many tables or even clearing it gently queues a long burst of seastar tasks #24814 service levels: cache is reloaded too often #25114 #23065 slow replace in the presence of a table with tablets #25163 utils::http::dns_connection_factory creates new system trust credential on each instance #24447 --- ### Page: https://forum.scylladb.com/t/multi-dc-clusters-version-compatibility-and-upgrading/5073 Title: Multi DC clusters. version compatibility and upgrading - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/multi-dc-clusters-version-compatibility-and-upgrading/5073 ## Headings Structure: H1: Multi DC clusters. version compatibility and upgrading H3: Related topics ## Main Content: H1: Multi DC clusters. version compatibility and upgrading H3: Related topics Originally from the User Slack @Yohan_Mok: Hi team, We have a couple questions about multi DC clusters. @dordor**:** Standard documentation is fine, it’s relatively easy to setup, you do need to read about the CL settings. It’s recommended to have the same version across DCs. @Yohan_Mok: We wanted to know if there were any documented limitations on version differences between DCs in the cluster. (eg. DC of version X is compatible up to version Y?) If this is documented like you say, is it ok if i ask for a link? > ps. cluster setup was not an issue + we are aware of CL settings and DCAwareRoundRobinPolicy @Felipe_Cardeneti_Mendes**:** @Yohan_Mok the general direction for single-DC applies. All nodes are part of a cluster irrespective of their placement. In other words: You can’t add a DC with a backleveled major version. As you upgrade one region, you should plan to also upgrade others. @Yohan_Mok: Thanks for the clarification! --- ### Page: https://forum.scylladb.com/t/release-scylladb-java-driver-3-11-5-8/5081 Title: [RELEASE] ScyllaDB Java Driver 3.11.5.8 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: Improvements Configuration Options: Added support for setting application name, version, and client ID, making it easier to identify clients in logs and metrics. Tablets Code: Cleaned up and optimized tablets code f… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-java-driver-3-11-5-8/5081 ## Headings Structure: H1: [RELEASE] ScyllaDB Java Driver 3.11.5.8 H3: Improvements H3: Fixes H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Java Driver 3.11.5.8 H3: Improvements H3: Fixes H3: Related topics Configuration Options: Added support for setting application name, version, and client ID, making it easier to identify clients in logs and metrics. Tablets Code: Cleaned up and optimized tablets code for better maintainability and performance. Zero-Token Nodes: Prevented logging of invalid rows for zero-token nodes. RackAwareRoundRobinPolicy: Fixed policy behavior to properly treat nodes from other racks as REMOTE. Detailed changelog is available here. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-290-2025-08-24/5082 Title: Last week in scylladb.git master (issue #290; 2025-08-24) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the f689d41747..87dd96f9a2 range are covered. There were 109 non-merge commits from 22 authors in that per… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-290-2025-08-24/5082 ## Headings Structure: H1: Last week in scylladb.git master (issue #290; 2025-08-24) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #290; 2025-08-24) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the f689d41747..87dd96f9a2 range are covered. There were 109 non-merge commits from 22 authors in that period. Some notable commits: Tables created with tablets can now be repaired using incremental repair. Incremental repair remembers what part of the tablet dataset was already repaired and avoids repairing it again. To avoid forgetting the repaired dataset, compaction cannot compact unrepaired and repaired sstables. Incremental repair is much faster than full repair. The redis protocol implementation, which was always very partial, was removed from the source tree. It is now possible to DROP ROLEs when using the saslauthd authenticator. There is now a warning when creating a keyspace where the replication factor does not match the rack count. Such keyspaces won’t work well with tablets. Alternator, ScyllaDB’s implementation of the DynamoDB API, now calculates Write Capacity Units (WCU) in a compatible manner compared to DynamoDB. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-291-2025-08-31/5084 Title: Last week in scylladb.git master (issue #291; 2025-08-31) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 87dd96f9a2..bc5773f777 range are covered. There were 82 non-merge commits from 13 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-291-2025-08-31/5084 ## Headings Structure: H1: Last week in scylladb.git master (issue #291; 2025-08-31) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #291; 2025-08-31) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 87dd96f9a2..bc5773f777 range are covered. There were 82 non-merge commits from 13 authors in that period. Some notable commits: There is now a centralized view building coordinator running on Raft group 0. It is responsible for coordination builds of new materialized views. Nodes will now reject user writes when nodes reach 98% storage utilization, while still allowing tablet migrations. This protects the cluster when scale-out was not able to keep up with data ingestion. The container image entry point learned new --dc and --rack options to set those parameters without bind-mounting the cassandra.rackdb file. When creating a vector index, CDC will automatically be enabled so that the vector store can stream vector updates from the database. The build system now uses precompiled headers to speed up compilation. A crash in commitlog when a mutation spans multiple segments has been fixed. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-2025-3-0/5085 Title: [RELEASE] ScyllaDB 2025.3.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of ScyllaDB 2025.3, a production-ready Short Term Support (STS) Minor Feature Release. More information on ScyllaDB’s Long Term Support (LTS) policy is available here.… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-2025-3-0/5085 ## Headings Structure: H1: [RELEASE] ScyllaDB 2025.3.0 H2: Related Links H2: New features H3: Native Backup H3: Alternator H3: Deployment options H3: More updates H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 2025.3.0 H2: Related Links H2: New features H3: Native Backup H3: Alternator H3: Deployment options H3: More updates H3: Related topics The ScyllaDB team is pleased to announce the release of ScyllaDB 2025.3, a production-ready Short Term Support (STS) Minor Feature Release. More information on ScyllaDB’s Long Term Support (LTS) policy is available here. The 2025.3 release adds new features like native backup and RHEL 10 support. Read more about ScyllaDB Upgrade from ScyllaDB 2025.2 to ScyllaDB 2025.3 ScyllaDB Enterprise customers are encouraged to upgrade to ScyllaDB 2025.3, and are welcome to contact our Support Team with questions. To get the most from ScyllaDB 2025.3, use ScyllaDB Manager 3.6 and later, and ScyllaDB Monitoring Stack 4.10 and later. The native backup feature delivers up to 15x faster backup performance compared to the previous rclone-based approach, with no impact on user queries.#24793 Introduced as experimental in ScyllaDB 2025.2, native backup is now production-ready. Previously, SSTable backups to S3 used the Scylla Manager Agent with rclone, managed via the Scylla Manager server. This release adds direct backup infrastructure between ScyllaDB and S3 while keeping the scheduling logic and API in Scylla Manager. Native restore support will be provided in a future release. To enable native backup in 2025.3, set up the S3 connectivity, by adding to each node scylla.yaml: S3 target bucket - you can use different bucket per region A max upload bandwidth upper limit stream_io_throughput_mb_per_sec, safe guards the backup rate from getting too high. The recommendation is to set it up as 75% of the max network throughput. A new method is added in Scylla Manager for the native backup. The available methods are: Rclone: use legacy rclone backup. Native: Backup directly from Scylla server to S3. Auto: try to use Native if possible: supported by scyllaDB version and back target (see bullet below) Rclone agent is always used to upload backup meta data. You should still configure the manager agent when using Native backup and restore. Native backup is only available for AWS. For other backup targets (such as GCP Cloud Storage and Azure Storage), ScyllaDB will continue to rely on the Manager Agent. Alternator, ScyllaDB’s implementation of the DynamoDB API, now supports a set of per-table metrics. #19824 Alternator, now supports time-to-live (TTL) attributes on tables running with tablets. #16567 AWS: This release adds support for the following I7i and I7ie instance types: i7i.large, i7i.xlarge, i7i.2xlarge and all and i7ie instances. The I7i and I7ie families are the next-generation successors to the i4i and i3en instances, delivering an improved price-to-performance ratio compared to earlier generations. Larger i7i instances will be supported in upcoming releases. image#709 GCP: The latest Z3 instances are now supported #728 image#728. Read more about ScyllaDB on Z3 in New Google Cloud Z3 Instances: Early Performance Benchmarks on ScyllaDB Show up to 24% Better Throughput, and Big ScyllaDB Performance Gains on Google Cloud’s New Smaller Z3 Instances ScyllaDB now tested on RHEL 10 scylla-pkg/#5222 Support for Ubuntu 20.04, which has reached end-of-life, has been removed. #24564 LWT: possible stale data read #24630. The issue is limited to rare cases where Replication Factor >= 4 , the most recent commit is stored (“learned”) only on one replica, and other replicas must have failed to receive an update last time LWT was executed for this partition key. An edge case with DECIMAL type parsing was fixed. #24581. For example, the following should have fail, but prior to this fix have not: CREATE TABLE IF NOT EXISTS test_table (pk int PRIMARY KEY, val decimal); INSERT INTO test_table (pk,val) values (12, 1.1e-2147483647); The data structure used for building mutations now has additional sanity checks for row clustering keys. Restore: Native restore does not finish after altering batch size #25262, #25453 S3 client: Refinement in how credential renewal and expiration #25044 The maximum length of keyspace, table, and view names was extended from 48 characters to 192 characters. #4480 Fixed an issue where the default role was re-created after a node restart. The role is now only re-created if no other superuser role exists. #24469 The native CQL transport now supports the metadata ID extension (from protocol version 5) that allows updating row metadata for prepared statements for SELECT * queries (when columns were added or removed) or SELECT udt queries (when the user defined type definition changed). Note a driver that supports the extension is required to make use of this. #20860 An issue in mapreduce_service caused by a possible race condition, where parallel aggregation may fail to include all values. #20662 ScyllaDB tracks the amount of memory in memtables that was spooled to an SSTable and tries to ensure new writes don’t consume memory faster than it is written to disk. A crash due to a rare edge case when tracking this memory caused an assertion `_mt._flushed_memory <= _mt.occupancy().total_space()’ failed during bulk ingestion #21413 There is now a queue for topology requests. Requests that cannot be processed in parallel will be queued one after the other. #16822 Empty clustering keys generated and spread freely in the system #24506 ScyllaDB now avoids large contiguous allocations during DESCRIBE statements with tables that have many (possibly deleted) columns. #24018 The repair small table optimization is used to speed up repairs on tables that are empty or almost empty (like many tables in system_distributed). On large clusters, the optimization could lead to an out-of-memory condition. #22244 A new schema application framework was merged, which will allow to improve atomicity when managing cluster metadata. #19649 #24531 The CQL binary protocol server now avoids exceptions when rejecting protocol versions it does not support. Unsupported protocol versions are common when drivers probe for server capability, and preventing exceptions improves robustness during connection storms. #24738 Background writes are now cancelled when a node is shut down, so shutdown completes more quickly. #23665 Remove throwing (some) protocol_exceptions when handling new connections #24567 #25272. avoid negative side effects of throwing these exceptions during a connection storm. Coredump during truncate: Data written after truncation time was incorrectly truncated #25013 Node crashes at shutdown with “Assertion this->_con->get()->sink_closed() failed” while restore in progress #25165 Tablets Repair: A node may crash If a node has many shards and many tablets on each shard #23632 Upgrade to raft topology may cause reads from system_distributed, leading to timeouts and defuncting group0 fiber #24963 HTTP TLS support now reuses credentials across connections, minimizing repeated disk loads. #24447 There is now support for converting CQL3 data type representations to a new byte-comparable representation. This is a step towards implementation of Trie sstable indexes in a follow-up release. #23541 Some stalls when managing schemas with thousands of tables were eliminated. #24815 Alternator, ScyllaDB’s implementation of the DynamoDB API, is now more careful to avoid large contiguous allocations during query/scan operations, as these can cause stalls. #23535 Tablets: After replacing a node, the load balancer will immediately check available space, to avoid the up to 60-second delay from regularly scheduled checks that can delay load balancing. #25163 service levels: cache is reloaded too often #25114, #23065 Slow replace in the presence of a table with tablets #25163 ScyllaDB can automatically parallelize some aggregation queries via an internal map/reduce service. This automatic parallelization is now optimized for tablets #21831 Lightweight transactions (LWT) now synchronize with tablet migrations. This is a step towards enabling LWT with tablets, which is yet not available in 2025.3. #24495 Repair: Tablet repair does not exclude global topology operations #24195. The fix delays repair until topology updates are completed. ScyllaDB now supports co-locating tablets of different tables. Co-located tablets are migrated, split, and merged. They will be used, in a follow-up release, for lightweight transactions, change data capture, and local materialized views. #17043 Tablet metadata can grow very large in large clusters. Its creation already avoided reactor stalls, and now its destruction as well. #24814 The nodetool refresh command gained a –skip-reshape switch. Reshaping can be an expensive operation and one may want to defer it. #24365 The nodetool repair command now rejects tablet keyspaces; those are repaired via a different command. #23660. To repair tablet keyspaces use nodetool cluster repair. Native Backup: The nodetool backup command learned the --move-files option which moves files instead of copying them. #24372 Using scylla sstable, the tool may crash in case of disengaged row #25325 Simplify the new Raft recovery procedure: simplify rolling restart with recovery_leader #25015 Introduced a queue for global topology requests, enabling multiple requests to run in parallel. #24293 An internal error in truncate was relaxed into a warning, after it was determined the condition can legitimately happen. #25173 #25013 Change severity of Failed force_keyspace_cleanup - gate_closed_exception error to WARNING #16732 seastar_memory - oversized allocation: in a cluster with 5000 tables #25040 S3 clients memory improvements, allowing operations to be cancelled when waiting for memory resources #25454 --- ### Page: https://forum.scylladb.com/t/how-do-i-update-extend-the-ttl-of-all-existing-data-by-x-time/5086 Title: How do I update/extend the TTL of all existing data by X time? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/how-do-i-update-extend-the-ttl-of-all-existing-data-by-x-time/5086 ## Headings Structure: H1: How do I update/extend the TTL of all existing data by X time? H3: Related topics ## Main Content: H1: How do I update/extend the TTL of all existing data by X time? H3: Related topics Originally from the User Slack @Saif_AliSaif_Ali**:** Hi everyone, I’m working with ScyllaDB and had a design/operations question: I’m currently writing data with a TTL of 90 days at insert time. Now I have a use case where I need to extend the TTL of all existing data by another 90 days. • Is there a way to do this in an optimized way at scale (without rewriting all rows manually)? • Or is the best practice to simply re-write the data with a new TTL? • The dataset size is about 500 GB, so I want to be careful about the approach. What strategy would you recommend for updating/extending TTLs on existing data, and how should I proceed with this? Thanks in advance for any pointer@Felipe_Cardeneti_Mendes! @Felipe_Cardeneti_Mendes**:** you need to rewrite the data, and add another 90d to their existing TTL. @Saif_Ali**:** SSTable count: 2354 SSTables in each level: [6/4, 0, 195/100, 0, 2153] Space used (live): 342.83 GiB Space used (total): 342.83 GiB Space used by snapshots (total): 0 bytes Off heap memory used (total): 106.78 MiB SSTable Compression Ratio: 0.558728 Number of partitions (estimate): 186070 Memtable cell count: 1149 Memtable data size: 68.54 MiB Memtable off heap memory used: 91.88 MiB Memtable switch count: 37983 Local read count: 18326724 Local read latency: 38.464 ms Local write count: 9350587021 Local write latency: 0.010 ms Pending flushes: 0 Percent repaired: 0.0 Bloom filter false positives: 1817 Bloom filter false ratio: 0.00571 Bloom filter space used: 434.8 KiB Bloom filter off heap memory used: 425.6 KiB Index summary off heap memory used: 14.49 MiB Compression metadata off heap memory used: 0 bytes Compacted partition minimum bytes: 51 Compacted partition maximum bytes: 268650950 Compacted partition mean bytes: 3917309 Average live cells per slice (last five minutes): 0.0 Maximum live cells per slice (last five minutes): 0 Average tombstones per slice (last five minutes): 0.0 Maximum tombstones per slice (last five minutes): 0 Dropped Mutations: 0 bytes These are my tablestats, is there a way for me to get an estimate of how many rows are there? for me to analyse how much time will this process take? @Guy: See here: https://forum.scylladb.com/t/get-the-approximate-number-of-rows-of-a-table/196 ScyllaDB Community NoSQL Forum: Get the approximate number of rows of a table? --- ### Page: https://forum.scylladb.com/t/last-3-weeks-in-scylla-cluster-tests-git-master-issue-106-2025-09-05/5089 Title: Last 3 weeks in scylla-cluster-tests.git master (issue #106; 2025-09-05) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last 3 weeks. Commits in the 6554c1a7…6b878721 range are covered. There were 70 non-merge commits from 15 authors in… Language: en Canonical URL: https://forum.scylladb.com/t/last-3-weeks-in-scylla-cluster-tests-git-master-issue-106-2025-09-05/5089 ## Headings Structure: H1: Last 3 weeks in scylla-cluster-tests.git master (issue #106; 2025-09-05) H3: Related topics ## Main Content: H1: Last 3 weeks in scylla-cluster-tests.git master (issue #106; 2025-09-05) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last 3 weeks. Commits in the 6554c1a7…6b878721 range are covered. There were 70 non-merge commits from 15 authors in that period. Some notable commits: The test_tester suite was rewritten to use pytester, improving isolation by running in a separate process, simplifying markers, and surfacing failing tests more clearly while removing the custom make_report hook. Manager test updates: Stress logs now include timestamps in both filenames and entries, enabling easier correlation with other logs and distinguishing parallel stress runs. README has been enhanced with guidance for running SCT locally with Docker (disabling Argus, logs directory, Docker backend limitations). Added native/rclone backup benchmarking under read/write stress, measuring S3 backup duration and stress latencies. Dependency and tooling updates: See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/dropping-a-table-does-not-reduce-disk-usage-why/5090 Title: Dropping a table does not reduce disk usage. Why? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I’m trying to free up disk space, but when I drop a table, it does not reduce the storage used by ScyllaDB. What’s happening? Language: en Canonical URL: https://forum.scylladb.com/t/dropping-a-table-does-not-reduce-disk-usage-why/5090 ## Headings Structure: H1: Dropping a table does not reduce disk usage. Why? H3: Related topics ## Main Content: H1: Dropping a table does not reduce disk usage. Why? H3: Related topics I’m trying to free up disk space, but when I drop a table, it does not reduce the storage used by ScyllaDB. What’s happening? What is happening is that ScyllaDB, by default, creates a snapshot of a table just before it is dropped, as a safety measure. This can be configured in scylla.yaml - there is an auto_snapshot parameter. When it is true (which it is by default), the snapshot is created before dropping the table. The snapshot itself is in the snapshots directory, under the table SSTable. For example, for dropped table users in keyspace mykeyspace: /var/lib/scylla/data/mykeyspace/users-bdba4e60f6d511e7a2ab000000000000/snapshots/1515678531438-users As the snapshot takes the same space as the dropped table, the disk usage remains the same, which is what you are seeing. You can clean snapshots by using nodetool clearsnapshot. Read more on snapshot and clearsnapshot. Learn more about Data Modeling in this ScyllaDB University course and database administration and configuration options in the ScyllaDB Operations course. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-292-2025-09-08/5091 Title: Last week in scylladb.git master (issue #292; 2025-09-08) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the bc5773f777..bb0255b2fb range are covered. There were 109 non-merge commits from 29 authors in that per… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-292-2025-09-08/5091 ## Headings Structure: H1: Last week in scylladb.git master (issue #292; 2025-09-08) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #292; 2025-09-08) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the bc5773f777..bb0255b2fb range are covered. There were 109 non-merge commits from 29 authors in that period. Some notable commits: A few minor incompatibilities in alternator’s GetRecords implementation were fixed. The peer IP mapping is now cached, reducing slowdowns in large cluster management. The byte-ordered comparable types now include collections. This is useful for clustering keys that are collection types for the upcoming Trie sstable indexes. The vector store client now disables Nagle’s algorithm to improve latency. LWT reads can now be retried internally. A crash during a vector index search if tracing is enabled was fixed. A crash during DROP TABLE is a tablet is concurrently cleaned was fixed. Object storage (backup/restore/tables on object storage) now support the GCP object storage protocol. The system.truncated table holds the last truncation time for any table that was truncated. It is now pruned of dropped tables. The replication factor can now be omitted from the CREATE KEYSPACE statement. It will default to replicating on every rack that has nodes (excluding zero-token nodes). The task manager now shows progress of compaction tasks. The DESCRIBE MATERIALIZED VIEW statement will now show the underlying materialized view of a secondary index. Profile-guided optimizations now train on authentication and counters. The authorization system now allows users who created a CDC LOG to SELECT from it. There is now a metric for memory used by the S3 driver. Hints are now sent to pending replicas (tablet migration targets or new nodes). See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-gocql-v1-15-3/5092 Title: [RELEASE]: GoCQL v1.15.3 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: Driver Release Summary: v1.15.3 Release link: Release v1.15.3 · scylladb/gocql · GitHub What’s Changed Fixes gocql.uuid: add IsEmpty() api (#515) Run control connection reconnect in background (#529) Make ClusterConfi… Language: en Canonical URL: https://forum.scylladb.com/t/release-gocql-v1-15-3/5092 ## Headings Structure: H1: [RELEASE]: GoCQL v1.15.3 H2: What’s Changed H3: Fixes H3: Improvements H3: Related topics ## Main Content: H1: [RELEASE]: GoCQL v1.15.3 H2: What’s Changed H3: Fixes H3: Improvements H3: Related topics Driver Release Summary: v1.15.3 Release link: Release v1.15.3 · scylladb/gocql · GitHub --- ### Page: https://forum.scylladb.com/t/release-scylla-doctor-v1-7/5093 Title: [RELEASE]: Scylla Doctor v1.7 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: Scylla Doctor v1.7 is released. Enhancements: Consistent topology checks utils.py: add SSL support for cqlsh commands utils.py: enhance cqlsh SSL handling with fallback mechanism Add raft topology debug rest collect… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-doctor-v1-7/5093 ## Headings Structure: H1: [RELEASE]: Scylla Doctor v1.7 H3: Related topics ## Main Content: H1: [RELEASE]: Scylla Doctor v1.7 H3: Related topics Scylla Doctor v1.7 is released. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-107-2025-09-12/5094 Title: Last week in scylla-cluster-tests.git master (issue #107; 2025-09-12) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the d14017d8…719aa45b range are covered. There were 20 non-merge commits from 11 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-107-2025-09-12/5094 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #107; 2025-09-12) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #107; 2025-09-12) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the d14017d8…719aa45b range are covered. There were 20 non-merge commits from 11 authors in that period. Some notable commits: A CI for validating xcloud provisioning was introduced, enabling automatic checks on PRs that touch xcloud-related machinery. The cassandra-harry image was updated to build with Java 21, fixing the “cgroup v2 NullPointerException in containers” error. Repository cleanup removed several long-unused items: CMakeLists.txt, tox.ini, unused tests, and queries-limits. Latte was bumped to 0.33.0-scylladb, adding the ability to choose serial consistency for LWT. Example: --serial-consistency=LOCAL_SERIAL --consistency=LOCAL_QUORUM . Performance regression tests for Alternator were added, covering read, write, and mixed scenarios for CQL (baseline), Alternator default, and Alternator with forced LWT. Scylla Cloud backend improvements: VPC peering support between Scylla Cloud clusters and SCT infrastructure and dynamic CIDR allocation for new clusters to avoid route conflicts. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/time-series-data-modelling-with-scylla/5095 Title: Time series data modelling with Scylla - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello! I am trying to learn better data modelling with Scylla and wanted to understand what would be the optimal setup for a time-series use-case..As an example, let’s say I have data coming in from multiple sensors (ea… Language: en Canonical URL: https://forum.scylladb.com/t/time-series-data-modelling-with-scylla/5095 ## Headings Structure: H1: Time series data modelling with Scylla H3: Related topics ## Main Content: H1: Time series data modelling with Scylla H3: Related topics I am trying to learn better data modelling with Scylla and wanted to understand what would be the optimal setup for a time-series use-case..As an example, let’s say I have data coming in from multiple sensors (each sensor has a unique ID) for a few hours everyday (at 20k rps).. I need to store the sensor data itself in a table and also store 10s, 1min, 10min, 1hour, 1d, 1w, 1month, 1year aggregates, (first, last, min, max, avg) – which the ingestor will compute and INSERT. For the raw sensor data, I think a good setup would be to have sensor_raw table with PRIMARY KEY ((sensor_id, date), timestamp) (order by desc) ? i could have a single aggregates table with PRIMARY KEY ((sensor_id, period), timestamp) (order by desc).. but i believe this would be bad since the partition will keep growing forever. I could have a single aggregates table with PRIMARY KEY ((sensor_id, period, date), timestamp) … this is probably better from a partition size perspective. but for higher periods like 1d, 1w, 1month, etc, client side logic would be very messy and i would have to force a single compaction strategy for all partitions.. Other option is to have multiple aggregates table like aggregates_ with PRIMARY KEY ((sensor_id, bucket), timestamp) structure where the bucket for lower timeframes is yyyy-mm-dd, and for higher timeframes like 1d, 1w is yyyy-mm etc.. This way, for lower timeframes, the compaction can be time-tiered. and for higher timeframes it can be size-tiered etc.. What is the best practice here? P.S.: This is a repost from Slack (general channel) For optimal time-series data modeling in Scylla handling sensor data at high write rates with multiple aggregation levels, best practices include: Avoid a single aggregates table partitioned only by (sensor_id, period) because partitions grow indefinitely. Using (sensor_id, period, date) as partition key is better for partition size management but complicates client logic and forces uniform compaction strategy. A recommended approach is to create multiple aggregate tables for different periods, e.g., aggregates_10s, aggregates_1min, aggregates_1d, each with a primary key like PRIMARY KEY ((sensor_id, bucket), timestamp). Define the bucket differently per aggregation period: for finer granularities like seconds or minutes, use daily buckets (e.g., yyyy-mm-dd), while for coarser granularities like daily, weekly, monthly, use larger buckets such as yyyy-mm or yyyy. This allows tailoring compaction strategies (time-tiered for frequent writes, size-tiered for less frequent ones) per aggregation table, optimizing read/write performance and data retention. In summary, multiple aggregation tables segmented by appropriate time buckets are preferable for time series in Scylla since they provide better partition size control, flexible compaction strategies, and cleaner client-side logic for higher-level aggregates Here’s what I ended up doing: For anything ≤ 1 min aggregation, I’m using monthly bucketed partitions (small tier). For our use case, since data input isn’t 24/7, each partition should stay under ~15 MB even at 5 s aggregation. From 1 min up to (but not including) 1 day, I use yearly buckets (medium tier). 1 day aggregations go into a century bucket with a hardcoded key 2000-01-01 (large tier). So, three tables total, with primary key ((sensor_id, aggregation_size, bucket), timestamp) I’m using TWCS for all bucket tables. The compaction interval is roughly aligned with the bucket period. Does that make sense, or would a daily compaction interval be better? Also, any tips for choosing the caching settings? Right now, I’ve set caching = { 'keys': 'ALL' }. Are there best practices for rows_per_partition (set to ALL or a specific number)? I couldn’t find clear details in the OSS docs. Our read pattern is usually “last 1000 rows.” For the 1min queries i am seeing high latencies .. I am wondering if this compaction setup is the problem. My exact schema looks something like this; It’s good that you are using TimeWindowCompactionStrategy (TWCS) and bucketing for partition size control. However, high query latencies for 1-minute aggregates can be due to several factors: 1. Compaction Window/Partition Alignment Your compaction window is set to 31 days but your buckets for 1-min aggregations are yearly—this mismatch might leave many SSTables open for reads, especially for recent data. Try aligning your compaction window to the granularity of your bucket (i.e., if your bucket is per year, consider setting compaction_window_size to 365, compaction_window_unit to DAYS for yearly buckets; for monthly buckets, use 30-31 days). This helps TWCS drop expired SSTables more efficiently and reduces SSTable fan-out on reads.​ If your partitions (with yearly buckets for 1-min data) grow too large, any read (such as “last 1000 rows”) may need to scan too many SSTables and tombstones. Monitor partition size with nodetool cfstats or Scylla Monitoring to see if you should reduce bucket duration even further. Setting 'rows_per_partition': '500' is a reasonable start for “last 1000 rows” access, but ALL isn’t always best for every read pattern if partitions get large. Experiment with increasing ‘rows_per_partition’ to match your most common query size or consider switching to partition-level caching if reads always load different ranges. 4. Speculative Retry and Latency Spikes 5. Bloom Filter and Read Statistics 6. SSTable Expiry and Query Filtering Summary Table: Possible Tuning Points Align the compaction window and bucket length. Check partition sizes and reduce the bucket if queries still slow down. Match caching rows to your “last N rows” queries. Monitor with Scylla Monitoring for hot partitions or pending compactions. Thank you for the detailed response. Really appreciate it! My 1minute aggregations are in monthly bucket only (small tier table with monthly is used for <=1minute aggregations). And their compaction is set to 31 Days. And other tables also have their compaction window set to align their bucketing window. But with cfstats i can see “high” number of SSTables (>100 is considered high i believe?).. For example, right now for the medium tier table (yearly bucket), I see SSTable count: 334; SSTables in each level: [334/4].. Caching, TTL & Partition Sizes Great questions—here’s my input on each: Compactions & SSTable Count Healthy SSTable Count: For TWCS tables, 100+ SSTables isn’t unusual for active, recent buckets, but sustained high counts (e.g., 334) may indicate compaction backlog, especially if data is uniformly loaded. It’s normal for older buckets (with mostly expired or cold data) to eventually compact down to fewer files. Try monitoring compaction throughput and queue length via Scylla Monitoring—what matters most is that compactions catch up regularly, especially for recent data. Manual Major Compaction: If major compaction reduces SSTable count sharply, frequent manual runs shouldn’t be needed unless you see prolonged backlog or latency spikes—ideally, TWCS should handle active buckets automatically. Guidelines: There isn’t a strict “healthy” number; aim for <100 SSTables per active bucket for most read patterns, but higher counts are okay if read latency, bloom filter hits, and compaction metrics are all healthy. Client vs. Server: Using client-side speculative retries (via Rust driver) is a solid strategy—this places low-latency control closer to your application. If you aren’t facing frequent server-side timeouts, consider disabling server-side speculative_retry and tuning this in your client instead. Metric Visibility: Scylla Monitoring tracks server-side speculative executions via the scylla_coordinator_speculative_retry metric. For most setups, watch for spikes in this metric that align with your app’s traffic and tail latency events. Caching, TTL, Partition Size Experimenting with rows_per_partition: Definitely keep adjusting—ideally, set this to match your “last N rows” query size for fastest cache-based lookup. TTL handling: It’s fine to avoid TTL if business archival needs require it; just ensure you don’t accumulate obsolete data in cold buckets. Partition Size: Reducing partition size to ~20 MB (from 100 MB) is excellent—you should see improved query performance and compaction health. I have applied all these changes (disable server-side speculative retry, tune caching, etc.) But i still see very big spikes in some cases. One suspicion I have right now is if the connection is getting overloaded – We have 600rps inbound api requests which turns into 1.6k rps actual queries per second to Scylla DB due to the bucketing. Is it possible that this is choking the single connection to each node that we have? There is option to set custom pool size, but the documentation says recommended is 1. SessionBuilder in scylla::client::session_builder - Rust This is the Scylla session setup I have: It is possible that the connection saturation and latency spikes you observe are related to the single connection per node setup in your session pool, which can become a bottleneck under high query rates. Although the official documentation recommends a pool size of 1, this is typically for lighter workloads. For higher throughput scenarios like yours—600 inbound API requests translating to 1.6k queries per second—it is recommended to increase the connection pool size to multiple connections per node to better distribute load and avoid connection choking. You can configure this in your Rust driver setup like so: Additionally, it is crucial to use the official Scylla Rust driver to ensure proper connection pooling, shard awareness, and optimized concurrency handling tailored for ScyllaDB’s architecture. Non-official or generic drivers may lack these optimizations, leading to bottlenecks. The official Scylla Rust driver is fully asynchronous, shard-aware, and designed for high-performance ScyllaDB access. You can find the official Scylla Rust driver repository here: https://github.com/scylladb/scylla-rust-driver Increasing the connection pool size beyond 1 is advisable to handle your high query rate and avoid choking a single connection. Verify you are using the official Scylla Rust driver to leverage the best connection management and performance features. Monitor client and server metrics to tune connection pool size and concurrency for your workload. This approach should help reduce latency spikes and improve throughput under load. If problems persist, sharing detailed metrics and client logs will help diagnose further. We are using the official driver. But the way to set pool size is different though. As per the docs here: GenericSessionBuilder in scylla::client::session_builder - Rust There is PerHost and PerShard pool sizes. Which one is preferred to be set? (We are using shard aware port) It is possible that the connection saturation and latency spikes you observe are related to the single connection per node setup in your session pool, which can become a bottleneck under high query rates. Although the official documentation recommends a pool size of 1, this is typically for lighter workloads. inbound API requests translating to 1.6k queries per second—it is recommended to increase the connection pool size to multiple connections per node to better distribute load and avoid connection choking. Increasing the connection pool size will allow for better concurrency and load balancing across the node, directly addressing the observed latency spikes crucial for your high-throughput time series application hmm thats pretty cool --- ### Page: https://forum.scylladb.com/t/release-scylladb-2025-1-7/5096 Title: [RELEASE] ScyllaDB 2025.1.7 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB 2025.1.7, a bug-fix production-ready patch release for ScyllaDB 2025.1 LTS Release. Note there is a new Short Term Support (STS) Feature release 2025.3. You are welcome to upgrade to… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-2025-1-7/5096 ## Headings Structure: H1: [RELEASE] ScyllaDB 2025.1.7 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 2025.1.7 H3: Related topics The ScyllaDB team announces ScyllaDB 2025.1.7, a bug-fix production-ready patch release for ScyllaDB 2025.1 LTS Release. Note there is a new Short Term Support (STS) Feature release 2025.3. You are welcome to upgrade to it for the latest and greatest features, or stay on the 2025.1 track for long term support. Upgrade from ScyllaDB Enterprise 2024.x to ScyllaDB 2025.1 Upgrade from ScyllaDB Open Source 6.2 to ScyllaDB Enterprise 2025.1.x The following issues are fixed in this release: service/qos: Modularize service level controller to avoid invalid access to a stopped auth::service #24792 Commitlog: segment_manager::discard_unused_segments - separate the modification of the vector (_segments) from actual releasing of objects #25709 Drop table during tablet cleanup initiated by a tablet migration may crash #25706 gossiper - race condition in gossip::apply_state_locally when receiving messages containing removed endpoints that no longer have a valid host id #25702, #25621 gossiper can reply with an empty self host id in `gossip_get_endpoint_states_response` during startup #25831 Tablets Repair: A node may crash If a node has many shards and many tablets on each shard #23632 Segmentation fault in scheduling group streaming during tablets streaming after node was decommissioned #23162 Improved handling of in-progress requests during node shutdown, ensuring requests are allowed to complete successfully. #24481 Removed a large allocation when loading SSTables #3335 raft_sys_table_storage: avoid an extra copy when deserializing log_entry #23903, seastar_memory - Removed an oversized allocation changing token_range_vectorto use chunked_vector #24156, #24115 storage_service::maybe_reconnect_to_preferred_ip() calls the gossiper::get_host_id() unnecessarily and can be passed directly as a parameter #25715 Optimized container footprint – The container image size has been minimized to improve efficiency #25479 Backport ‘reactor: Update seastar submodule to get more info and robustness on segfault’ scylladb#25681 db/hints: Improve logs #25466 Debug logs: deletion_time formatter prints wrong results #25556 db/commitlog: Extend error messages for corrupted data #25459 Warn users about using RF!=Racks with Tablets Keyspaces #23330 storage_service/get_natural_endpoints has no way to pass a key containing a colon character #24829 Alternator, ScyllaDB’s implementation of the DynamoDB API, is now more careful to avoid large contiguous allocations during query/scan operations, as these can cause stalls. #23535 Improve performance of worst case scenario in auth::passwords::detail::hash_with_salt #24524 Improve performance for partial repair with a node down and with many tablets #22413 storage_service: on_change: frequent system.peers reloads with many peers may cause high CPU usage #25660 system_keyspace::drop_truncation_rp_records can cause unbound parallelism #25682 --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-293-2025-09-14/5097 Title: Last week in scylladb.git master (issue #293; 2025-09-14) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the bb0255b2fb..5307d1b9a8 range are covered. There were 69 non-merge commits from 19 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-293-2025-09-14/5097 ## Headings Structure: H1: Last week in scylladb.git master (issue #293; 2025-09-14) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #293; 2025-09-14) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the bb0255b2fb..5307d1b9a8 range are covered. There were 69 non-merge commits from 19 authors in that period. Some notable commits: Gossiper was hardened against a race condition that could leave empty host entries, that could later lead to a crash. The REST API for repair can now select incremental mode as an option. When tablets are split, a compaction process is initiated to break apart sstables that span the new tablet boundary. This compaction is now prepared for a tablet merge to happen before the split compaction is complete. The scylla sstable tool now uses the more modern UUID generations rather than numeric generations. The CQL binary protocol server now avoids exceptions when processing protocol errors; these happen when drivers negotiate the supported protocol version with the server. Avoiding exceptions reduces CPU use during connection storms. The nodetool stop cleanup command did not stop all cleanup tasks; only the ones currently running. Pending cleanup compactions would still run. It now abort all pending tasks. Vector indexes are now versioned. This is used to synchronize metadata between the database and the vector search nodes. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/adding-a-node-to-an-existing-cluster-node-stuck-in-joining-status/5098 Title: Adding a node to an existing cluster, node stuck in joining status - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/adding-a-node-to-an-existing-cluster-node-stuck-in-joining-status/5098 ## Headings Structure: H1: Adding a node to an existing cluster, node stuck in joining status H3: Related topics ## Main Content: H1: Adding a node to an existing cluster, node stuck in joining status H3: Related topics Originally from the User Slack @Shivaprasad_BhatShivaprasad_Bhat**:** Hello.. I added a node to an existing Scylla cluster.. Running nodetool status on any of the cluster nodes shows all nodes as UN .. It’s been ~~4 hours and all i see is this. not even sure if it’s working (we have ~~90GB of data). Any th@aviughts here? @avi**:** Check the Advanced dashboard, look for Streaming I/O bandwidth and CPU, it indicates bandwidth for streaming But if it’s just 90GB, it’s probably stuck, what version a@Shivaprasad_Bhate you running? @Shivaprasad_Bhat**:** i let it run and went to sleep. it was done by morning. might have been because of the small node size we hav@avi. we are running scylla 6.0.0 @avi**:** It shou@Shivaprasad_Bhatd have completed in a few minutes @Shivaprasad_Bhat**:** We had a cluster of 4 nodes with 2 cores each (i4i.large).. I have upgraded this to 3 node cluster with 4 cores each (i4i.xlarge).. Could this time be related to having very large partitions? we have time-series sort of data but partitions are not time bucketed so they are ever growing (many partitions are around 70MB – checked from the large_paritions system table). Also, the compaction is set to STCS which is also not ideal for time-series@avi.. would that compaction affect the time as well? @avi**:** 70MB is not very large, and should not affect steaming performance@Shivaprasad_Bhat STCS is not ideal but should also not be a problem. @Shivaprasad_Bhat**:** Ohh. Overall, I have seen that whenever I try to add a node, compaction runs for a few hours first And then streaming starts which also takes a few hours. Not sure if it was because we had 2 core nodes befo@avie which was not enough (during this, cpu is maxed out on all nodes). @avi**:** C@Shivaprasad_Bhatmpacting 70GB should take ~15 minutes, not hours What hardware is this? @Shivaprasad_Bhat**:** i4i.large AWS. Nitro S@aviD. 2 vCPU, 16GB memory Any pointers on what other factors would impa@Shivaprasad_Bhatt compaction time? @avi**:** It should run at 10-40 MB/s/shard, depending on the data model @Shivaprasad_Bhat**:** I’m suspecting an issue with the data model itself (we also see random latency spikes which are increasing with increasing data size). It’s pure time series data but with STCS, and no time bucketing. https://scylladb-users.slack.com/archives/C2NLNBXLN/p1757748173227739 The use case is similar to this one. Let me know your thoughts / suggestions on this as well. I’m going through the Scylla data modeling course as well @avi**:** If you have tiny cells and/or huge amounts of tombstones, it can slow down compaction. You can try moving to i7i instances which have faster CPUs --- ### Page: https://forum.scylladb.com/t/release-scylladb-2025-3-1/5099 Title: [RELEASE] ScyllaDB 2025.3.1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of ScyllaDB 2025.3.1, a production-ready path release for ScyllaDB 2025.3 Feature Release. Related Links Read more about ScyllaDB Get ScyllaDB 2025.3 Upgrade… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-2025-3-1/5099 ## Headings Structure: H1: [RELEASE] ScyllaDB 2025.3.1 H2: Related Links H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 2025.3.1 H2: Related Links H3: Related topics The ScyllaDB team is pleased to announce the release of ScyllaDB 2025.3.1, a production-ready path release for ScyllaDB 2025.3 Feature Release. Read more about ScyllaDB Upgrade from ScyllaDB 2025.2 to ScyllaDB 2025.3 The following issues are fixed in this release: --- ### Page: https://forum.scylladb.com/t/time-to-live-ttl-expiring-data-new-explanation-video/5101 Title: Time to Live (TTL), expiring data new explanation video - University and Training - ScyllaDB Community NoSQL Forum Meta Description: Check out this new video explaining how to use TTL in ScyllaDB, by @Attila_Toth . If you have any questions, you can ask here. Language: en Canonical URL: https://forum.scylladb.com/t/time-to-live-ttl-expiring-data-new-explanation-video/5101 ## Headings Structure: H1: Time to Live (TTL), expiring data new explanation video H3: Related topics ## Main Content: H1: Time to Live (TTL), expiring data new explanation video H3: Related topics Check out this new video explaining how to use TTL in ScyllaDB, by @Attila_Toth . If you have any questions, you can ask here. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-108-2025-09-19/5103 Title: Last week in scylla-cluster-tests.git master (issue #108; 2025-09-19) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 6340efcd…a7af7bda range are covered. There were 40 non-merge commits from 12 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-108-2025-09-19/5103 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #108; 2025-09-19) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #108; 2025-09-19) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 6340efcd…a7af7bda range are covered. There were 40 non-merge commits from 12 authors in that period. Some notable commits: Updated scylladb/gemini Docker tag to v2.1.4. Added Azure Key Management Service (KMS) integration: Healthcheck is now wrapped with adaptive timeout, making healthcheck duration visible in Argus Results (closes #9170). Scylla-Cloud support improvements: New longevity test configuration for LWT with tablets: switched from c-s to latte; added a loader script with LWT queries and deletions to trigger tablet merge/split. Added more actions log lines to cover all disruptions. Performance tests use default cassandra-stress image. 2025.3 and 2025.4 releases were added to performance trigger matrix. Enabled scylla_doctor checks in non-root artifact test. Some are not working due missing root permissions and were disabled. Added AGENTS.md to improve Copilot/LLM Agents experience with a codebase overview and common commands. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/release-scylla-operator-1-18-1/5104 Title: [RELEASE] Scylla Operator 1.18.1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: We are pleased to announce the release of Scylla Operator 1.18.1. Release highlights: Fix potential data races in RenderTemplate function (#2967,@czeslavo) Update prometheus-operator types to a pre-v0.86.0 commit (695… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-operator-1-18-1/5104 ## Headings Structure: H1: [RELEASE] Scylla Operator 1.18.1 H3: Related topics ## Main Content: H1: [RELEASE] Scylla Operator 1.18.1 H3: Related topics We are pleased to announce the release of Scylla Operator 1.18.1. For the full list of changes, please refer to the release notes. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-294-2025-09-21/5105 Title: Last week in scylladb.git master (issue #294; 2025-09-21) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 5307d1b9a8..1690e5265a range are covered. There were 170 non-merge commits from 33 authors in that per… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-294-2025-09-21/5105 ## Headings Structure: H1: Last week in scylladb.git master (issue #294; 2025-09-21) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #294; 2025-09-21) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 5307d1b9a8..1690e5265a range are covered. There were 170 non-merge commits from 33 authors in that period. Some notable commits: Change Data Capture (CDC) now works with tablet-enabled tables. There is a new automatically created service level named driver. It is used during driver connections and for control connections. This reduces workload disruptions during connection storms. An sstable Bloom filter is built with an estimate of the number of partitions it will hold. If the estimate turns out to be incorrect, we rebuild the Bloom filter in order not to waste memory. We now avoid the rebuild if the memory wasted is low enough to be ignored. After streaming or repairing data, the portion of the row cache affected is invalidated since it no longer reflects the underlying sstables. We now invalidate at partition granularity rather than token-range granularity, resulting in increased cache efficiency, particularly after repair. The nodetool cluster repair command gained an --incremental-mode option. Gossiper operations are now enforced to run in the gossip scheduling group, even if invoked from other components. The CREATE KEYSPACE ks statement can now be executed without any optional clauses; it will create a keyspace using NetworkTopologyStrategy with replication to all racks in all datacenters. The table used for the Raft log now has caching disabled. There are now metrics for S3 prefetches. Lightweight transactions (LWT) now implement fencing, which prevents old requests that were send using an old version of the topology from being incorrectly applied onto a new topology. This is a step for implementing LWT on tablets. The system.clients columns reporting encrypted connections (SSL/TLS) are now filled in. The tablet load balancer now considers dead nodes in its calculations. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/request-for-guidance-using-scylladb-for-api-response-caching-poc-in-progress/5106 Title: Request for Guidance: Using ScyllaDB for API Response Caching (POC in Progress) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Dear ScyllaDB Team, I am currently working on a POC (Proof of Concept) using ScyllaDB on-prem trial version as a caching layer for some of our APIs, such as transaction history , account summary etc. The goal is to redu… Language: en Canonical URL: https://forum.scylladb.com/t/request-for-guidance-using-scylladb-for-api-response-caching-poc-in-progress/5106 ## Headings Structure: H1: Request for Guidance: Using ScyllaDB for API Response Caching (POC in Progress) H3: Related topics ## Main Content: H1: Request for Guidance: Using ScyllaDB for API Response Caching (POC in Progress) H3: Related topics I am currently working on a POC (Proof of Concept) using ScyllaDB on-prem trial version as a caching layer for some of our APIs, such as transaction history , account summary etc. The goal is to reduce load on our primary database and serve frequently accessed responses with low latency. As part of this POC, I want to validate the best approach and architecture to utilize ScyllaDB effectively for this scenario. Below are some design options I am considering, and I would appreciate your input or alternative recommendations. Please feel free to suggest best alternatives, your guidance would be very valuable: Request Hash–based Caching Store API responses against a unique request hash. Concern: If requests grow into millions, ScyllaDB may accumulate old unused data. Question: Should we rely on TTL for automatic cleanup, or are there better strategies for eviction and cache invalidation? Save request parameters directly as columns with the API response. Concern: If new request parameters are introduced later, schema evolution may be required. Question: Is this a recommended approach for caching APIs where input parameters may evolve over time? We would highly value your guidance on: As this is part of our ongoing POC, your suggestions will be critical in shaping our final architecture and ensuring that we leverage ScyllaDB’s strengths effectively. Looking forward to your recommendations. Best regards, Dadasaheb What is the primary database you’re using? Some users started out using ScyllaDB as a caching layer, but later switched the entire database to ScyllaDB to get better results. This webinar is a useful resource. A replication factor of three works for many use cases. Regarding sizing, it really depends, what is your payload size? This database and cache internals blog post might also be relevant, and maybe can add more info. --- ### Page: https://forum.scylladb.com/t/exception-while-connecting-to-scylladb-with-ssl/5108 Title: Exception while connecting to ScyllaDB with SSL - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I’m trying to connect with our on prem ScyllaDB using ScyllaDBCSharpDriver via our .net 8 web API in visual studio, but i’m getting this exception. Cassandra.NoHostAvailableException: All hosts tried for query failed (… Language: en Canonical URL: https://forum.scylladb.com/t/exception-while-connecting-to-scylladb-with-ssl/5108 ## Headings Structure: H1: Exception while connecting to ScyllaDB with SSL H3: Related topics ## Main Content: H1: Exception while connecting to ScyllaDB with SSL H3: Related topics I’m trying to connect with our on prem ScyllaDB using ScyllaDBCSharpDriver via our .net 8 web API in visual studio, but i’m getting this exception. Cassandra.NoHostAvailableException: All hosts tried for query failed (tried x.x.x.x:9042: TimeoutException ‘The timeout period elapsed prior to completion of SSL authentication operation.’; y.y.y.y:9042: TimeoutException ‘The timeout period elapsed prior to completion of SSL authentication operation.’) The exception is thrown at cluster.Conntect The port 9042, which you point the driver at, is normally configured for nonencrypted communication. The encrypted counterpart is 9142. Please try specifying 9142 as the connection port in the driver and let me know if it works. Hi Thanks a lot for the reply. However, I was able to figure this out. There was an issue with the certificate i used. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-109-2025-09-26/5109 Title: Last week in scylla-cluster-tests.git master (issue #109; 2025-09-26) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the ffcf305e…27ef7a07 range are covered. There were 9 non-merge commits from 5 authors in that… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-109-2025-09-26/5109 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #109; 2025-09-26) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #109; 2025-09-26) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the ffcf305e…27ef7a07 range are covered. There were 9 non-merge commits from 5 authors in that period. Some notable commits: Vector-store is now supported in AWS backend. Local dev environment setup readme was updated covering uv installation and configuration. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-21/5113 Title: [RELEASE] ScyllaDB Enterprise 2024.1.21 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB Enterprise 2024.1.21, a bug-fix production-ready patch release for ScyllaDB Enterprise 2024.1 Long Term Support (LTS) Release. Note there is a later, better, LTS (Long term Support) R… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-enterprise-2024-1-21/5113 ## Headings Structure: H1: [RELEASE] ScyllaDB Enterprise 2024.1.21 H3: Related Links H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Enterprise 2024.1.21 H3: Related Links H4: Fixed Issue with an source available reference: H3: Related topics The ScyllaDB team announces ScyllaDB Enterprise 2024.1.21, a bug-fix production-ready patch release for ScyllaDB Enterprise 2024.1 Long Term Support (LTS) Release. Note there is a later, better, LTS (Long term Support) Release 2025.1, and a Feature Release 2025.3. You are encouraged to upgrade in coordination with the ScyllaDB Support team. --- ### Page: https://forum.scylladb.com/t/incremental-backup-repair-compaction-and-disk-space/5114 Title: Incremental backup, repair, compaction and disk space - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/incremental-backup-repair-compaction-and-disk-space/5114 ## Headings Structure: H1: Incremental backup, repair, compaction and disk space H3: Related topics ## Main Content: H1: Incremental backup, repair, compaction and disk space H3: Related topics Originally from the User Slack @Mahdi_KamaliMahdi_Kamali**:** Does incremental backup include the data received in a repair operation? What about new SSTables generated in com@doraction? @dor**:** yes, data written from repair is written as a standard new sstable (mixed with other data that came in to tha@Mahdi_Kamali memtable) @Mahdi_Kamali**:** Ni@dore! W@dorat about compaction? @dor @dor**:** If you turn on incremental backup, it copy-on-write every sstable, including all new s@Mahdi_Kamalitables created by@dorcompaction @Mahdi_Kamali**:** Are you sure?! @dor Therefore running compaction will fill the disk space. Because new large SSTables will be created and that will almost duplicate the data. Am I right? @dor**:** Compactions removes data eventually too --- ### Page: https://forum.scylladb.com/t/scylladb-cloud-presents-the-new-x-cloud-beta/5116 Title: ScyllaDB Cloud presents the new X Cloud (Beta) - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: I am excited to announce the next step in our journey toward an amazingly fast, truly elastic, and cost-efficient ScyllaDB Cloud. The new X Cloud product builds on the technical achievements of ScyllaDB and ScyllaDB Clo… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-cloud-presents-the-new-x-cloud-beta/5116 ## Headings Structure: H1: ScyllaDB Cloud presents the new X Cloud (Beta) H3: Tablets Technology H3: Faster Infrastructure H3: Autonomous Scaling H3: Availability H3: Still Beta H3: Related topics ## Main Content: H1: ScyllaDB Cloud presents the new X Cloud (Beta) H3: Tablets Technology H3: Faster Infrastructure H3: Autonomous Scaling H3: Availability H3: Still Beta H3: Related topics I am excited to announce the next step in our journey toward an amazingly fast, truly elastic, and cost-efficient ScyllaDB Cloud. The new X Cloud product builds on the technical achievements of ScyllaDB and ScyllaDB Cloud engineering teams, delivering excellent scalability and efficiency. Harnessing the full potential of ScyllaDB tablets to deliver: We have optimized and parallelized the cluster deployment process to significantly reduce provisioning time. The data is distributed much faster, and the new nodes start serving in minutes. With the new X Cloud Beta, we introduce autonomous scaling, a capability that continuously keeps clusters running at optimal storage efficiency, based on the data from the cluster monitoring. Clusters scale in and out automatically to maintain up to 90% disk utilization, ensuring efficient resource use and cost control. Customers can still reserve a minimum vCPU or storage capacity to ensure clusters never scale below a defined baseline—even if idle—so they’re ready for any planned workload spikes. By combining these technologies, your cluster can now scale just-in-time, in precise increments, and be ready to serve requests from all nodes within minutes—while also optimizing storage utilization to save you even more. For more information, I recommend reading the following X Cloud articles: Introducing ScyllaDB X Cloud: A (Mostly) Technical Overview ScyllaDB X Cloud: An Inside Look with Avi Kivity (Part 1) Alan Shimel and Dor Laor on Database Elasticity, ScyllaDB X Cloud - ScyllaDB X Cloud Beta is available today in ScyllaDB Cloud for all customer accounts. If you need documentation, you can always find it here: X Cloud Documentation. We strongly recommend using X Cloud Beta only for testing and evaluation, not for production workloads. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-110-2025-10-03/5117 Title: Last week in scylla-cluster-tests.git master (issue #110; 2025-10-03) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 677f7e3d…d07d943b range are covered. There were 9 non-merge commits from 7 authors in that… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-110-2025-10-03/5117 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #110; 2025-10-03) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #110; 2025-10-03) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 677f7e3d…d07d943b range are covered. There were 9 non-merge commits from 7 authors in that period. Some notable commits: Refactored MV/SI creation utilities, consolidating materialized view and secondary index helpers in one place. Introduced new nemesis disrupt_kill_mv_building_coordinator. The nemesis locates the topology leader running the coordinator, kills Scylla there to trigger leader re-election, and asserts that MV creation successfully completes afterwards. Latte bumped to 0.40.0-scylladb, adding support for business-logic-side “data validation” in rune scripts. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/removing-and-adding-a-column-with-the-same-name-getting-an-error/5118 Title: Removing and adding a column with the same name, getting an error - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/removing-and-adding-a-column-with-the-same-name-getting-an-error/5118 ## Headings Structure: H1: Removing and adding a column with the same name, getting an error H3: Related topics ## Main Content: H1: Removing and adding a column with the same name, getting an error H3: Related topics Originally from the User Slack @KishoreKishore**:** Hi, We recently added a new User-Defined Type (UDT) column to one of our tables, but subsequently dropped it. Upon attempting to re-add the same column, we encountered the following error: Please advise on the root cause and suggest a solution? @Botond_Dénes: The solution is to use a different name for the recreated column. @Kishore**:** okay, thanks. --- ### Page: https://forum.scylladb.com/t/last-fortnight-in-scylladb-git-master-issue-295-2025-10-05/5119 Title: Last fortnight in scylladb.git master (issue #295; 2025-10-05) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last two week. Commits in the 1690e5265a..20aeed1607 range are covered. There were 266 non-merge commits from 40 authors in that… Language: en Canonical URL: https://forum.scylladb.com/t/last-fortnight-in-scylladb-git-master-issue-295-2025-10-05/5119 ## Headings Structure: H1: Last fortnight in scylladb.git master (issue #295; 2025-10-05) H3: Related topics ## Main Content: H1: Last fortnight in scylladb.git master (issue #295; 2025-10-05) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last two week. Commits in the 1690e5265a..20aeed1607 range are covered. There were 266 non-merge commits from 40 authors in that period. Some notable commits: The tablet load balancer now runs in the maintenance/streaming scheduling group, rather than the gossip group. This reduces impact on node failure detection and Raft which also run in the gossip group. The CREATE TABLE IF NOT EXISTS no longer fails if CDC is specified. There are now Lua scripts for examining purgable tobmstone and write-time histograms in sstables. The batchlog mechanism will now drop batches for a table that was itself dropped. The tablet scheduler can schedule migrations in different racks independently. It is now more careful not to generate conflicting tablet migration plans. The tablet scheduler will no longer attempt expensive cross-rack migrations if the cluster is known to have exactly one tablet replica per rack. This is much faster. The sstable scrubber now handles malformed sstables better. We now store more complete schemas in sstables, so data recovery from a stray sstable is easier. There is now a standalone scylla sstable upgrade command, which can be used to rewrite sstables using different versions or options. It is now possible to set the default sstable compression options via configuration. Alternator, ScyllaDB’s implementation of the DynamoDB API, now caches parsed expressions to reduce per-request overhead. There is a new sstable version, ms. This is like me but with Trie primary indexes. The index is more compact and can yield faster lookups. It is not yet enabled by default. ScyllaDB now formats the data filesystem using 4k block size. We previously used 1k block size to work around a kernel deficiency, which has since been fixed. The nodetool getendpoints command now works with keys containing : for more cases. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-2025-2-3/5120 Title: [RELEASE] ScyllaDB 2025.2.3 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB 2025.2.3, a bug-fix production-ready patch release for ScyllaDB 2025.2 Feature Release. Note there is a new Short Term Support (STS) Feature release 2025.3. You are welcome to upgrad… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-2025-2-3/5120 ## Headings Structure: H1: [RELEASE] ScyllaDB 2025.2.3 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 2025.2.3 H3: Related topics The ScyllaDB team announces ScyllaDB 2025.2.3, a bug-fix production-ready patch release for ScyllaDB 2025.2 Feature Release. Note there is a new Short Term Support (STS) Feature release 2025.3. You are welcome to upgrade to it for the latest and greatest features. The following issues are fixed in this release: --- ### Page: https://forum.scylladb.com/t/checklist-of-scylladb-files-to-run-with-antivirus/5121 Title: Checklist of scylladb files to run with antivirus - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 2025.3.1 #Cluster size: standalone os (RHEL/CentOS/Ubuntu/AWS AMI): Ubuntu (Azure VM) Hi , I am facing issue while running scylladb with sentinelone antivirus. Sentinelone is … Language: en Canonical URL: https://forum.scylladb.com/t/checklist-of-scylladb-files-to-run-with-antivirus/5121 ## Headings Structure: H1: Checklist of scylladb files to run with antivirus H2: Essential ScyllaDB File and Directory Exclusions H2: Additional Recommendations H2: Exclusion Principles for Sentinelone H2: Advisory Notes H3: Related topics ## Main Content: H1: Checklist of scylladb files to run with antivirus H2: Essential ScyllaDB File and Directory Exclusions H2: Additional Recommendations H2: Exclusion Principles for Sentinelone H2: Advisory Notes H3: Related topics Installation details #ScyllaDB version: 2025.3.1 #Cluster size: standalone os (RHEL/CentOS/Ubuntu/AWS AMI): Ubuntu (Azure VM) Hi , I am facing issue while running scylladb with sentinelone antivirus. Sentinelone is killing the scylla process. I have allowed most of the scylla files which are required to run the scylla , but still getting issue . screenshot shared below. Can i get any checklist of scylla files to allow ,so that scylla can run. Below is a comprehensive checklist based on official ScyllaDB security and requirements docs, as well as database best practices for AV exclusions. For a standalone ScyllaDB deployment on Ubuntu, you should exclude these from antivirus real-time scanning and mitigation, especially with Sentinelone: ScyllaDB server binary (default location): ScyllaDB system service scripts (typically): Data directories (most critical, required for disk access and integrity): /var/lib/scylla/data/ /var/lib/scylla/commitlog/ /var/lib/scylla/hints/ /var/lib/scylla/saved_caches/ Runtime directories and logs: /tmp/scylla* (used for temporary files and socket communication) Whitelist any custom install location if you use non-default paths for binaries or configs. For upgrade scripts or web installer, also exclude: Do NOT exclude entire /var/ or /usr/ unless the path is dedicated to ScyllaDB. Use path-based exclusions for all listed directories and binaries. If Sentinelone flagged specific binaries, exclude the SHA1 hash as a targeted file exclusion.​ Avoid excluding the entire system; only necessary paths/binaries to minimize risk.​ After adding exclusions, restart Sentinelone services and verify that ScyllaDB processes (scylla, scylla-server) can start and run without interruption. Running a database engine with antivirus, even with exclusions, can sometimes interfere with performance and stability. Ensure excluded directories (especially /var/lib/scylla/) are never quarantined or locked.​ Always keep ScyllaDB up to date for best security alongside active antivirus. Summary Table: ScyllaDB Antivirus Exclusion Checklist Ensure AV exclusions are correctly configured in the Sentinelone console, either by hash or directory path, as appropriate. Thanks a lot Gabriel. --- ### Page: https://forum.scylladb.com/t/how-to-migrate-data-from-bitnami-scylladb-to-scylla-in-kubernetes/5126 Title: How to migrate data from bitnami ScyllaDB to scylla in Kubernetes? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 6.2.0 #Cluster size: 3 nodes os (RHEL/CentOS/Ubuntu/AWS AMI): Kubernetes We had used bitnami ScyllaDB in a project. Now we want to use scylla-operator and Kubernetes. How can … Language: en Canonical URL: https://forum.scylladb.com/t/how-to-migrate-data-from-bitnami-scylladb-to-scylla-in-kubernetes/5126 ## Headings Structure: H1: How to migrate data from bitnami ScyllaDB to scylla in Kubernetes? H2: 1. Understand the setup H2: 2. Preferred migration method: sstableloader H2: 3. Alternate approach: Scylla Manager Migrator H2: 4. Configuration considerations H2: 5. Optional: Full export/import H3: Related topics ## Main Content: H1: How to migrate data from bitnami ScyllaDB to scylla in Kubernetes? H2: 1. Understand the setup H2: 2. Preferred migration method: sstableloader H2: 3. Alternate approach: Scylla Manager Migrator H2: 4. Configuration considerations H2: 5. Optional: Full export/import H3: Related topics Installation details #ScyllaDB version: 6.2.0 #Cluster size: 3 nodes os (RHEL/CentOS/Ubuntu/AWS AMI): Kubernetes We had used bitnami ScyllaDB in a project. Now we want to use scylla-operator and Kubernetes. How can we migrate data from bitnami ScyllaDB to our new cluster? If you’re currently running Bitnami’s ScyllaDB Helm deployment and plan to migrate to an official Scylla cluster managed by Scylla Operator, the process is straightforward once you view Bitnami as simply “another ScyllaDB cluster.” The operator just changes how the nodes are managed — not the actual data format — so standard Scylla migration tools apply. Bitnami’s ScyllaDB chart wraps official Scylla images with its own init containers, config defaults, and host volume mounts. The Scylla Operator instead deploys native ScyllaCluster custom resources that manage lifecycle (scaling, rolling upgrades, etc.) automatically. Both store data in compatible SSTable formats, so direct data-level migration is possible.​ You can migrate from Bitnami to Operator-managed Scylla by exporting SSTables and streaming them into the new cluster: On each Bitnami Scylla node (Pod), take a snapshot of all keyspaces: Copy the snapshot data from /var/lib/scylla/data/ to an intermediate host or storage volume (or mount it via NFS). On a machine with scylla-tools-core installed (you can use a temporary pod for this in Kubernetes), run: sstableloader -d This will stream the SSTables into your new Scylla Operator cluster while preserving schema and data.​ Verify data consistency with nodetool status and application-level checks. This method works well for clusters of any size and is version-agnostic between 6.x-compatible builds. If you already use Scylla Manager, the scylla-manager-migrator tool can handle streaming between two live clusters (Bitnami and Operator-managed) automatically. It uses snapshots and repair-safe streaming internally.​ Ensure both clusters use the same replication settings before loading data. Verify that endpoint_snitch, compaction_strategy, and keyspace definitions match. Deploy the new cluster first via Helm + Scylla Operator chart per Scylla Operator docs before migrating.​ If using PersistentVolumeClaims, make sure the volume capacity is sufficient to hold both existing and incoming data. For smaller datasets, you can use standard cqlsh export/import as a lightweight alternative: cqlsh -e "DESCRIBE KEYSPACE your_keyspace" > schema.cql cqlsh -e "COPY your_table TO 'data.csv' WITH HEADER=TRUE;" Then recreate schema and reimport on the new cluster: cqlsh -f schema.cql COPY your_table FROM 'data.csv' WITH HEADER=TRUE; This method is slower but simple for non-production migrations. Bitnami ScyllaDB → Scylla Operator (Kubernetes) migration can be done cleanly with: nodetool snapshot + sstableloader (recommended for production) or scylla-manager-migrator (automated streaming) or cqlsh COPY (for small dev/test datasets) Take snapshots on Bitnami nodes Deploy new Scylla Operator cluster Stream data into the target cluster Once complete, decommission the Bitnami deployment and update your application’s contact points. cc @mflendrich & @Yehuda_Lebi who could provide any additional insights, if needed Thank you very much for the procedure. We will follow it and let you know the results. its ineed helped thanks!! --- ### Page: https://forum.scylladb.com/t/scylladb-cloud-vector-search-beta-is-here/5127 Title: ScyllaDB Cloud Vector Search Beta is here! - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: :rocket: ScyllaDB Cloud Vector Search Beta is here! We’re excited to announce the Early Access Program (EAP) for ScyllaDB Cloud Vector Search - built on the ScyllaDB core architecture to deliver ultra-low-millisecond la… Language: en Canonical URL: https://forum.scylladb.com/t/scylladb-cloud-vector-search-beta-is-here/5127 ## Headings Structure: H1: ScyllaDB Cloud Vector Search Beta is here! H3: Related topics ## Main Content: H1: ScyllaDB Cloud Vector Search Beta is here! H3: Related topics ScyllaDB Cloud Vector Search Beta is here! We’re excited to announce the Early Access Program (EAP) for ScyllaDB Cloud Vector Search - built on the ScyllaDB core architecture to deliver ultra-low-millisecond latency and massive throughput for AI and real-time workloads. Be among the first to explore Vector Search, try it on your own workloads (for free, before GA), and help shape the product with your feedback. We’ll schedule a quick kickoff (demo + Q&A) You’ll get early access and direct support from our team Ready to dive in? → Join the Early Access Program ScyllaDB Cloud Vector Search Beta is here! We’re excited to announce the Early Access Program (EAP) for ScyllaDB Cloud Vector Search - built on the ScyllaDB core architecture to deliver ultra-low-millisecond latency and massive throughput for AI and real-time workloads. Be among the first to explore Vector Search, try it on your own workloads (for free, before GA), and help shape the product with your feedback. We’ll schedule a quick kickoff (demo + Q&A) You’ll get early access and direct support from our team Ready to dive in? → Join the Early Access Program You can read more about ScyllaDB’s Vector Search in this great blog post. Also, a few of the talks in the upcoming P99 Conf deal with AI and Vector Search, you can save your (free) spot here. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-111-2025-10-10/5128 Title: Last week in scylla-cluster-tests.git master (issue #111; 2025-10-10) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 64038966…4db66ea8 range are covered. There were 10 non-merge commits from 7 authors in tha… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-111-2025-10-10/5128 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #111; 2025-10-10) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #111; 2025-10-10) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 64038966…4db66ea8 range are covered. There were 10 non-merge commits from 7 authors in that period. Some notable commits: Updated scylla-bench to v0.3.0. introducing mixed workload support. All relevant jobs from branch-perf-v15 were moved to branch-perf-v17. The branch-perf-v15 folder was removed. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-296-2025-10-12/5130 Title: Last week in scylladb.git master (issue #296; 2025-10-12) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 20aeed1607..8cd9f5d271 range are covered. There were 88 non-merge commits from 17 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-296-2025-10-12/5130 ## Headings Structure: H1: Last week in scylladb.git master (issue #296; 2025-10-12) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #296; 2025-10-12) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 20aeed1607..8cd9f5d271 range are covered. There were 88 non-merge commits from 17 authors in that period. Some notable commits: The replication factor can now be specified as a list of racks, rather than a number. This allows a keyspace to be migrated off a particular rack before it is decommissioned. Creating materialized views in tablet keyspaces now requires that the replication factor match the number of racks. The BYPASS CACHE query option now bypasses the cache for index reads using Trie indexes. The setup tool now checks if 4k block sizes are preferred by the disk, even if it advertises 512 byte physical sector sizes, to compensate for disks that misreport the physical sector size. The data filesystem will now be mounted with the lazytime option, to reduce metadata updates. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/getting-cdc-to-run-in-real-time-and-window-sizes/5133 Title: Getting CDC to run in real time and window sizes - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/getting-cdc-to-run-in-real-time-and-window-sizes/5133 ## Headings Structure: H1: Getting CDC to run in real time and window sizes H3: Related topics ## Main Content: H1: Getting CDC to run in real time and window sizes H3: Related topics Originally from the User Slack @Kỳ_Trần: Hello everyone, I’m having an issue with ScyllaDB’s CDC — it’s not running in real time. It seems to be processing in batches about every 60 seconds. I tried to fix it to make it real time again, but I haven’t been able to. I’d really appreciate any help. @Felipe_Cardeneti_Mendes**:** See window sizes https://github.com/scylladb/scylla-cdc-source-connector/blob/master/README.md#advanced-configuration-parameters @Kỳ_Trần: thank you verymuch --- ### Page: https://forum.scylladb.com/t/release-scylladb-2025-1-9/5134 Title: [RELEASE] ScyllaDB 2025.1.9 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB 2025.1.9, a bug-fix production-ready patch release for ScyllaDB 2025.1 LTS Release. Note there is a new Short Term Support (STS) Feature release 2025.3. You are welcome to upgrade to… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-2025-1-9/5134 ## Headings Structure: H1: [RELEASE] ScyllaDB 2025.1.9 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 2025.1.9 H3: Related topics The ScyllaDB team announces ScyllaDB 2025.1.9, a bug-fix production-ready patch release for ScyllaDB 2025.1 LTS Release. Note there is a new Short Term Support (STS) Feature release 2025.3. You are welcome to upgrade to it for the latest and greatest features, or stay on the 2025.1 track for long term support. The following issues are fixed in this release: Enhancements & Correctness --- ### Page: https://forum.scylladb.com/t/release-scylladb-2025-3-2/5135 Title: [RELEASE] ScyllaDB 2025.3.2 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of ScyllaDB 2025.3.2, a production-ready path release for ScyllaDB 2025.3 Feature Release. Related Links Read more about ScyllaDB Get ScyllaDB 2025.3 Upgrade from Sc… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-2025-3-2/5135 ## Headings Structure: H1: [RELEASE] ScyllaDB 2025.3.2 H2: Related Links H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 2025.3.2 H2: Related Links H3: Related topics The ScyllaDB team is pleased to announce the release of ScyllaDB 2025.3.2, a production-ready path release for ScyllaDB 2025.3 Feature Release. The following issues are fixed in this release: Enhancements & Correctness --- ### Page: https://forum.scylladb.com/t/how-to-trigger-cleanup-tasks-sequentially-instead-of-in-parallel-on-1-12/5136 Title: How to trigger cleanup tasks sequentially instead of in parallel on 1.12 - Kubernetes Operator - ScyllaDB Community NoSQL Forum Meta Description: Hi, I’m using Scylla Operator 1.12. After scaling, cleanup jobs seem to run simultaneously on all nodes, which puts significant pressure on the cluster. Wonder is there a way to run these jobs sequentially instead of in … Language: en Canonical URL: https://forum.scylladb.com/t/how-to-trigger-cleanup-tasks-sequentially-instead-of-in-parallel-on-1-12/5136 ## Headings Structure: H1: How to trigger cleanup tasks sequentially instead of in parallel on 1.12 H3: Related topics ## Main Content: H1: How to trigger cleanup tasks sequentially instead of in parallel on 1.12 H3: Related topics Hi, I’m using Scylla Operator 1.12. After scaling, cleanup jobs seem to run simultaneously on all nodes, which puts significant pressure on the cluster. Wonder is there a way to run these jobs sequentially instead of in parallel? @Yehuda_Lebi take a look… The Scylla Operator automatically kicks off cleanup jobs after a ScyllaCluster scale event. It creates per-node Jobs in parallel, there’s no user knob to serialize them today. Thanks for the reply. Not sure but I feel such an option might be necessary, as the cluster could be under heavy load overall if the dataset is large. Thanks for the feedback. I’ve opened Add option to control cleanup parallelism · Issue #3015 · scylladb/scylla-operator · GitHub as a tracker feature request. --- ### Page: https://forum.scylladb.com/t/p99-conference-is-happening-next-week/5137 Title: P99 Conference is happening next week! - Announcements - ScyllaDB Community NoSQL Forum Meta Description: P99 Conference is happening next week! This technical conference (free + virtual) is designed for engineers who care about P99 percentiles and high-performance, low-latency applications. It’ll feature speakers from comp… Language: en Canonical URL: https://forum.scylladb.com/t/p99-conference-is-happening-next-week/5137 ## Headings Structure: H1: P99 Conference is happening next week! H3: Related topics ## Main Content: H1: P99 Conference is happening next week! H3: Related topics P99 Conference is happening next week! This technical conference (free + virtual) is designed for engineers who care about P99 percentiles and high-performance, low-latency applications. It’ll feature speakers from companies like Uber, Microsoft, Meta, Bloomberg, SAP, Netflix, ShareChat, PayPal, Disney and many more. See the full agenda here. I hope to see you there! --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-111-2025-10-17/5142 Title: Last week in scylla-cluster-tests.git master (issue #111; 2025-10-17) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the a20dd652…f1941804 range are covered. There were 30 non-merge commits from 9 authors in tha… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-111-2025-10-17/5142 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #111; 2025-10-17) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #111; 2025-10-17) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the a20dd652…f1941804 range are covered. There were 30 non-merge commits from 9 authors in that period. Some notable commits: Latte bumped to 0.40.1-scylladb, bringing data‑validation fixes and the ability to generate 32‑bit floats in rune scripts. Also added support for specifying user/password in stress commands via --user /--password needed by SLA testing. Provision tests added for clusters with a Vector Store node for AWS and Docker backends. Support for performance v15/v16 branches was fully removed. Scylla Cloud backend now supports the ‘release:latest’ tag to deploy the latest supported released version. Add a weekly provision test for xcloud backend on both AWS and GCE to spot backend issues early. After months of stable operation in tier1, logs transport default was switched to Vector. Updated cassandra-stress to v3.18.2, bringing driver v3.11.5.8 and fixing a case where thread errors incorrectly yielded exit code 0. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-2025-3-2/5147 Title: [RELEASE] ScyllaDB 2025.3.2 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of ScyllaDB 2025.3.2, a production-ready path release for ScyllaDB 2025.3 Feature Release. Related Links Read more about ScyllaDB Get ScyllaDB 2025.3 Upgrade from Sc… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-2025-3-2/5147 ## Headings Structure: H1: [RELEASE] ScyllaDB 2025.3.2 H2: Related Links H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 2025.3.2 H2: Related Links H3: Related topics The ScyllaDB team is pleased to announce the release of ScyllaDB 2025.3.2, a production-ready path release for ScyllaDB 2025.3 Feature Release. The following issues are fixed in this release: Enhancements & Correctness --- ### Page: https://forum.scylladb.com/t/whats-the-correct-model-schema-for-my-use-case/5148 Title: Whats the correct model schema for my use case? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello, I am the owner of a roblox extension and I had recently encounter dificulties with ScyllaDB that made me consider switching to MongoDB. My application is stable right now but I want to know what I did wrong in my … Language: en Canonical URL: https://forum.scylladb.com/t/whats-the-correct-model-schema-for-my-use-case/5148 ## Headings Structure: H1: Whats the correct model schema for my use case? H3: Related topics ## Main Content: H1: Whats the correct model schema for my use case? H3: Related topics Hello, I am the owner of a roblox extension and I had recently encounter dificulties with ScyllaDB that made me consider switching to MongoDB. My application is stable right now but I want to know what I did wrong in my context. I have the following database: My queries are to SELECT based solely on the following fields server_id , place_id and last_updated. I had created secondary queries for those fields like this: await client.execute(` CREATE INDEX IF NOT EXISTS game_servers_new_last_updated_idx ON game_servers_new (last_updated ); `); My issue is with this query: SELECT game_id, active_players, last_updated FROM games WHERE last_updated < ? LIMIT ? ALLOW FILTERING So my question is: Even if i have secondary index on last_updated why does it says that my query can only run if ALLOW FILTERING is present? Where is my mistake? How should be my data constructed? To answer your question about the correct model schema for your ScyllaDB use case and why your query requires ALLOW FILTERING even with a secondary index: Summary of the issue: You created a table with server_id as the primary key, and a secondary index on last_updated. The database says you must use ALLOW FILTERING. In ScyllaDB (and Cassandra), secondary indexes only allow efficient equality queries (WHERE last_updated = ?), not range queries (WHERE last_updated < ?). For range queries on a secondary index (<, >, etc.), ScyllaDB does a full table scan behind the scenes and applies the filter after collecting the data. This operation is risky and may be expensive, so the system forces you to use ALLOW FILTERING to acknowledge the load and risk.​ SELECT...WHERElast_updated = ?-- Fast with index SELECT...WHERElast_updated < ? ALLOW FILTERING-- Slow, requires filter How to fix/model properly: If you need to frequently query by ranges (e.g., last_updated < ?), you should change your schema so that last_updated is part of your primary key (partition key or clustering key). For example: CREATE TABLEgames ( game_idtext, last_updated timestamp, ... PRIMARY KEY(game_id, last_updated) ) Or, if you want to efficiently select by last_updated only: Use a materialized view with last_updated as the base (ScyllaDB supports materialized views natively). CREATEMATERIALIZEDVIEWgames_by_last_updatedAS SELECT*FROMgames WHERElast_updated IS NOT NULL PRIMARY KEY(last_updated, game_id); SELECT...FROMgames_by_last_updatedWHERElast_updated < ? Secondary Index Limitations: Secondary indexes work for equality. Range queries over secondary indexes will almost always require ALLOW FILTERING, indicating a potentially slow, unoptimized query. Best Practice: Model your schema so that your most common query fields are part of your primary key (partition/clustering), or use materialized views if you need alternative fast access patterns.​ You are a genius. I didnt know that. I perfectly got it now. I have 1 more easy question about the data modeling: server_id is a random uuid, and it shall be unique place_id is a number Multiple server ids can have the same place_id. My question is the following: Why did you opted for PRIMARY KEY(game_id, last_updated) ) instead of PRIMARY KEY(server_id, last_updated) ) My queries are the following: server_id - exact query place_id - exact query only last_updated is a range query Therefore Is it better for my db to be structured like this? : PRIMARY KEY(server_id, last_updated) ) So we have unique server id and secondary index on place_id where exact querys will work by default? Again thank you for your much appreciated help. Great to hear you got the concept, thanks for the compliments, it’s just about knowing how the DB operates and proper modeling Regarding your question about the choice of primary key — whether to use PRIMARY KEY (game_id, last_updated) or PRIMARY KEY (server_id, last_updated) — here are some considerations: Since server_id is a unique random UUID and you will query by it with exact matches, using PRIMARY KEY (server_id, last_updated) is a good design choice. It ensures rows are uniquely identified by server_id and locally clustered by last_updated for efficient range queries per server. In your case, game_id vs server_id as partition key depends on which identifier you query mostly and which uniquely identifies the row. If your intent is to store data keyed by servers (each server having multiple records distinguished by timestamp), then (server_id, last_updated) makes more sense. The schema should reflect your most common query patterns and uniqueness constraints. If you frequently query by server_id and last_updated ranges, structure your table with PRIMARY KEY (server_id, last_updated). Use a secondary index or materialized view to support queries by place_id since it is not part of the primary key. This schema aligns well with your use case of unique servers queried by exact ID and timestamp ranges, plus location filtering via secondary indexing. Thank you so much. Now everything makes way more sense to me. I think I will get back to scylla now that I know what my mistake was. It had incredible performance but I got upset since I didnt understeand what I did wrong. Have a wonderfull day! Glad to help and good luck! Let us know if you have further questions, you can also reach-out on our Scylla Users slack - http://slack.scylladb.com/ And for other ways, take a look at this post - Engaging with the ScyllaDB Community I also tried right now to use a materialized view but I got: [CENTRALIZED CACHE] Failed to refill server cache: Only EQ and IN relation are supported on the partition key (unless you use the token() function or ALLOW FILTERING) ResponseError: Only EQ and IN relation are supported on the partition key (unless you use the token() function or ALLOW FILTERING) at FrameReader.readError (/root/BetterBloxAPI/node_modules/cassandra-driver/lib/readers.js:389:17) at Parser.parseBody (/root/BetterBloxAPI/node_modules/cassandra-driver/lib/streams.js:209:66) at Parser._transform (/root/BetterBloxAPI/node_modules/cassandra-driver/lib/streams.js:152:10) at Transform._write (node:internal/streams/transform:171:8) at writeOrBuffer (node:internal/streams/writable:572:12) at _write (node:internal/streams/writable:501:10) at Writable.write (node:internal/streams/writable:510:10) at Protocol.ondata (node:internal/streams/readable:1009:22) at Protocol.emit (node:events:518:28) at addChunk (node:internal/streams/readable:561:12) { info: ‘Represents an error message from the server’, code: 8704, query: ‘SELECT server_id, place_id, last_updated, datacenter \n’ + ’ FROM game_servers_by_last_updated \n’ + ’ WHERE last_updated < ?’ } I really dont understeand why this doesnt work… Edit: I have structured my materialized view as PRIMARY KEY (server_id, last_updated); The main table only has PRIMARY KEY (server_id) in order to automatically filter duplicates with index on place_id This made all my queries work. Hopefully I didnt made a mistake Based on ScyllaDB best practices and your handling of the secondary index issue: If your changes remove the need for ALLOW FILTERING by designing queries around primary and clustering keys or by modeling data to access patterns, your solution is aligned with recommended ScyllaDB schema strategies. Secondary indexes are best used for low-cardinality, infrequently updated fields; for anything dynamic or high-volume, it’s better to rely on a schema designed for efficient partition and clustering key reads.​ If your approach only needs secondary indexes for rare access paths and keeps core queries fast, this is considered acceptable and in line with current ScyllaDB guidance. If you encounter further performance issues or need a more opinionated schema review, please share your exact table definitions and query patterns—happy to offer more targeted advice! cc @GuyCarmin if you have further insights. --- ### Page: https://forum.scylladb.com/t/release-scylla-monitoring-stack-4-12-1/5149 Title: [RELEASE] Scylla Monitoring Stack 4.12.1 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of ScyllaDB Monitoring Stack 4.12.0 + 4.12.1 Due to an issue with 4.12.0 we jump to 4.12.1 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylla-monitoring-stack-4-12-1/5149 ## Headings Structure: H1: [RELEASE] Scylla Monitoring Stack 4.12.1 H3: Related topics ## Main Content: H1: [RELEASE] Scylla Monitoring Stack 4.12.1 H3: Related topics The ScyllaDB team is pleased to announce the release of ScyllaDB Monitoring Stack 4.12.0 + 4.12.1 Due to an issue with 4.12.0 we jump to 4.12.1 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB based on Prometheus and Grafana. ScyllaDB Monitoring Stack 4.12.1 supports: ScyllaDB 2024.1, 2025.1, 2025.2, 2025.3, and the upcoming 2025.4 release This release introduces an updated look to the overview dashboard and a new dashboard for the upcoming vector search. Download ScyllaDB Monitoring Stack 4.12.1 ScyllaDB Monitoring Stack Docs Upgrade from ScyllaDB Monitoring 3.x to 4.y Upgrade from ScyllaDB Monitoring 4.x to 4.y Version updates for ScyllaDB Monitoring Stack 4.12.0 Prometheus upgraded to version 3.6.0 Grafana upgraded to version 12.2.0 New Information in ScyllaDB Dashboards Overview Dashboard Change Cluster status is now displayed in a panel showing the cluster status in text and a panel that shows the number of active nodes (and if any: joining, leaving, and unreachable). #2593 The new service groups table summarizes the available service groups, and act as a quick navigation #2592 The number of cores has been added to the node table #2591 The DC disk usage graph shows all the nodes in (vs average) that DC #2666 The manager panel was removed from the top row #2594 The cluster name is now shown (not just ID) when available #2589 Vector Search dashboard The upcoming ScyllaDB Vector search service runs on a dedicated node. The new vector-search dashboard monitors both the vector search service and the node physical statistics. Node metrics include: Disk, Memory and CPU usage General status indicates if it’s available or not Per index information combines: The index size (number of vectors in the index) Index build rate (how many new vectors per second are added to the index) Request latencies in Average time, P95, and P99 Detailed Dashboard Change Keyspace Dashboard change Tablets load balancer, make the title and description clearer #2640 LWT panels are not displayed well (full label data makes them hard to read). #2637 Write AWait panel - tool tip is not formatted properly. #2625 Update the hint panel to the new metric name #2619 The -c option in start-all.sh does not support values with spaces #2602 The device drop-down in the OS Dashboard should only show node_exporter devices #2628 Limit the SG dropdown (Service Level Group) in the Detailed dashboard to: user-facing service groups #2511 When adding the vector-search service, add the node_exporter targets for that node #2636 Read Grafana server metrics, enable alerts on Grafana server failure #2632 Allow Vector service to work with Prometheus directory binding #2667 Documentation: The docker-compose.example.yml is automatically updated on every release #2627 Documentation: Add usage examples for the -c option in start-all.sh #2601 Enable experimental-promql-functions for Prometheus #2629 --- ### Page: https://forum.scylladb.com/t/how-do-i-change-the-replication-factor-from-3-to-2-in-a-running-cluster-should-i-move-nodes-between-racks/5151 Title: How do I change the replication factor from 3 to 2 in a running cluster? Should I move nodes between racks? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/how-do-i-change-the-replication-factor-from-3-to-2-in-a-running-cluster-should-i-move-nodes-between-racks/5151 ## Headings Structure: H1: How do I change the replication factor from 3 to 2 in a running cluster? Should I move nodes between racks? H3: Related topics ## Main Content: H1: How do I change the replication factor from 3 to 2 in a running cluster? Should I move nodes between racks? H3: Related topics Originally from the User Slack @SudeepSudeep**:** I have 12 nodes with replication factor of 3 and 3 racks. I want to change the replication factor to 2 because my cluster is running out of disk space. What is the best approach? after changing the replication factor, does it make sense to move 4 nodes from rack c and move it to rack a and rack @Felipe_Cardeneti_Mendes? @Felipe_Cardeneti_Mendes**:** well, best is to bootstrap new nodes so you can make more disk space, instead of reducing RF and risk losing quorum / read stale data that said, if you want to do this, then so long you’re aware of the side-effects, the general doc is https://docs.scylladb.com/manual/stable/kb/rf-increase.html which is also relevant to Decreasing it. How to Safely Increase the Replication Factor | ScyllaDB Docs remember to @Sudeepleanup after the change. @Sudeep**:** thank you. It’s okay to keep 3 @Felipe_Cardeneti_Mendesacks after replication change? @Felipe_Cardeneti_Mendes**:** yes if you remove the “third” rack, then mathematics will hit y@Sudeepu and you’ll run close to out of disk space again :^) @Sudeep**:** one more question? is it okay to run nodetool cleanup nodes at the same time. This cluster won’t serve any traffic. @Felipe_Cardeneti_Mendes**:** yes --- ### Page: https://forum.scylladb.com/t/release-scylladb-cdc-rust-0-5-0/5152 Title: [RELEASE] ScyllaDB CDC Rust 0.5.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce scylla-cdc-rust v0.5.0, a library that makes it easy to develop Rust applications that consume data from a Scylla Change Data Capture Log. In this version: Support for CDC in k… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-cdc-rust-0-5-0/5152 ## Headings Structure: H1: [RELEASE] ScyllaDB CDC Rust 0.5.0 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB CDC Rust 0.5.0 H3: Related topics The ScyllaDB team is pleased to announce scylla-cdc-rust v0.5.0, a library that makes it easy to develop Rust applications that consume data from a Scylla Change Data Capture Log. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-113-2025-10-24/5154 Title: Last week in scylla-cluster-tests.git master (issue #113; 2025-10-24) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 0fb16a68…9434f481 range are covered. There were 24 non-merge commits from 12 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-113-2025-10-24/5154 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #113; 2025-10-24) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #113; 2025-10-24) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 0fb16a68…9434f481 range are covered. There were 24 non-merge commits from 12 authors in that period. Some notable commits: Test pipelines no longer require extra quotes when passing dict-like values via extra_environment_variables, preventing accidental parsing as strings instead of dicts. To tackle sporadic freezes in results processing, HDR histogram analysis was fixed. New test for Scylla Manager Alternator restore, including YCSB-based verification and hooks for indexes and special table names. Mimics the flow for CQL based tests. Monitoring was moved to version 4.12.1. Llatte will no longer run redundant ‘latte schema’ commands. Most Manager pipelines switch to Ubuntu24 as the main distro in make Ubuntu24 a main distro for Manager tests. Gemini image was bumped to v2.1.5 bringing stability fixes and reduced false error cases. Scylla.yaml config now have rf_rack_valid_keyspaces set to True by default to allow MV testing with tablets. In case of errors in creating keyspaces, adjust racks properly or set this value to False for given test. We updated the base versions rules for rolling upgrade tests by supporting upgrades from any supported STS up to the last LTS. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/migrating-from-vnodes-to-tablets-lightweight-transactions-lwt-materialized-views-mv-support/5155 Title: Migrating from vNodes to Tablets, Lightweight Transactions (LWT), Materialized Views (MV) support - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/migrating-from-vnodes-to-tablets-lightweight-transactions-lwt-materialized-views-mv-support/5155 ## Headings Structure: H1: Migrating from vNodes to Tablets, Lightweight Transactions (LWT), Materialized Views (MV) support H3: Related topics ## Main Content: H1: Migrating from vNodes to Tablets, Lightweight Transactions (LWT), Materialized Views (MV) support H3: Related topics Originally from the User Slack @Mikael_HedbergMikael_Hedberg**:** Hello! I’ve been working on a product where I think scylla is a really nice fit. How should I think about the upcoming tablets? I am relying on LWTs and MVs at the moment which is not supported yet with tablets. How much of a hassle is it to go from vnodes to tablets onc@Felipe_Cardeneti_Mendes its out? @Felipe_Cardeneti_Mendes**:** Hi Mikael, currently you would need to migrate data to a tablets enabled keyspace, for example using the ScyllaDB Migrator and doing dual writes on your application and that should be it. As a side-note, LWT and MV support is close. We recently wired support for CDC. And we are also thinking (not there yet) of wa@Mikael_Hedbergs to make the transition easier. @Mikael_Hedberg**:** Do you have a timeframe for the 2025.4.0 release? @Felipe_Cardeneti_Mendes**:** It’s been branched, so somewhere late this month/early next . --- ### Page: https://forum.scylladb.com/t/should-disable-cache-for-a-1-row-partition-modeling/5157 Title: Should disable cache for a "1-row-partition" modeling? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi everyone, I have a table where each partition contains only a single row — for example, using a uuid as the partition key: CREATE TABLE user_cache ( user_id uuid, timestamp timestamp, data text ); The f… Language: en Canonical URL: https://forum.scylladb.com/t/should-disable-cache-for-a-1-row-partition-modeling/5157 ## Headings Structure: H1: Should disable cache for a "1-row-partition" modeling? H3: Related topics ## Main Content: H1: Should disable cache for a "1-row-partition" modeling? H3: Related topics I have a table where each partition contains only a single row — for example, using a uuid as the partition key: The frequency of the acess on the same partition is unpredictable. I wonder if in this case I should disable cache in WITH caching = {'keys': 'NONE', 'rows_per_partition': 'NONE'}, in order to avoid store a tons of reference in cache, I’m afraid of huge consume of RAM in future, as much as the data increases. --- ### Page: https://forum.scylladb.com/t/segfaults-on-old-scylladb-cluster-how-can-i-upgrade-it/5159 Title: Segfaults on old ScyllaDB cluster, how can I upgrade it? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/segfaults-on-old-scylladb-cluster-how-can-i-upgrade-it/5159 ## Headings Structure: H1: Segfaults on old ScyllaDB cluster, how can I upgrade it? H3: Related topics ## Main Content: H1: Segfaults on old ScyllaDB cluster, how can I upgrade it? H3: Related topics Originally from the User Slack @SabinSabin**:** I have really old Scylla Server (Scylla version 4.2.1-0.20201108.4fb8ebccff) which is “running” at work. Recently, I have encountered issues where there’s random segfaults on shard X. How can I debug and possibly fix it? I want to migrate it to newer version but this issue is holding me back I am mostly seeing these on the journalctl for Scylla Server This crashes the service and the kernel on… Ubuntu 18.04 TL@avi @avi**:** It is highly recommended to upgrade both the server OS and the ScyllaDB soft@Sabinare @Sabin**:** If I create a new instance with recent version of ScyllaDB, would it be able to join the@avicluster? @avi**:** It’s not recommended, especially for such old versions Upgrade in place, one version at a time, until you reach a supported version, and don’t let it lag i@Sabin the future @Sabin**:** I am worried about the data. Also, I don’t have much experience wi@avih the DB itself. @avi**:** Take a backup. Leaving it like that will cause it to rot until one day yo@Sabin cannot recover it. @Sabin**:** Is there a way to fix this error apart from upgrading? I have backup but recoveri@avig it takes way too long. @avi**:** Try to decode it on @Sabintp://backtrace.scylladb.com @Sabin**:** I wonder if I can do rolling up@avirade of ScyllaDB from 4.x to 6.x @avi**:** No, we only test one minor version at a time @Sabin**:** That will take time. The node is now up as the actual issue was with hardware (faulty RAM sticks). I will try to get the data off the nodes and try to restore it on newer version on another cluster. --- ### Page: https://forum.scylladb.com/t/release-scylladb-manager-3-7-0/5160 Title: [RELEASE] ScyllaDB Manager 3.7.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces the release of ScyllaDB Manager 3.7.0, a production-ready minor release of the stable 3.7 branch. ScyllaDB Manager is a centralized cluster administration and recurrent tasks automation tool. … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-manager-3-7-0/5160 ## Headings Structure: H1: [RELEASE] ScyllaDB Manager 3.7.0 H3: Alternator backup and restore (#4509) H3: Incremental repair for tablet keyspaces (#4591) H3: Healthcheck intervals (#4445) H3: Compatibility with ScyllaDB 2025.4 H3: Bug fixes and other improvements H3: Upgrade to the new release H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Manager 3.7.0 H3: Alternator backup and restore (#4509) H3: Incremental repair for tablet keyspaces (#4591) H3: Healthcheck intervals (#4445) H3: Compatibility with ScyllaDB 2025.4 H3: Bug fixes and other improvements H3: Upgrade to the new release H3: Related topics The ScyllaDB team announces the release of ScyllaDB Manager 3.7.0, a production-ready minor release of the stable 3.7 branch. ScyllaDB Manager is a centralized cluster administration and recurrent tasks automation tool. This release focuses on the support for Alternator clusters backup and restore tasks and integration with the upcoming ScyllaDB 2025.4 release. Below are the changes in this release. Even though Alternator can be thought of as just an alternative frontend to CQL, it comes with different tricks and considerations. Before ScyllaDB Manager 3.7.0 release, backing up and restoring Alternator clusters worked in the exact same way as for the CQL clusters. This worked in most scenarios, but it didn’t cover all corner cases where manual alterations would be needed. ScyllaDB Manager 3.7.0 provides improvements to Alternator backup and restore tasks, so that manual alterations are no longer needed. In order for those improvements to take place, users need to add Alternator credentials for managed clusters to ScyllaDB Manager. Alternator credentials are handled in the same way as the regular CQL credentials and can be added with: sctool cluster update --alternator-access-key-id --alternator-secret-access-key. Both Alternator and CQL credentials can be linked to the same underlying CQL role (see authentication in Alternator for more context). Note that both CQL and Alternator credentials are required for backup and restore tasks in Alternator clusters. Restoring contents of Alternator Secondary Indexes has a slight limitation. When restore task is completed, restored LSIs are available, yet they are still in the process of applying view updates. This means that querying LSIs right after restore might return partial results. When all updates are applied, LSIs will return full results. The same applies to GSIs for ScyllaDB versions older than 2025.1. ScyllaDB 2025.4 adds incremental repair feature for tablet keyspaces. It allows for skipping the repair of already repaired sstables and hence decreases repair time and impact on the cluster. By default, all repair tasks use incremental repair, but it is also possible to schedule full repair with the new sctool repair --incremental-mode flag. ScyllaDB Manager 3.7.0 increases the default healthcheck tasks interval to 1 minute to reduce healthcheck noise in bigger, multi-datacenter clusters. It also makes it possible to configure healthcheck tasks interval from scylla-manager.yaml config. These are the adjustments needed for achieving compatibility with ScyllaDB 2025.4 release: ScyllaDB customers are encouraged to upgrade to ScyllaDB Manager 3.7.0 in coordination with the ScyllaDB support team. The new release includes upgrades of both ScyllaDB Manager Server and Agent. Scylla Manager 3.7.0 supports the following ScyllaDB releases You can install and run ScyllaDB Manager on Kubernetes using ScyllaDB Operator. More here. --- ### Page: https://forum.scylladb.com/t/cant-get-host-id-to-ip-mapping-error/5161 Title: Can't get host id to ip mapping error - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello, I am new to scylladb. I set up all the prerequisite on gke using GitOps (kubectl) | ScyllaDB Docs and all the component are up and running. Now I am following the instruction on ScyllaClusters | ScyllaDB Docs to d… Language: en Canonical URL: https://forum.scylladb.com/t/cant-get-host-id-to-ip-mapping-error/5161 ## Headings Structure: H1: Can't get host id to ip mapping error H3: Related topics ## Main Content: H1: Can't get host id to ip mapping error H3: Related topics Hello, I am new to scylladb. I set up all the prerequisite on gke using GitOps (kubectl) | ScyllaDB Docs and all the component are up and running. Now I am following the instruction on ScyllaClusters | ScyllaDB Docs to deploy a cluster. Only the difference is the region names (I am using us-central-1 instead of us-east-1) and the cluster/dc name. The pod for the first region started without issue and the second region one did not pass the readiness the pod log shows the following messag Do you have any recommendation what to look for to resolve the issue? For your information I would appreciate it any help. I have been stuck with this for a while. Thanks in advance seems like you are trying to create a multi-DC cluster - as you mentioned 1st region and 2nd region? if so the 2nd region needs to be able to access the seed host/IP from the 1st region to connect. do you have 2 k8s clusters? my example here may help you. GitHub - tluck/Scylla-K8s-Example: Scylla K8s Example - End-to-End deployment. let me know happy to get this sorted! Thanks for your response seems like you are trying to create a multi-DC cluster No. The following is the copy of the example I refer to and there is only one DC. It creates statefulset per zone and they work as a single DC (single region). Thanks for sharing your repo. I will check it out and compare with what I did. Thanks again --- ### Page: https://forum.scylladb.com/t/how-much-disk-space-required-when-changing-compaction-strategy/5162 Title: How much disk space required when changing compaction strategy? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/how-much-disk-space-required-when-changing-compaction-strategy/5162 ## Headings Structure: H1: How much disk space required when changing compaction strategy? H3: Related topics ## Main Content: H1: How much disk space required when changing compaction strategy? H3: Related topics Originally from the User Slack @SudeepSudeep**:** If I wanted to change compaction strategy for a table how much disk space is required? I have a table that is 250GB. Does that mean I need to have at least another 250GB of free space when changing compaction strategy? @Felipe_Cardeneti_Mendes**:** Oftentimes no. A 100% space amplification is more of a STCS thing. That said, changing the strategy does involve recompacting your entire dataset, so you definitely want some breathing room. The disk space required depends on the specific compaction strategy. You can learn more about that in the Compaction Strategies ScyllaDB University lesson. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-114-2025-10-31/5163 Title: Last week in scylla-cluster-tests.git master (issue #114; 2025-10-31) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 99b56777…f6ed66b5 range are covered. There were 40 non-merge commits from 15 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-114-2025-10-31/5163 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #114; 2025-10-31) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #114; 2025-10-31) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 99b56777…f6ed66b5 range are covered. There were 40 non-merge commits from 15 authors in that period. Some notable commits: Switch rolling-upgrade tests to Rocky Linux 10, replacing CentOS Stream 9 as GCP deprecates CentOS. Upgrade latency regression now sends Latte results (before/after upgrade) to Argus. Comprehensive GitHub Copilot instructions were added, covering commit format, pre-commit checks, test layout, backend labels, and manual testing notes. Latte updated to v0.41.0, adding support for the Vector CQL data type. Vector Search node deployment on Scylla Cloud for new clusters; Beta limits apply (restricted instance types, one VS node per AZ, no multi-DC). Can reuse existing Scylla Cloud clusters previously deployed by SCT. Install and delete Vector Search nodes in already running clusters (not only at creation time). Added tests for Scylla Manager restore object_storage_method, covering both Native and Rclone methods. cassandra-stress updated to v3.18.3, switching to RANDOM replica ordering for rack-aware LB and creating user schemas with QUORUM consistency. Added HDR analysis support for c-s user profile workloads. New utility for deeper HDR analysis, running build_histograms_summary_with_interval at shorter intervals to pinpoint latency spikes. scylla-bench updated to v0.3.1, fixing missing co-fixed-read and co-fixed-write tags in mixed-mode HDR histograms. SCT events are sent to Argus in realtime (with rate limiter for the same events in a short time period). Elastic Cloud space reclamation tests at 90% utilization: validate reclamation via truncate, drop, and TTL expiry (checking metrics); tests run periodically on demand . Move select jobs to AWS i8ge.xlarge (Ubuntu2205-arm artifact and longevity-twcs-48h.yaml), offering similar price and disk size with better performance. Manager 3.7 becomes the default: add repos, bump agent to 3.7.0, update Docker image and Jenkins pipelines, refresh unit tests and docs; remove Manager 3.4. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/error-when-trying-to-restore-all-keyspaces-from-backup-using-alternator-dynamodb-version-issue/5164 Title: Error when trying to restore all keyspaces from Backup, using Alternator (DynamoDB), version issue? - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/error-when-trying-to-restore-all-keyspaces-from-backup-using-alternator-dynamodb-version-issue/5164 ## Headings Structure: H1: Error when trying to restore all keyspaces from Backup, using Alternator (DynamoDB), version issue? H3: Related topics ## Main Content: H1: Error when trying to restore all keyspaces from Backup, using Alternator (DynamoDB), version issue? H3: Related topics Originally from the User Slack @Phil_GebhardtPhil_Gebhardt**:** Hi everyone, I’m new to ScyllaDB, and I’m trying to restore all keyspaces from a backup, but receiving an error I’m having trouble getting past. This is how I’m attempting the backup using ScyllaDB Manager’s sctool When I try this, I receive: I’ve tried all sorts of different inputs for --keyspace with no luck. Any help would be greatly app@Felipe_Cardeneti_Mendeseciated! @Felipe_Cardeneti_Mendes**:** hmm this is a ScyllaDB API (as in port 10000 through the reserve proxy) call. I wonder if the scylla version is incompatible with the request issued by the Manager? Though I dont re@Phil_Gebhardtall seeing this endpoint change @Phil_Gebhardt**:** Thanks for looking at this Felipe, I suspected versions as well but according to the compatibility matrix, I think what I am running should be compatible. I should also mention that I am using alternator. I did just find this epic #4509. Is it possible that my restore problems stem @Felipe_Cardeneti_Mendesrom being on a version before these ch@Phil_Gebhardtnges? @Felipe_Cardeneti_Mendes**:**@Felipe_Cardeneti_Mendesgood catch, so that explains it. @Phil_Gebhardt**:** Thanks for your eyes on this @Felipe_Cardeneti_Mendes**:** IIRC alternator follows a specific naming convention which was known to fail when calling the API. btw you miiiight be able to get it running with a recent Manager - we literally stopped updating the compatibility list and running regression tests on older releases as – well, we no longer support them docker is probably your best friend here just to assess it quickly :)) @Phil_Gebhardt**:** thanks for the info, I might give that a try before upgrading --- ### Page: https://forum.scylladb.com/t/upcoming-change-to-the-system-users-structure-no-action-required/5165 Title: Upcoming change to the System Users Structure (No action required) - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: As part of our always ongoing efforts to strengthen ScyllaDB Cloud’s security and compliance, we are updating the structure of the system-managed service users. No action is required from you. The change will take effec… Language: en Canonical URL: https://forum.scylladb.com/t/upcoming-change-to-the-system-users-structure-no-action-required/5165 ## Headings Structure: H1: Upcoming change to the System Users Structure (No action required) H3: Related topics ## Main Content: H1: Upcoming change to the System Users Structure (No action required) H3: Related topics As part of our always ongoing efforts to strengthen ScyllaDB Cloud’s security and compliance, we are updating the structure of the system-managed service users. No action is required from you. The change will take effect starting November 5, 2025, and will be transparent to customers. There will be no downtime involved. After the change is applied, you might start seeing the new service users in the database audit logs. The new structure provides more granular and easily auditable access for our internal service operations, maintenance, and monitoring. It will ensure even stricter adherence to compliance framework recommendations such as SOC 2, PCI DSS, and ISO 27001. For additional details, please refer to our documentation portal. Service Users in ScyllaDB Cloud If you have any questions or require further information about this update, please don’t hesitate to contact our support team. (support@scylladb.com) Thank You, ScyllaDB Cloud Team --- ### Page: https://forum.scylladb.com/t/release-scylladb-2025-3-3/5167 Title: [RELEASE] ScyllaDB 2025.3.3 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of ScyllaDB 2025.3.3, a production-ready path release for ScyllaDB 2025.3 Feature Release. Related Links Read more about ScyllaDB Get ScyllaDB 2025.3 Upgrade from Sc… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-2025-3-3/5167 ## Headings Structure: H1: [RELEASE] ScyllaDB 2025.3.3 H2: Related Links H3: Deployment options H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 2025.3.3 H2: Related Links H3: Deployment options H3: Related topics The ScyllaDB team is pleased to announce the release of ScyllaDB 2025.3.3, a production-ready path release for ScyllaDB 2025.3 Feature Release. The following issues are fixed in this release: Enhancements & Correctness --- ### Page: https://forum.scylladb.com/t/release-python-driver-3-29-5/5168 Title: [RELEASE] Python Driver 3.29.5 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: Driver Release Summary: 3.29.5 Release Link: Release 3.29.5 · scylladb/python-driver · GitHub Bugs fixed Fix dc aware and rack aware policies initialization by (#578) Fix TokenAwarePolicy to work properly with tab… Language: en Canonical URL: https://forum.scylladb.com/t/release-python-driver-3-29-5/5168 ## Headings Structure: H1: [RELEASE] Python Driver 3.29.5 H3: Related topics ## Main Content: H1: [RELEASE] Python Driver 3.29.5 H4: Bugs fixed H4: Improvements H4: Environment & Compatibility Updates H3: Related topics Driver Release Summary: 3.29.5 Release Link: Release 3.29.5 · scylladb/python-driver · GitHub Fix dc aware and rack aware policies initialization by (#578) Fix TokenAwarePolicy to work properly with tablets when used in ExecutionProfile (#579) Update TokenAwarePolicy.make_query_plan to schedule to replicas first (#574) Smaller local/peers queries (#528) Fix metadata request timeout (#539) --- ### Page: https://forum.scylladb.com/t/release-scylladb-rust-driver-1-4-0/5169 Title: [RELEASE] ScyllaDB Rust Driver 1.4.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce ScyllaDB Rust Driver 1.4.0, an asynchronous CQL driver for Rust, optimized for Scylla, but also compatible with Apache Cassandra! Some interesting statistics: over 5.200k down… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-rust-driver-1-4-0/5169 ## Headings Structure: H1: [RELEASE] ScyllaDB Rust Driver 1.4.0 H2: Changes H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Rust Driver 1.4.0 H2: Changes H3: Related topics The ScyllaDB team is pleased to announce ScyllaDB Rust Driver 1.4.0, an asynchronous CQL driver for Rust, optimized for Scylla, but also compatible with Apache Cassandra! Some interesting statistics: New features / enhancements: Internal cleanups/refactors: CI / developer tool improvements: Congrats to all contributors and thanks everyone for using our driver! The source code of the driver can be found here: The official crates.io registry entry is here: Thank you for your attention, please do not hesitate to contact us if you have any questions, issues, feature requests, or are simply interested in our driver! Contributors since the last release: --- ### Page: https://forum.scylladb.com/t/does-disk-space-free-automatically-after-reducing-rf-and-running-repair/5170 Title: Does disk space free automatically after reducing RF and running repair? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: We reduced RF from 3→2 using ALTER KEYSPACE and ran nodetool repair on all 3 nodes. Will disk space be freed automatically through compaction? Language: en Canonical URL: https://forum.scylladb.com/t/does-disk-space-free-automatically-after-reducing-rf-and-running-repair/5170 ## Headings Structure: H1: Does disk space free automatically after reducing RF and running repair? H3: Related topics ## Main Content: H1: Does disk space free automatically after reducing RF and running repair? H3: Related topics We reduced RF from 3→2 using ALTER KEYSPACE and ran nodetool repair on all 3 nodes. Will disk space be freed automatically through compaction? need to run nodetool cleanup keyspace if we are reducing the RF. --- ### Page: https://forum.scylladb.com/t/s3-backed-storage-experimentation/5171 Title: S3 backed storage experimentation - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 2025.3.3 #Cluster size: 1 node os (RHEL/CentOS/Ubuntu/AWS AMI): Ubuntu Hello, I wanted to test out the experimental S3 backed storage for keyspaces but I am a bit confused abo… Language: en Canonical URL: https://forum.scylladb.com/t/s3-backed-storage-experimentation/5171 ## Headings Structure: H1: S3 backed storage experimentation H3: Related topics ## Main Content: H1: S3 backed storage experimentation H4: docs/dev/object_storage.md H3: Related topics Installation details #ScyllaDB version: 2025.3.3 #Cluster size: 1 node os (RHEL/CentOS/Ubuntu/AWS AMI): Ubuntu I wanted to test out the experimental S3 backed storage for keyspaces but I am a bit confused about the configuration part. I don’t have AWS and I am not at all familiar with how it works (which may explain the confusion). I have my own S3 provider with a custom domain. My buckets are available through this provider using ONLY virtual-hosted-style URLs e.g https://./. Does Scylla support this ? I don’t see any bucket field in the object_storage_endpoints option in scylla.yaml. The bucket is defined on the keyspace definition, something like So, once you define your endpoint in the scylla.yaml you can create a keyspace which will be using this endpoint. As of the endpoint you mentioned you can try and do the following Try to leave the bucket empty. If it doesnt work for you and you can tolerate additional level of directories created try to do the following For the reference on how to configure object storage --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-115-2025-11-07/5172 Title: Last week in scylla-cluster-tests.git master (issue #115; 2025-11-07) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 51a61c3c…dfeb826e range are covered. There were 7 non-merge commits from 5 authors in that… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-115-2025-11-07/5172 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #115; 2025-11-07) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #115; 2025-11-07) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 51a61c3c…dfeb826e range are covered. There were 7 non-merge commits from 5 authors in that period. Some notable commits: We’ve cleaned up outdated RBNO cases and old “feature” tests which were already reimplemented as Nemesis or not needed. Use java 25 LTS image for localhost java containers as java23 image is deprecated. All driver performance tests were moved under branch-perf-v17 where other perf tests already reside. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/performance-when-changing-a-cluster-size-with-vnodes-and-with-tablets/5173 Title: Performance when changing a cluster size with vNodes and with Tablets - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/performance-when-changing-a-cluster-size-with-vnodes-and-with-tablets/5173 ## Headings Structure: H1: Performance when changing a cluster size with vNodes and with Tablets H3: Related topics ## Main Content: H1: Performance when changing a cluster size with vNodes and with Tablets H3: Related topics Originally from the User Slack @FRED_FAT: Hi, does anyone compared the rebalance performance between vNode and tablet? vNode only suggest to add new node one by one and tablet could support parallel add @Felipe_Cardeneti_Mendes**:** you mean purely data rebalancing time? Yes, tablets wins by a very large margin as both data gets and nodes are distributed/bootstrapped in parallel --- ### Page: https://forum.scylladb.com/t/last-month-in-scylladb-git-master-issue-297-2025-11-10/5175 Title: Last month in scylladb.git master (issue #297; 2025-11-10) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 8cd9f5d271..cdba3bebda range are covered. There were 377 non-merge commits from 42 authors in that per… Language: en Canonical URL: https://forum.scylladb.com/t/last-month-in-scylladb-git-master-issue-297-2025-11-10/5175 ## Headings Structure: H1: Last month in scylladb.git master (issue #297; 2025-11-10) H3: Related topics ## Main Content: H1: Last month in scylladb.git master (issue #297; 2025-11-10) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 8cd9f5d271..cdba3bebda range are covered. There were 377 non-merge commits from 42 authors in that period. Some notable commits: CQL batchlog replays are now sent with consistency level EACH_QUORUM instead of ALL to increase their availability. When alternator uses Paxos for writes, it now also uses Paxos to request record pre-images for change data capture (CDC). This increases consistency of CDC records. There is now an experimental flag for enabling strongly consistent tables, marking the beginning of active development. ScyllaDB can now read and write to Google Cloud Storage via its native protocol rather than S3 emulation. ScyllaDB now authenticates vector store connections to the main database. Alternator, ScyllaDB’s implementation of the DynamoDB API, now supports the userIdentity field for records deleted via time-to-live (TTL) and reported via the Streams API. Previously, tablet merges with active materialized views were disallowed. They are now allowed when the cluster topology allows it (when the number of replicas is equal to the number of racks). A race condition between tablet spit and load-and-stream (importing sstables) was fixed. Alternator, ScyllaDB’s implementation of the DynamoDB API, now skips materialized view building when indexes are created for empty tables. The nodetool scrub command can now drop sstables which cannot be recovered. The scylla sstable command can now create sstables offline using CQL statements. Change Data Capture will now garbage-collect unnecessary CDC streams when it is used with tablets. Materialized view update concurrency control was improved, reducing the probability of flooding the system with materialized view updates. The maximum record size of trace records is now configurable. Alternator, ScyllaDB’s implementation of the DynamoDB API, now has better protection against oversize requests. We are now more careful updating the repair time after batchlog failure replay. The is important for tombstone_gc = repair mode. Change Data Capture (CDC) pre-images now work for tables without a clustering key. Alternator, ScyllaDB’s implementation of the DynamoDB API, generated incorrect records for TTL expiration events. This is now fixed. When disabling autocompaction, we no longer abort an ongoing major compaction (as it is a user-requested compaction, not automatic). We now automatically expand a numeric replication factor (‘my_dc’: 3) to a list of racks ('my_dc: [‘rackA’, ‘rackB’, ‘rackC’]) to persist the choice of racks for replication, allowing new racks to be added later. The compaction controller will now compute (and emit a metric for) the compaction backlog, allowing operators to estimate the effect of the automatic compaction controller even when it is not enabled. The default sstable compression algorithm was changed to LZ4WithDictsCompressor as it delivers better compression. There is a new nodetool excludenode command that allows marking a node as permanently down. The removenode operation sometimes cannot be used if it would violate replication invariants. Tables with counter columns are now supported with tablets. Authentication can now work in an unenforcing mode (warnings only) to test the effects of changing auth configuration. There is now a metric for time spent on tablet repair. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/error-on-repair-malformed-sstable-exception-after-topology-change/5176 Title: Error on repair, malformed sstable exception, after topology change - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/error-on-repair-malformed-sstable-exception-after-topology-change/5176 ## Headings Structure: H1: Error on repair, malformed sstable exception, after topology change H3: Related topics ## Main Content: H1: Error on repair, malformed sstable exception, after topology change H3: Related topics Originally from the User Slack @Jake_LaFountainJake_LaFountain**:** I was running a repair on 6.2.1-0.20241106.a3a0ffbcd015 and we got this. I’ve never seen this before and can’t find any results on Google either. Is there any way to re@Felipe_Cardeneti_Mendesolve this? @Felipe_Cardeneti_Mendes**:** hmmm interesting - I think this is the closest I could find (and if its that we only really saw it on failing tests - not manifesting out in the wild) https://github.com/scylladb/scylladb/commit/1d0c6aa26f105056aab001022a0f9487850af16b malformed_sstable_exception is effectively corruption (and sadly we had other issues in the past). If the disks are fine, since this affect system keyspace (a node-local one), perhaps replacing the faulty node could help. As in particular there should be noth@Jake_LaFountainng to repair due to LocalStrategy @Jake_LaFountain**:** Interesting! Some additional context: we had moved from a cluster of 5 (RF - 2) to a cluster of 4 to upgrade some drives. Adding the node back in has proved to be quite a struggle but during this last week we were able to successfully cleanup and repair the keyspace. We wanted to go up to a RF of 3, so we increased and repaired again and ran into this error on the new node. I’ll check into the actual disks but I doubt it’s the issue here unfortunately Replacing the node is possible I suppose but we did just do that a few days ago. Will follow up here. @Felipe_Cardeneti_Mendes**:** yeah, sorry about that. I suppose given the same keyspace and table are involved in the above commit and your particular problem that it might be related as both the source pull request (see #21558) mention topology changes and TTL expiration as a trigger. From what I can see the fix made it to 6.2.3 in 933ec7c so you may want to try that out - after you fix the already malformed sstable situation and before the wanderer maintainers call me out I shall diligently write down a reminder 6.2 is no longer supported, and ask for an upgrade :)) --- ### Page: https://forum.scylladb.com/t/release-gocql-v1-17-0/5177 Title: [RELEASE]: GoCQL v1.17.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: Driver Release Summary: v1.17.0 Release link: Release v1.17.0 · scylladb/gocql · GitHub What’s Changed Tests Skip TestRecreateSchema/UDT on 2024.1 (#600) Simplify tablets tests (#603) Fixes Fix inserting time.Time{}… Language: en Canonical URL: https://forum.scylladb.com/t/release-gocql-v1-17-0/5177 ## Headings Structure: H1: [RELEASE]: GoCQL v1.17.0 H2: What’s Changed H3: Tests H3: Fixes H3: Improvements H3: Related topics ## Main Content: H1: [RELEASE]: GoCQL v1.17.0 H2: What’s Changed H3: Tests H3: Fixes H3: Improvements H3: Related topics Driver Release Summary: v1.17.0 Release link: Release v1.17.0 · scylladb/gocql · GitHub Full Changelog: Comparing v1.16.1...v1.17.0 · scylladb/gocql · GitHub --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-116-2025-11-14/5178 Title: Last week in scylla-cluster-tests.git master (issue #116; 2025-11-14) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the e4596550…681aae9e range are covered. There were 40 non-merge commits from 12 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-116-2025-11-14/5178 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #116; 2025-11-14) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #116; 2025-11-14) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the e4596550…681aae9e range are covered. There were 40 non-merge commits from 12 authors in that period. Some notable commits: SkipPerIssues now recognizes dtest skip labels in addition to SCT ones (e.g., sct-2025.1-skip and dtest/2025.1-skip), smoothing cross-project skip handling. Rolling upgrade pipeline received two updates: ARM AMI tests moved to i8g.2xlarge, and the Jenkins job now accepts extra_environment_variables for more flexible runs. Tier1 coverage on AWS was adjusted by moving several cases to the i7i family. Nemesis flow was hardened by removing a dead node before running repair to avoid repair target creation errors. Cloud automation: ScyllaDB Cloud clusters can now be cleaned up based on a keep tag (e.g., names with -keep-Xh), and docs were updated with xcloud backend configuration details. Docker nodes will automatically use Google Container Registry mirrors to speed up and stabilize image pulls. Backtrace decoding became more efficient: we added selective decoding with stall filtering and enabled it across all performance tests (decode crashes/errors but skip reactor stalls). Dependencies and tooling were refreshed: requirements were fixed to use lz4 for python-driver compression and the scylla-driver was bumped to v3.29.5. Hydra gained a utility to find AMI equivalents across regions/architectures, easing cross-Region test setup (see docs/faq.md). On GCE, we enabled Cloud KMS encryption with key rotation, added external-runner credential retrieval, and improved certificate handling when the scylla user is absent. Test configs were also tidied up to stop using deprecated prepared loaders. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/disabling-backups-in-scylladb-manager/5179 Title: Disabling backups in ScyllaDB Manager - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/disabling-backups-in-scylladb-manager/5179 ## Headings Structure: H1: Disabling backups in ScyllaDB Manager H3: Related topics ## Main Content: H1: Disabling backups in ScyllaDB Manager H3: Related topics Originally from the User Slack @Brad_PetersBrad_PetersBrad_PetersBrad_Peters**:** @Patrick_Bossman @Patrick_Bossman Great chat at Kubecon, following up on disabling backups in Scylla manager@Brad_PetersBrad_Peters @Patrick_Bossman: Hi @Brad_Peters Here are the commands to see a backup, disable it, see it disabled, and re-enable it: When it’s disabled, you need to use -a to see it, and it will have an asterick in front of it, which means the task is disabled. The cli portion of scylla manager is very useful. The --keyspace argument can be used with glob patterns to include/exclude objects, if you only want to exclude certain objects. https://manager.docs.scylladb.com/stable/sctool/backup.html Backup | ScyllaDB Docs @Brad_Peters**:** Thanks Patrick --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-298-2025-11-16/5180 Title: Last week in scylladb.git master (issue #298; 2025-11-16) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the cdba3bebda..3672715211 range are covered. There were 112 non-merge commits from 23 authors in that per… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-298-2025-11-16/5180 ## Headings Structure: H1: Last week in scylladb.git master (issue #298; 2025-11-16) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #298; 2025-11-16) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the cdba3bebda..3672715211 range are covered. There were 112 non-merge commits from 23 authors in that period. Some notable commits: Alternator, ScyllaDB’s implementation of the DynamoDB API, now uses tablets by default instead of vnodes. This greatly improves scale-out performance. The bundled Prometheus node_exporter package was updated to version 1.10.2 to solve security issues. The AWS Key Management Service (KMS) integration now retries certain API calls to improve robustness. The CQL layer now correctly deserializes vectors of collections. Note that these data types are not useful in practice. Native restore from backup (where ScyllaDB directly accesses object storage) can now restore to only the primary replica, leaving full replication for later. The view builder now backs off when view building fails due to RPC errors, reducing log spam. Bootstrap and decommission now use regular streaming rather than repair-based node operations (RBNO) by default due to problems with performance and some stability issues. The vector store client, used for vector index searches, now backs off failed nodes and fails over to live nodes. The topology coordinator now includes joining nodes (not completely bootstrapped) in topology barriers, to improve robustness. Streaming and repair now verify sstable digest and checksums. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-monitoring-4-12-2/5182 Title: [RELEASE] ScyllaDB Monitoring 4.12.2 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of ScyllaDB Monitoring Stack 4.12.2 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB based on Prometheus and Grafana. ScyllaDB Monitoring Sta… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-monitoring-4-12-2/5182 ## Headings Structure: H1: [RELEASE] ScyllaDB Monitoring 4.12.2 H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Monitoring 4.12.2 H3: Related topics The ScyllaDB team is pleased to announce the release of ScyllaDB Monitoring Stack 4.12.2 ScyllaDB Monitoring Stack is an open-source stack for monitoring ScyllaDB based on Prometheus and Grafana. ScyllaDB Monitoring Stack 4.12.2 supports: ScyllaDB 2024.1, 2025.1, 2025.2, 2025.3, and the upcoming 2025.4 release This is a patch release with the following bug fixes Add support for Docker 29.x Fix the aggregate formula for large partitions Download ScyllaDB Monitoring Stack 4.12.2 ScyllaDB Monitoring Stack Docs Upgrade from ScyllaDB Monitoring 3.x to 4.y Upgrade from ScyllaDB Monitoring 4.x to 4.y --- ### Page: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-117-2025-11-21/5186 Title: Last week in scylla-cluster-tests.git master (issue #117; 2025-11-21) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 466a75db…f2ba4869 range are covered. There were 37 non-merge commits from 13 authors in th… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylla-cluster-tests-git-master-issue-117-2025-11-21/5186 ## Headings Structure: H1: Last week in scylla-cluster-tests.git master (issue #117; 2025-11-21) H3: Related topics ## Main Content: H1: Last week in scylla-cluster-tests.git master (issue #117; 2025-11-21) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the 466a75db…f2ba4869 range are covered. There were 37 non-merge commits from 13 authors in that period. Some notable commits: Gradual throughput growth via latte $throttle is now supported in PerformanceRegressionPredefinedStepsTest , allowing the same scenario to be driven by latte in addition to cassandra-stress. Serverless cleanup: removed the deprecated cloud_config option to align with the upcoming driver change. Upgrade tests gained a switch to disable Gemini during rolling upgrades; the new run_gemini_in_rolling_upgrade defaults to false for safer baseline runs. CLI: xcloud provider support was added to list-resources, including filtering by test-id/user and nicer table output across lab and staging. AWS provisioner now provides verbose error logging for spot/fleet failures with codes and state hints to speed up debugging. Rolling-upgrade pipelines were updated to replace debian-12 with debian-13, keeping test images current. Latte bumped to 0.42.0-scylladb, bringing Rust driver 1.4.1, an OCI image, and helpers for large data files (useful for vector search scenarios). Tester gained S3 artifact download functionality via the download_from_s3 config and a helper method to fetch required artifacts. Vector Search: a latte-based sanity test and script were added to create the schema, load data, and validate ANN results. GCE upgrade jobs were migrated to z3-highmem-8-highlssd instances for more consistent performance and lower costs. Housekeeping: updated CODEOWNERS to reflect current infra ownership. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-operator-1-19-0/5188 Title: [RELEASE] ScyllaDB Operator 1.19.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of ScyllaDB Operator 1.19.0. ScyllaDB Operator is an open-source project that helps you run ScyllaDB on Kubernetes by managing ScyllaDB clusters deployed to Kubern… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-operator-1-19-0/5188 ## Headings Structure: H1: [RELEASE] ScyllaDB Operator 1.19.0 H1: Multi-tenant monitoring with Prometheus and OpenShift support H1: Sensitive information excluded from must-gather H1: Configuration of kernel parameters (sysctl) H1: Topology change operations synchronisation H1: Other notable changes H2: Deprecation of ScyllaDBMonitoring components’ exposeOptions H2: Dependency updates H1: Upgrade instructions H1: Supported versions H1: Getting started with ScyllaDB Operator H1: Related links H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Operator 1.19.0 H1: Multi-tenant monitoring with Prometheus and OpenShift support H1: Sensitive information excluded from must-gather H1: Configuration of kernel parameters (sysctl) H1: Topology change operations synchronisation H1: Other notable changes H2: Deprecation of ScyllaDBMonitoring components’ exposeOptions H2: Dependency updates H1: Upgrade instructions H1: Supported versions H1: Getting started with ScyllaDB Operator H1: Related links H3: Related topics The ScyllaDB team is pleased to announce the release of ScyllaDB Operator 1.19.0. ScyllaDB Operator is an open-source project that helps you run ScyllaDB on Kubernetes by managing ScyllaDB clusters deployed to Kubernetes and automating tasks related to operating a ScyllaDB cluster, like installation, vertical and horizontal scaling, as well as rolling upgrades. ScyllaDB Operator 1.19 brings new features, stability improvements and documentation updates. ScyllaDB Operator monitoring uses Prometheus (an industry-standard cloud-native monitoring system) for metric collection and aggregation. Up until now, running a fresh, clean instance of Prometheus for every ScyllaDB cluster was the only supported way. We coined the term “Managed mode” for this architecture (because, in that case, ScyllaDB Operator would manage the Prometheus deployment): Please refer to the new ScyllaDB Monitoring overview and Setting up ScyllaDB Monitoring documents to learn more about the new mode and how to set up ScyllaDBMonitoring with an existing Prometheus instance. The Setting up ScyllaDB Monitoring on OpenShift guide offers guidance on how to set up User Workload Monitoring (UWM) for ScyllaDB in OpenShift. That being said, our experience shows that cluster administrators prefer closer control over the monitoring stack than offered with the Managed mode. For this reason we intend to standardize on using External in the long run. Therefore, the Managed mode remains supported, but is being deprecated and will be removed in a future Operator version. If you are an existing user, please consider deploying your own Prometheus using the Prometheus Operator platform guide and switching from Managed to External. ScyllaDB Operator comes with an embedded tool (called must-gather) that helps preserve the configuration (Kubernetes objects) and runtime state (ScyllaDB node logs, gossip information, nodetool status, etc.) in a convenient archive, allowing comparative analysis and troubleshooting with a holistic, reproducible view. As of ScyllaDB Operator 1.19, must-gather comes with a new setting --exclude-resource that serves as an additional guardrail preventing accidental inclusion of sensitive information – covering Secrets and SealedSecrets by default. Users can specify additional types to be restricted from capturing, or override the defaults by setting the --include-sensitive-resources flag. Please see the Gathering data with must-gather guide for more information. ScyllaDB nodes require kernel parameter (sysctl) configuration for optimal performance and stability – ScyllaDB Operator 1.19 improves the API to do that. Before 1.19, it has been possible to configure these parameters through v1.ScyllaCluster’s .spec.sysctls. Experience has indicated that this wasn’t the right place in the API for a setting that affects entire Kubernetes nodes. Therefore, ScyllaDB Operator 1.19 lets you configure sysctls through v1alpha1.NodeConfig for a range of Kubernetes nodes at once by matching the specified placement rules using a label-based selector. Please refer to the Configuring kernel parameters (sysctls) section of the documentation to learn how to configure the sysctl values recommended for production-grade ScyllaDB deployments. With the introduction of sysctl to NodeConfig, the legacy way of configuring sysctl values through v1.ScyllaCluster’s .spec.sysctls is now deprecated. ScyllaDB requires that when a new node is added to a cluster, no existing nodes are down. ScyllaDB Operator 1.19 addresses this issue by extending ScyllaDB Pods for newly joining nodes with a barrier blocking the ScyllaDB container from starting until the preconditions for bootstrapping a new node are met. This feature is opt-in in ScyllaDB Operator 1.19. You can enable it by setting the --feature-gates=BootstrapSynchronisation=true command-line argument to ScyllaDB Operator. This feature supports ScyllaDB 2025.2 and newer. If you are running a multi-datacenter ScyllaDB cluster (multiple ScyllaCluster objects bound together with external seeds), you are still required to verify the preconditions yourself before initiating any topology changes, because the synchronisation only occurs on the level of an individual ScyllaCluster. Please refer to Synchronising bootstrap operations in ScyllaDB for more information. With the introduction of support for external Prometheus instances, ScyllaDB Operator 1.19 makes a step towards reducing ScyllaDBMonitoring’s complexity by deprecating exposeOptions in both ScyllaDBMonitoring’s Prometheus and Grafana components. The use of exposeOptions is limited as it provides no way to configure an Ingress that will terminate TLS, which is likely the most common approach in production. As an alternative, this release introduces a more pragmatic and flexible approach by simply documenting how the components’ corresponding Services can be exposed and giving you the flexibility to do just what your use case requires. Please refer to the Exposing Grafana page in ScyllaDB Operator’s documentation to learn how to expose Grafana deployed by ScyllaDBMonitoring using a self-managed Ingress resource. The deprecated ScyllaDBMonitoring’s exposeOptions will be removed in a future Operator version. This release also includes regular updates of ScyllaDB Monitoring and the packaged dashboards to support the latest ScyllaDB releases (4.11.1->4.12.1, #3031), as well as its dependencies: Grafana (12.0.2->12.2.0) and Prometheus (v3.5.0->v3.6.0). For more changes and details, check out the GitHub release notes. For instructions on upgrading ScyllaDB Operator to 1.19, please refer to the Upgrading Scylla Operator section of the documentation. Your feedback is always welcome! Feel free to open an issue or reach out on the #scylla-operator channel in ScyllaDB User Slack. The ScyllaDB Operator Team --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-299-2025-11-24/5189 Title: Last week in scylladb.git master (issue #299; 2025-11-24) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 3672715211..724dc1e582 range are covered. There were 119 non-merge commits from 28 authors in that per… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-299-2025-11-24/5189 ## Headings Structure: H1: Last week in scylladb.git master (issue #299; 2025-11-24) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #299; 2025-11-24) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 3672715211..724dc1e582 range are covered. There were 119 non-merge commits from 28 authors in that period. Some notable commits: The view build worker builds materialized views from data staged by repair and sstable import. It now supports tablets migrating within a node. Audit can now simultaneously write to both syslog and a table. Batchlog replay is now more careful when updating the last replay marker. Automatic cleanup now allows the operator to override its decisions. There is now a metric showing uncompressed sstable file size, making it easier to gauge the benefit of compression. The load-and-stream feature is now more careful to select the correct sstables to stream. Tablet merging (when a table shrinks in size) can now happen concurrently with repair. It is now possible to configure internode compression only between racks. This is useful in avoiding public cloud inter-availability-zone charges while not compressing intra-availability-zone traffic. The vector search client can now use HTTPS to connect to vector search nodes. The default sstable format for new clusters is now ms, using Trie indexes. File based streaming (used for tablet migration) now validates the checksum and digest. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/trying-to-run-a-simple-benchmark-using-docker-and-cassandra-stress-tool/5192 Title: Trying to run a simple benchmark using docker and cassandra stress tool - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: latest, using docker run #Cluster size: Ubuntu 22.04 LTS, single node 8 GB RAM, 2 vcpu (small instance for MVP). I started scylla using docker: docker run --name some-scylla -… Language: en Canonical URL: https://forum.scylladb.com/t/trying-to-run-a-simple-benchmark-using-docker-and-cassandra-stress-tool/5192 ## Headings Structure: H1: Trying to run a simple benchmark using docker and cassandra stress tool H3: Related topics ## Main Content: H1: Trying to run a simple benchmark using docker and cassandra stress tool H3: Related topics Installation details #ScyllaDB version: latest, using docker run #Cluster size: Ubuntu 22.04 LTS, single node 8 GB RAM, 2 vcpu (small instance for MVP). I started scylla using docker: docker run --name some-scylla --hostname some-scylla -d scylladb/scylla --smp 1 Here is the running container: The stress test version and command (and error below): cassandra-stress write n=100000 -node 172.17.0.2 There is an option to select the protocol version, it defaults to the latest NEWEST_SUPPORTED version: --- ### Page: https://forum.scylladb.com/t/release-added-support-for-aws-i8g-and-i8ge-instance-families/5193 Title: [RELEASE] Added support for AWS i8g and i8ge instance families - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: We are happy to announce a new update to the ScyllaDB Cloud service. New features and enhancements: We’ve added support for the AWS i8g and i8ge instance families. These Graviton-based instances deliver strong price/pe… Language: en Canonical URL: https://forum.scylladb.com/t/release-added-support-for-aws-i8g-and-i8ge-instance-families/5193 ## Headings Structure: H1: [RELEASE] Added support for AWS i8g and i8ge instance families H3: Supported Regions H3: Related topics ## Main Content: H1: [RELEASE] Added support for AWS i8g and i8ge instance families H3: Supported Regions H4: i8g available in: H4: i8ge available in: H3: Related topics We are happy to announce a new update to the ScyllaDB Cloud service. New features and enhancements: We’ve added support for the AWS i8g and i8ge instance families. These Graviton-based instances deliver strong price/performance for storage-intensive workloads, giving customers more flexibility to optimize their deployments for cost efficiency, performance, or a balanced mix of both. How they compare to previous generations: i8g vs i7i and i4i: The AWS i8g family provides improved price/performance over both i7i and i4i by combining Graviton4 processors with next-generation Nitro improvements. Compared to i7i, i8g typically delivers similar or higher throughput at a lower hourly cost. Compared to i4i, it offers a substantial uplift in compute efficiency and storage bandwidth, making it a compelling upgrade path for customers seeking better TCO. i8ge vs i7ie and i3en: The AWS i8ge family builds on i8g with enhanced network capabilities and more consistent I/O performance. It provides stronger throughput than i7ie and represents a modern, more efficient alternative to i3en. This makes i8ge well-suited for high-throughput, latency-sensitive workloads. --- ### Page: https://forum.scylladb.com/t/memtable-and-cache-behavior/5194 Title: Memtable and cache behavior - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/memtable-and-cache-behavior/5194 ## Headings Structure: H1: Memtable and cache behavior H3: Related topics ## Main Content: H1: Memtable and cache behavior H3: Related topics Originally from the User Slack @Mahdi_KamaliMahdi_KamaliMahdi_KamaliMahdi_Kamali**:** When a memtable is being flushed, while some rows have been updated, what happens to those rows in cache? When rows in cache will be invalidated and when will be merged? @avi**:** If the partition exists in cache, rows in the memtable being flushed will be merged into cache --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-300-2025-11-30/5195 Title: Last week in scylladb.git master (issue #300; 2025-11-30) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 724dc1e582..44c605e59c range are covered. There were 63 non-merge commits from 27 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-300-2025-11-30/5195 ## Headings Structure: H1: Last week in scylladb.git master (issue #300; 2025-11-30) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #300; 2025-11-30) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 724dc1e582..44c605e59c range are covered. There were 63 non-merge commits from 27 authors in that period. Some notable commits: Nodes that are excluded from the cluster now receive a notification about it, to improve their diagnostics. Tablets of co-located tables are no longer repaired individually. The scylla sstable tool now has a dump-schema command. One can now limit the maximum shares generated by the compaction controller to reduce autocompaction impact on performance. The tablet scheduler moves the two halves of to-be-merged tablets together. It now avoids conflicting migration plans when doing so. Error handling for service level manipulations from CQL was improved. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/why-is-my-aws-alb-routing-traffic-to-unhealthy-targets-even-though-health-checks-look-normal/5197 Title: Why is my AWS ALB routing traffic to unhealthy targets even though health checks look normal? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: I’m running an application behind an AWS Application Load Balancer with two EC2 instances in an Auto Scaling group. Both instances show as healthy in the Target Group health checks, but I’m seeing inconsistent behavior … Language: en Canonical URL: https://forum.scylladb.com/t/why-is-my-aws-alb-routing-traffic-to-unhealthy-targets-even-though-health-checks-look-normal/5197 ## Headings Structure: H1: Why is my AWS ALB routing traffic to unhealthy targets even though health checks look normal? H3: Related topics ## Main Content: H1: Why is my AWS ALB routing traffic to unhealthy targets even though health checks look normal? H3: Related topics I’m running an application behind an AWS Application Load Balancer with two EC2 instances in an Auto Scaling group. Both instances show as healthy in the Target Group health checks, but I’m seeing inconsistent behavior where the ALB still routes requests to an instance that is clearly failing at the application level (timeouts, 500s, etc.). I’ve already checked: Health check path is correct (/health) Security groups allow ALB → EC2 traffic Application logs show intermittent failures but the health endpoint still returns 200 Is there a scenario where the ALB keeps routing traffic to a target even though the application behind it isn’t responding properly? Should I tighten the health check settings, or is there another configuration I might be missing? Would appreciate guidance from anyone who has dealt with similar behavior. --- ### Page: https://forum.scylladb.com/t/stuck-topology-request-after-increasing-replication-factor/5202 Title: Stuck topology request after increasing replication factor - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 6.2 #Cluster size: 30 nodes x 1 dc #OS: Ubuntu 24.04 Hey everybody, I encounter some issues with a stuck topology request after increasing the replication factor. The query wh… Language: en Canonical URL: https://forum.scylladb.com/t/stuck-topology-request-after-increasing-replication-factor/5202 ## Headings Structure: H1: Stuck topology request after increasing replication factor H3: Related topics ## Main Content: H1: Stuck topology request after increasing replication factor H3: Related topics Installation details #ScyllaDB version: 6.2 #Cluster size: 30 nodes x 1 dc #OS: Ubuntu 24.04 I encounter some issues with a stuck topology request after increasing the replication factor. The query which caused it was:ALTER KEYSPACE uniprottry WITH replication = { ‘class’ : ‘NetworkTopologyStrategy’, ‘cubimedrubdc’ : 3} cqlsh:uniprottry> select * from system.topology_requests ; cqlsh:uniprottry> select * from system.topology where global_topology_request_id = 75105bf2-cb73-11f0-d73b-a33c60dc186d allow filtering; here: Stuck topology request after increasing RF · GitHub It look suspicously like: , but I have no ghost nodes.Nodetool is reporting all nodes are up normal (UN) on every node. Increasing the repliction factor to 2, a few days earlier worked flawlessly. Is there something I can do to get rid of these stuck topology requests? It look suspicously like: Sorry, the link was missing As time is rather limited and everything I tried to solve this issue was unsuccessfull, I moved the data to a new cluster. --- ### Page: https://forum.scylladb.com/t/why-are-we-seeing-cl-one-queries-the-application-team-is-using-a-different-consistency-level/5204 Title: Why are we seeing CL=ONE queries? The application team is using a different consistency level - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/why-are-we-seeing-cl-one-queries-the-application-team-is-using-a-different-consistency-level/5204 ## Headings Structure: H1: Why are we seeing CL=ONE queries? The application team is using a different consistency level H3: Related topics ## Main Content: H1: Why are we seeing CL=ONE queries? The application team is using a different consistency level H3: Related topics Originally from the User Slack @KishoreKishore**:** Hi, We have a Scylla 5.0.13 cluster in production and are intermittently seeing CL=ONE queries, but the application team reports they aren’t issuing them. This occurred only in one DC. During the same period, we observed spikes in client connections, active SSTable reads, and queued SSTable reads. Any thoughts on what could cause this if it’s not coming from the applicati@Felipe_Cardeneti_Mendesn? @Felipe_Cardeneti_Mendes**:** Driver internal queries, or using Cassandra role to authenticate are common reasons @Kishore**:** got it. Thanks. --- ### Page: https://forum.scylladb.com/t/devopsdays-tel-aviv-workshop-december-11th/5205 Title: DevOpsDays Tel Aviv Workshop, December 11th - University and Training - ScyllaDB Community NoSQL Forum Meta Description: This Thursday, Eyal Shpin and I will host a workshop, Database Performance at Scale. Come say hi, also stop by our booth for some cool swag. Language: en Canonical URL: https://forum.scylladb.com/t/devopsdays-tel-aviv-workshop-december-11th/5205 ## Headings Structure: H1: DevOpsDays Tel Aviv Workshop, December 11th H3: Related topics ## Main Content: H1: DevOpsDays Tel Aviv Workshop, December 11th H3: Related topics This Thursday, Eyal Shpin and I will host a workshop, Database Performance at Scale. Come say hi, also stop by our booth for some cool swag. --- ### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-301-2025-12-07/5206 Title: Last week in scylladb.git master (issue #301; 2025-12-07) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 44c605e59c..47efbdffbc range are covered. There were 72 non-merge commits from 26 authors in that peri… Language: en Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-301-2025-12-07/5206 ## Headings Structure: H1: Last week in scylladb.git master (issue #301; 2025-12-07) H3: Related topics ## Main Content: H1: Last week in scylladb.git master (issue #301; 2025-12-07) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 44c605e59c..47efbdffbc range are covered. There were 72 non-merge commits from 26 authors in that period. Some notable commits: Alternator, ScyllaDB’s implementation of the DynamoDB API, now accepts compressed requests. This can help reduce network transfer costs in some cloud environments. The vector search component now ensures high availability during request timeouts. Authentication now uses a shard-local cache for authentication data. This increases performance during connection storms. The PRUNE MATERIALIZED VIEW statement, used to remove materialized view rows that no longer correspond to base table rows, can now have user-defined concurrency via a USING CONCURRENCY clause. A crash when a table using dictionary compression exceeded 8TB uncompressed per node was fixed. Encryption-at-rest supports IPv6 for KMIP. There are now checksums for all sstable components, reducing silent data corruption probability. Support for legacy schema support, obsoleted i 2017, was removed. The Raft failure detector was optimized for larger clusters. The row cache is updated every time a memtable is flushed. It will now avoid stalls when a range tombstone that covers many rows is merged. This scenario is common for Raft. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/release-scylladb-2025-3-5/5207 Title: [RELEASE] ScyllaDB 2025.3.5 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of ScyllaDB 2025.3.5, a production-ready patch release for ScyllaDB 2025.3 Feature Release. Related Links Read more about ScyllaDB Get ScyllaDB 2025.3 Upgrad… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-2025-3-5/5207 ## Headings Structure: H1: [RELEASE] ScyllaDB 2025.3.5 H2: Related Links H2: Bug Fixes H3: Change Data Capture (CDC) H3: Cloud/Connectivity H3: Operations/Management H3: Stability/Reliability H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 2025.3.5 H2: Related Links H2: Bug Fixes H3: Change Data Capture (CDC) H3: Cloud/Connectivity H3: Operations/Management H3: Stability/Reliability H3: Related topics The ScyllaDB team is pleased to announce the release of ScyllaDB 2025.3.5, a production-ready patch release for ScyllaDB 2025.3 Feature Release. Read more about ScyllaDB Upgrade from ScyllaDB 2025.2 to ScyllaDB 2025.3 The following issues are fixed in this release. Critical errors due to a malformed SSTable exception Issue: Critical errors (sstables::malformed_sstable_exception) were occurring because a column was reported as missing in the current schema for the cdc_log table. This could happen when recreating a column too soon. Fix: Added a check to prevent recreating a column too soon, and the logic was updated to set the column drop timestamp in the future to prevent the schema mismatch. scylladb#26340, scylladb#27036 Notification about expiring ERM held for too long was broken Issue: The system failed to properly notify when an Effective Replication Map (ERM) token was held for too long after its expiry. Fix: The notification logic for the expiring ERM held for too long was corrected. scylladb#27141, scylladb#27275 EC2 metadata querying should use back-off for “service unavailable” Issue: When querying EC2 metadata (used by AWS KMS), “service unavailable” responses (e.g., HTTP 503 errors) were not handled with a retry mechanism. Fix: The KMS host was updated to include the HTTP error code in KMS errors, and an exponential backoff-retry mechanism was added specifically for 503 errors. scylladb#27062, scylladb#27063 S3 client error handling for transient network errors Issue: The S3 client was not classifying all transient network errors as retryable, leading to unnecessary failures. Fix: Error handling for the S3 client was extended to correctly classify additional transient network errors as retryable. scylladb#27349, scylladb#27390 Automatic cleanup improvements Issue: Automatic cleanup logic was limited and lacked user-facing controls. Fix: Automatic cleanup was improved to allow a node to opt out of automatic cleanup. This update also introduced a RESTful API to reset the cleanup needed flag, and a nodetool cluster cleanup command to run cleanup on all dirty nodes. scylladb#26866, scylladb#27093 Maintenance mode functionality was broken Issue: Maintenance mode was non-functional, and the related test (test_maintenance_mode) did not perform as expected. Fix: The service QoS was updated to fall back to the default scheduling group when using the maintenance socket, restoring maintenance mode functionality. scylladb#26816, scylladb#27039 More logging for load_new_sstables/download_new_sstables Issue: The logging output for load_new_sstables and download_new_sstables was insufficient, lacking logging of all option values. Fix: The functions were updated to log all option values used during execution, and additional logging was added to streaming operations. scylladb#27299, scylladb#27341 Node locator missing `_excluded` field in operations Issue: The node locator was not preserving the _excluded field in clone() and omitting it from the verbose formatter. Fix: The locator logic was updated to preserve and include the _excluded field in all necessary places. scylladb#27290 Conflicting tablet migrations in the scheduler Issue: The tablet scheduler could emit conflicting migrations for the same tablet in different DCs or conflicting inter-node and intra-node migrations, resulting in incorrect reads. Fix: The scheduler logic was updated to prevent emitting conflicting migrations in the plan and during merge colocation. scylladb#26038, scylladb#27304, scylladb#26048, scylladb#27312, scylladb#27330 Load-and-stream with tablets failing with “Unable to load SSTable” Issue: Load-and-stream operations with tablets would sometimes fail with an “Unable to load SSTable” error. Fix: Synchronization logic was added to the sstables_loader to prevent bypassing synchronization when the topology is busy. scylladb#22707, scylladb#26730 Multiple oversized memory allocation errors with Vnodes Issue: Creating thousands of tables with Vnodes could lead to multiple seastar_memory - oversized allocation errors. Fix: The issue was resolved by changing the internal type of a table metadata variable. scylladb#26787, scylladb#27198 Node coredumped after tablet cleanup log line Issue: A node could coredump after logging that tasks were stopped for compactions due to tablet cleanup. Fix: The replica logic was updated to fail a timed-out single-key read on a cleaned-up tablet replica. scylladb#26229, scylladb#27155 Premature break causes SSTables to be skipped during streaming Issue: A premature loop break in the tablet_sstable_streamer::stream function was causing SSTables to be unexpectedly skipped. Fix: The loop break condition in tablet_sstable_streamer::stream was fixed. scylladb#26979, scylladb#27153 Race condition between tablet split and load-and-stream Issue: A race condition could occur between the tablet split process and the load-and-stream operation. Fix: Synchronization logic was implemented to correctly synchronize tablet split and load-and-stream. scylladb#26455, scylladb#26648 --- ### Page: https://forum.scylladb.com/t/release-scylladb-2025-1-10/5208 Title: [RELEASE] ScyllaDB 2025.1.10 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces ScyllaDB 2025.1.10, a bug-fix production-ready patch release for ScyllaDB 2025.1 LTS Release. Note that there is a new Short-Term Support (STS) Feature release, version 2025.3. You are welcom… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-2025-1-10/5208 ## Headings Structure: H1: [RELEASE] ScyllaDB 2025.1.10 H2: Related Links H2: Bug Fixes H3: Authentication H3: CDC (Change Data Capture) H3: Documentation H3: Repair H3: Operations H3: Reliability H3: Tooling/Automation H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 2025.1.10 H2: Related Links H2: Bug Fixes H3: Authentication H3: CDC (Change Data Capture) H3: Documentation H3: Repair H3: Operations H3: Reliability H3: Tooling/Automation H3: Related topics The ScyllaDB team announces ScyllaDB 2025.1.10, a bug-fix production-ready patch release for ScyllaDB 2025.1 LTS Release. Note that there is a new Short-Term Support (STS) Feature release, version 2025.3. You are welcome to upgrade to it for the latest and greatest features, or stay on the 2025.1 track for long-term support. Upgrade from ScyllaDB Enterprise 2024.x to ScyllaDB 2025.1 Upgrade from ScyllaDB Open Source 6.2 to ScyllaDB Enterprise 2025.1.x The following issues are fixed in this release. Issue: CQL connections were accepted in the main scheduling group and remained there if authentication was not enabled. Fix: Calls to update the scheduling group are now made for non-authenticated connections. (scylladb#26040, scylladb#26581) Issue: Service levels on Raft experienced local reads with timeouts from authentication tables. Fix: A long timeout is now set for authentication queries during the Service Level cache update. scylladb#25290 Issue: Documentation was missing explicit support for Debian 12. Fix: Support for Debian 12 has been added to the documentation. scylladb#26640 Issue: The documentation was missing information on the --list-active-releases option for the Web Installer. Fix: The documentation has been updated to include --list-active-releases for the Web Installer. scylladb#26688 Issue: An incorrect order of parameters was being logged in the repair log. Fix: The order of uuid and nodes_down in the repair log has been corrected. scylladb#26536 Issue: The node operations progress metric was not always resetting to 100% when the operation was complete. Fix: Node operations progress is now always reset to 100% upon completion. scylladb#26193 Issue: Maintenance mode was broken, causing test_maintenance_mode to fail. Fix: The system now falls back to the default scheduling group when using the maintenance socket, resolving the issue. scylladb#26816 Issue: The command to opt out a node from automatic cleanup was missing. Fix: Support to allow a node to opt out of automatic cleanup has been added. scylladb#26866 --- ### Page: https://forum.scylladb.com/t/two-new-lessons-on-scylladb-university-vector-search-and-tablets-elasticity-and-x-cloud/5209 Title: Two new lessons on ScyllaDB University, Vector Search and Tablets, Elasticity, and X Cloud - University and Training - ScyllaDB Community NoSQL Forum Meta Description: Two new lessons are now available on ScyllaDB University! 1. Vector Search: learn how to get started with ScyllaDB Vector Search. You will also understand how to efficiently store, index, and query vectors, enabling you… Language: en Canonical URL: https://forum.scylladb.com/t/two-new-lessons-on-scylladb-university-vector-search-and-tablets-elasticity-and-x-cloud/5209 ## Headings Structure: H1: Two new lessons on ScyllaDB University, Vector Search and Tablets, Elasticity, and X Cloud H3: Related topics ## Main Content: H1: Two new lessons on ScyllaDB University, Vector Search and Tablets, Elasticity, and X Cloud H3: Related topics Two new lessons are now available on ScyllaDB University! 1. Vector Search: learn how to get started with ScyllaDB Vector Search. You will also understand how to efficiently store, index, and query vectors, enabling you to build high-performance vector search applications. 2. Tablets, Elasticity, and X Cloud: you’ll learn about Tablets, Elasticity, and X Cloud. ScyllaDB distributes data by splitting tables into tablets. Each tablet has its replicas on different nodes, depending on the RF (Replication Factor). Each partition of a table is deterministically mapped to a single tablet. When you query or update the data, ScyllaDB can quickly identify the tablet that stores the relevant partition. Using the new tablets’ replication architecture significantly improves elasticity. Data is dynamically redistributed as the workload and topology evolve. New nodes can be spun up in parallel and adapt to the load in near real-time. Based on the above, X Cloud is an elastic database that supports variable/unpredictable workloads with consistent low latency, which translates to low costs. Each of the lessons includes some theory, practical points, and hands-on labs. Any feedback or something you’d like to see? Share it here. --- ### Page: https://forum.scylladb.com/t/release-scylladb-cpp-rs-driver-0-6-0/5210 Title: [RELEASE] ScyllaDB CPP RS Driver 0.6.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce ScyllaDB CPP RS Driver 0.6.0, an API-compatible rewrite of GitHub - scylladb/cpp-driver: Scylla C/C++ Driver as a wrapper for the Rust driver. It will fully replace the CPP driver… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-cpp-rs-driver-0-6-0/5210 ## Headings Structure: H1: [RELEASE] ScyllaDB CPP RS Driver 0.6.0 H2: Changes H2: Changes H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB CPP RS Driver 0.6.0 H2: Changes H2: Changes H3: Related topics The ScyllaDB team is pleased to announce ScyllaDB CPP RS Driver 0.6.0, an API-compatible rewrite of GitHub - scylladb/cpp-driver: Scylla C/C++ Driver as a wrapper for the Rust driver. It will fully replace the CPP driver, which is already reaching its End of Life. The current driver version should be considered Beta. Some minor features still need to be included. See Limitations section in README.md. The underlying Rust driver used version: 4f02969823f7b07b0d6b007a11863fe4aac95821 (not yet released commit from main branch). Big changes on our way to 1.0: New features / enhancements: CI / developer tool improvements: Congrats to all contributors and thanks everyone for using our driver! The source code of the driver can be found here: Contributions are most welcome! Thank you for your attention, please do not hesitate to contact us if you have any questions, issues, feature requests, or are simply interested in our driver! Contributors since the last release: The ScyllaDB team is pleased to announce ScyllaDB CPP RS Driver 0.6.0, an API-compatible rewrite of GitHub - scylladb/cpp-driver: Scylla C/C++ Driver as a wrapper for the Rust driver. It will fully replace the CPP driver, which is already reaching its End of Life. The current driver version should be considered Beta. Some minor features still need to be included. See Limitations section in README.md. The underlying Rust driver used version: 4f02969823f7b07b0d6b007a11863fe4aac95821 (not yet released commit from main branch). Big changes on our way to 1.0: New features / enhancements: CI / developer tool improvements: Congrats to all contributors and thanks everyone for using our driver! The source code of the driver can be found here: Contributions are most welcome! Thank you for your attention, please do not hesitate to contact us if you have any questions, issues, feature requests, or are simply interested in our driver! Contributors since the last release: --- ### Page: https://forum.scylladb.com/t/why-are-my-lightweight-transactions-lwts-not-working/5211 Title: Why are my lightweight transactions (LWTs) not working? - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello all! Lightweight transactions do not produce the expected changes in my ScyllaDB instance, and I’m at a loss as to why. The operations seem to complete successfully, as implied by the value of the [applied] column… Language: en Canonical URL: https://forum.scylladb.com/t/why-are-my-lightweight-transactions-lwts-not-working/5211 ## Headings Structure: H1: Why are my lightweight transactions (LWTs) not working? H3: Related topics ## Main Content: H1: Why are my lightweight transactions (LWTs) not working? H3: Related topics Lightweight transactions do not produce the expected changes in my ScyllaDB instance, and I’m at a loss as to why. The operations seem to complete successfully, as implied by the value of the [applied] column in the result set each of them returns. However, the contents of the singular row each such operation targets remains unchanged. Here is the schema for the table which the queries target: Here is the target row in the table (CSV): Here is the LWT query: And here is the result of the LWT query: Subsequent SELECT queries targeting the same row do not reflect the changes contained in the LWT. What could the issue be? I consulted Gemini on the matter, and it gave me the following possibilities: Installation details #ScyllaDB version: 2025.1.2-0.20250422.502c62d91d48 #Cluster size: 1 node os (RHEL/CentOS/Ubuntu/AWS AMI): MacOS (database is run through Docker Desktop) --- ### Page: https://forum.scylladb.com/t/deleting-old-sstables-after-ttl-change/5212 Title: Deleting old sstables after ttl change - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Installation details #ScyllaDB version: 5.2.19 (the update is planned after OS update ) #Cluster size: 9 nodes os (RHEL/CentOS/Ubuntu/AWS AMI): CentOS 8 (mostly) So I added a ttl for a table(TWCS) and want to manuall… Language: en Canonical URL: https://forum.scylladb.com/t/deleting-old-sstables-after-ttl-change/5212 ## Headings Structure: H1: Deleting old sstables after ttl change H3: Related topics ## Main Content: H1: Deleting old sstables after ttl change H3: Related topics Installation details #ScyllaDB version: 5.2.19 (the update is planned after OS update ) #Cluster size: 9 nodes os (RHEL/CentOS/Ubuntu/AWS AMI): CentOS 8 (mostly) So I added a ttl for a table(TWCS) and want to manually delete old data that was saved without any ttl. I’ve seen recommendations about writing all not expired data into new table but I really don’t want to do that since my ttl is big and most of the data is not yet expired - I would end up using x2 disk during this migration (i think ?) I was thinking about locating sstables that are fully expired (but don’t have ttl) with sstablemetadata and deleting them on disk. Is this possible ? The way twcs is described it shpud not be a problem. Data outside of ttl is rarely read and I dont care if a few read request in old time windows will fail but I plan to disable read repairs during this cleanup. Don’t know if I should remove a node from cluster during this cleanup anyway ? --- ### Page: https://forum.scylladb.com/t/logged-batch-atomicity/5214 Title: Logged BATCH atomicity - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hi everyone, I’m using denormalized tables with different partition keys (e.g., orders_by_order_id, orders_by_customer_id by customer_id, orders_by_merchant_id by merchant_id). When I write to all three tables using a … Language: en Canonical URL: https://forum.scylladb.com/t/logged-batch-atomicity/5214 ## Headings Structure: H1: Logged BATCH atomicity H3: My Questions: H3: Related topics ## Main Content: H1: Logged BATCH atomicity H3: My Questions: H3: Related topics I’m using denormalized tables with different partition keys (e.g., orders_by_order_id, orders_by_customer_id by customer_id, orders_by_merchant_id by merchant_id). When I write to all three tables using a logged BATCH: This pattern is also mentioned “ScyllaDB in action”, chapter 6.4. The batch will be executed against 3 different tables with 3 different partition keys. Will the batch log ensure ALL three inserts eventually succeed, or can some succeed while others fail permanently? --- ### Page: https://forum.scylladb.com/t/release-scylladb-manager-3-8-0/5215 Title: [RELEASE] ScyllaDB Manager 3.8.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team announces the release of ScyllaDB Manager 3.8.0, a production-ready minor release of the stable 3.8 branch. ScyllaDB Manager is a centralized cluster administration and recurrent tasks automation tool. … Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-manager-3-8-0/5215 ## Headings Structure: H1: [RELEASE] ScyllaDB Manager 3.8.0 H3: Dedicated tablet repair task (#4644) H3: New general repair task flags (#4645, #4708) H3: Recommendations for general and tablet repair tasks H3: Extended control over cluster suspend and resume (#4647) H3: Recommendations for suspending cluster to perform topology changes H3: Cleanup snapshots from interrupted backup (#4648) H3: Bug fixes and other improvements H3: Upgrade to the new release H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB Manager 3.8.0 H3: Dedicated tablet repair task (#4644) H3: New general repair task flags (#4645, #4708) H3: Recommendations for general and tablet repair tasks H3: Extended control over cluster suspend and resume (#4647) H3: Recommendations for suspending cluster to perform topology changes H3: Cleanup snapshots from interrupted backup (#4648) H3: Bug fixes and other improvements H3: Upgrade to the new release H3: Related topics The ScyllaDB team announces the release of ScyllaDB Manager 3.8.0, a production-ready minor release of the stable 3.8 branch. ScyllaDB Manager is a centralized cluster administration and recurrent tasks automation tool. This release focuses on handling ScyllaDB Manager tasks during cluster topology changes. Below are the changes in this release. ScyllaDB Manager 3.8.0 introduces new tablet repair task. This lightweight task is resilient to ongoing topology changes and uses ScyllaDB’s 2025.4 incremental repair feature, allowing it to run more frequently with minimal overhead. It can be scheduled with sctool repair tablet command. Apart from the tablet repair task, ScyllaDB Manager 3.8.0 extends general repair task flags with: It’s recommended to schedule both tablet and general repair tasks with: This solution splits a single general repair task into two separate ones, but it also brings additional benefits: Note that such a schedule is also recommended for clusters with user data replicated only with tablet or only with vnode keyspaces, since ScyllaDB system keyspaces can be replicated with either of them. The ScyllaDB Manager 3.8.0 release aims to make suspending and resuming the cluster both easier and safer with the following features: ScyllaDB Manager tasks assume stable cluster topology during their execution, with the exception of the new tablet repair task described above. It forces administrators to suspend ScyllaDB Manager tasks for the time of changing cluster topology. Moreover, tasks incompatible with topology changes need to be started from scratch afterwards. Otherwise, unexpected errors may occur. It’s recommended to suspend the cluster with sctool suspend --allow-task-type=tablet_repair --no-continue and resume it with sctool resume --start-tasks --start-tasks-missed-activation --no-continue. This solution brings the following benefits: In more urgent cases (e.g., cluster resize caused by the need to add more storage), it is advised to use --suspend-policy=stop_running_tasks so that the urgent operation can be performed without delay. In less urgent cases (e.g., cluster resize caused by too much free storage), it is advised to use fail_if_running_task to make sure that suspend won’t interrupt and lose the progress of currently running tasks. As mentioned above, ScyllaDB Manager 3.8.0 automatically removes snapshots from nodes’ disks left by the interrupted backups when --no-continue is used on suspend. Since this is a generally useful feature, it’s also possible to run such cleanup with sctool backup delete local-snapshots command. Fix ScyllaDB Manager Agent choice of pinned CPU (#4627) Use ScyllaDB side default incremental repair mode by default (#4683) Set query idempotency allowing for retries when querying local ScyllaDB Manager database (#4705) Fix filtering of task names containing slashes in sctool (#4653) ScyllaDB customers are encouraged to upgrade to ScyllaDB Manager 3.8.0 in coordination with the ScyllaDB support team. The new release includes upgrades of both ScyllaDB Manager Server and Agent. Scylla Manager 3.8.0 supports the following ScyllaDB releases You can install and run ScyllaDB Manager on Kubernetes using ScyllaDB Operator. More here. --- ### Page: https://forum.scylladb.com/t/last-4-weeks-in-scylla-cluster-tests-git-master-issue-118-2025-12-19/5217 Title: Last 4 weeks in scylla-cluster-tests.git master (issue #118; 2025-12-19) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the d5b33de6…b8bf758c range are covered. There were 100 non-merge commits from 20 authors in t… Language: en Canonical URL: https://forum.scylladb.com/t/last-4-weeks-in-scylla-cluster-tests-git-master-issue-118-2025-12-19/5217 ## Headings Structure: H1: Last 4 weeks in scylla-cluster-tests.git master (issue #118; 2025-12-19) H3: Related topics ## Main Content: H1: Last 4 weeks in scylla-cluster-tests.git master (issue #118; 2025-12-19) H3: Related topics This short report brings to light some interesting commits to scylla-cluster-tests.git master from the last week. Commits in the d5b33de6…b8bf758c range are covered. There were 100 non-merge commits from 20 authors in that period. Some notable commits: xcloud backend capabilities grew across the board, with GCP/GCE support for Vector Search tests, Manager node operations over SSH, SSH access and log collection for VS nodes, and log collection from the Manager node. Vector logging now filters noise on the client side, removing compaction/repair/streaming and other low-value entries before they reach SCT to lower memory and network overhead. Out-of-space error handling is workarounded with new nemesis ENOSPC selector. Added a comprehensive out-of-space prevention test suite covering write rejections, restarts near thresholds, index creation, compactions, repairs, decommissioning, and RF increases. Rolling-upgrade on ARM AMIs now forces iotune to compensate for older versions lacking i8g defaults, a safeguard to keep until 2026.1 becomes LTS. Performance testing gained a new “read disk only” scenario via test_read_disk_only_gradual_increase_load, focused on 100% cache misses with gradual load growth. Gemini adopts client-side timestamps with USING TIMESTAMP so operations are applied in the same order even when time skew between nodes is big. SkipPerIssue now accepts Jira references, both short jira:KEY and full URLs. Hydra moved to Python 3.14 on Debian Trixie. Cql-stress Docker tag was bumped to v0.2.5 bringing rust driver v1.3.1 and other dependency updates. The latte has been upgraded to 0.43.0-scylladb bringing performance improvements and functions for prepare phase. Also added schema-scaling rune script that can create thousands of tables with controlled cache hit/miss and flexible workloads. Repository formatting standardized on Ruff, starting with the switch to ruff format, followed by a treewide reformat and a dedicated .git-blame-ignore-revs entry to hide the mega-format commit in blame. A new quarantined folder isolates known-broken cases without removing them from the tree. CI was unblocked by bumping oracle cluster to 2024.1 in Gemini tests, replacing missing 2022.1 AMIs. See you in the next issue of last week in scylla-cluster-tests.git master! --- ### Page: https://forum.scylladb.com/t/release-new-billing-tab-for-active-clusters-in-scylladb-cloud/5219 Title: [RELEASE] New Billing Tab for Active Clusters in ScyllaDB Cloud - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: We’re happy to announce a new update to the ScyllaDB Cloud service that improves billing transparency for customers running active clusters. New features and enhancements We’ve added a new Billing tab for active cluster… Language: en Canonical URL: https://forum.scylladb.com/t/release-new-billing-tab-for-active-clusters-in-scylladb-cloud/5219 ## Headings Structure: H1: [RELEASE] New Billing Tab for Active Clusters in ScyllaDB Cloud H3: New features and enhancements H3: What’s included H3: Related topics ## Main Content: H1: [RELEASE] New Billing Tab for Active Clusters in ScyllaDB Cloud H3: New features and enhancements H3: What’s included H3: Related topics We’re happy to announce a new update to the ScyllaDB Cloud service that improves billing transparency for customers running active clusters. We’ve added a new Billing tab for active clusters, allowing customers to easily review how their clusters are currently being billed, broken down per data center. With this update, customers can clearly see whether a cluster is billed on demand or covered by one or more contract line items, all in one place. The new Billing tab is available for all active clusters and can be accessed via Cluster → Billing in the ScyllaDB Cloud UI. You can also view all of your contract line items (both active and expired) from the main Billing page, accessible via the top-right menu → Billing, making it easier to review contracts across clusters. --- ### Page: https://forum.scylladb.com/t/last-fortnight-in-scylladb-git-master-issue-302-2025-12-21/5220 Title: Last fortnight in scylladb.git master (issue #302; 2025-12-21) - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last two weeks. Commits in the 47efbdffbc..f65db4e8eb range are covered. There were 208 non-merge commits from 37 authors in tha… Language: en Canonical URL: https://forum.scylladb.com/t/last-fortnight-in-scylladb-git-master-issue-302-2025-12-21/5220 ## Headings Structure: H1: Last fortnight in scylladb.git master (issue #302; 2025-12-21) H3: Related topics ## Main Content: H1: Last fortnight in scylladb.git master (issue #302; 2025-12-21) H3: Related topics This short report brings to light some interesting commits to scylladb.git master from the last two weeks. Commits in the 47efbdffbc..f65db4e8eb range are covered. There were 208 non-merge commits from 37 authors in that period. Some notable commits: Tablet repair now reports its progress via the task manager. Metrics exposed by Prometheus node_exporter now include the ethtool collector. The frozen toolchain was rebased to Fedora 43 with clang 21.1. Authentication now uses a password hash implementation that runs on the reactor but avoids stalls. This prevents a CPU bottleneck when running on non-reactor CPUs. Batchlog replay performance was improved. There is a new system.client_routes table, intended to help clients determine routes to nodes behind a reverse proxy (e.g. AWS privatelink and similar). The ALTER KEYSPACE statement can now convert a keyspace with a numeric replication factor to an explicit rack list. Tablets will be migrated to the specified racks. See you in the next issue of last week in scylladb.git master! --- ### Page: https://forum.scylladb.com/t/configuring-kernel-parameters-for-scylla/5221 Title: Configuring kernel parameters for Scylla - ScyllaDB - ScyllaDB Community NoSQL Forum Meta Description: Hello! Is there a current article on configuring kernel parameters for Scylla? I thought the scylla_setup and scylla_prepare scripts automatically configure the kernel, but, for example, these scripts don’t add HugePage… Language: en Canonical URL: https://forum.scylladb.com/t/configuring-kernel-parameters-for-scylla/5221 ## Headings Structure: H1: Configuring kernel parameters for Scylla H3: Related topics ## Main Content: H1: Configuring kernel parameters for Scylla H3: Related topics Is there a current article on configuring kernel parameters for Scylla? I thought the scylla_setup and scylla_prepare scripts automatically configure the kernel, but, for example, these scripts don’t add HugePages, which are necessary for most databases. --- ### Page: https://forum.scylladb.com/t/release-scylladb-2025-4-0/5222 Title: [RELEASE] ScyllaDB 2025.4.0 - Release Notes - ScyllaDB Community NoSQL Forum Meta Description: The ScyllaDB team is pleased to announce the release of ScyllaDB 2025.4, a production-ready ScyllaDB Short Term Support (STS) Minor Feature Release. More information on ScyllaDB’s Long Term Support (LTS) policy is avail… Language: en Canonical URL: https://forum.scylladb.com/t/release-scylladb-2025-4-0/5222 ## Headings Structure: H1: [RELEASE] ScyllaDB 2025.4.0 H2: Related Links H2: New Features H3: Vector Search H3: Tablets: Alternator H3: Tablets: Support for Materialized Views and Secondary Indexes H3: Tablets: Support for Lightweight Transactions (LWT) H3: Tablets: Support for Change Data Capture (CDC) H3: Tablets: Driver Support H3: Out-of-space Prevention Guardrail H3: Default Replication Factor H3: Default CREATE KEYSPACE Syntax H3: Add Azure Key Vault as an Encryption at Rest Key Provider H3: Alternator Improvements H3: Default SSTable compression algorithm H3: Trie-Based SSTable Index Format H3: Deployment options H2: More Updates H3: Reference H3: Related topics ## Main Content: H1: [RELEASE] ScyllaDB 2025.4.0 H2: Related Links H2: New Features H3: Vector Search H3: Tablets: Alternator H3: Tablets: Support for Materialized Views and Secondary Indexes H3: Tablets: Support for Lightweight Transactions (LWT) H3: Tablets: Support for Change Data Capture (CDC) H3: Tablets: Driver Support H3: Out-of-space Prevention Guardrail H3: Default Replication Factor H3: Default CREATE KEYSPACE Syntax H3: Add Azure Key Vault as an Encryption at Rest Key Provider H3: Alternator Improvements H3: Default SSTable compression algorithm H3: Trie-Based SSTable Index Format H3: Deployment options H2: More Updates H4: Stability H4: Tablets H4: Enhancements & Correctness H4: Optimization H4: Performance H4: Change Data Capture (CDC) H4: Tools and API H4: CQL H4: Log and traces H4: Monitoring H4: Setup H3: Reference H3: Related topics The ScyllaDB team is pleased to announce the release of ScyllaDB 2025.4, a production-ready ScyllaDB Short Term Support (STS) Minor Feature Release. More information on ScyllaDB’s Long Term Support (LTS) policy is available here. Highlights of the 2025.4 release include: Tablets now support Materialized Views (MV), Secondary Indexes (SI), Change Data Capture (CDC), and Lightweight Transactions (LWT). This fully bridges the previous feature gap between Tablets and vNodes. ScyllaDB Vector Search is now available (in GA), introducing native low-latency Approximate Nearest Neighbor similarity search (ANN) through CQL. See the Getting Started Guide and try it out. Alternator fully supports Tablets by following the tablets_mode_for_new_keyspaces configuration flag, except for the still-experimental Streams. The new Trie-based index format improves indexing efficiency. New deployment options with i8g and i8ge show significant performance advantages over i4i, i3en, as well as i7i and i7ie. ScyllaDB 2025.4 introduces native Vector Search to power AI-driven applications. By integrating vector indexing directly into the ScyllaDB ecosystem, teams can now perform similarity searches without moving data to a separate vector database. Vector Search is currently available only in ScyllaDB Cloud, the fully managed ScyllaDB service. For more information, see: This release introduces support for Change Data Capture with Tablet-enabled tables. #22576 Previously, CDC was supported only for vnode-based tables. In this release, ScyllaDB adds full CDC support for Tablet-enabled tables. This allows users to track inserts, updates, and deletes in tablet-based keyspaces. With Tablet-based CDC: Enabling CDC: CDC can be enabled on a table at creation or afterwards: Enabling it on an existing table: ALTER TABLE my_table WITH cdc = {'enabled': true}; Once CDC is enabled, all changes to the table will automatically be logged in the corresponding CDC log table. Tablets support and performance optimizations have been added to the drivers. See the documentation to identify which driver versions include Tablets support. Improved Alternator to fully align with DynamoDB getRecords behavior, ensuring consistent results. #6931 Performance Improvement Introduced caching of expressions in requests, as parsing expressions can be slow. Performance gains depend on expression type and complexity, with observed single-node throughput improvements of ~7–15% in example workloads. #5023 Add Support for Per-Table Metrics in Alternator. #19824 ScyllaDB 2025.4 introduces a new SSTable Index implementation based on a Tries index. This feature improves the performance of modification operations and the performance of data lookup (reads) in many cases. Trie-Based SSTable Index is disabled by default in 2025.4. The Trie-based index is not enabled by default. It can be enabled by setting the sstable_format parameter in the scylla.yaml file to ms. This release introduces a trie-based index format designed to improve performance through faster lookups and more efficient memory usage. The default SSTable format remains “me”, the same index type used in previous versions. The new format, “ms”, is similar to “me” but uses Trie-based indexes. It provides a more compact index structure and typically faster lookups. #25626 When set to ms, newly created SSTables will use the new format, while existing SSTables will continue using their current index. To convert existing SSTables to the new format, rewrite them using nodetool upgradesstables. AWS: This release expands support to include all I7i and I7ie instance types, in addition to the previously supported i7i.large, i7i.xlarge, i7i.2xlarge, and all i7ie instances. The I7i and I7ie families offer an improved price-to-performance ratio compared to previous generations. AWS: This release adds support for i8g and ig8e instance families, which provide a better price-performance ratio compared to x86-based instances. Enabling Alternator streams on a table with a very long name crashes ScyllaDB. #24598 Nodes now become eligible to be Raft voters earlier during bootstrap, closing a gap that could expose clusters to loss-of-quorum. #24420 Encryption using AWS Key Management Service (KMS) can now use externally-provided credentials. This makes recovery tasks simpler. #22470 Cleanup of Key Management Interoperability Protocol (KMIP) connections, used by encryption, is more robust, avoiding TLS errors. #24873 A large remote procedure call (RPC) hash table is now sized in advance to prevent resizing from causing latency jumps. #24660 #24217 Improved handling of in-progress requests during node shutdown, ensuring requests are allowed to complete successfully. #24481 Drop table during tablet cleanup initiated by a tablet migration may crash. #25706 Raft: Atomic schema updates (COW for schema state). Schema changes are now applied to an in-memory copy and made visible atomically, ensuring each Raft command’s modifications are isolated. This simplifies reasoning about state, avoids cross-command dependencies, and improves restart performance by reducing repeated full-state copies. #19649 Commitlog: segment_manager::discard_unused_segments - separate the modification of the vector (_segments) from actual releasing of objects. #25709 A deadlock when removing a failed node in a cluster with materialized views was fixed. #24807 service/qos: Modularize service level controller to avoid invalid access to a stopped auth::service. #24792 Fix stability issue with KMS caching testing. #24574 Alternator: character 0x255 in base64 causes an out-of-bounds read. #25701 Alternator: expression parsing memory leak in specific situations. 25878 auth: default user creation may cause a crash if done during upgrade to raft topology. #24975 load_balancer: std::out_of_bounds when decommissioning with empty nodes. #26203 Reworked group0 Raft server shutdown to avoid use-after-free issues during drain or shutdown. Introduced abort_and_drain() to safely stop background tasks while keeping the server alive, and moved final destruction to abort_and_destroy() for proper cleanup. #24625 Coredump during truncate: Data written after truncation time was incorrectly truncated. #25013 Coredump right after seed node decommission. #23911 Potential data race in utils::alien_worker. #24751 Gossip: Failed to add server: “No ip address for … when one is expected”. #23407 Gossiper - race condition in gossip::apply_state_locally when receiving messages containing removed endpoints that no longer have a valid host id. #25702, #25621 Gossiper can reply with an empty self- host IDid in `gossip_get_endpoint_states_response` during startup. #25831 Node shutdown is stuck waiting for batchlog manager drain. #24599 Non-RBNO streaming may hang on failure if IO permits are deadlocked. #24925 Potential node crash when doing base-view pairing after decreasing RF. #21492 token_metadata_ptr may be destroyed without gentle cleaning. #13381 Unregister raft_topology_get_cmd_status on shutdown. #24910 Raft topology: make the voter handler consider only group 0 members. #26321 service/qos: set long timeout for auth queries on SL cache update. #25290 Race condition between tablet split and load-and-stream. #26455 s3_client: parse multipart response XML defensively. #25009 Raft topology: fix group0 tombstone GC in the Raft-based recovery procedure. #26534 Fixed an issue in tablet migration rollback where view-building tasks were recreated for the target instead of the source replica. The rollback logic now correctly restores tasks on the migration source. #26825 Fixed a synchronization issue between tablet split and load-and-stream that occurred when using Gossip topology. #22707 Failed to boot because of a missing TOC for SSTable. #25919 Tablets: stop storage group on deallocation #24857 #24828. This can cause errors when a node is restarted during tablet cleanup and other use cases. Hints may get lost during a node replace. #24980 Tablets: repair: tablet repair isn’t safe in case of topology coordinator failover. #23318 When tablets are split, a compaction process is initiated to break apart SSTables that span the new tablet boundary. This compaction is now prepared for a tablet merge to happen before the split compaction is complete. #24153 API: updated range_to_endpoint_map to return tablet ranges consistent with system.tablets for tablet-enabled keyspaces, ensuring correct token boundaries and replica mappings. #26331 The load balancer now tracks migration badness, measuring how a tablet’s move affects table balance on the source and destination. Previously, badness for source and destination was combined using std::max(), which ignored negative values (good migrations). This could result in incorrect migration decisions. The computation has been corrected to preserve the actual badness values. #26091 Improved compaction task progress tracking by introducing an ‘expected total workload’ calculation. This ensures progress is reported as (sum of child task progresses) / expected total workload, reducing fluctuations in the total workload. #8392, #6406, #7845. The batchlog mechanism will now drop batches for a table that was itself dropped. #24806 Hints are now sent to pending replicas, such as tablet migration targets or new nodes. #19835 Introduced view_building_coordinator, a cluster-wide coordinator for building tablet-based views. It splits work into tablet-level tasks, is resilient to RPC failures and Raft leader changes, adapts to tablet operations, and safely handles staging SSTables to ensure consistent view updates. #22288, #19149, #21564, #17603, #22586, #18826, #23930 The container image entry point now supports --dc and --rack options, allowing these parameters to be set without bind-mounting the cassandra.rackdb file. #23423 The system will no longer create SSTables with numeric generation numbers (only UUIDs). It can still read such SSTables. #24248 Lightweight transactions (LWT) now implement fencing, which prevents old requests that were sent using an old version of the topology from being incorrectly applied to a new topology. This is a step for implementing LWT on tablets. #22332 PRUNE MATERIALIZED VIEW statement can miss some ghost rows. #25655 Fixed scrub failures caused by concurrent compactions deleting SSTables before checksum verification. #23363 Changed ignore_nodes from IPs to host IDs to streamline removenode operations. #26249, #25958 The CREATE TABLE IF NOT EXISTS no longer fails if CDC is specified. #26142 There are now Lua scripts for examining purgeable tombstones and write-time histograms in SSTables. #26062 Unauthorized connections now switch to the sl:default scheduling group for consistency with authenticated users without an assigned service level. #26040 CDC: Set specific tombstone_gc when creating log tables to ensure consistent behavior with other schema entities and avoid confusing output of DESCRIBE. #25187 Compaction: reorder operations in compaction_manager::stop to ensure that all compaction executors are stopped and prevent assertion failures. #25806 Repair: always set node ops progress to 100% on completion to prevent lingering incomplete status after errors. #26193 encryption::encryption_file_io_extension::wrap_sink: removed default case in component_type switch and explicitly handle all values to enforce consideration of new components. #23724 Improved S3 client error handling in chunked_download_source to prevent callback-related errors from propagating. Disabled retries on Seastar’s side to ensure data integrity and avoid multiple downloads of the same range. #25043 DB/View: fixed cross-shard access by wrapping shared_sstable in foreign_ptr within view_building_worker, ensuring safe transfer of staging SSTables between shards. #25859 Tasks: updated make_and_start_task to return task::impl instead of task to allow direct access to implementation details, enabling features like returning values from tasks. #22146 Replica: fixed race between table drop and merge completion that could cause false compaction checks. The fix makes truncate ignore stopped groups during compaction verification, preventing failures when merging and dropping overlap. #25551 SSTables: updated make_entry_descriptor regex to be non-greedy, ensuring correct keyspace and table matching in snapshot directories. #25242 Storage Service: drain the view builder before group0 to ensure proper coordination of view building operations. #25096 HTTP: updated query parameter usage to the new interface. Accessing parameters directly is deprecated; use {get,set}_query_param() to handle multiple values per key. #26023 Utils/Stall-Free: fixed detection of clear_gently for const payload types in smart/shared pointer containers (e.g., foreign_ptr), ensuring objects are gently cleared before destruction while still disallowing direct calls on const objects. #24605, #25026 s3_client: fix `when` condition to prevent infinite locking. #26497 s3_client: track memory starvation in background filling fiber. #26465 s3: Fix chunked download source metrics calculations. #25875 Improved efficiency of stopping table-specific compactions by introducing a filter function to the compaction manager. #25846, #26082 raft topology: disable schema pulls in the Raft-based recovery procedure. #26569 Access tablet map through table->effective_replication_map() on SSTable discovery. #26403 Materialized view creation: improved handling of cases where a reader returns no partitions for a partial range. #26635 The row cache is able to purge expired tombstones in order to improve the performance of reads that later touch the same key. To prevent data resurrection, it checks memtables for overlapping data. It now avoids checking memtables for which it can prove there is no overlapping order data, reducing false positives and increasing the number of tombstones purged. #24962 The small table repair optimization is used when bootstrapping or repairing tables with little data, like system_tracing tables. It now optimizes token range calculations for larger clusters. #24817 An SSTable Bloom filter is built with an estimate of the number of partitions it will hold. If the estimate turns out to be incorrect, we rebuild the Bloom filter in order not to waste memory. We now avoid the rebuild if the memory wasted is low enough to be ignored. #25464, #25468 After streaming or repairing data, the portion of the row cache affected is invalidated since it no longer reflects the underlying SSTables. We now invalidate at partition granularity rather than token-range granularity, resulting in increased cache efficiency, particularly after repair. #9136 Gossiper operations are now enforced to run in the gossip scheduling group, even if invoked from other components. #25907 The table used for the Raft log now has caching disabled. #26027 The tablet load balancer now considers dead nodes in its calculations. #24485 The SSTable scrubber now handles malformed SSTables better. #19059 More comprehensive schema information is now stored in SSTables, making data recovery from an orphaned SSTable easier. #24187 ScyllaDB now formats the data filesystem using 4k block size. We previously used 1k block size to work around a kernel deficiency, which has since been fixed. #25441 Dropped tables should be removed from system.truncated. #25683 raft_topology: updated remove node logic to improve concurrency by checking raft_topology_change_enabled() on shard0 and routing execution accordingly; allows concurrent remove node operations on Raft-enabled clusters. #24737 storage_service::maybe_reconnect_to_preferred_ip() calls the gossiper::get_host_id() unnecessarily and can be passed directly as a parameter. #25715 Alternator validation of the table name on ordinary read/write requests is done only if the table lookup fails. This provides a small optimization. #12538 The main data structure holding vnode tokens was changed to avoid large allocations, which can produce stalls on very large clusters. #24876 Repair now sends smaller messages for partition differences with many small rows (e.g., tombstone-only rows), reducing stalls and improving efficiency. #24808 storage_service: on_change: frequent system.peers reloads with many peers may cause high CPU usage. #25660 Improve performance of worst case scenario in auth::passwords::detail::hash_with_salt. #24524, #13136 A large remote procedure call (RPC) hash table is now sized in advance to prevent resizing from causing latency jumps, #24660, #24217 Alternator is now more careful to avoid large contiguous allocations during query/scan operations, as these can cause stalls. #23535 Boot-time stall while creating CDC generation. #24522 The tablet load balancer now runs in the maintenance/streaming scheduling group, rather than the gossip group. This reduces the impact on node failure detection and Raft, which also run in the gossip group. #26037 The tablet scheduler will no longer attempt expensive cross-rack migrations if the cluster is known to have exactly one tablet replica per rack. This is much faster. #26016 Alternator, ScyllaDB’s implementation of the DynamoDB API, now caches parsed expressions to reduce per-request overhead. #25855 Large allocation when converting a batch statement to a mutation for batchlog. #24809 Locator:util: optimized describe_ring by caching per-endpoint info, reducing CPU time by ~20%. #24887 Fix oversized allocation in Paxos under pressure. #25559 Service/Tablet Allocator: replaced large contiguous vector in make_repair_plan() with a chunked vector to prevent stalls from large allocations. #24713 system_keyspace::drop_truncation_rp_records can cause unbound parallelism. #25682 The setup tool now checks if 4k block sizes are preferred by the disk, even if it advertises 512 byte physical sector sizes, to compensate for disks that misreport the physical sector size. #25315 Alternator Streams: fixed PutItem to emit a single INSERT or MODIFY event, matching DynamoDB Streams behavior. Previously, each PutItem produced both REMOVE and MODIFY events. The fix uses tombstones for :attrs and regular columns to generate the correct CDC event. #6930, #24991 Alternator: fixed handling of UpdateItem with a combination of legacy AttributeUpdates and ReturnValues=ALL_OLD. #25894 Alternator: improved error reporting for non-existent resources. Operations like TagResource now return ResourceNotFoundException with a clear message when an ARN points to a table that doesn’t exist, instead of the previous vague AccessDeniedException. This matches DynamoDB behavior. #26179 Alternator: simplify std::views::transform usage. Replace lambdas that extract a class member with pointer-to-member syntax, making the code more concise. #25012 Alternator Streams: refactored std::views::transform usage to remove side effects. Side effects inside transforms could run multiple times depending on the algorithm, so appending to the column vector is now done before the transform to ensure it happens exactly once. #25011 describe_multi_item previously misinterpreted its last argument, causing a double conversion that undercounted RCU. The fix removes the extra conversion and updates the API. Tests pass against DynamoDB, ensuring correct RCU calculation. #25847 Alternator, ScyllaDB’s implementation of the DynamoDB API, now skips materialized view building when indexes are created for empty tables. #26615 Authentication can now work in an unenforcing mode (warnings only) to test the effects of changing the auth configuration. #25308, #25457 Alternator: DescribeTable may return an incorrect RANGE key for GSIs. #5320 It’s possible to manually drop columns from CDC log. #24643 Node crashes when dropping a column while writing to it with CDC. #24952 A new REST API endpoint allows dropping quarantined SSTables to reclaim their storage. SSTables are quarantined automatically when corruption is detected. #19061 Fixed CDC permission checks so users with Select rights on the base table can now query CDC tables. #19798 Warn users about using RF!=Racks with Tablets Keyspaces. #23330 It is now possible to DROP ROLEs when using the saslauthd authenticator. #25571 Alternator: now calculates Write Capacity Units (WCU) in a compatible manner compared to DynamoDB. #24436 Internode communication errors are now reported to the user during LWT transaction failures. #25318 storage_service/get_natural_endpoints has no way to pass a key containing a colon character. #24829, #16596 Alternator now supports writing to system tables, in addition to reading them. #12348 The SscyllaDB SSTable tool now uses the more modern UUID generations rather than numeric generations. #25166 The nodetool stop cleanup command did not stop all cleanup tasks; only the ones currently running. Pending cleanup compactions would still run. It now aborts all pending tasks. #20823 The system.clients columns reporting encrypted connections (SSL/TLS) are now filled in. #9216 Refusing tablet table repairs via /storage_service/repair_async: The previous API is unsafe for tablet tables and may cause crashes . A new repair API, /storage_service/tablets/repair, is now available for tablets. Once all users (manager and nodetool) have migrated to the new API, the old API will no longer allow repairing tablet tables. #23008 Alternator: Improve error message for overflowing list indexes. #25947 There is now a standalone SscyllaDB SSsstable upgrade command, which can be used to rewrite SSTables using different versions or options. #26109 repair: to_repair_rows_on_wire stalls destroying input list #24725 DESCRIBE MATERIALIZED VIEW now shows the actual view definition (as a commented CREATE MATERIALIZED VIEW), instead of a misleading CREATE INDEX statement. #24610 Table Compression: ScyllaDB’s compression DDL property lets you set the algorithm and chunk size per table. By default, LZ4Compressor with a 4 KiB chunk size is used for both user and system tables. #25195 CQL3/ResultSet: set GLOBAL_TABLES_SPEC in query responses when all columns belong to the same table to reduce metadata overhead. #17788 The system.clients virtual table now lists active Alternator requests in addition to CQL connections. #24993 Compaction start/end messages were demoted to DEBUG level, as they were deemed too noisy. #24949 Reduced log severity in dirty_memory_manager::flush_one – The severity level of logs has been lowered for seastar::named_gate_closed_exception, as this flush failure is expected when the corresponding compaction group has already been stopped. #25037 Incoherent commas in AUTH audit log. #24410 Removed unnecessary logs when aborting tablet splits on node stop. #24850 Fixed SSTable streaming logs during tablet migration to correctly show the list of files sent, preventing confusion from previously empty file lists. #25830 Streaming: Enclose potential throws in try block and ensure sink close before logging. #25591, #25497 db/hints: Improve logs. #25466 Tablets: demoted sstables_repaired_at log to debug to reduce verbosity on large tablet counts. #25926 Fixed separator formatting from “ ,” to “,“ for improved log readability. #23883 Debug logs: deletion_time formatter prints wrong results. #25556 Change the error message of colocated repair to show table names instead of table IDs. #26567 Incorrect order of parameters in repair log. #26536 Compaction: demote normal compaction start/end log messages to debug level. #24949 Compaction: Enhanced trace logging in get_max_purgeable_timestamp. Now includes keyspace and table names in log messages, and adds trace logs for early return cases when tombstone_gc is disabled or gc_check_only_compacting_sstables is enabled. #24914. Improved log clarity: When compactions are aborted, the log message now specifies the exact user-triggered reason (“truncate”, “cleanup”, “rewrite”, or “split”) instead of the generic “user-triggered operation”. #25136 Avoid setting the request_type field for truncation commands unless the topology_global_request_queue feature is enabled, ensuring compatibility with older nodes that don’t recognize global request types. #24731, #24716 There are now metrics for S3 prefetches. #25876 Added general metrics for each index. #25970 S3 Client: Memory Usage Metric – providing visibility into backup and restore memory consumption. #25864, #25769 New metric for time spent on tablet repair. #26505 scylla_sysconfig_setup prints warnings in 2025.2.0. #24915 Docker: Container image lacks the PS command used by ScyllaDB scripts. #24827 --- ### Page: https://forum.scylladb.com/t/performance-issues-running-scylladb-in-a-vm/5225 Title: Performance issues, running ScyllaDB in a VM - ScyllaDB - ScyllaDB Community NoSQL Forum Language: en Canonical URL: https://forum.scylladb.com/t/performance-issues-running-scylladb-in-a-vm/5225 ## Headings Structure: H1: Performance issues, running ScyllaDB in a VM H3: Related topics ## Main Content: H1: Performance issues, running ScyllaDB in a VM H3: Related topics Originally from the User Slack @hamonicahamonica**:** I am planning to store observability data using ScyllaDB 6.2, but I am experiencing lower-than-expected performance and would like to request technical guidance. Under normal conditions, writing 10,000–20,000 rows per second is not a problem. These rows represent what I call root spans. However, each root span has child spans beneath it—anywhere from fewer than 10 to as many as 100–300 child spans. Assuming an average of 150 child spans per root span, I am observing significant disk I/O wait during writes. I created a span_trace table as shown below. The data is primarily grouped by trace_id. However, if I include trace_id as part of the partition key, I understand that this can lead to an excessive number of partitions, which may also be problematic. I am trying to understand how to improve write performance under this model. I have experimented with bucketing created_at at various granularities—5 seconds, 10 seconds, and currently 1 minute and 5 minute buckets—but I am still testing to find the optimal configuration. To clarify, the 10k–20k rows/sec figure represents an upper-bound scenario. At present, the actual workload is closer to 1,000–2,000 TPS (rows per second). OS : Rocky-Linux-9 CPU: Architecture: x86_64 CPU op-mode(s): 32-bit, 64-bit Address sizes: 52 bits physical, 57 bits virtual Byte Order: Little Endian CPU(s): 8 On-line CPU(s) list: 0-7 Vendor ID: AuthenticAMD BIOS Vendor ID: QEMU Model name: AMD EPYC 9224 24-Core Processor @dor**:** Well, you run in a VM. It’s ok but you need to make sure that the VM cpus/memory aren’t taken by other apps on the host. Ideally pin them. In addition, you must run scylla setup at the beginning ---