### Page: https://forum.scylladb.com/t/about-the-announcements-category/1
Title: About the Announcements category - Announcements - ScyllaDB Community NoSQL Forum
Meta Description: News and information related to products, the ScyllaDB community, events, webinars, and so on
Language: en
Canonical URL: https://forum.scylladb.com/t/about-the-announcements-category/1
## Headings Structure:
H1: About the Announcements category
H3: Related topics
## Main Content:
H1: About the Announcements category
H3: Related topics
News and information related to products, the ScyllaDB community, events, webinars, and so on
---
### Page: https://forum.scylladb.com/t/about-the-database-community-category/3
Title: About the Database Community category - Database Community - ScyllaDB Community NoSQL Forum
Meta Description: Have fun and engage with your fellow community members.
Feel free to introduce yourself here, add feature requests, and provide feedback. Also, the place to post ScyllaDB related job opportunities.
How are you using S…
Language: en
Canonical URL: https://forum.scylladb.com/t/about-the-database-community-category/3
## Headings Structure:
H1: About the Database Community category
H3: Related topics
## Main Content:
H1: About the Database Community category
H3: Related topics
Have fun and engage with your fellow community members.
Feel free to introduce yourself here, add feature requests, and provide feedback. Also, the place to post ScyllaDB related job opportunities.
How are you using ScyllaDB?
---
### Page: https://forum.scylladb.com/t/welcome-to-the-scylladb-community-nosql-forum/7
Title: Welcome to the ScyllaDB Community NoSQL Forum - Database Community - ScyllaDB Community NoSQL Forum
Meta Description: Explore NoSQL topics as you learn from the ScyllaDB community
Before posting, please search and make sure the question hasn't been asked before. To ask a question, [log in](https://forum.scylladb.com/login) first.
A fe…
Language: en
Canonical URL: https://forum.scylladb.com/t/welcome-to-the-scylladb-community-nosql-forum/7
## Headings Structure:
H1: Welcome to the ScyllaDB Community NoSQL Forum
H2: Explore NoSQL topics as you learn from the ScyllaDB community
H3: Related topics
## Main Content:
H1: Welcome to the ScyllaDB Community NoSQL Forum
H2: Explore NoSQL topics as you learn from the ScyllaDB community
H3: Related topics
A few helpful resources:
Also, here’s a blog with details on why we launched this forum and how it relates to Slack and our other community channels: Introducing the ScyllaDB Community Forum - ScyllaDB
---
### Page: https://forum.scylladb.com/t/about-the-scylladb-category/12
Title: About the ScyllaDB category - ScyllaDB - ScyllaDB Community NoSQL Forum
Meta Description: For general questions related to ScyllaDB products. Topics include Troubleshooting, Benchmarks, Data modeling, Drivers, 3rd party integrations, etc.
Language: en
Canonical URL: https://forum.scylladb.com/t/about-the-scylladb-category/12
## Headings Structure:
H1: About the ScyllaDB category
H3: Related topics
## Main Content:
H1: About the ScyllaDB category
H3: Related topics
For general questions related to ScyllaDB products. Topics include Troubleshooting, Benchmarks, Data modeling, Drivers, 3rd party integrations, etc.
Hi,
I studied you peoples forum and earlier also checked one webinar, any chance to join you people to learn something more.
---
### Page: https://forum.scylladb.com/t/about-the-university-and-training-category/13
Title: About the University and Training category - University and Training - ScyllaDB Community NoSQL Forum
Meta Description: For topics regarding specific ScyllaDB University courses, lessons and training events.
Language: en
Canonical URL: https://forum.scylladb.com/t/about-the-university-and-training-category/13
## Headings Structure:
H1: About the University and Training category
H3: Related topics
## Main Content:
H1: About the University and Training category
H3: Related topics
For topics regarding specific ScyllaDB University courses, lessons and training events.
---
### Page: https://forum.scylladb.com/t/scylladb-summit-2023-call-for-speakers-is-open/30
Title: ScyllaDB Summit 2023 Call for Speakers is open! - Announcements - ScyllaDB Community NoSQL Forum
Meta Description: Call for Speakers for ScyllaDB Summit 2023 is now open! If you have war stories of deployments or awesome tools and integrations we’d love to hear!
Language: en
Canonical URL: https://forum.scylladb.com/t/scylladb-summit-2023-call-for-speakers-is-open/30
## Headings Structure:
H1: ScyllaDB Summit 2023 Call for Speakers is open!
H3: Related topics
## Main Content:
H1: ScyllaDB Summit 2023 Call for Speakers is open!
H3: Related topics
Call for Speakers for ScyllaDB Summit 2023 is now open! If you have war stories of deployments or awesome tools and integrations we’d love to hear!
---
### Page: https://forum.scylladb.com/t/high-performance-nosql-masterclass-register-now-for-09-nov-2022/31
Title: High Performance NoSQL Masterclass — register now for 09 Nov 2022 - Announcements - ScyllaDB Community NoSQL Forum
Meta Description: Register now for our latest Masterclass, and read the agenda. Hosted by ScyllaDB and our friends at Pythian.
If you’re registered, sound off below!
Language: en
Canonical URL: https://forum.scylladb.com/t/high-performance-nosql-masterclass-register-now-for-09-nov-2022/31
## Headings Structure:
H1: High Performance NoSQL Masterclass — register now for 09 Nov 2022
H3: Related topics
## Main Content:
H1: High Performance NoSQL Masterclass — register now for 09 Nov 2022
H3: Related topics
Register now for our latest Masterclass, and read the agenda. Hosted by ScyllaDB and our friends at Pythian.
If you’re registered, sound off below!
---
### Page: https://forum.scylladb.com/t/engaging-with-the-scylladb-community/32
Title: Engaging with the ScyllaDB Community - Database Community - ScyllaDB Community NoSQL Forum
Meta Description: There are already a number of ways to engage with the ScyllaDB Community. Let’s just go over a few:
ScyllaDB User Slack (http://slack.scylladb.com/)
ScyllaDB User Google Group “scylladb-users” (https://groups.google.co…
Language: en
Canonical URL: https://forum.scylladb.com/t/engaging-with-the-scylladb-community/32
## Headings Structure:
H1: Engaging with the ScyllaDB Community
H3: Related topics
## Main Content:
H1: Engaging with the ScyllaDB Community
H3: Related topics
There are already a number of ways to engage with the ScyllaDB Community. Let’s just go over a few:
So why would we want another? Because users appreciate choice, and different people like to engage via different media.
That said, there are usually strengths and weaknesses to different media which would gravitate certain conversations to different platforms. Online forums are nothing new, but community forums, discussion boards, persist because they have certain advantages:
If people are happy to be using their existing media of choice to communicate with the ScyllaDB community and are getting the answers they need, then please stick with it! These forums are our way of adding more choice, and to hopefully to get better re-use of the knowledge you all share with the community day in and day out.
As the Director of Technical Advocacy here at ScyllaDB, we hope to use this forum to communicate with you more actively, interactively and directly. A big shout-out to Guy Shtub who championed the project into fruition!
Thanks Peter!
A bit more about when to use slack and when to use this forum:
Community Forum
This forum is a great place to ask questions, get detailed answers, have in-depth discussions, and search for all the previously answered questions.
Benefits:
User Slack
Our community Slack channel provides real-time chat with ScyllaDB experts and other ScyllaDB users.
Benefits:
Great work!
Lets send it to the open world…
An official Discord server would be so cool tohave for Scylla.
---
### Page: https://forum.scylladb.com/t/say-hello-how-are-you-using-scylladb/35
Title: Say Hello, How are you using ScyllaDB? - Database Community - ScyllaDB Community NoSQL Forum
Meta Description: Welcome to the forum! Please go ahead and introduce yourself.
ScyllaDB is used for many different use cases.
Any interesting projects you’d like to share with your fellow community members?
Are there any tips/tricks y…
Language: en
Canonical URL: https://forum.scylladb.com/t/say-hello-how-are-you-using-scylladb/35
## Headings Structure:
H1: Say Hello, How are you using ScyllaDB?
H3: Related topics
## Main Content:
H1: Say Hello, How are you using ScyllaDB?
H3: Related topics
Welcome to the forum! Please go ahead and introduce yourself.
ScyllaDB is used for many different use cases.
Any interesting projects you’d like to share with your fellow community members?
Are there any tips/tricks you learned along the way?
How did you first hear about ScyllaDB?
I’ll go first
My name is Guy Shtub, and I’m Head of Training at ScyllaDB.
I joined Scylla over four years ago to build ScyllaDB University.
My tip is that we have lots of topics covered in ScyllaDB University and in the Documentation.
If you have any questions or issues, you can ask them here. Also, feel free to suggest new topics/lessons/hands-on labs!
Hiii
My name is Mahdi Fardkohan and I have been working as a DBA & BI Engineer in a software company for several years. I’m so passionate about Data Mesh Architecture, and I found ScyllaDB when I was researching high-performance open-source distributed databases for Big Data to use in a Retail startup and I was amazed by its capabilities.
I’m honored to be here with great people like you in the ScyllaDB community, and I’m also very grateful to the venerable Scylla and its wonderful experts who put this power and knowledge into the hands of the world.
Thanks for your kind feedback, Mahdi!
My name is Yaniv and I’m the VP R&D of Scylla, very happy to be here and read the forum notes!
My name is Diego and I’m the support engineer at Scylla, interested in the forum notes, thanks!
Hi, my name is Fabio, and I’m a software engineer in testing leader.
I’m into details, and passionate about performance, so happy to be here.
Hi Everyone!
Great to be here!
This is my first time ever to join a cool community like ScyllaDB
My name is Trinh, an expat working in Germany as Data Engineer. I have no experience with NoSQL in general, but using SQL and Distributed DB (ex: HDFS) a lot. I’d like to fullfill my knowledge in this area so that I can complete my picture of Data industry. Happy to meet you in my Twitter https://twitter.com/bkincities
Hi,
My name is Raya and I am a a software engineer in testing manager at ScyllaDB.
Thrilled to be here
Hi,
My name is Orane Gabrielovitch, and I’m the Israel office admin.
Happy to be here
Hi all, I’m happy to be part of our ScyllaDB community
Inna B.
Operations & Admin Manager at ScyllaDB
Hello ScyllaDB community. I’m happy to be here. I don’t use ScyllaDB but I’m curious to learn more.
Hi there. I’m Suraj Vijayan, Undergraduate student from India. I’m a Full Stack web and app developer (started coding at 16) and currently into Rust and Cassandra. Learnt Cassandra (through DataStax workhops and their academy) and got certified as Developer Associate.
I had always been looking into Discord’s tech stack and came to know the fact that they moved from Cassandra to ScyllaDB. Though I studied a lot about it in blogs, I so curios (excited too like - “Whoa!! Damn! Awesome!! why not give a try?” ). And this how I ended here. Looking forward to learn ScyllaDB.
Talking about projects I have a made a lot, but the one I like a lot is my portfolio site. As its an full stack site and uses Cassandra (AstraDB) to collect website stats (yes a lot, no google analytics), people can comment, like and share my works and achievments. Why not give a try to it - surajvijayan.me and also an honorable mention to my AstraDB Navigator - A low code environment to manage you AstraDB instances.
I guess this is an pretty huge message. but yeah thank you if you came this far.
Hi,
My name is Clive Attard, and I’m a software engineer learning new technologies.
Excited to be here
Hello, This is kartikeya. Currently started using scylla db in our org
Hi, I am Abdulshakur. Excited to learn Scylladb
Hi, I am Gary, looking forward to learn Scylladb!
Hi , My name is Amir , I am a Software engineer.
Hi Scylladb community
I’m Charan working as a Database reliability engineer. Currently I’m learning and implementing scylla on my environments. Happy to connect here.
---
### Page: https://forum.scylladb.com/t/about-the-knowledge-base-category/40
Title: About the Knowledge Base category - Knowledge Base - ScyllaDB Community NoSQL Forum
Meta Description: Topics for understanding and troubleshooting ScyllaDB. These are frequently asked questions and general topics.
Language: en
Canonical URL: https://forum.scylladb.com/t/about-the-knowledge-base-category/40
## Headings Structure:
H1: About the Knowledge Base category
H3: Related topics
## Main Content:
H1: About the Knowledge Base category
H3: Related topics
Topics for understanding and troubleshooting ScyllaDB. These are frequently asked questions and general topics.
---
### Page: https://forum.scylladb.com/t/what-is-the-difference-between-clustering-primary-partition-and-composite-or-compound-keys-in-scylladb/41
Title: What is the difference between Clustering, Primary, Partition, and Composite (or Compound) Keys in ScyllaDB? - Knowledge Base - ScyllaDB Community NoSQL Forum
Meta Description: In ScyllaDB (and Apache Cassandra for that matter) A Primary Key is defined within a table. It is one or more columns used to identify a row. All tables must include a definition for a Primary Key. For example, in the ta…
Language: en
Canonical URL: https://forum.scylladb.com/t/what-is-the-difference-between-clustering-primary-partition-and-composite-or-compound-keys-in-scylladb/41
## Headings Structure:
H1: What is the difference between Clustering, Primary, Partition, and Composite (or Compound) Keys in ScyllaDB?
H3: Related topics
## Main Content:
H1: What is the difference between Clustering, Primary, Partition, and Composite (or Compound) Keys in ScyllaDB?
H3: Related topics
In ScyllaDB (and Apache Cassandra for that matter) A Primary Key is defined within a table. It is one or more columns used to identify a row. All tables must include a definition for a Primary Key. For example, in the table:
The Primary Key is a single column – the pet_chip_id. If a Primary Key is made up of a single column, it is called a Simple Primary Key.
It’s also possible to define the Primary Key to include more than one column, in which case it is called a Composite (or Compound) key. For example:
In this case, the first part of the Primary Key is called the Partition Key (pet_chip_id in the above example) and the second part is called the Clustering Key (time).
The Partition Key is responsible for data distribution across the nodes. It determines which node will store a given row. It can be one or more columns.
The Clustering Key is responsible for sorting the rows within the partition. It can be zero or more columns.
If a table has more than one column defined as the Primary Key, for example:
In this case, the Partition Key includes two columns: pet_chip_id and time, and the Clustering Key is pet_name. Every query must include all the columns defined in the Partition Key (pet_chip_id and time) in this case.
Look at another example:
If there is more than one column in the Clustering Key (pet_name and heart_rate in the example above), the order of these columns defines the clustering order. For a given partition, all the rows are physically ordered inside ScyllaDB by the clustering order. This order determines what select queries you can efficiently run on this partition.
In this example, the ordering is first by pet_name and then by heart_rate.
In addition to the Partition Key columns, a query may include the Clustering Key. If it does include the Clustering Key columns, they must be used in the same order as they were defined.
Additional Resources:
---
### Page: https://forum.scylladb.com/t/how-many-connections-should-i-open-from-each-scylladb-client-application/43
Title: How many connections should I open from each ScyllaDB client application? - Knowledge Base - ScyllaDB Community NoSQL Forum
Meta Description: As a rule of thumb, for Scylla’s best performance, each client needs at least 1-3 connections per Scylla core. For example, in a cluster with three nodes, each node with 16 cores, each client application should open 32 (…
Language: en
Canonical URL: https://forum.scylladb.com/t/how-many-connections-should-i-open-from-each-scylladb-client-application/43
## Headings Structure:
H1: How many connections should I open from each ScyllaDB client application?
H3: Related topics
## Main Content:
H1: How many connections should I open from each ScyllaDB client application?
H3: Related topics
As a rule of thumb, for Scylla’s best performance, each client needs at least 1-3 connections per Scylla core. For example, in a cluster with three nodes, each node with 16 cores, each client application should open 32 (2x16) connections to each Scylla node.
Additional Resources:
---
### Page: https://forum.scylladb.com/t/whats-the-best-way-to-count-scan-a-table-in-scylladb-full-table-scan/48
Title: What's the best way to count / scan a table in ScyllaDB? (full table scan) - Knowledge Base - ScyllaDB Community NoSQL Forum
Meta Description: If you wish to perform a full table scan, please look into the following blog posts on how to do an efficient full table scan:
The theory behind an efficient full table scan Blog by Avi (our CTO)
Follow up Blog with co…
Language: en
Canonical URL: https://forum.scylladb.com/t/whats-the-best-way-to-count-scan-a-table-in-scylladb-full-table-scan/48
## Headings Structure:
H1: What's the best way to count / scan a table in ScyllaDB? (full table scan)
H3: Related topics
## Main Content:
H1: What's the best way to count / scan a table in ScyllaDB? (full table scan)
H3: Related topics
If you wish to perform a full table scan, please look into the following blog posts on how to do an efficient full table scan:
Note: There’s also a python version of that efficient full table scan script.
---
### Page: https://forum.scylladb.com/t/scylladb-version-upgrade/49
Title: ScyllaDB Version Upgrade - Knowledge Base - ScyllaDB Community NoSQL Forum
Meta Description: Can I upgrade from ScyllaDB Version X to version Y?
Language: en
Canonical URL: https://forum.scylladb.com/t/scylladb-version-upgrade/49
## Headings Structure:
H1: ScyllaDB Version Upgrade
H3: Related topics
## Main Content:
H1: ScyllaDB Version Upgrade
H3: Related topics
Can I upgrade from ScyllaDB Version X to version Y?
If you’re using ScyllaDB Cloud, you don’t have to upgrade, as it’s a fully managed service. It deploys the latest ScyllaDB Enterprise version, and all upgrades are performed by ScyllaDB.
If you’re using ScyllaDB Open Source or Enterprise, you can upgrade to a newer version following these upgrade guides:
A ScyllaDB upgrade is a rolling procedure - it does not require full cluster shutdown and is performed without any downtime or disruption of service.
To ensure a successful upgrade and avoid breaking anything, you should perform your upgrades consecutively - to each successive version. For example, to upgrade from version 4.4 to 5.0, you should first upgrade ScyllaDB to version 4.5, next from 4.5 to 4.6, and finally, from 4.6 to 5.0, without skipping any major version.
---
### Page: https://forum.scylladb.com/t/scylladb-is-using-up-all-of-my-memory-what-s-going-on/50
Title: ScyllaDB is using up all of my memory. What’s going on? - Knowledge Base - ScyllaDB Community NoSQL Forum
Meta Description: ScyllaDB is designed to use all the memory it has and to put it to good use. Most notably to cache data. By default, when ScyllaDB starts up, it inspects the node’s hardware configuration and claims all memory to itself,…
Language: en
Canonical URL: https://forum.scylladb.com/t/scylladb-is-using-up-all-of-my-memory-what-s-going-on/50
## Headings Structure:
H1: ScyllaDB is using up all of my memory. What’s going on?
H3: Related topics
## Main Content:
H1: ScyllaDB is using up all of my memory. What’s going on?
H3: Related topics
ScyllaDB is designed to use all the memory it has and to put it to good use. Most notably to cache data. By default, when ScyllaDB starts up, it inspects the node’s hardware configuration and claims all memory to itself, leaving some reserve for the operating system. The assumption is that ScyllaDB does not run in a shared environment.
If for some reason you do want to give scyllaDB less memory, say for testing or development, you can do so:
Additional Resources:
---
### Page: https://forum.scylladb.com/t/scylladb-university-live-1st-of-december/52
Title: ScyllaDB University LIVE - 1st of December - Announcements - ScyllaDB Community NoSQL Forum
Meta Description: The next ScyllaDB University LIVE training event will take place on the 1st of December. It’s a half-day free online training with some of our best engineers and experts. Learn more here, hope to see you there.
Language: en
Canonical URL: https://forum.scylladb.com/t/scylladb-university-live-1st-of-december/52
## Headings Structure:
H1: ScyllaDB University LIVE - 1st of December
H3: Related topics
## Main Content:
H1: ScyllaDB University LIVE - 1st of December
H3: Related topics
The next ScyllaDB University LIVE training event will take place on the 1st of December. It’s a half-day free online training with some of our best engineers and experts. Learn more here, hope to see you there.
Check out my blog post ScyllaDB University LIVE, Fall 203: From Getting Started to Expert Tips & Tricks for more details, see you there!
This is happening tomorrow (Wednesday). You can still save your spot here.
---
### Page: https://forum.scylladb.com/t/i-m-having-performance-issues-what-should-i-do/53
Title: I’m having performance issues. What should I do? - Knowledge Base - ScyllaDB Community NoSQL Forum
Meta Description: ScyllaDB auto-tunes for optimal performance, however, users still need to apply best practices in order to get the most out of their ScyllaDB deployments.
Lower than expected performance can be a result of many factors,…
Language: en
Canonical URL: https://forum.scylladb.com/t/i-m-having-performance-issues-what-should-i-do/53
## Headings Structure:
H1: I’m having performance issues. What should I do?
H3: Related topics
## Main Content:
H1: I’m having performance issues. What should I do?
H3: Related topics
ScyllaDB auto-tunes for optimal performance, however, users still need to apply best practices in order to get the most out of their ScyllaDB deployments.
Lower than expected performance can be a result of many factors, from hardware (storage, CPU, network) to data modeling to the application layer.
As a first step, make sure you have ScyllaDB Monitoring in place. Looking at the monitoring dashboards is the best way to look for bottlenecks and understand the cause of the issues. Here are some other tips:
Additional Resources:
---
### Page: https://forum.scylladb.com/t/scylladb-monitoring-has-no-data-in-the-charts/54
Title: ScyllaDB Monitoring has no data in the charts - Knowledge Base - ScyllaDB Community NoSQL Forum
Meta Description: The most common reason for seeing no data about your cluster after installing the ScyllaDB Monitoring stack is an issue with the connection to Prometheus. To check this:
Login to the Prometheus console by pointing your…
Language: en
Canonical URL: https://forum.scylladb.com/t/scylladb-monitoring-has-no-data-in-the-charts/54
## Headings Structure:
H1: ScyllaDB Monitoring has no data in the charts
H3: Related topics
## Main Content:
H1: ScyllaDB Monitoring has no data in the charts
H3: Related topics
The most common reason for seeing no data about your cluster after installing the ScyllaDB Monitoring stack is an issue with the connection to Prometheus. To check this:
Other things to check:
Make sure you are not using the local network for the local IP range when using Docker containers. By default, the local IP range (127.0.0.X) is inside the Docker container. If you are trying to connect to a target via the local IP range from a Docker container, you need to use the -l flag to enable the local network stack.
Verify that Prometheus is pointing to the correct target by checking prometheus/scylla_servers.yml.
Make sure that your dashboard and Scylla versions are aligned. If, for example, you are running Scylla 5.1, you can specify a specific version with the -v flag when starting the monitoring stack: ./start-all.sh -v 5.1
Additional Resources:
---
### Page: https://forum.scylladb.com/t/when-to-use-filtering-and-when-not/55
Title: When to use filtering - and when not - Knowledge Base - ScyllaDB Community NoSQL Forum
Meta Description: Sometimes you want to be able to query by different columns, but you’re not interested in creating secondary indexes. Filtering is one more way of allowing such queries. The mechanism is really simple – the coordinator w…
Language: en
Canonical URL: https://forum.scylladb.com/t/when-to-use-filtering-and-when-not/55
## Headings Structure:
H1: When to use filtering - and when not
H3: Related topics
## Main Content:
H1: When to use filtering - and when not
H3: Related topics
Sometimes you want to be able to query by different columns, but you’re not interested in creating secondary indexes. Filtering is one more way of allowing such queries. The mechanism is really simple – the coordinator will fetch all of the results specified by the key restrictions, and then filter out rows that do not match the rest of the restrictions.
But, there’s a catch. Filtering can be very performance-heavy, it can even result in fetching all rows from the table and then filter out just a few rows. Because of that, queries that involve filtering must be explicitly allowed to do so, by appending CQL ALLOW FILTERING keyword to each such query.
Here are some popular resources about CQL ALLOW FILTERING
---
### Page: https://forum.scylladb.com/t/how-do-you-use-scylladb-with-spring-boot/56
Title: How do you use ScyllaDB with Spring Boot? - Knowledge Base - ScyllaDB Community NoSQL Forum
Meta Description: This blog explains how to use Spring Boot apps with ScyllaDB for time series data, taking advantage of shard-aware drivers and prepared statements: Using Spring Boot, ScyllaDB and Time Series Data - ScyllaDB
Language: en
Canonical URL: https://forum.scylladb.com/t/how-do-you-use-scylladb-with-spring-boot/56
## Headings Structure:
H1: How do you use ScyllaDB with Spring Boot?
H3: Related topics
## Main Content:
H1: How do you use ScyllaDB with Spring Boot?
H3: Related topics
This blog explains how to use Spring Boot apps with ScyllaDB for time series data, taking advantage of shard-aware drivers and prepared statements: Using Spring Boot, ScyllaDB and Time Series Data - ScyllaDB
---
### Page: https://forum.scylladb.com/t/about-the-blog-posts-category/58
Title: About the Blog Posts category - Blog Posts - ScyllaDB Community NoSQL Forum
Meta Description: (Replace this first paragraph with a brief description of your new category. This guidance will appear in the category selection area, so try to keep it below 200 characters.) Use the following paragraphs for a longer description, or to establish category guidelines or rules: Why should people use this category? What is it for? How exactly is this different than the other categories we already have? What should topics in this category generally contain? Do we need this category? Can...
Language: en
Canonical URL: https://forum.scylladb.com/t/about-the-blog-posts-category/58
## Headings Structure:
H1: About the Blog Posts category
H3: Related topics
## Main Content:
H1: About the Blog Posts category
H3: Related topics
(Replace this first paragraph with a brief description of your new category. This guidance will appear in the category selection area, so try to keep it below 200 characters.)
Use the following paragraphs for a longer description, or to establish category guidelines or rules:
Why should people use this category? What is it for?
How exactly is this different than the other categories we already have?
What should topics in this category generally contain?
Do we need this category? Can we merge with another category, or subcategory?
Honestly, this category feels too vague. before making a new one, ask yourself: who is it for, what belongs here, and what doesn’t.
most new categories fail because they overlap with existing ones or nobody posts there.
If you can’t clearly say “post this here, not there,” just use tags or a subcategory instead. keeps the forum clean and less confusing.
---
### Page: https://forum.scylladb.com/t/how-scylladb-helped-an-adtech-company-focus-on-core-business/60
Title: How ScyllaDB Helped an AdTech Company Focus on Core Business - Blog Posts - ScyllaDB Community NoSQL Forum
Meta Description: [image]
Language: en
Canonical URL: https://forum.scylladb.com/t/how-scylladb-helped-an-adtech-company-focus-on-core-business/60
## Headings Structure:
H1: How ScyllaDB Helped an AdTech Company Focus on Core Business
H3: How ScyllaDB Helped an AdTech Company Focus on Core Business
H3: Related topics
## Main Content:
H1: How ScyllaDB Helped an AdTech Company Focus on Core Business
H3: How ScyllaDB Helped an AdTech Company Focus on Core Business
H3: Related topics
AdTech innovator GumGum wanted to escape the maintenance issues connected with Apache Cassandra. This podcast shares how they moved to a database-as-a-service (ScyllaDB Cloud DBaaS).
---
### Page: https://forum.scylladb.com/t/experimental-features/61
Title: Experimental features - Knowledge Base - ScyllaDB Community NoSQL Forum
Meta Description: Experimental features are still under development and are not stable enough to be used in production. Their design is not finalized, and their API will likely change, breaking backward or forward compatibility. Experimen…
Language: en
Canonical URL: https://forum.scylladb.com/t/experimental-features/61
## Headings Structure:
H1: Experimental features
H3: Related topics
## Main Content:
H1: Experimental features
H3: Related topics
Experimental features are still under development and are not stable enough to be used in production. Their design is not finalized, and their API will likely change, breaking backward or forward compatibility. Experimental features are available for ScyllaDB Open Source users to test and provide feedback.
ScyllaDB Enterprise and ScyllaDB Cloud do not support experimental features.
To list all the experimental features available in your ScyllaDB version, run scylla --help.
To enable an experimental feature, you can do one of the following:
---
### Page: https://forum.scylladb.com/t/which-driver-should-i-use/62
Title: Which driver should I use? - ScyllaDB - ScyllaDB Community NoSQL Forum
Meta Description: Which driver should I use? Can I use Apache Cassandra drivers?
Language: en
Canonical URL: https://forum.scylladb.com/t/which-driver-should-i-use/62
## Headings Structure:
H1: Which driver should I use?
H3: Related topics
## Main Content:
H1: Which driver should I use?
H3: Related topics
Which driver should I use? Can I use Apache Cassandra drivers?
ScyllaDB is compatible with Apache Cassandra on the protocol level, so that every Cassandra Driver will work with Scylla out of the box.
That being said, for many drivers (Java, Python, Go, Rust…) there are Scylla forks that take advantage of Scylla features, like shared per core, to allow better performance.
It is recommended to use this fork when available.
See Scylla CQL Drivers | Scylla Docs for a list of drivers.
---
### Page: https://forum.scylladb.com/t/database-security-authentication-authorization-rba-encyrption-and-audit/63
Title: Database Security: Authentication, Authorization, RBA, Encyrption and Audit - ScyllaDB - ScyllaDB Community NoSQL Forum
Meta Description: We’re concerned with data breaches. How secure is ScyllaDB?
Language: en
Canonical URL: https://forum.scylladb.com/t/database-security-authentication-authorization-rba-encyrption-and-audit/63
## Headings Structure:
H1: Database Security: Authentication, Authorization, RBA, Encyrption and Audit
H3: Scylla Security Checklist | Scylla Docs
H3: Related topics
## Main Content:
H1: Database Security: Authentication, Authorization, RBA, Encyrption and Audit
H3: Scylla Security Checklist | Scylla Docs
H3: Related topics
We’re concerned with data breaches. How secure is ScyllaDB?
Security is a big topic that includes many features like:
Encryption (on rest, in transit), Authentication and Authorization, role-based access control, auditing, and more.
An excellent place to start is Scylla Security Checklist:
Scylla is an Apache Cassandra-compatible NoSQL data store that can handle 1 million transactions per second on a single server.
---
### Page: https://forum.scylladb.com/t/can-the-compression-strategy-be-changed-from-stcs-to-ics-live-or-is-a-node-restart-required/64
Title: Can the compression strategy be changed from STCS to ICS live, or is a node restart required? - ScyllaDB - ScyllaDB Community NoSQL Forum
Meta Description: Yes, you can change the compaction strategy live - without restarting any nodes.
Note that there might be increased compaction activity for a while after you make the change, so schedule it to off-peak hours if possible…
Language: en
Canonical URL: https://forum.scylladb.com/t/can-the-compression-strategy-be-changed-from-stcs-to-ics-live-or-is-a-node-restart-required/64
## Headings Structure:
H1: Can the compression strategy be changed from STCS to ICS live, or is a node restart required?
H3: Related topics
## Main Content:
H1: Can the compression strategy be changed from STCS to ICS live, or is a node restart required?
H3: Related topics
Yes, you can change the compaction strategy live - without restarting any nodes.
Note that there might be increased compaction activity for a while after you make the change, so schedule it to off-peak hours if possible.
See How to Change Compaction Strategy for details.
---
### Page: https://forum.scylladb.com/t/scylladb-5-1-release-candidate/65
Title: ScyllaDB 5.1 release candidate - Announcements - ScyllaDB Community NoSQL Forum
Meta Description: The Scylla team is pleased to announce ScyllaDB Open Source 5.1 RC4, a Release Candidate for the Scylla Open Source 5.1 minor release.
We encourage you to run ScyllaDB 5.1 release candidates on your test environments; t…
Language: en
Canonical URL: https://forum.scylladb.com/t/scylladb-5-1-release-candidate/65
## Headings Structure:
H1: ScyllaDB 5.1 release candidate
H3: Related topics
## Main Content:
H1: ScyllaDB 5.1 release candidate
H3: Related topics
The Scylla team is pleased to announce ScyllaDB Open Source 5.1 RC4, a Release Candidate for the Scylla Open Source 5.1 minor release.
We encourage you to run ScyllaDB 5.1 release candidates on your test environments; this will help ensure that an upgrade to ScyllaDB 5.1 General Availability will proceed smoothly with your workload. Use the release candidate with caution; RC4 is not production-ready yet. You can help stabilize Scylla Open Source 5.1 by reporting bugs here.
Only the last two minor releases of the ScyllaDB Open Source project are supported. Once Scylla Open Source 5.1 is officially released, ScyllaDB Open Source 5.1 and 5.0 will be supported, and ScyllaDB 4.6 will be retired.
For a complete description of ScyllaDB 5.1 see ScyllaDB 5.1.
ScyllaDB 5.1 RC1 , RC2, RC3
Get ScyllaDB Open Source 5.1 (under “More Versions” for each distro)
Upgrade from ScyllaDB 5.0 to 5.1
Updates and bug fixes since 5.1 RC3
Stability: Memtable flush aborts the node if it fails due to out-of storage space #11245
The Scylla team is pleased to announce ScyllaDB Open Source 5.1 RC4, a Release Candidate for the Scylla Open Source 5.1 minor release.
We encourage you to run ScyllaDB 5.1 release candidates on your test environments; this will help ensure that an upgrade to ScyllaDB 5.1 General Availability will proceed smoothly with your workload. Use the release candidate with caution; RC4 is not production-ready yet. You can help stabilize Scylla Open Source 5.1 by reporting bugs here.
Only the last two minor releases of the ScyllaDB Open Source project are supported. Once Scylla Open Source 5.1 is officially released, ScyllaDB Open Source 5.1 and 5.0 will be supported, and ScyllaDB 4.6 will be retired.
For a complete description of ScyllaDB 5.1 see ScyllaDB 5.1.
ScyllaDB 5.1 RC1 , RC2, RC3
Get ScyllaDB Open Source 5.1 (under “More Versions” for each distro)
Upgrade from ScyllaDB 5.0 to 5.1
Updates and bug fixes since 5.1 RC3
The Scylla team is pleased to announce ScyllaDB Open Source 5.1 RC5, a Release Candidate for the Scylla Open Source 5.1 minor release.
We encourage you to run ScyllaDB 5.1 release candidates on your test environments; this will help ensure that an upgrade to ScyllaDB 5.1 General Availability will proceed smoothly with your workload. Use the release candidate with caution; RC5 is not production-ready yet. You can help stabilize Scylla Open Source 5.1 by reporting bugs here.
Only the last two minor releases of the ScyllaDB Open Source project are supported. Once Scylla Open Source 5.1 is officially released, ScyllaDB Open Source 5.1 and 5.0 will be supported, and ScyllaDB 4.6 will be retired.
For a complete description of ScyllaDB 5.1 see ScyllaDB 5.1.
ScyllaDB 5.1 RC1 , RC2, RC3, RC4
Get ScyllaDB Open Source 5.1 (under “More Versions” for each distro)
Upgrade from ScyllaDB 5.0 to 5.1
Updates and bug fixes since 5.1 RC4
---
### Page: https://forum.scylladb.com/t/sharechat-s-path-to-high-performance-nosql-q-a-with-geetish-nayak/67
Title: ShareChat’s Path to High-Performance NoSQL: Q&A with Geetish Nayak - Blog Posts - ScyllaDB Community NoSQL Forum
Meta Description: [800x400-blog-sharechat-webinar-ondemand]
Language: en
Canonical URL: https://forum.scylladb.com/t/sharechat-s-path-to-high-performance-nosql-q-a-with-geetish-nayak/67
## Headings Structure:
H1: ShareChat’s Path to High-Performance NoSQL: Q&A with Geetish Nayak
H3: ShareChat’s Path to High-Performance NoSQL: Q&A with Geetish Nayak
H3: Related topics
## Main Content:
H1: ShareChat’s Path to High-Performance NoSQL: Q&A with Geetish Nayak
H3: ShareChat’s Path to High-Performance NoSQL: Q&A with Geetish Nayak
H3: Related topics
How India's social media unicorn achieves microsecond P99 latency with 1.2M op/sec – for 180M monthly active users expecting real-time engagement with 2.5B posts per month.
---
### Page: https://forum.scylladb.com/t/scylladb-support/69
Title: ScyllaDB Support - ScyllaDB - ScyllaDB Community NoSQL Forum
Meta Description: How does ScyllaDB Support work? Is it 24/7?
Language: en
Canonical URL: https://forum.scylladb.com/t/scylladb-support/69
## Headings Structure:
H1: ScyllaDB Support
H3: Related topics
## Main Content:
H1: ScyllaDB Support
H3: Related topics
How does ScyllaDB Support work? Is it 24/7?
ScyllaDB protects every customer with professional, top-notch technical support.
24/7 and around the globe our professional team is ready to deliver the knowledge and expertise needed to resolve any issue, answer any questions, and meet our customer satisfaction commitment.
OSS users will get community support, official documentation and free access to ScyllaDB university.
Scylla Enterprise and Scylla Cloud customers subscriptions include a ScyllaDB Enterprise license, tested and certified binaries, software updates, hot fixes, and technical NoSQL support.
More information can be found here
---
### Page: https://forum.scylladb.com/t/migrating-to-scylladb-from-cassandra-mongodb-dynamodb/70
Title: Migrating to ScyllaDB from Cassandra MongoDB DynamoDB - ScyllaDB - ScyllaDB Community NoSQL Forum
Meta Description: What does the typical migration process from (Cassandra/DynamoDB/Mongo) look like? How long does it take?
Language: en
Canonical URL: https://forum.scylladb.com/t/migrating-to-scylladb-from-cassandra-mongodb-dynamodb/70
## Headings Structure:
H1: Migrating to ScyllaDB from Cassandra MongoDB DynamoDB
H3: Related topics
## Main Content:
H1: Migrating to ScyllaDB from Cassandra MongoDB DynamoDB
H3: Related topics
What does the typical migration process from (Cassandra/DynamoDB/Mongo) look like? How long does it take?
This is a very wide question (you can learn more in a Scylla University dedicated lesson).
You need to consider whether you’ll be doing a Hot/Live or a Cold/Offline migration.
There’s the live traffic aspect and the historical data aspect.
If you’ll be doing a HOT migration, here is the sequence you should follow:
Now let’s talk about historical data migration.
The following tools can be used to migrate historical data from Cassandra:
The following tools can be used to migrate historical data from DynamoDB to ScyllaDB (CQL API / DynamoDB compatible API, a.k.a Alternator):
In order to migrate from MongoDB you’ll 1st need to design your schema / data model. A nice tool that can help you with that is Hackolade
As for data migration tools:
Thanks @Tomersan! Very helpful.
I’ll add that if you’re doing a migration, get in touch with us, we’re happy to help.
Also, in the upcoming LIVE event, we’ll have a dedicated session on migrating from DynamoDB.
Hate to necro this thread but it’s been about a year - has any of the information above changed substantially? We’re considering moving from our Cass 3.11 cluster (12 nodes, 2 DCs) to Scylla 5.1. Any recommendations or tips are welcome.
No substantial change that I’m aware of.
Later today, we’ll host another LIVE training event, and there will be a specific session on migrating from DynamoDB to ScyllaDB, including a hands-on example. Hope to see you there.
@Guy Can I see recordings of such events? If not, think about it, it might be useful for newcomers.
These events are available live only.
We’re working on updating the Migration lesson on ScyllaDB University.
Also, stay tuned for the next training event (we have one every few weeks) as well as for the upcoming ScyllaDB Summit.
---
### Page: https://forum.scylladb.com/t/scylladb-and-large-partitions/71
Title: ScyllaDB and Large Partitions - ScyllaDB - ScyllaDB Community NoSQL Forum
Meta Description: Are large partitions still an issue and if so, how can I deal with it?
What is the maximal partition size?
Language: en
Canonical URL: https://forum.scylladb.com/t/scylladb-and-large-partitions/71
## Headings Structure:
H1: ScyllaDB and Large Partitions
H3: Related topics
## Main Content:
H1: ScyllaDB and Large Partitions
H3: Related topics
Are large partitions still an issue and if so, how can I deal with it?
What is the maximal partition size?
Large partitions are an anti-pattern in Scylla - one should aim to have data spread evenly among the partitions in the cluster.
The presence of large partitions create issues such as higher shard latency (because it has a hot spot for data) or oversized allocation warnings/errors on logs.
In any case, you can use Scylla’s system.large_partitions virtual table (doc) to help you track them, along with log lines that indicate their occurrence.
Check this Scylla University lesson covering this topic.
As to how avoid them on the first place, you should wisely plan your data modeling in a way that leads to a high cardinality distribution of data among partitions. Check this blog post for some ideas around that.
---
### Page: https://forum.scylladb.com/t/is-scylladb-right-for-my-application-scylladbs-sweet-spot/72
Title: Is ScyllaDB Right for My Application? ScyllaDB's Sweet Spot - ScyllaDB - ScyllaDB Community NoSQL Forum
Meta Description: This is a question that I often get:
I’m currently evaluating different databases for my application. What is the sweet spot for ScyllaDB?
Language: en
Canonical URL: https://forum.scylladb.com/t/is-scylladb-right-for-my-application-scylladbs-sweet-spot/72
## Headings Structure:
H1: Is ScyllaDB Right for My Application? ScyllaDB's Sweet Spot
H3: Related topics
## Main Content:
H1: Is ScyllaDB Right for My Application? ScyllaDB's Sweet Spot
H3: Related topics
This is a question that I often get:
I’m currently evaluating different databases for my application. What is the sweet spot for ScyllaDB?
I’d say it’s a need for HIGH volume throughput with HIGH cardinality and LOW (15ms or even single digit) tail (p95/99) latency.
Many times our I/O scheduler will do most of the prioritization for you.
Application not optimized enough? reach out to our Solution architects for advise on better data modeling.
Want more optimization tips? Make sure to check Scylla Monitoring:
Have cardinality problems? check this out.
I agree with what Tomer wrote about performance (High throughput, low latency).
I’d also add to the sweet spot High Availability and Big Data:
There are other reasons that teams choose ScyllaDB:
I’d recommend a relational database and not ScyllaDB if:
I’m going to add a few more elements that make for a sweet spot for ScyllaDB:
Transactional (OLTP) vs. Analytical (OLAP) — while ScyllaDB does have workload prioritization for balancing various workloads on the same cluster, and certain kinds of analytics can be run on ScyllaDB, we’re far more focused on transactional/operational workloads. ScyllaDB is a row-oriented data store, vs. a columnar database that is designed for analytics.
Single-digit Millisecond P99 Latencies — ScyllaDB is optimized for locally-attached NVMe SSD. While we have a built-in row-based in-memory cache, we’re not a pure-play in-memory database or data grid. So you will find that ScyllaDB is faster than a database connected to block storage, but not as expensive as a RAM-based system. It works in a “goldilocks” zone optimizing performance and price.
High Throughput — This is a subjective term, but generally ScyllaDB is the database to look at when you scale to tens of thousands, hundreds of thousands and millions of operations per second (OPS). ScyllaDB uses immutable LSM tree based storage, so it is optimized for fast writes. And because it has a built-in row-based cache we are also good for fast reads.
Multi-datacenter Replication — ScyllaDB automagically does replication across multiple sites, so if you are looking to deploy your data in the region of your users we will take care of that distribution for you.
You don’t need to do all four of these. Any one of these make ScyllaDB a good fit. But the more of these criteria apply to your use case, the more you should be adding ScyllaDB to a technology short list for consideration.
This page provides some additional guidance: Fit - ScyllaDB
---
### Page: https://forum.scylladb.com/t/scylladb-on-aws-one-large-node-or-more-smaller-ones/73
Title: ScyllaDB on AWS, One Large Node or More Smaller Ones? - ScyllaDB - ScyllaDB Community NoSQL Forum
Meta Description: I’m looking into running a ScyllaDB cluster on AWS. How do I determine sizing?
Is it better to run a few large instances or many smaller ones?
Language: en
Canonical URL: https://forum.scylladb.com/t/scylladb-on-aws-one-large-node-or-more-smaller-ones/73
## Headings Structure:
H1: ScyllaDB on AWS, One Large Node or More Smaller Ones?
H3: On-Demand Webinar: Does it still make sense to do Big Data with Small Nodes?
H3: Related topics
## Main Content:
H1: ScyllaDB on AWS, One Large Node or More Smaller Ones?
H3: On-Demand Webinar: Does it still make sense to do Big Data with Small Nodes?
H3: Related topics
I’m looking into running a ScyllaDB cluster on AWS. How do I determine sizing?
Is it better to run a few large instances or many smaller ones?
Larger nodes are easier to manage and their resources aggregate much more performance and speed. Think of applying a patch or performing a rolling restart in a hundred small nodes, also the number of failures would be higher so you’d expect more often maintenance.
A disadvantage would be: since for AWS the storage size increases proportionately the larger the instance type, more powerful instances would become an expensive option in cases when the user doesn’t have a huge dataset. But still it may be recommended depending on the throughput and latency requirements.
A benchmark would help to find the sweet spot, as well as other details that can be found in a nicely explicative Webinar about this exact subject:
In the world of Big Data, scaling out is the norm. However, many Big Data deployments are trapped in a sea of small box clusters. Join us to learn the pros and cons of large nodes, and explore why people resist using big machines.
I recommend watching it!
---
### Page: https://forum.scylladb.com/t/scylladb-cloud-or-enterprise/74
Title: ScyllaDB Cloud or Enterprise? - ScyllaDB - ScyllaDB Community NoSQL Forum
Meta Description: What are the benefits of using Scylla Cloud vs. Scylla Enterprise?
Language: en
Canonical URL: https://forum.scylladb.com/t/scylladb-cloud-or-enterprise/74
## Headings Structure:
H1: ScyllaDB Cloud or Enterprise?
H3: Related topics
## Main Content:
H1: ScyllaDB Cloud or Enterprise?
H3: Related topics
What are the benefits of using Scylla Cloud vs. Scylla Enterprise?
By choosing ScyllaDB Cloud you’ll get the latest Enterprise version of Scylla DB as a service. Our experts will take care of everything, leaving you to focus on your data and applications.
With flexible pricing and sizing choices, ScyllaDB Cloud may be used to suit your needs right away and grow alongside you in the future. Our solutions team will continuously check and manage your database to make sure it is always fully operational and operating at peak efficiency.
You can read more here
---
### Page: https://forum.scylladb.com/t/what-are-the-different-scylladb-flavors-when-should-i-use-each-one/75
Title: What are the different ScyllaDB flavors? When should I use each one? - ScyllaDB - ScyllaDB Community NoSQL Forum
Meta Description: ScyllaDB comes in three different variants:
Open Source: a free, open-source, community-supported version. This is licensed under the AGPL. It offers great performance, a low node count, reduced complexity, and a low …
Language: en
Canonical URL: https://forum.scylladb.com/t/what-are-the-different-scylladb-flavors-when-should-i-use-each-one/75
## Headings Structure:
H1: What are the different ScyllaDB flavors? When should I use each one?
H3: Related topics
## Main Content:
H1: What are the different ScyllaDB flavors? When should I use each one?
H3: Related topics
ScyllaDB comes in three different variants:
Open Source: a free, open-source, community-supported version. This is licensed under the AGPL. It offers great performance, a low node count, reduced complexity, and a low total cost of ownership. If you’re technical, know your way around source code, and like to be the first one to work with the latest features, this might be the version for you.
Enterprise: This variant is based on the ScyllaDB open-source project. It includes tested and certified production-ready binaries, software updates, and hotfixes. Additionally, it includes mission-critical technical support that guarantees that you have access to the engineers who developed ScyllaDB. ScyllaDB Enterprise has a commercial license.
This version also includes the ScyllaDB Manager, which provides centralized cluster administration and recurrent task automation, automation of periodic repair, as well as other features which are exclusive to ScyllaDB Enterprise customers.
Cloud: This is the fastest and most affordable NoSQL database as a service. You get fully managed ScyllaDB Enterprise clusters. Using ScyllaDB Cloud spares your team from database administrative tasks, giving you access to ScyllaDB clusters with automatic backup, repairs, performance optimization, security hardening, and 24/7 maintenance and support. This is offered for a fraction of the cost of other database as a service offerings, without any vendor lock-in.
---
### Page: https://forum.scylladb.com/t/difference-between-reshape-and-compaction/76
Title: Difference between reshape and compaction - ScyllaDB - ScyllaDB Community NoSQL Forum
Meta Description: What is the difference between reshape and compaction in scylla? Even after checking the related documentation, I don’t quite understand.
Is it correct for reshape to do the following?
reshape: To make compaction work…
Language: en
Canonical URL: https://forum.scylladb.com/t/difference-between-reshape-and-compaction/76
## Headings Structure:
H1: Difference between reshape and compaction
H3: ScyllaDB Open Source 4.2
H3: Related topics
## Main Content:
H1: Difference between reshape and compaction
H3: ScyllaDB Open Source 4.2
H3: Related topics
What is the difference between reshape and compaction in scylla? Even after checking the related documentation, I don’t quite understand.
ScyllaDB Open Source 4.2 provides new features: a far more efficient binary search algorithm, and improvements to our DynamoDB compatible interface.
Is it correct for reshape to do the following?
reshape: To make compaction work well per shard, reshape sstables to fit critica before bootstrap. In the case of stcs, set sstable count for 4 to hit min_threshold , and when using LCS adjusts the sstable size to 160mb.
If the node is restarted after draining it, the reshape operation occurs before node join, and compaction seems to operate after node join. Why does this work like that?
Reshape is a Rewrite of a set of SSTables to satisfy a compaction strategy’s criteria. For example, restoring data from an old backup or before the strategy update.
Reshape happens when the data is wildly out-of-shape and regular compaction would require a lot of work to get it back into shape. It is an offline operation which allows the system to devote all resources to the work.
Regular compaction happens when the data gets somewhat out of shape (e.g. a new sstable is added when a memtable is flushed). It’s an online operation that requires mild resource use.
---
### Page: https://forum.scylladb.com/t/gui-visualization-for-scylladb/77
Title: GUI Visualization for ScyllaDB - ScyllaDB - ScyllaDB Community NoSQL Forum
Meta Description: Hello, Team!
I am a Scylladb newbie. Does ScyllaDB have a GUI?
Language: en
Canonical URL: https://forum.scylladb.com/t/gui-visualization-for-scylladb/77
## Headings Structure:
H1: GUI Visualization for ScyllaDB
H3: Related topics
## Main Content:
H1: GUI Visualization for ScyllaDB
H3: Related topics
Hello, Team!
I am a Scylladb newbie. Does ScyllaDB have a GUI?
You can manage ScyllaDB with the ScyllaDB Manager.
Also have a look at the Monitoring Stack.
If you are talking about the tools such as pgAdmin (for PostgreSQL) or SQL Server Management Studio, then this may help:
I am using JetBrains Rider for my .NET C# code development on my Linux desktop, and I was surprised to find out that that IDE’s support for databases is almost on par with dedicated and paid-for tools, while it’s only a plugin to Rider.
I am working with DDL and DML code for my ScyllaDB using JetBrains Rider.
It’s a paid-for tool, but I already have it anyway, so I might as well use it for the DB work.
I also tried their database tool (called “DataGrip”), but I did not find any functionality in it which would be useful for my needs and that is not already in the Rider plugin.
DBeaver now supports ScyllaDB in the Enterprise version.
More than just a GUI, it’s a universal database manager for SQL and NoSQL databases.
You can read more about it in this blog post. It provides step-by-step instructions on how to connect DBeaver to ScyllaDB.
Once connected, users can run CQL queries and view result tables directly within DBeaver. If you try this out, please share your experience here.
---
### Page: https://forum.scylladb.com/t/running-repair-after-changing-the-replication-factor/78
Title: Running Repair after changing the Replication Factor - ScyllaDB - ScyllaDB Community NoSQL Forum
Meta Description: After I change the replication settings of a keyspace, should I run a repair on all nodes? Or will Scylla rescale the whole structure on its own?
Language: en
Canonical URL: https://forum.scylladb.com/t/running-repair-after-changing-the-replication-factor/78
## Headings Structure:
H1: Running Repair after changing the Replication Factor
H3: Related topics
## Main Content:
H1: Running Repair after changing the Replication Factor
H3: Related topics
After I change the replication settings of a keyspace, should I run a repair on all nodes? Or will Scylla rescale the whole structure on its own?
ScyllaDB doesn’t automatically stream data to new replicas after changing the RF. You have to run a repair for that to happen.
---
### Page: https://forum.scylladb.com/t/kubernetes-memory-limit-for-scylladb-pod/79
Title: Kubernetes memory limit for ScyllaDB pod - ScyllaDB - ScyllaDB Community NoSQL Forum
Meta Description: I set up Kubernetes for a limit of 64GB RAM for a Scylla pod. What if Scylla hits 64GB? Will my system crash? Or will it auto-free memory?
Also, if I increase the memory limit to 96GB, will this improve performance?
Language: en
Canonical URL: https://forum.scylladb.com/t/kubernetes-memory-limit-for-scylladb-pod/79
## Headings Structure:
H1: Kubernetes memory limit for ScyllaDB pod
H3: Related topics
## Main Content:
H1: Kubernetes memory limit for ScyllaDB pod
H3: Related topics
I set up Kubernetes for a limit of 64GB RAM for a Scylla pod. What if Scylla hits 64GB? Will my system crash? Or will it auto-free memory?
Also, if I increase the memory limit to 96GB, will this improve performance?
ScyllaDB grabs all available memory and manages it internally.
It is completely normal to have ScyllaDB at the memory limit. This is by design.
Regarding increasing the memory limit, of course, the more memory, the better. The excess memory – that is not strictly needed for normal operations – is used to cache data from the disk, greatly increasing access speed for cached rows. More memory also allows for supporting more storage (ScyllaDB aims for a 1:100 memory:disk ratio).
---
### Page: https://forum.scylladb.com/t/running-a-large-number-of-keyspaces-in-a-cluster/80
Title: Running a large number of keyspaces in a cluster - ScyllaDB - ScyllaDB Community NoSQL Forum
Meta Description: What’s the operational experience like with large numbers of keyspaces in a cluster? Are there any functional or operational limits to be aware of?
Language: en
Canonical URL: https://forum.scylladb.com/t/running-a-large-number-of-keyspaces-in-a-cluster/80
## Headings Structure:
H1: Running a large number of keyspaces in a cluster
H3: Related topics
## Main Content:
H1: Running a large number of keyspaces in a cluster
H3: Related topics
What’s the operational experience like with large numbers of keyspaces in a cluster? Are there any functional or operational limits to be aware of?
It’s recommended to limit the number of keyspaces/tables in a single cluster to a few hundreds.
Using more can be tricky in some cases as Scylla needs many files, and Memtable flush is per table. Additionally, the amount of RAM available and how homogenous those keyspaces are homogenous also have an impact.
---
### Page: https://forum.scylladb.com/t/calculating-per-record-memory-in-a-cluster/81
Title: Calculating "per record" memory in a cluster - ScyllaDB - ScyllaDB Community NoSQL Forum
Meta Description: I have been tasked to come up with “some number” of how much memory it takes to store one record in Scylla (and related overhead if any). We had commands in Redis Cache that produced these metrics. Is that even possible …
Language: en
Canonical URL: https://forum.scylladb.com/t/calculating-per-record-memory-in-a-cluster/81
## Headings Structure:
H1: Calculating "per record" memory in a cluster
H3: Related topics
## Main Content:
H1: Calculating "per record" memory in a cluster
H3: Related topics
I have been tasked to come up with “some number” of how much memory it takes to store one record in Scylla (and related overhead if any). We had commands in Redis Cache that produced these metrics. Is that even possible to do in Scylla? I have scyllatop running and also the monitoring stack.
One way to accomplish this task is to create a million “average” records, execute “nodetool flush” to make sure they get flushed to disk (which will also place them in cache), then look at the cache metrics and divide the cache memory usage by the total rows in cache.
Thanks, a couple of follow-up questions. When you say “cache metrics,” do you imply ones provided by scyllatop ? I presume this is where I get cache memory usage from? Also, would that work in a clustered environment (I have 8 nodes spread over four DCs). I know cache is over cluster, but what happens if not all records are in the cache? How would I find out how many have been put into cache?
The metrics are available by scyllatop, though a nicer way to see them is with Prometheus or Grafana (see scylla-monitoring.git). The metrics contain the number of partitions in cache, rows in cache, and memory in cache, so from there you compute any statistic you want.
About the rows not in cache, well they don’t have any cache footprint.
---
### Page: https://forum.scylladb.com/t/which-cassandra-version-is-scylladb-it-compatible-with/84
Title: Which Cassandra version is ScyllaDB it compatible with? - ScyllaDB - ScyllaDB Community NoSQL Forum
Meta Description: ScyllaDB is a drop-in replacement for Apache Cassandra 3.11, but it has some additional features from Apache Cassandra 4.0.
See ScyllaDB and Apache Cassandra Compatibility for details.
Language: en
Canonical URL: https://forum.scylladb.com/t/which-cassandra-version-is-scylladb-it-compatible-with/84
## Headings Structure:
H1: Which Cassandra version is ScyllaDB it compatible with?
H3: Related topics
## Main Content:
H1: Which Cassandra version is ScyllaDB it compatible with?
H3: Related topics
ScyllaDB is a drop-in replacement for Apache Cassandra 3.11, but it has some additional features from Apache Cassandra 4.0.
See ScyllaDB and Apache Cassandra Compatibility for details.
---
### Page: https://forum.scylladb.com/t/data-model-with-a-lot-of-empty-columns-collections/88
Title: Data model with a lot of empty columns, collections - ScyllaDB - ScyllaDB Community NoSQL Forum
Meta Description: Hello, I’m trying to find a good practice for column vs. Map column. For example, if I have 200 columns in a table and I usually use only 30% of them, am I better off using a Map columns (collection) instead of having em…
Language: en
Canonical URL: https://forum.scylladb.com/t/data-model-with-a-lot-of-empty-columns-collections/88
## Headings Structure:
H1: Data model with a lot of empty columns, collections
H3: Related topics
## Main Content:
H1: Data model with a lot of empty columns, collections
H3: Related topics
Hello, I’m trying to find a good practice for column vs. Map column. For example, if I have 200 columns in a table and I usually use only 30% of them, am I better off using a Map columns (collection) instead of having empty columns? I’ve read here that storage is not affected but memory is. Any advice on this?
Yes, storage is not affected by empty columns. They are simply not stored if empty. It is similar in memory.
It used to be that we used a different container for columns storage based on the number of columns in the schema: we used a vector (very efficient lookups but empty columns also take memory) by default and switched to a set for larger column counts (less efficient lookups but empty columns don’t use memory). We now uniformly switched to a compact radix tree, which should also not use any memory for empty columns.
So overall, I think you are better off with columns. Although if you are not yet using clustering keys, you might consider refactoring your schema, so some of these maybe-empty columns are separate rows.
Thank you for your answer! We are effectively organizing our schema to migrate our data from PostgreSQL.
Can you elaborate a little more about the idea of an empty column as a separate row? Like giving me an example and why it will be better that way.
I meant that if there is a pattern of some columns being empty in certain partitions, you can maybe organize your schema such that these are separate clustering rows instead, part of the same partition. Can’t really write an example without knowing more about the schema. But having partitions with lots of columns is also completely fine.
thanks for the awesome information.
---
### Page: https://forum.scylladb.com/t/running-nodetool-upgradesstables-after-version-upgrade/89
Title: Running nodetool upgradesstables after version upgrade - ScyllaDB - ScyllaDB Community NoSQL Forum
Meta Description: After an upgrade from 4.6 to 5.0, do we need to run the nodetool upgradesstables ?
Language: en
Canonical URL: https://forum.scylladb.com/t/running-nodetool-upgradesstables-after-version-upgrade/89
## Headings Structure:
H1: Running nodetool upgradesstables after version upgrade
H3: Related topics
## Main Content:
H1: Running nodetool upgradesstables after version upgrade
H4: apache/cassandra/blob/trunk/src/java/org/apache/cassandra/io/sstable/format/big/BigFormat.java#L369
H3: Related topics
After an upgrade from 4.6 to 5.0, do we need to run the nodetool upgradesstables ?
The short answer is no.
The long answer is: even if a new ScyllaDB version introduces a new sstable version, you don’t have to manually upgrade sstables via nodetool upgradesstables ,this will happen automatically as the sstables are compacted, and in general new sstables are written in the new version.
Gradually all sstables are migrated to the new version. This might take some time, but this is fine. There is no rush to get all sstables to the same version.
In some cases one might be in a hurry to get some of the optimizations a new sstables format version has.
This is a quick summary of version content:
For example me adds host id information to sstable, I’m not sure it’s for preformece optimizations
---
### Page: https://forum.scylladb.com/t/point-in-time-recovery-in-scylladb/90
Title: Point in Time Recovery in ScyllaDB - ScyllaDB - ScyllaDB Community NoSQL Forum
Meta Description: User Question:
I’m a n00b in ScyllaDB and reading the docs I understand that there is no such thing as Point-In-Time-Recovery. So you can only restore whatever your last backup has. Am I right or did I misunderstand thi…
Language: en
Canonical URL: https://forum.scylladb.com/t/point-in-time-recovery-in-scylladb/90
## Headings Structure:
H1: Point in Time Recovery in ScyllaDB
H3: Related topics
## Main Content:
H1: Point in Time Recovery in ScyllaDB
H3: Related topics
User Question:
I’m a n00b in ScyllaDB and reading the docs I understand that there is no such thing as Point-In-Time-Recovery. So you can only restore whatever your last backup has. Am I right or did I misunderstand things ?
Answer:
Correct. It is not something to worry about, however, as even if you lose a full AZ you can still serve requests and have your data fully considering you issue queries with LOCAL_QUORUM. And - of course - you can always go multi-DC for a true DR scenario.
User Question:
Indeed, but I am thinking about the scenario where the application has a new release and there is some bug or some operator executes the wrong DDL and blows things up.
If that happened, I guess the only way to restore things up to the point of failure would be by replaying the transactions (Kafka) or by doing a PITR with an RDBMS and then dump all the data again onto ScyllaDB.
Am I wrong ?
Answer:
As you are correct in the first question, then of course you are also correct in the last one.
You can always snapshot before such changes. And for TRUNCATE/DROP table statements, Scylla will always snapshot the table in question (unless you told it not to) when these DDL statements are run
---
### Page: https://forum.scylladb.com/t/lost-connection-to-remote-peer-exception-when-querying-data-synchronously/91
Title: “Lost connection to remote peer” exception when querying data synchronously - ScyllaDB - ScyllaDB Community NoSQL Forum
Meta Description: User Question:
I met “Lost connection to remote peer” exception when query data from scylla synchronously which will break the main client process. At the time of exception happened, seems one node became down but other…
Language: en
Canonical URL: https://forum.scylladb.com/t/lost-connection-to-remote-peer-exception-when-querying-data-synchronously/91
## Headings Structure:
H1: “Lost connection to remote peer” exception when querying data synchronously
H3: Related topics
## Main Content:
H1: “Lost connection to remote peer” exception when querying data synchronously
H3: Related topics
User Question:
I met “Lost connection to remote peer” exception when query data from scylla synchronously which will break the main client process. At the time of exception happened, seems one node became down but other nodes were working well. Do I need to switch to asynchronous access? CQL deriver version is 4.14.1
Answer:
Well, it seems inflight queries threw an exception which you didn’t handle. As a result, your main thread died.
Handle it and retry the failed queries, and it should work picking up other coordinator nodes.
---
### Page: https://forum.scylladb.com/t/counter-columns-alternative-and-overcoming-limitations/94
Title: Counter columns alternative and overcoming limitations - ScyllaDB - ScyllaDB Community NoSQL Forum
Meta Description: Question:
According to the official documentation, among the restrictions on counter columns in ScyllaDB are:
The only other columns in a table with a counter column can be columns of the primary key (which cannot be u…
Language: en
Canonical URL: https://forum.scylladb.com/t/counter-columns-alternative-and-overcoming-limitations/94
## Headings Structure:
H1: Counter columns alternative and overcoming limitations
H3: Related topics
## Main Content:
H1: Counter columns alternative and overcoming limitations
H3: Related topics
According to the official documentation, among the restrictions on counter columns in ScyllaDB are:
In my table design I have hundreds of rows where each row needs to be linked to a counter column.
The table has the following structure:
Where K is the partition key, C is the cluster key, V is a value and COUNTER is a counter column.
Generally, the above table is queried as follows:
This results in about 500 rows being returned per query. As seen, each row is linked to a COUNTER column. Specifically, each combination of K1, K2, C1, C2, C3 is linked to a different COUNTER value.
How am I suppose to model this table, if the COUNTER has to be moved to an entirely different table?
If I understand correctly, the only way to do this, is to define another table → table_counter without any of the values (V):
However, I have several issues with this approach:
1) It seems extremely inelegant to break up a logically cohesive table like this
2) It means that whenever I want to execute the previous query I would essentially need to execute two queries, instead of one, just to get the counter information linked to each K1 K2 C1 C2 C3 combination
3) I would also be forced to combine the results of the two queries above into a single data structure on the client side (for it to be useful)
Is this correct? If yes, is there an alternative to COUNTER column where I could add the COUNTER column to the first table?
One approach I was thinking of is to use a regular INTEGER as the counter column. Whenever the counter column needs to be updated I can read the current counter (integer) value and increment it on the client side and then write the new value back to the database. I understand that I won’t be protected from concurrent reads/writes so that if two clients read the counter value at the same time, and they both increment/update it, one write (update) will be lost (e.g, only the last write will be preserved). However, I can live with the occasional lost write as we are not tracking anything critical where a single (or handful) of lost writes will make a major difference. I also understand that this would require a read and then a write (two operations) each time I want to update the counter column, however, it allows me to keep the counter column as part of the original table as well as reduce the querying from two tables to one. Additionally, there would be no need to combine the results of querying two tables on the client side with this design. Seems more efficient than using a counter column and much more elegant.
Is this approach a viable alternative to the COUNTER column? Are there any pitfalls I missed? Is there another approach that might work better in my example?
*The question was asked on Stack Overflow by S.O.S
Answer:
An Alternative to the “classic” DRDT-inspired (but not quite) counter column is to use Scylla’s lightweight transactions (LWT) - basically your counter becomes a normal integer column, which you can read normally if you wish, but writes use a conditional update (UPDATE … IF …). For example to modify value V and also increment the counter you can:
This pattern is known as “optimistic locking” - step 1-2 are “optimistic” in assuming that the new item they build will be able to be written, but if some other concurrent update beat us, step 1-2 will need to be repeated. This is to contrast with pessimistic locking approaches, where the client “takes a lock”, and only after knowing it is holding the lock, bothers to calculate the new value of the item (step 2).
LWT is much more powerful than counters, and you can do with it a lot more than you can do with counters. Reads can be as efficient as regular reads, but note that writes do become slower. Scylla is working on a next-generation LWT implementation based on Raft (the current implementation is based on Paxos), so you can expect improvements in LWT write performance in the future.
*The answer was provided on Stack Overflow by Nadav Har’El
Thanks for your excellent response. Just to be clear, using LWT still requires the client to submit two transactions - one to read the current value and update the value on the client’s side and a second transaction to write the new value to db. In other words, it’s not possible to combine reading and updating into a single transaction with LWT. In other words, LWT only ensures that a concurrent write is not lost but it still requires two transactions from the client. In a case where I don’t mind losing the occasional write a regular integer column without LWT would suffice. Do you concur?
Also, when would you recommend using LWT instead of custom COLUMN counter? Only when we expect few writes on the column or always? Thanks!
CQL does not currently have the syntax to increment a non-counter column - doing
UPDATE ... X = X + 1
will be an error if X is not a counter column. Both with LWT and without LWT. So you need to do a separate read - even though the underlying LWT implementation could have handle an atomic increment just fine without a read. Remember, though, that if the counter is not stand-alone and is used for optimistic locking (as I explained above) you would need that extra read anyway. And yes, if you don’t care about write isolation and missed increments, you can just use an unsafe read+write.
---
### Page: https://forum.scylladb.com/t/check-if-an-item-is-inside-a-list-or-collection/95
Title: Check if an item is inside a list or collection - ScyllaDB - ScyllaDB Community NoSQL Forum
Meta Description: Question:
I’d like to check if an item is inside a list. How can I do that?
Language: en
Canonical URL: https://forum.scylladb.com/t/check-if-an-item-is-inside-a-list-or-collection/95
## Headings Structure:
H1: Check if an item is inside a list or collection
H3: Related topics
## Main Content:
H1: Check if an item is inside a list or collection
H3: Related topics
Question:
I’d like to check if an item is inside a list. How can I do that?
Answer:
To do that, you use the CONTAINS operator.
This operator can be used on lists, sets, and maps.
In the case of maps, CONTAINS applies to the map values.
Also, see the Documentation.
---
### Page: https://forum.scylladb.com/t/is-there-an-api-for-scylla-nodetool/96
Title: Is there an API for Scylla Nodetool? - ScyllaDB - ScyllaDB Community NoSQL Forum
Meta Description: Question:
Is there an API for nodetool? Especially nodetool tablestats?
I saw there is scylladb/api/api-doc at master · scylladb/scylladb · GitHub. Is this the right place to look for APIs?
*The question was asked on …
Language: en
Canonical URL: https://forum.scylladb.com/t/is-there-an-api-for-scylla-nodetool/96
## Headings Structure:
H1: Is there an API for Scylla Nodetool?
H3: Related topics
## Main Content:
H1: Is there an API for Scylla Nodetool?
H4: scylladb/scylla-tools-java/blob/master/src/java/org/apache/cassandra/tools/nodetool/TableStats.java
H4: scylladb/scylla-tools-java/blob/master/src/java/org/apache/cassandra/tools/nodetool/stats/TableStatsHolder.java#L117
H4: scylladb/scylla-jmx/blob/master/src/main/java/org/apache/cassandra/db/ColumnFamilyStore.java
H3: Related topics
Question:
Is there an API for nodetool? Especially nodetool tablestats?
I saw there is scylladb/api/api-doc at master · scylladb/scylladb · GitHub. Is this the right place to look for APIs?
*The question was asked on Stack Overflow by SilentCanon
Answer:
The Scylla server indeed has a REST API and it’s documentation is at the URL you point to. You can find a Swagger UI when you start the Scylla server as well: https://docs.scylladb.com/operating-scylla/rest/.
But, please note that nodetool, for example, does not use the API directly. Instead, it talks to the Scylla JMX proxy, which is a Java process that implements Cassandra-compatible JMX API. You can still use the REST API directly, but you have to figure out the mapping between the JMX operations and the REST API yourself.
For something like nodetool tablestats, the first step is to check what JMX APIs nodetool uses:
The command delegates to TableStatsHolder class:
which uses the ColumnFamilyStoreMBean JMX API for querying table statistics.
You can find the implementation of the JMX API in the scylla-jmx project by finding the ColumnFamilyStore class (without the MBean suffix):
From that class, you can see that, for example, the ColumnFamilyStore.getSSTableCountPerLevel() method delegates to the column_family/sstables/per_level/
REST API URL.
*The answer was provided on Stack Overflow by Pekka Enberg
Hey, thanks for the answer but i’d like to ask more. If you know, or could point me at the direction of how to divide the actual scylla and the scylla-jmx. Meaning we wanted to get rid of cassandra on hosts just because it uses java, and it’s a huge package almost 30mb in size, while our goal was to make a tiniest distro possible. Is it possible to, let’s say, start scylla on the host and jmx nodetool services in a container on the same host, so API ports are reachable from inside the container? This way we would have no need for java in the initial distro and therefore on hosts
We have not done that, but we’ll be looking at this direction in the future - makes total sense to separate them.
Thank you for the quick response. Is there an approximate date you may look in to it, maybe in your backlog? So i would know when to get back to the though of replacing cassandra again?
It’s not urgent. Besides being large and perhaps the need to update more often, it’s not a hassle. Splitting it has consequences that we’ll have to deal with. For example, running multiple containers side-by-side.
Have heard you. Thanks anyway!
---
### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-156-2022-11-27/113
Title: Last week in scylladb.git master (issue #156; 2022-11-27) - Database Community - ScyllaDB Community NoSQL Forum
Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 41629e97de…996eac9569 range are covered.
There were 90 non-merge commits from 12 authors in that perio…
Language: en
Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-156-2022-11-27/113
## Headings Structure:
H1: Last week in scylladb.git master (issue #156; 2022-11-27)
H3: Related topics
## Main Content:
H1: Last week in scylladb.git master (issue #156; 2022-11-27)
H3: Related topics
This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 41629e97de…996eac9569 range are covered.
There were 90 non-merge commits from 12 authors in that period. Some notable commits:
A regression where CQL ignored some WHERE clause components when a multi-column restriction ((col_a, col_b) < (0, 0)) was present was fixed.
The bundled cqlsh now uses the ScyllaDB Python driver (rather than the generic Cassandra driver) and supports Scylla Cloud connection bundles.
When a query completes a page, ScyllaDB caches the query activity as an inactive read. When the client requests the next page, ScyllaDB re-activates the read and continues where it left off. A bug in this mechanism that could cause crashes has been fixed.
There is now documentation for the (mostly automatic) procedure to upgrade a cluster to use Raft, and for the manual procedure to recover in case of problems.
A bug in alternator WHERE condition for global secondary index range key, that does not appear to have any user visible impact, has been fixed.
The task manager is now aware of repair-related tasks.
See you in the next issue of last week in scylladb.git master!
---
### Page: https://forum.scylladb.com/t/release-scylladb-4-6-10/115
Title: [RELEASE] ScyllaDB 4.6.10 - Release Notes - ScyllaDB Community NoSQL Forum
Meta Description: The ScyllaDB team announces ScyllaDB Open Source 4.6.10, a bugfix release of the ScyllaDB 4.6 stable branch.
Please note the latest ScyllaDB stable release is ScyllaDB 5.0, and you are encouraged to upgrade to it.
Scyl…
Language: en
Canonical URL: https://forum.scylladb.com/t/release-scylladb-4-6-10/115
## Headings Structure:
H1: [RELEASE] ScyllaDB 4.6.10
H3: Related topics
## Main Content:
H1: [RELEASE] ScyllaDB 4.6.10
H3: Related topics
The ScyllaDB team announces ScyllaDB Open Source 4.6.10, a bugfix release of the ScyllaDB 4.6 stable branch.
Please note the latest ScyllaDB stable release is ScyllaDB 5.0, and you are encouraged to upgrade to it.
ScyllaDB Open Source 4.6.10, like all past and future 4.x.y releases, is backward compatible and supports rolling upgrades.
Issue fixed in this release:
---
### Page: https://forum.scylladb.com/t/scylladb-university-live-nosql-roundtable-avi-kivity-dor-laor-tzach-livyatan/116
Title: ScyllaDB University Live NoSQL Roundtable: Avi Kivity, Dor Laor & Tzach Livyatan - Blog Posts - ScyllaDB Community NoSQL Forum
Meta Description: [800x400-blog-scylladb-u-roundtable-fall-22]
Language: en
Canonical URL: https://forum.scylladb.com/t/scylladb-university-live-nosql-roundtable-avi-kivity-dor-laor-tzach-livyatan/116
## Headings Structure:
H1: ScyllaDB University Live NoSQL Roundtable: Avi Kivity, Dor Laor & Tzach Livyatan
H3: ScyllaDB University Live NoSQL Roundtable: Avi Kivity, Dor Laor & Tzach...
H3: Related topics
## Main Content:
H1: ScyllaDB University Live NoSQL Roundtable: Avi Kivity, Dor Laor & Tzach Livyatan
H3: ScyllaDB University Live NoSQL Roundtable: Avi Kivity, Dor Laor & Tzach...
H3: Related topics
A look at the expert NoSQL roundtables that are a part of every ScyllaDB University LIVE event – including the fall session on December 1.
---
### Page: https://forum.scylladb.com/t/scylladb-return-inconsistent-data-after-node-full-rebuild/119
Title: Scylladb return inconsistent data after node full rebuild - Database Community - ScyllaDB Community NoSQL Forum
Meta Description: Hi! I am testing node rebuild after loosing the data volume in docker (like described here Rebuild a Node After Losing the Data Volume | Scylla Docs).
My steps
Create new cluster with static ip’s(3 DC, 2 nodes in each…
Language: en
Canonical URL: https://forum.scylladb.com/t/scylladb-return-inconsistent-data-after-node-full-rebuild/119
## Headings Structure:
H1: Scylladb return inconsistent data after node full rebuild
H3: Related topics
## Main Content:
H1: Scylladb return inconsistent data after node full rebuild
H3: Related topics
Hi! I am testing node rebuild after loosing the data volume in docker (like described here Rebuild a Node After Losing the Data Volume | Scylla Docs).
Image version: scylladb/scylla:4.6.8
Hello, scylla replace operations fetches data from nodes in the same DC. In case RF = 1, it is expected the replacing node will miss some of the data after replace. You can run a cross DC repair to fix the data. In addition, I would recommend using RF more than 1 per DC.
---
### Page: https://forum.scylladb.com/t/i-wonder-how-much-cpu-share-each-workload-prioritization-type-oltp-olap-of-scylladb-has/121
Title: I wonder how much cpu share each Workload Prioritization type (oltp, olap) of scylladb has - ScyllaDB - ScyllaDB Community NoSQL Forum
Meta Description: There are two workload types in scylladb other than the default value, but it doesn’t seem to be specified in the documentation how much cpu share this option has.
interactive - workload sensitive to latency, expected…
Language: en
Canonical URL: https://forum.scylladb.com/t/i-wonder-how-much-cpu-share-each-workload-prioritization-type-oltp-olap-of-scylladb-has/121
## Headings Structure:
H1: I wonder how much cpu share each Workload Prioritization type (oltp, olap) of scylladb has
H3: Related topics
## Main Content:
H1: I wonder how much cpu share each Workload Prioritization type (oltp, olap) of scylladb has
H4: scylladb/scylladb/blob/master/docs/dev/service_levels.md
H4: scylladb/scylladb/blob/b551cd254c6daf07014267950fcdd17aaf19acf0/transport/server.cc#L590-L598
H3: Related topics
There are two workload types in scylladb other than the default value, but it doesn’t seem to be specified in the documentation how much cpu share this option has.
Even if I check the related PR, it is not specified how much cpu share the corresponding workload type takes.
May I know the relevant code?
Hi @Stewart_Han ,
It is not about cpu shares, the workload type is used by ScyllaDB to decide on the proper action during
excesive load on the system.
For batch workloads , where the concurrency is bounded, we can throttle the load by delaying
answers to the client, this in turn will make him send less requests per second for each thread.
For interactive workload, throteling doesn’t make much sense since the requests are not queued on
the client side so the only resort is to fail early and signal to the client that we are overloaded (OverloadedException), this behaviour is a little bit speculative since we need to predict in advance our ability to serve a specific requests withing the timeout limits.
Here is a pointer to the code where we drop requests for interactive workload if we predict that we will
not be able to meet the timeout:
If the workload is not interactive scylla will just continue normally, assuming that the workload will converge to the maximal throughput possible.
thanks for your reply!
---
### Page: https://forum.scylladb.com/t/meet-scylladb-s-new-vp-of-r-d-yaniv-kaul/123
Title: Meet ScyllaDB’s New VP of R&D, Yaniv Kaul - Blog Posts - ScyllaDB Community NoSQL Forum
Meta Description: [800x400-blog-intro-yaniv-kaul (1)]
Language: en
Canonical URL: https://forum.scylladb.com/t/meet-scylladb-s-new-vp-of-r-d-yaniv-kaul/123
## Headings Structure:
H1: Meet ScyllaDB’s New VP of R&D, Yaniv Kaul
H3: Meet ScyllaDB’s New VP of R&D, Yaniv Kaul
H3: Related topics
## Main Content:
H1: Meet ScyllaDB’s New VP of R&D, Yaniv Kaul
H3: Meet ScyllaDB’s New VP of R&D, Yaniv Kaul
H3: Related topics
Get to know Yaniv Kaul, who just joined ScyllaDB as VP of Research and Development, in this quick Q & A. You'll hear about his fascination with tackling complex challenges: from distributed system engineering to photographing bees.
Does anyone here have questions for Yaniv?
---
### Page: https://forum.scylladb.com/t/release-scylla-5-1-0-part-1/126
Title: [RELEASE] Scylla 5.1.0 - part 1 - Release Notes - ScyllaDB Community NoSQL Forum
Meta Description: The ScyllaDB team is pleased to announce the release of ScyllaDB Open Source 5.1, a production-ready release of our open-source NoSQL database.
ScyllaDB 5.1 introduces Partition level rate limit, Distributed select coun…
Language: en
Canonical URL: https://forum.scylladb.com/t/release-scylla-5-1-0-part-1/126
## Headings Structure:
H1: [RELEASE] Scylla 5.1.0 - part 1
H2: New Features
H3: Distributed SELECT COUNT(*)
H3: Limit partition access rate
H3: Load and stream
H3: Materialized Views: Prune
H3: Alternator updates
H3: Materialized Views: Synchronous Mode - Experimental
H3: Performance: Eliminate exceptions from the read and write path
H3: Raft Updates
H3: Web Assembly (WASM) based UDA/UDF - Experimental
H1: [RELEASE] Scylla 5.1.0 - part 2
H2: Updates in this Release
H3: Deployment and Packaging
H3: CQL API updates
H3: Stability and Performance Improvements
H3: Tooling
H3: Storage
H3: Configuration
H3: Monitoring and Tracing
H3: Bug Fixes
H3: Related topics
## Main Content:
H1: [RELEASE] Scylla 5.1.0 - part 1
H2: New Features
H3: Distributed SELECT COUNT(*)
H3: Limit partition access rate
H3: Load and stream
H3: Materialized Views: Prune
H3: Alternator updates
H3: Materialized Views: Synchronous Mode - Experimental
H3: Performance: Eliminate exceptions from the read and write path
H3: Raft Updates
H3: Web Assembly (WASM) based UDA/UDF - Experimental
H1: [RELEASE] Scylla 5.1.0 - part 2
H2: Updates in this Release
H3: Deployment and Packaging
H3: CQL API updates
H3: Stability and Performance Improvements
H3: Tooling
H3: Storage
H3: Configuration
H3: Monitoring and Tracing
H3: Bug Fixes
H3: Related topics
The ScyllaDB team is pleased to announce the release of ScyllaDB Open Source 5.1, a production-ready release of our open-source NoSQL database.
ScyllaDB 5.1 introduces Partition level rate limit, Distributed select count, and more functional, performance and stability improvements.
Only the last two minor releases of the ScyllaDB Open Source project are supported. As ScyllaDB Open Source 5.1 is officially released, only ScyllaDB Open Source 5.1 and ScyllaDB 5.0 will be supported; ScyllaDB 4.6 will be retired. Users are encouraged to upgrade to the latest and greatest release. Upgrade instructions are here.
Get ScyllaDB Open Source 5.1 as binary packages (RPM/DEB), AWS AMI, GCP Image and Docker image
Upgrade from ScyllaDB 5.0 to ScyllaDB 5.1
ScyllaDB will now automatically run SELECT COUNT(*) statements on all nodes and all shards in parallel, which brings a considerable speedup, even 100X in larger clusters.
This feature is limited to queries that do not use GROUP BY or filtering.
The implementation includes a new level of coordination.
A Super-Coordinator node splits aggregation queries into sub-queries, distributes them across some group of coordinators, and merges results.
Like a regular coordinator, the Super-Coordinator is a per operation function.
A 3 node cluster setup on powerful desktops (3x32 vCPU)
Filled the cluster with ~2 * 10^8 rows using scylla-bench and run:
time cqlsh --request-timeout=3600 -e “select count(*) from scylla_bench.test using timeout 1h;”
Before Distributed Select: 68s
After Distributed Select: 2s
You can disable this feature by setting enable_parallelized_aggregation config parameter to false.
It is now possible to limit read rates and writes rates into a partition with a new WITH per_partition_rate_limit clause for the CREATE TABLE and ALTER TABLE statements. This is useful to prevent hot-partition problems when high rate reads or writes are bogus (for example, arriving from spam bots). #4703
Limits are configured separately for reads and writes. Some examples:
ALTER TABLE t WITH per_partition_rate_limit = {
‘max_reads_per_second’: 100,
‘max_writes_per_second’: 200
Limit reads only, no limit for writes:
ALTER TABLE t WITH per_partition_rate_limit = {
‘max_reads_per_second’: 200
This feature extends nodetool refresh to allow loading arbitrary sstables that do not belong to a particular node into the cluster. It loads the sstables from disk, calculates the data’s owning nodes, and automatically streams the data to the owning nodes. In particular this is useful when restoring a cluster from backup.
For example, say the old cluster has 6 nodes and the new cluster has 3 nodes.
One can copy the sstables from the old cluster to the new nodes and trigger the load and stream process.
This can make restores and migrations much easier:
Load_and_stream option also updates the relevant Materialized Views #9205
curl -X POST "http://{ip}:10000/storage_service/sstables/{keyspace}?cf={table}&load_and_stream=true
Note there is an open bug, #282 , for the Nodetool refresh --load-and-stream operation. Until it is fixed, use the REST API above.
A new CQL extension PRUNE MATERIALIZED VIEW statement can now be used to remove inconsistent rows from materialized views. A special statement is dedicated for pruning ghost rows from materialized views.
A ghost row is an inconsistency issue which manifests itself by having rows in a materialized view which do not correspond to any base table rows. Such inconsistencies should be prevented altogether and ScyllaDB strives to avoid them, but if they happen, this statement can be used to restore a materialized view
to a fully consistent state without rebuilding it from scratch.
PRUNE MATERIALIZED VIEW my_view;
PRUNE MATERIALIZED VIEW my_view WHERE token(v) > 7 AND token(v) < 1535250;
PRUNE MATERIALIZED VIEW my_view WHERE v = 19;
Alternator, ScyllaDB’s implementation of the DynamoDB API, now provides the following improvements:
There is now a synchronous mode for materialized views. In ordinary, asynchronous mode materialized views operations return before the view is updated. In synchronous mode materialized views operations do not return until the view is updated. This enhances consistency but reduces availability as in some situations all nodes might be required to be functional.
CREATE MATERIALIZED VIEW main.mv
AS SELECT * FROM main.t
WITH synchronous_updates = true;
ALTER MATERIALIZED VIEW main.mv WITH synchronous_updates = true;
When a coordinator times out, it generates an exception which is then caught in a higher layer and converted to a protocol message. Since exceptions are slow, this can make a node that experiences timeouts become even slower. To prevent that, the coordinator write path and read path has been converted not to use exceptions for timeout cases, treating them as another kind of result value instead. Further work on the read path and on the replica reduces the cost of timeouts, so that goodput is preserved while a node is overloaded.
Improvement results below
While strong consistent schema management remains experimental in 5.1, the work on the Raft consensus algorithm continues toward more use cases, such as safe topology updates, improving traceability and stability. Here are selected updates in this release:
ScyllaDB 5.1 brings experimental support for Wasm-based User Defined Functions (UDFs) and User Defined Aggregates (UDAs).
The CQL syntax is compatible with Apache Cassandra. Examples:
CREATE FUNCTION sample ( arg int ) …;
CREATE FUNCTION sample ( arg text ) …;
A full example of using Rust to create a UDF will be shared soon.
To enable WASM UDF in ScyllaDB 5.1:
–enable-user-defined-functions true --experimental-features udf
experimental_features:
Issues fixed in this release:
More update in Part 2 below!
A list of CQL bug fix and extensions:
The LIKE operator on descending order clustering keys now works. #10183
ScyllaDB would incorrectly use an index with some IN queries, leading to incorrect results. This is now fixed.
In CREATE AGGREGATE statements, the INITCOND and FINALFUNC clauses are now optional (defaulting to NULL and the identity function respectively).
CREATE KEYSPACE now has a WITH STORAGE clause, allowing to customize where data is stored. For now, this is only a placeholder for future extensions.
When talking to drivers using the older v3 protocol, ScyllaDB did not serialize timeout exceptions correctly, resulting in the driver complaining about protocol violations. This is now fixed. #5610
The CQL grammar was relaxed to allow bind markers in collection literals, e.g. UPDATE tab SET my_set = { ?, ‘foobar’, :variable }.
ScyllaDB now validates collections for NULLs more carefully. #10580
After this change, the following query
INSERT INTO ks.t (list_column) VALUES (?);
And the driver sending a list with null inside as the bound value, something like [1, 2, null, 4] Would result in an invalid_request_exception instead of an ugly marshaling error.
The tool can be used to list the different API functions and their parameters, and to print detailed help for each function.
Then, when invoking any function, scylla-api-cli performs basic validation on the function arguments and prints the result to the standard output. Note that json results msy be pretty-printed using commonly available command line utilities. It is recommended to use scylla-api-cli for interactive usage of the REST API over plain http tools, like curl, to prevent human errors.
It is now possible to limit, and control in real time, the bandwidth of streaming and compaction.
These and more configuration updates below:
Below are a list of monitoring and tracing related work in this release:
For a full list of fixed issues see git log and 5.1 release candidates notes.
---
### Page: https://forum.scylladb.com/t/release-scylladb-5-0-5/128
Title: [RELEASE] ScyllaDB 5.0.5 - Release Notes - ScyllaDB Community NoSQL Forum
Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.0.5, a bugfix release of the ScyllaDB 5.0 stable branch. ScyllaDB Open Source 5.0.5, like all past and future 5.x.y releases, is backward compatible and supports rolling…
Language: en
Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-0-5/128
## Headings Structure:
H1: [RELEASE] ScyllaDB 5.0.5
H3: Related topics
## Main Content:
H1: [RELEASE] ScyllaDB 5.0.5
H3: Related topics
The ScyllaDB team announces ScyllaDB Open Source 5.0.5, a bugfix release of the ScyllaDB 5.0 stable branch. ScyllaDB Open Source 5.0.5, like all past and future 5.x.y releases, is backward compatible and supports rolling upgrades.
Note that the latest stable branch of ScyllaDB is 5.1, and you are encouraged to upgrade.
Issues fixed in this release:
ScyllaDB 5.0.5 was acutally released at Oct, resent by mistake.
---
### Page: https://forum.scylladb.com/t/what-table-options-should-be-used-for-fast-writes-reads-and-no-deletes/129
Title: What table options should be used for fast writes & reads and no deletes - Database Community - ScyllaDB Community NoSQL Forum
Meta Description: ( New user to Scylladb … excuse my naivety )
I am going to use ScyllaDB cluster ( single DC ) for key value pair , key will be the queue-id and value will be the job data . Key size 32 bytes , values varying avg 50kb -…
Language: en
Canonical URL: https://forum.scylladb.com/t/what-table-options-should-be-used-for-fast-writes-reads-and-no-deletes/129
## Headings Structure:
H1: What table options should be used for fast writes & reads and no deletes
H3: Related topics
## Main Content:
H1: What table options should be used for fast writes & reads and no deletes
H3: Related topics
( New user to Scylladb … excuse my naivety )
I am going to use ScyllaDB cluster ( single DC ) for key value pair , key will be the queue-id and value will be the job data . Key size 32 bytes , values varying avg 50kb - max 10 Mb
There will by 800 million writes every day at peaks of 4 k writes per second ( which might grow)
No batch inserts all single . No updates . Reads will happen exactly once per record
To avoid any tombstones I will use 1 table per day and drop entire tables after 2 days
There are my doubts
What table type should I used . I think Caching enabled=true ?
What is the ideal concurrency I should design my writers and readers ?
If there are no updates or deletes can I use this to speed up my read/writes by setting up some config ?
You can use default table,
however I suggest to use TimeWindow Compaction Strategy instead of dropping and creating the table with TTL of X days (make sure you properly then set the window unit in TWCS, don’t create within your TTL more than 12-13 windows please)
then tombstones won’t be an issue, Scylla will effectively throw them away (thanks to TTL and windows)
However a warning here is - you won’t overwrite old data here (outside of current window), if you will, then this effective removing of old data will get broken.
(and in such case default ICS with TTL should also work with either SAG or periodic major compaction)
I don’t know what will be your latency SLAs, but 50k-10M payloads are huuuge range, where the rows around 50k will be quite fast, but processing of 10M payload might bottleneck the cpu(shard).
Limits we suggest to keep are here: scylladb/config.cc at master · scylladb/scylladb · GitHub , so for you if we assume single row partition, then it’s about cell size, which is 1 MB, you will have 10MB, so I’d expect not single digit ms latencies, but worst case 10x more
Above largely depends on how many reads per second will you do and how many cpus will be there (and how good will be your PK distribution).
Concurrency depends on distribution and how many cpus you will have and how many reads/writes (with ideally percentiles for that data, since that payload size range of yours is big). Some guidance is in Sizing Up Your ScyllaDB Cluster - ScyllaDB , or check sizing calculators (take them as guidance, not as a rule of thumb, cassandra-stress is your best friend here to see how much 1 cpu will be able to handle with your RF). You can also read Maximizing Performance via Concurrency While Minimizing Timeouts in Distributed Databases - ScyllaDB .
If there are no updates and deletes, then this is perfect TTL + TWCS situation assuming you really want to drop all your data older than 2 days. And that is basically your best tuning - Compaction | ScyllaDB Docs
(but do check other strategies, too)
---
### Page: https://forum.scylladb.com/t/release-scylladb-5-0-6/133
Title: [RELEASE] ScyllaDB 5.0.6 - Release Notes - ScyllaDB Community NoSQL Forum
Meta Description: The ScyllaDB team announces ScyllaDB Open Source 5.0.6, a bugfix release of the ScyllaDB 5.0 stable branch. ScyllaDB Open Source 5.06, like all past and future 5.x.y releases, is backward compatible and supports rolling …
Language: en
Canonical URL: https://forum.scylladb.com/t/release-scylladb-5-0-6/133
## Headings Structure:
H1: [RELEASE] ScyllaDB 5.0.6
H3: Related topics
## Main Content:
H1: [RELEASE] ScyllaDB 5.0.6
H3: Related topics
The ScyllaDB team announces ScyllaDB Open Source 5.0.6, a bugfix release of the ScyllaDB 5.0 stable branch. ScyllaDB Open Source 5.06, like all past and future 5.x.y releases, is backward compatible and supports rolling upgrades.
Note that the latest stable branch of ScyllaDB is 5.1, and you are encouraged to upgrade.
Issues fixed in this release:
---
### Page: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-157-2022-12-04/134
Title: Last week in scylladb.git master (issue #157; 2022-12-04) - Database Community - ScyllaDB Community NoSQL Forum
Meta Description: This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 996eac9569…b551cd254c range are covered.
There were 149 non-merge commits from 19 authors in that peri…
Language: en
Canonical URL: https://forum.scylladb.com/t/last-week-in-scylladb-git-master-issue-157-2022-12-04/134
## Headings Structure:
H1: Last week in scylladb.git master (issue #157; 2022-12-04)
H3: Related topics
## Main Content:
H1: Last week in scylladb.git master (issue #157; 2022-12-04)
H3: Related topics
This short report brings to light some interesting commits to scylladb.git master from the last week. Commits in the 996eac9569…b551cd254c range are covered.
There were 149 non-merge commits from 19 authors in that period. Some notable commits:
A crash was fixed during an illegal lightweight transaction INSERT with NULL clustering key
Alternator, ScyllaDB’s implementation of the DynamoDB API, now considers the time-to-live implementation stable and no longer experimental.
Transient errors in Alternator TTL scanning (where alternator looks for expired items) no longer cause Alternator to abort the scan; instead it continues.
ScyllaDB caches rows and (since 4.6) index entries in a single unified cache. It was observed that in some small-partition workloads index caching causes a performance regression, so index caching is now disabled by default. It can still be enabled for workloads that benefit from it. We plan to re-enable it when the regression is fixed.
The bundled cqlsh now considers system_distributed_everywhere a system keyspace.
Last week’s change to the scylla driver in cqlsh was reverted, as it causes regressions around the USE statement.
Evaluation of Boolean binary operators (e.g. “=”) has been refactored to use the same expression evaluation code as other expressions, paving the way for relaxation of the CQL grammar to be more similar to SQL.
WebAssembly (WASM) has been re-enabled for aarch64 (ARM).
The topology management code is more relaxed about unknown endpoints to prevent crashes in tests that check for edge cases. This fixes a recent regression.
Hinted handoff now checks that a node exists in topology before doing anything; this helps with a recent regression due to topology refactoring.
In SELECT JSON statements, column names can be given aliases (just as with traditional SELECT). However, ScyllaDB ignored those aliases. It will now honor them.
The container (docker) image now uses the C locale to reduce image size.
The Raft protocol implementation now supports changing a node’s IP address without changing its identity.
Raft failure detection now uses a separate RPC verb from gossip failure detection.
The task manager now controls repair tasks with shard granularity.
A crash in the task manager related to repair tasks has been fixed.
A crash while fetching repaid ids from the repair history table was fixed.
See you in the next issue of last week in scylladb.git master!
---
### Page: https://forum.scylladb.com/t/io-uring-what-is-it/136
Title: Io_uring: what is it? - ScyllaDB - ScyllaDB Community NoSQL Forum
Meta Description: I’ve recently been seeing io_uring come up in locations. Understanding it has something to do with input and output, what exactly is it?
*The question was originally asked on Stack Overflow by Amani
Language: en
Canonical URL: https://forum.scylladb.com/t/io-uring-what-is-it/136
## Headings Structure:
H1: Io_uring: what is it?
H3: Related topics
## Main Content:
H1: Io_uring: what is it?
H3: Related topics
I’ve recently been seeing io_uring come up in locations. Understanding it has something to do with input and output, what exactly is it?
*The question was originally asked on Stack Overflow by Amani
io_uring is a (new as of mid-2019) Linux kernel interface to efficiently allows you to send and receive data asynchronously. It was originally designed to target block devices and files but has since gained the ability to work with things like network sockets.
Unlike something like epoll(), it is built around a completion model rather than a readiness model. This is desirable because other operating systems have used the completion model successfully for some time. io_uring provides something competitive and complete for Linux without the drawbacks the previous Linux AIO interface has.
The author of io_uring has written a PDF document titled Efficient IO with io_uring, which technically discusses its usage. A gentler introduction is provided by the Lord of the io_uring guide. You can read ScyllaDB developer Glauber Costa proselytize it in How io_uring and eBPF Will Revolutionize Programming in Linux. Lastly, LWN.net has written about io_uring many times.
*The answer was originally provided on Stack Overflow by Anon
---
### Page: https://forum.scylladb.com/t/what-are-the-differences-between-column-families-in-cassandras-data-model-compared-to-bigtable/137
Title: What are the differences between column families in Cassandra's data model compared to Bigtable? - ScyllaDB - ScyllaDB Community NoSQL Forum
Meta Description: I am learning about Cassandra’s data model and its relation to Bigtable but have some things I still don’t understand regarding the Column Family concept.
Is the column-family based data model of Cassandra the same as t…
Language: en
Canonical URL: https://forum.scylladb.com/t/what-are-the-differences-between-column-families-in-cassandras-data-model-compared-to-bigtable/137
## Headings Structure:
H1: What are the differences between column families in Cassandra's data model compared to Bigtable?
H3: Related topics
## Main Content:
H1: What are the differences between column families in Cassandra's data model compared to Bigtable?
H3: Related topics
I am learning about Cassandra’s data model and its relation to Bigtable but have some things I still don’t understand regarding the Column Family concept.
Is the column-family based data model of Cassandra the same as the column-family based data model of Google’s BigTable?
Firstly I’ve read the Bigtable paper, including the part about its data model, that is, how data is stored. As far as I understood, each table in Bigtable relies on a multi-dimensional sparse map with the dimensions row, column, and time. The map is sorted by rows. Columns can be grouped with the name convention family:qualifier to a column family. Therefore, a single row can contain multiple-column families.
Although it is stated that Cassandra relies on the Bigtable data model, I read multiple times that in Cassandra, a column family contains multiple rows and is, to some extent, comparable to a table in relational data stores. Isn’t this contrary to Bigtable’s approach, where a row could contain multiple column families? What comes first, the column family or row? Are these concepts even comparable?
*Based on a question originally asked on Stack Overflow by OxideNt
When Cassandra started, its data model was indeed based on BigTable’s. A row of data could include any number of columns, each of these columns has a name and a value. A row could have a thousand different columns, and a different row could have a thousand other columns - rows do not have to have the same columns. Such a database is called “schema-less”, because there is no schema that each row needs to adhere to.
But Toto, we’re not in Kansas anymore - and Cassandra’s model changed in focus (though not in essence) since, and I’ll try to explain how and why:
As Cassandra matured, its developers started to realize that schema-less isn’t as great as they once thought it was. Schemas are valuable in ensuring application correctness. Moreover, one doesn’t normally get to 1000 columns in a single row just because there are 1000 individually-named fields in one record. Rather, the more common case is that the record actually contains 200 entries, each with 5 fields. The schema should fix these 5 fields that every one of these entries should have, and what defines each of these separate entries is called a “clustering key”. So around the time of Cassandra 0.8, these ideas were introduced to Cassandra as the “CQL” (Cassandra Query Language).
For example, in CQL, one declares that a column-family (which was dutifully renamed “table”) has a schema, with a known list of fields:
This schema says that each wide row in the table (now, in modern Cassandra, this was renamed a “partition”) with the key “groupname” is a possibly long list of users, each with username, email, and age fields. The first name in the “PRIMARY KEY” specifier is the partition key (it determines the key of the wide rows), and the second is called the clustering key (it determines the key of the small rows that together make up the wide rows).
Despite the new CQL dressup, Cassandra continued to implement these new concepts using the good-old-BigTable-wide-row-without-schema implementation. For example, consider that our data has a group “mygroup” with two people, (john, john@somewhere.com, 27) and (joe, joe@somewhere.com, 38). Cassandra adds the following four column names->values to the wide row:
Note how we ended up with a wide row with 4 columns - 2 non-key fields per row (email and age), multiplied by the number of rows in the partition (2). The clustering key field “username” no longer appears anywhere as the value, but rather as part of the column’s name! So If we have two username values “john” and “joe”, We have some columns prefixed “john” and some columns prefixed “joe”, and when we read the column “joe:email” we know this is the value of the email field of the row which has username=joe.
Cassandra still has this internal duality - converting the user-facing CQL rows and clustering keys into old-style wide rows. Previously, Cassandra’s on-disk format known as “SSTables” was still schema-less and used composite names as shown above for column names. I wrote a detailed description of the SSTable format on Scylla’s site SSTables Data File · scylladb/scylladb Wiki · GitHub (Scylla is a more efficient C++ re-implementation of Cassandra to which I contribute). However, column names are very inefficient in this format so Cassandra, in version 3.0, switched to a different file format, which for the first time, accepts clustering keys and schema-full rows as first-class citizens. This was the last nail in the coffin of the schema-less Cassandra from 13 years ago. Cassandra is now schema-full, all the way.
*Based on an answer originally on Stack Overflow by Nadav Har’El
---
### Page: https://forum.scylladb.com/t/error-message-key-cartesian-product-size-is-greater-than-maximum/138
Title: Error Message: key cartesian product size {} is greater than maximum {} - ScyllaDB - ScyllaDB Community NoSQL Forum
Meta Description: I’m seeing this error while reading data from ScyllaDB RequestHandler: ip:9042 replied with server error (clustering key cartesian product size 600 is greater than maximum 100), defuncting connection.
Any idea?
*Based …
Language: en
Canonical URL: https://forum.scylladb.com/t/error-message-key-cartesian-product-size-is-greater-than-maximum/138
## Headings Structure:
H1: Error Message: key cartesian product size {} is greater than maximum {}
H3: Related topics
## Main Content:
H1: Error Message: key cartesian product size {} is greater than maximum {}
H3: Related topics
I’m seeing this error while reading data from ScyllaDB RequestHandler: ip:9042 replied with server error (clustering key cartesian product size 600 is greater than maximum 100), defuncting connection.
Any idea?
*Based on a question originally asked on Stack Overflow by rohan-vadje
This error is returned to prevent too large restriction sets from being generated, which may put a strain on your server. If you’re aware of the risks and know a reasonable upper bound of the number of restrictions for your queries, you can manually change the maximum in scylla.yaml, e.g. max_clustering_key_restrictions_per_query: 650. Note, however, that this option has a warning in its description, and it should be acknowledged:
In particular, setting this flag above a couple of hundred is risky - 600 should be alright, but at this point, you could also consider rephrasing your query so that they have less values in their IN restrictions - perhaps splitting some queries into multiple smaller ones?
Source from Scylla tracker: https://github.com/scylladb/scylla/pull/4797
*Based on an answer originally on Stack Overflow by Piotr Sarna
---
### Page: https://forum.scylladb.com/t/scylladbs-read-path-vs-cassandras-and-performance-when-using-hdd-vs-ssd/139
Title: ScyllaDB's read path vs. Cassandra's and performance when using HDD vs SSD - ScyllaDB - ScyllaDB Community NoSQL Forum
Meta Description: What is the difference between ScyllaDB’s read path and Cassandra’s read path? When I stress test Cassandra and ScyllaDB, ScyllaDB’s read performance is poorer by five times compared to Cassandra’s using 16 cores and a n…
Language: en
Canonical URL: https://forum.scylladb.com/t/scylladbs-read-path-vs-cassandras-and-performance-when-using-hdd-vs-ssd/139
## Headings Structure:
H1: ScyllaDB's read path vs. Cassandra's and performance when using HDD vs SSD
H3: Related topics
## Main Content:
H1: ScyllaDB's read path vs. Cassandra's and performance when using HDD vs SSD
H3: Related topics
What is the difference between ScyllaDB’s read path and Cassandra’s read path? When I stress test Cassandra and ScyllaDB, ScyllaDB’s read performance is poorer by five times compared to Cassandra’s using 16 cores and a normal HDD.
I expect better read performance on ScyllaDB compared to Cassandra when using a normal HDD.
Can someone please confirm, if it’s possible to achieve better read performance using a normal HDD?
If so, what changes are required to Scylla’s config? Please guide me!
*Based on a question originally asked on Stack Overflow by sateesh
There can be various reasons why you are not getting the most out of your Scylla Cluster.
I strongly recommend reading the following docs that will provide you with more insights:
*Based on an answer on Stack Overflow by TomerSan
Both Cassandra and ScyllaDB utilize the same disk storage architecture (LSM). That means that they have relatively the same disk access patterns because the algorithms are largely the same. The LSM trees were built with the idea in mind that it is not necessary to do instant in-place updates. It consists of immutable data buckets that are large continuous pieces of data on disk. That means less random IO, and more sequential IO for which the HDD works great (not counting utilized parallelism by modern database implementations).
All the above means that the difference that you see is not induced by the difference in how those databases use a disk. It must be related to the configuration differences and what happens underneath. Maybe ScyllaDB tries to utilize more parallelism or more aggressively do compaction. It depends.
To be able to say anything specific, please share your tests, envs, and configurations.
*Based on an answer on Stack Overflow by Ivan Prisyazhnyy
Both databases use an LSM tree, but Scylla has thread-per-core architecture on top plus we use O_Direct while C* uses the page cache. Scylla also has a sophisticated IO scheduler that makes sure not to overload the disk, and thus scylla_setup runs a benchmark automatically to tune. Check your output of it in io.conf.
There are far more things to review, an more information from you is required to look into it. In general, Scylla should perform better in this case as well, but your disk is likely to be the bottleneck in both cases.
*Based on an answer on Stack Overflow by dor laor
Some other responses focused on write performance, but this isn’t what you asked about - you asked about reads.
Uncached read performance on HDDs is bound to be poor in Cassandra and Scylla, because reads from disk each require several seeks on the HDD, and even the best HDD cannot do more than, say, 200 of those seeks per second. Even with a RAID of several of these disks, you will rarely be able to do more than, say, 1000 requests per second. Since a modern multi-core can do orders of magnitude more CPU work than 1000 requests per second, in both Scylla and Cassandra cases, you’ll likely see free CPU. So Scylla’s main benefit of using much less CPU per request will not even matter when the disk is the performance bottleneck. In such cases, I would expect Scylla’s and Cassandra’s performance (I am assuming that you’re measuring throughput when you talk about performance?) should be roughly the same.
If still, you’re seeing better throughput from Cassandra than Scylla, several details may explain why, beyond the general client misconfiguration issues raised in other responses:
I don’t know which of these differences - or something else - is causing the performance of your use-case to be lower in Scylla, but please keep in mind that whatever you fix, your performance is always going to be bad with HDDs. With SDDs, we’ve measured in the past more than a million random-access read requests per second on a single node. HDDs cannot come to anything close. If you really need optimum performance or performance per dollar, SDDs are the way to go.
*Based on an answer on Stack Overflow by Nadav Har’El
---
### Page: https://forum.scylladb.com/t/column-design-and-udt-recommendations-for-a-specific-problem/141
Title: Column design and UDT recommendations for a specific problem - ScyllaDB - ScyllaDB Community NoSQL Forum
Meta Description: Hi, I’m learning about Cassandra/Scylla and I have a question. Lets say I’m storing books and each book has N fields, like title, alt title, synopsis, etc… Each of this fields is a string and has an associated language (…
Language: en
Canonical URL: https://forum.scylladb.com/t/column-design-and-udt-recommendations-for-a-specific-problem/141
## Headings Structure:
H1: Column design and UDT recommendations for a specific problem
H3: Related topics
## Main Content:
H1: Column design and UDT recommendations for a specific problem
H3: Related topics
Hi, I’m learning about Cassandra/Scylla and I have a question. Lets say I’m storing books and each book has N fields, like title, alt title, synopsis, etc… Each of this fields is a string and has an associated language (or “unknown”/null). What would be the most efficient way to store this kind of data? I was thinking about two solutions: The first one consists of a UDT column with: title map>, desc map>... the key is the language, for example “en” or “English” and the set is all variants (a language may have a couple). I would add a field into the UDT for every possible field I can think of (title and desc are 2 of 16) like the example below:
raw_scraps.fields( title frozen