# Scylla 5.2 Load and Stream

**URL:** <https://forum.scylladb.com/t/scylla-5-2-load-and-stream/991>\
**Category:** ScyllaDB\
**Tags:** open-source\
**Created:** [November 9, 2023, 12:52pm UTC](https://forum.scylladb.com/t/scylla-5-2-load-and-stream/991 "2023-11-09T12:52:48Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![ericjones2172](https://avatars.discourse-cdn.com/v4/letter/e/9d8465/32.png) [@ericjones2172](https://forum.scylladb.com/u/ericjones2172)\
**Post date:** [November 9, 2023, 12:52pm UTC](https://forum.scylladb.com/t/scylla-5-2-load-and-stream/991/1 "2023-11-09T12:52:48Z")

</div>

Hello,

I am trying to understand this feature. [Nodetool refresh | ScyllaDB Docs](https://opensource.docs.scylladb.com/stable/operating-scylla/nodetool-commands/refresh.html#load-and-stream)

What I don’t understand is if you are going to from 6 nodes(RF=3) to 4 nodes(RF=2), do you need to need to load data from all 6 nodes even if the replication factor in 6 node cluster is 3? If we do need to load data from all 6 nodes into 4 node cluster, are there any risk of running out of space in the new cluster?

---

<div class="post-metadata">

**Author:** ![felipemendes](https://sea2.discourse-cdn.com/flex016/user_avatar/forum.scylladb.com/felipemendes/32/11_2.png) [@felipemendes](https://forum.scylladb.com/u/felipemendes)\
**Post date:** [November 10, 2023, 3:51am UTC](https://forum.scylladb.com/t/scylla-5-2-load-and-stream/991/2 "2023-11-10T03:51:05Z")

</div>

Excellent question. I addressed Load and Stream specifics in [https://www.scylladb.com/2023/09/18/5-more-intriguing-scylladb-capabilities-you-might-have-overlooked/](https://www.scylladb.com/2023/09/18/5-more-intriguing-scylladb-capabilities-you-might-have-overlooked/) , so you may also want to check on that.

> if you are going to from 6 nodes(RF=3) to 4 nodes(RF=2), do you need to need to load data from all 6 nodes even if the replication factor in 6 node cluster is 3?

Copying all 6 nodes indeed seem an overkill. But the real answer is that it depends.

Are you dual-writing to both clusters? Do you expect all data present in the source cluster to match its target? Also, Do you use NetworkTopologyStrategy and spread the data to 3 AZs?

If yes, then you can start dual-writing, run a repair job and once that repair job finishes snapshot your data from a single AZ and copy it over. Both cluster should be in sync afterwards.

> If we do need to load data from all 6 nodes into 4 node cluster, are there any risk of running out of space in the new cluster?

All SSTable data is going to get streamed to its replicas, so you may want to let compaction pick up as you go through each Load and Stream step.

---

<div class="post-metadata">

**Author:** ![tzach](https://sea2.discourse-cdn.com/flex016/user_avatar/forum.scylladb.com/tzach/32/172_2.png) [@tzach](https://forum.scylladb.com/u/tzach)\
**Post date:** [November 13, 2023, 10:35am UTC](https://forum.scylladb.com/t/scylla-5-2-load-and-stream/991/3 "2023-11-13T10:35:16Z")

</div>

@felipemendes good input. Should we add it to the docs?

---

<div class="post-metadata">

**Author:** ![ericjones2172](https://avatars.discourse-cdn.com/v4/letter/e/9d8465/32.png) [@ericjones2172](https://forum.scylladb.com/u/ericjones2172)\
**Post date:** [November 13, 2023, 1:48pm UTC](https://forum.scylladb.com/t/scylla-5-2-load-and-stream/991/4 "2023-11-13T13:48:52Z")

</div>

Do I need to disable compaction during load and stream? And then enable after each load and stream is completed.

---

<div class="post-metadata">

**Author:** ![felipemendes](https://sea2.discourse-cdn.com/flex016/user_avatar/forum.scylladb.com/felipemendes/32/11_2.png) [@felipemendes](https://forum.scylladb.com/u/felipemendes)\
**Post date:** [November 13, 2023, 9:20pm UTC](https://forum.scylladb.com/t/scylla-5-2-load-and-stream/991/5 "2023-11-13T21:20:52Z")

</div>

Not really. You may want to disable tombstones from getting compacted though, in case their gc\_grace\_seconds happen to expire. You can do so by setting `tombstone_gc` to repair. See [Preventing Data Resurrection with Repair Based Tombstone Garbage Collection - ScyllaDB](https://www.scylladb.com/2022/06/30/preventing-data-resurrection-with-repair-based-tombstone-garbage-collection/)

@tzach , that’s definitely a good idea. 🙂 Ping me if anything
