> For the complete documentation index, see [llms.txt](https://umbertogriffo.gitbook.io/apache-spark-best-practices-and-tuning/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://umbertogriffo.gitbook.io/apache-spark-best-practices-and-tuning/dataframe/joining-a-large-and-a-medium-size-dataset.md).

# Joining a large and a small Dataset

A technique to improve the performance is analyzing the **DataFrame** size to get the best join strategy.

If the smaller DataFrame is small enough to fit into the memory of each worker, we can turn **ShuffleHashJoin** or **SortMergeJoin** into a **BroadcastHashJoin**. In broadcast join, the smaller DataFrame will be broadcasted to all worker nodes. Using the BROADCAST hint guides Spark to broadcast the smaller DataFrame when joining them with the bigger one:

```scala
largeDf.join(smallDf.hint("broadcast"), Seq("id"))
```

This way, the larger DataFrame does not need to be shuffled at all.

> *Recently Spark has increased the maximum size for the broadcast table from 2GB to 8GB. Thus, it is not possible to broadcast tables which are greater than 8GB.*
