mirror of
https://github.com/Microsoft/sql-server-samples.git
synced 2025-12-08 14:58:54 +00:00
44 lines
2.7 KiB
Markdown
44 lines
2.7 KiB
Markdown
# SQL Server big data clusters
|
|
|
|
## Pre-requisites
|
|
1. Kubernetes cluster configuration & Kubectl command-line utility
|
|
2. Curl utility
|
|
3. Sqlcmd and bcp utility (Installation instructions [here for Linux](https://docs.microsoft.com/en-us/sql/linux/sql-server-linux-setup-tools?view=sql-server-ver15) and [here for Windows](https://www.microsoft.com/en-us/download/details.aspx?id=53591))
|
|
4. Azure Data Studio or SQL Server Management Studio
|
|
5. SQL Server 2019 big data cluster
|
|
|
|
Installation instructions for SQL Server 2019 big data cluster can be found [here](https://docs.microsoft.com/en-us/sql/big-data-cluster/deployment-guidance?view=sql-server-2017).
|
|
|
|
## Samples Setup
|
|
|
|
**Before you begin**, download the sample database [backup file](https://sqlchoice.blob.core.windows.net/sqlchoice/static/tpcxbb_1gb.bak) and save it locally. Run the CMD script called *bootstrap-sample-db.cmd* or the shell script *bootstrap-sample-db.sh* depending on your platform. This script will restore the database on the SQL Master instance, execute the *bootstrap-sample-db.sql* script, create the database objects needed, export the web_clickstreams & inventory tables to CSV file, and upload the web_clickstreams CSV file to HDFS inside the SQL Server 2019 big data cluster.
|
|
|
|
__[data-pool](data-pool/)__
|
|
|
|
### Data ingestion using Spark
|
|
Connect to the master instance in your SQL Server big data cluster and the SQL Server big data cluster endpoint, and follow the steps in *data-pool/data-ingestion-spark.sql*.
|
|
|
|
### Data ingestion using sql
|
|
Connect to the master instance in your SQL Server big data cluster and execute the steps in *data-pool/data-ingestion-sql.sql*.
|
|
|
|
__[data-virtualization](data-virtualization/)__
|
|
|
|
### External table over HDFS
|
|
Connect to the master instance in your SQL Server big data cluster and execute the steps in *data-virtualization/external-table-hdfs.sql*.
|
|
|
|
### External table over Oracle
|
|
To execute this sample script, you will need following:
|
|
1. Oracle instance and credentials
|
|
1. Create inventory table in Oracle using [data-virtualization/inventory-oracle.sql](data-virtualization/inventory-oracle.sql/) script
|
|
1. Import the inventory.csv file generated by the bootstrap-sample-db script to a table in Oracle
|
|
|
|
Connect to the master instance in your SQL Server big data cluster and execute the steps in *data-virtualization/external-table-oracle.sql*.
|
|
|
|
__[machine-learning](machine-learning/)__
|
|
|
|
### SQL Server ML Services on master instance
|
|
Connect to the master instance in your SQL Server big data cluster and execute the steps in *machine-learning/sql/book-category-r-ml.sql*.
|
|
|
|
### Spark ML
|
|
Connect to the SQL Server big data cluster endpoint, and run the notebook files *machine-learning/spark/1-data-prep.ipynb* and *machine-learning/spark/2-build-ml-model.ipynb* cell by cell.
|