mirror of
https://github.com/Microsoft/sql-server-samples.git
synced 2025-12-08 14:58:54 +00:00
Initial samples for SQL Server 2019 big data cluster
Demonstrates various functionality in big data cluster.
This commit is contained in:
@@ -0,0 +1,44 @@
|
||||
# SQL Server big data clusters
|
||||
|
||||
## Pre-requisites
|
||||
1. Kubernetes cluster configuration & Kubectl command-line utility
|
||||
2. Curl utility
|
||||
3. Sqlcmd utility
|
||||
4. Bcp utility
|
||||
5. Azure Data Studio or SQL Server Management Studio
|
||||
6. SQL Server 2019 big data cluster
|
||||
|
||||
Installation instructions for SQL Server 2019 big data cluster can be found [here](https://docs.microsoft.com/en-us/sql/big-data-cluster/deployment-guidance?view=sql-server-2017).
|
||||
|
||||
## Samples Setup
|
||||
|
||||
**Before you begin**, download the sample database [backup file](https://sqlchoice.blob.core.windows.net/sqlchoice/static/tpcxbb_1gb.bak) and save it locally. Run the CMD script called *bootstrap-sample-db.cmd* or the shell script *bootstrap-sample-db.sh* depending on your platform. This script will restore the database on the SQL Master instance, execute the *bootstrap-sample-db.sql* script, create the database objects needed, export the web_clickstreams & inventory tables to CSV file, and upload the web_clickstreams CSV file to HDFS inside the SQL Server 2019 big data cluster.
|
||||
|
||||
__[data-pool](data-pool/)__
|
||||
|
||||
### Data ingestion using Spark
|
||||
Connect to the master instance in your SQL Server big data cluster and the SQL Server big data cluster endpoint, and follow the steps in *data-pool/data-ingestion-spark.sql*.
|
||||
|
||||
### Data ingestion using sql
|
||||
Connect to the master instance in your SQL Server big data cluster and execute the steps in *data-pool/data-ingestion-sql.sql*.
|
||||
|
||||
__[data-virtualization](data-virtualization/)__
|
||||
|
||||
### External table over HDFS
|
||||
Connect to the master instance in your SQL Server big data cluster and execute the steps in *data-virtualization/external-table-hdfs.sql*.
|
||||
|
||||
### External table over Oracle
|
||||
To execute this sample script, you will need following:
|
||||
1. Oracle instance and credentials
|
||||
1. Create inventory table in Oracle using [data-virtualization/inventory-oracle.sql](data-virtualization/inventory-oracle.sql/) script
|
||||
1. Import the inventory.csv file generated by the bootstrap-sample-db script to a table in Oracle
|
||||
|
||||
Connect to the master instance in your SQL Server big data cluster and execute the steps in *data-virtualization/external-table-oracle.sql*.
|
||||
|
||||
__[machine-learning](machine-learning/)__
|
||||
|
||||
### SQL Server ML Services on master instance
|
||||
Connect to the master instance in your SQL Server big data cluster and execute the steps in *machine-learning/sql/book-category-r-ml.sql*.
|
||||
|
||||
### Spark ML
|
||||
Connect to the SQL Server big data cluster endpoint, and run the notebook files *machine-learning/spark/1-data-prep.ipynb* and *machine-learning/spark/2-build-ml-model.ipynb* cell by cell.
|
||||
Reference in New Issue
Block a user