From 274577212bd9e435ca1bb25a4053842bf44bb7b9 Mon Sep 17 00:00:00 2001 From: ShivSood Date: Fri, 21 Feb 2020 13:29:50 -0800 Subject: [PATCH] first draft : connector with AD mode --- .../mssql_spark_connector_ad_pyspark.ipynb | 385 ++++++++++++++++++ 1 file changed, 385 insertions(+) create mode 100644 samples/features/sql-big-data-cluster/spark/data-virtualization/mssql_spark_connector_ad_pyspark.ipynb diff --git a/samples/features/sql-big-data-cluster/spark/data-virtualization/mssql_spark_connector_ad_pyspark.ipynb b/samples/features/sql-big-data-cluster/spark/data-virtualization/mssql_spark_connector_ad_pyspark.ipynb new file mode 100644 index 00000000..ffdfe758 --- /dev/null +++ b/samples/features/sql-big-data-cluster/spark/data-virtualization/mssql_spark_connector_ad_pyspark.ipynb @@ -0,0 +1,385 @@ +{ + "metadata": { + "kernelspec": { + "name": "pysparkkernel", + "display_name": "PySpark" + }, + "language_info": { + "name": "pyspark", + "mimetype": "text/x-python", + "codemirror_mode": { + "name": "python", + "version": 2 + }, + "pygments_lexer": "python2" + } + }, + "nbformat_minor": 2, + "nbformat": 4, + "cells": [ + { + "cell_type": "markdown", + "source": [ + "# Using MsSQLSparkConnector with Integrated AD Auth\r\n", + "\r\n", + "This sample shows how to use the MsSQLSparkConnector with integrated AD Auth when using use principal and keytab instead of username/password. \r\n", + "\r\n", + "## PreReq\r\n", + "-------\r\n", + "- SQL Server 2019 big data cluster is deployed with AD\r\n", + "- Have access to AD controller to create keytabs that we will use in this sample. \r\n", + "- Download [AdultCensusIncome.csv]( https://amldockerdatasets.azureedge.net/AdultCensusIncome.csv ) to your local machine. Upload this file to hdfs folder named *spark_data*. \r\n", + "- The sample uses a SQL database *spark_sql_db* to create/update tables. The database needs to be created before the sample is run.\r\n", + " " + ], + "metadata": { + "azdata_cell_guid": "2757df21-8174-4bb1-a52c-66eaf94f6b96" + } + }, + { + "cell_type": "markdown", + "source": [ + "\r\n", + "" + ], + "metadata": { + "azdata_cell_guid": "48874729-541f-4b01-888f-3fdd2aeb59da" + } + }, + { + "cell_type": "markdown", + "source": [ + "# Creating KeyTab file\r\n", + "The following section shows how to generate principal and keytab. This assumes you have a SS19 Big Data Cluster installed with Windows AD contoller for domain AZDATA.LOCAL. One of the users is testusera1@AZDATA.LOCAL and the user is part of Domain Admin group.\r\n", + "\r\n", + "## Create KeyTab file using ktpass\r\n", + "1. Login to the Windows AD controller with user1 credentials.\r\n", + "2. Open command prompt in Administrator mode.\r\n", + "3. Use ktpass to create a key tab. Refer [here](https://docs.microsoft.com/en-us/windows-server/administration/windows-commands/ktpass) for documentation on using ktpass. \r\n", + "\r\n", + "```sh\r\n", + "ktpass -out testusera1.keytab -mapUser testusera1@AZDATA.LOCAL -pass -mapOp set +DumpSalt -crypto AES256-SHA1 -ptype KRB5_NT_PRINCIPAL -princ testusera1@AZDATA.LOCAL\r\n", + "```\r\n", + "\r\n", + "The command above should generate a keytab file named testusera1.keytab. Transfer this file to hdfs folder in Big Data Cluster. In this sample we transfer the file to /user/testusera1/testusera1.keytab\r\n", + "\r\n", + "## Create KeyTab file using kinit\r\n", + "\r\n", + "If you are on a linux machine kinit can be used as follows to create keytab. Note that you linux machine shoud be connected to the domain controler.\r\n", + "\r\n", + "``` sh\r\n", + "ktutil\r\n", + "ktutil : add_entry -password -p testusera1@AZDATA.LOCAL -k 1 -e arcfour-hmac-md5\r\n", + "Password for testusera1@myDomain:\r\n", + "ktutil : add_entry -password -p testusera1@AZDATA.LOCAL -k 1 -e des-cbc-md4\r\n", + "ktutil : wkt testusera1@AZDATA.keytab \r\n", + "```\r\n", + "\r\n", + "``` sh\r\n", + "## Check if keytab generated properly. Any error implies that keytab is not generated right.\r\n", + "kinit -kt testusera1.keytab testusera1@AZDATA.LOCAL\r\n", + "```\r\n", + "\r\n", + "Load Keytab to HDFS for use\r\n", + "\r\n", + "```sh\r\n", + "hadoop fs -mkdir -p /user/testusera1/\r\n", + "hadoop fs -copyFromLocal -f testusera1.keytab /user/testusera1/testusera1.keytab\r\n", + "```\r\n", + "\r\n", + "\r\n", + "\r\n", + " " + ], + "metadata": { + "azdata_cell_guid": "5e733ddd-cbf8-4aa1-9f97-9fff40c28332" + } + }, + { + "cell_type": "markdown", + "source": [ + "# Create a database user\r\n", + "\r\n", + "``` sql\r\n", + "IF NOT EXISTS (select name from sys.server_principals where name='AZDATA.LOCAL\\testusera1')\r\n", + "BEGIN\r\n", + " CREATE LOGIN [AZDATA.LOCAL\\testusera1] FROM WINDOWS\r\n", + "END\r\n", + "\r\n", + "ALTER SERVER ROLE dbcreator ADD MEMBER [AZDATA.LOCAL\\testusera1]\r\n", + "GRANT VIEW SERVER STATE to [AZDATA.LOCAL\\testusera1]\r\n", + "\r\n", + "# Create a database named \"spark_mssql_db\"\r\n", + "IF NOT EXISTS (SELECT * FROM sys.databases WHERE name = N'spark_mssql_db')\r\n", + " CREATE DATABASE spark_mssql_db\r\n", + "```" + ], + "metadata": { + "azdata_cell_guid": "453e7b2f-e590-4b95-9fbd-5dc8b9d1f02c" + } + }, + { + "cell_type": "markdown", + "source": [ + "# Configure Spark applicaion to point to the key tab file\r\n", + "Here we configure spark to use the keytab file once the keytab is created and uploaded to HDFS (/user/testusera1/testusera1.keytab). \r\n", + "Note the usage of \"spark.files\" : \"/user/testusera1/testusera1.keytab\". As a result of this configuration Spark driver distributes the file to all executors. \r\n", + "\r\n", + "Run the cell below to start the spark application.\r\n", + "" + ], + "metadata": { + "azdata_cell_guid": "8ecbc487-996a-4bd1-b4f0-bcbf439decea" + } + }, + { + "cell_type": "code", + "source": [ + "%%configure -f\r\n", + "{\"conf\": {\r\n", + " \"spark.files\" : \"/user/testusera1/testusera1.keytab\",\r\n", + " \"spark.executor.memory\": \"4g\",\r\n", + " \"spark.driver.memory\": \"4g\",\r\n", + " \"spark.executor.cores\": 2,\r\n", + " \"spark.driver.cores\": 1,\r\n", + " \"spark.executor.instances\": 4\r\n", + " }\r\n", + "}" + ], + "metadata": { + "azdata_cell_guid": "32acad02-f758-4a0a-a4e3-e64a52948986", + "tags": [] + }, + "outputs": [ + { + "output_type": "stream", + "name": "stdout", + "text": "Starting Spark application\n" + }, + { + "output_type": "display_data", + "data": { + "text/plain": "", + "text/html": "\n
IDYARN Application IDKindStateSpark UIDriver logCurrent session?
12application_1581458669418_0041pysparkidleLinkLink
" + }, + "metadata": {} + }, + { + "output_type": "display_data", + "data": { + "text/plain": "FloatProgress(value=0.0, bar_style='info', description='Progress:', layout=Layout(height='25px', width='50%'),…", + "application/vnd.jupyter.widget-view+json": { + "version_major": 2, + "version_minor": 0, + "model_id": "ae4ec82192fb473e8b77e4065f4c1927" + } + }, + "metadata": {} + }, + { + "output_type": "stream", + "name": "stdout", + "text": "SparkSession available as 'spark'.\n" + }, + { + "output_type": "display_data", + "data": { + "text/plain": "", + "text/html": "Current session configs: {'conf': {'spark.files': '/user/user1/user1.keytab', 'spark.executor.memory': '4g', 'spark.driver.memory': '4g', 'spark.executor.cores': 2, 'spark.driver.cores': 1, 'spark.executor.instances': 4}, 'kind': 'pyspark'}
" + }, + "metadata": {} + }, + { + "output_type": "display_data", + "data": { + "text/plain": "", + "text/html": "\n
IDYARN Application IDKindStateSpark UIDriver logCurrent session?
6application_1581458669418_0035pysparkidleLinkLink
" + }, + "metadata": {} + } + ], + "execution_count": 8 + }, + { + "cell_type": "markdown", + "source": [ + "# Read CSV into a data frame\r\n", + "In this step we read the CSV into a data frame. This dataframe would then be written to SQL table using MSSQL Spark Connector \r\n", + "\r\n", + "\r\n", + "" + ], + "metadata": { + "azdata_cell_guid": "ed8b58e0-3607-4a71-8dc8-034bc0180ee4" + } + }, + { + "cell_type": "code", + "source": [ + "#spark = SparkSession.builder.getOrCreate()\r\n", + "sc.setLogLevel(\"INFO\")\r\n", + "\r\n", + "#Read a file and then write it to the SQL table\r\n", + "datafile = \"/spark_data/AdultCensusIncome.csv\"\r\n", + "df = spark.read.format('csv').options(header='true', inferSchema='true', ignoreLeadingWhiteSpace='true', ignoreTrailingWhiteSpace='true').load(datafile)\r\n", + "df.show(5)\r\n", + "\r\n", + "\r\n", + "#Process this data. Very simple data cleanup steps. Replacing \"-\" with \"_\" in column names\r\n", + "columns_new = [col.replace(\"-\", \"_\") for col in df.columns]\r\n", + "df = df.toDF(*columns_new)\r\n", + "df.show(5)\r\n", + "" + ], + "metadata": { + "azdata_cell_guid": "813bbfa3-2613-45dd-9556-94faba602977" + }, + "outputs": [ + { + "output_type": "stream", + "name": "stdout", + "text": "Starting Spark application\n" + }, + { + "output_type": "display_data", + "data": { + "text/plain": "", + "text/html": "\n
IDYARN Application IDKindStateSpark UIDriver logCurrent session?
7application_1581458669418_0036pysparkidleLinkLink
" + }, + "metadata": {} + }, + { + "output_type": "display_data", + "data": { + "text/plain": "FloatProgress(value=0.0, bar_style='info', description='Progress:', layout=Layout(height='25px', width='50%'),…", + "application/vnd.jupyter.widget-view+json": { + "version_major": 2, + "version_minor": 0, + "model_id": "52afbfffaf3545e28a2d09864d1c4ced" + } + }, + "metadata": {} + }, + { + "output_type": "stream", + "name": "stdout", + "text": "SparkSession available as 'spark'.\n" + }, + { + "output_type": "display_data", + "data": { + "text/plain": "FloatProgress(value=0.0, bar_style='info', description='Progress:', layout=Layout(height='25px', width='50%'),…", + "application/vnd.jupyter.widget-view+json": { + "version_major": 2, + "version_minor": 0, + "model_id": "7e03eb6a81cd41e4878f12df5ee31dc7" + } + }, + "metadata": {} + }, + { + "output_type": "stream", + "name": "stdout", + "text": "+---+----------------+------+---------+-------------+------------------+-----------------+-------------+-----+------+------------+------------+--------------+--------------+------+\n|age| workclass|fnlwgt|education|education-num| marital-status| occupation| relationship| race| sex|capital-gain|capital-loss|hours-per-week|native-country|income|\n+---+----------------+------+---------+-------------+------------------+-----------------+-------------+-----+------+------------+------------+--------------+--------------+------+\n| 39| State-gov| 77516|Bachelors| 13| Never-married| Adm-clerical|Not-in-family|White| Male| 2174| 0| 40| United-States| <=50K|\n| 50|Self-emp-not-inc| 83311|Bachelors| 13|Married-civ-spouse| Exec-managerial| Husband|White| Male| 0| 0| 13| United-States| <=50K|\n| 38| Private|215646| HS-grad| 9| Divorced|Handlers-cleaners|Not-in-family|White| Male| 0| 0| 40| United-States| <=50K|\n| 53| Private|234721| 11th| 7|Married-civ-spouse|Handlers-cleaners| Husband|Black| Male| 0| 0| 40| United-States| <=50K|\n| 28| Private|338409|Bachelors| 13|Married-civ-spouse| Prof-specialty| Wife|Black|Female| 0| 0| 40| Cuba| <=50K|\n+---+----------------+------+---------+-------------+------------------+-----------------+-------------+-----+------+------------+------------+--------------+--------------+------+\nonly showing top 5 rows\n\n+---+----------------+------+---------+-------------+------------------+-----------------+-------------+-----+------+------------+------------+--------------+--------------+------+\n|age| workclass|fnlwgt|education|education_num| marital_status| occupation| relationship| race| sex|capital_gain|capital_loss|hours_per_week|native_country|income|\n+---+----------------+------+---------+-------------+------------------+-----------------+-------------+-----+------+------------+------------+--------------+--------------+------+\n| 39| State-gov| 77516|Bachelors| 13| Never-married| Adm-clerical|Not-in-family|White| Male| 2174| 0| 40| United-States| <=50K|\n| 50|Self-emp-not-inc| 83311|Bachelors| 13|Married-civ-spouse| Exec-managerial| Husband|White| Male| 0| 0| 13| United-States| <=50K|\n| 38| Private|215646| HS-grad| 9| Divorced|Handlers-cleaners|Not-in-family|White| Male| 0| 0| 40| United-States| <=50K|\n| 53| Private|234721| 11th| 7|Married-civ-spouse|Handlers-cleaners| Husband|Black| Male| 0| 0| 40| United-States| <=50K|\n| 28| Private|338409|Bachelors| 13|Married-civ-spouse| Prof-specialty| Wife|Black|Female| 0| 0| 40| Cuba| <=50K|\n+---+----------------+------+---------+-------------+------------------+-----------------+-------------+-----+------+------------+------------+--------------+--------------+------+\nonly showing top 5 rows" + } + ], + "execution_count": 5 + }, + { + "cell_type": "markdown", + "source": [ + "# Write and READ to/from SQL Table ( using Integrated Auth)\r\n", + "- Write dataframe to SQL table to Master instance\r\n", + "- Read SQL Table to Spark dataframe\r\n", + "\r\n", + "In both scenarions here we use integrated auth with principal\\keytab file rather than username\\password of the user." + ], + "metadata": { + "azdata_cell_guid": "a6afceb2-6fbc-435b-af88-e9f5cc784f5d" + } + }, + { + "cell_type": "code", + "source": [ + "#Write from Spark to SQL table using MSSQL Spark Connector\r\n", + "print(\"MSSQL Spark Connector write(overwrite) start \")\r\n", + "\r\n", + "servername = \"jdbc:sqlserver://master-p-svc:1433\"\r\n", + "dbname = \"mssql_spark_db\"\r\n", + "security_spec = \";integratedSecurity=true;authenticationScheme=JavaKerberos;\"\r\n", + "url = servername + \";\" + \"databaseName=\" + dbname + security_spec\r\n", + "\r\n", + "dbtable = \"AdultCensus_test\"\r\n", + "principal = \"testusera1@AZDATA.LOCAL\"\r\n", + "keytab = \"/user/testusera1/testusera1.keytab\" \r\n", + "\r\n", + "try:\r\n", + " df.write \\\r\n", + " .format(\"com.microsoft.sqlserver.jdbc.spark\") \\\r\n", + " .mode(\"overwrite\") \\\r\n", + " .option(\"url\", url) \\\r\n", + " .option(\"dbtable\", dbtable) \\\r\n", + " .option(\"principal\", principal) \\\r\n", + " .option(\"keytab\", keytab) \\\r\n", + " .save()\r\n", + "except ValueError as error :\r\n", + " print(\"MSSQL Spark Connector write(overwrite) failed\", error)\r\n", + "\r\n", + "print(\"MSSQL Connector write(overwrite) done \")" + ], + "metadata": { + "azdata_cell_guid": "b851fe61-6e85-4e46-a20f-4063fc6586e0" + }, + "outputs": [ + { + "output_type": "display_data", + "data": { + "text/plain": "FloatProgress(value=0.0, bar_style='info', description='Progress:', layout=Layout(height='25px', width='50%'),…", + "application/vnd.jupyter.widget-view+json": { + "version_major": 2, + "version_minor": 0, + "model_id": "644bfe3ba48142f38f52717cab9eb19d" + } + }, + "metadata": {} + }, + { + "output_type": "stream", + "name": "stdout", + "text": "Use MSSQL connector to write to master SQL instance \nMSSQL Connector write(overwrite) succeeded" + } + ], + "execution_count": 7 + }, + { + "cell_type": "code", + "source": [ + "#Read from SQL table using MSSQ Connector\r\n", + "print(\"MSSQL Spark Connector read start \")\r\n", + "jdbcDF = spark.read \\\r\n", + " .format(\"com.microsoft.sqlserver.jdbc.spark\") \\\r\n", + " .option(\"url\", url) \\\r\n", + " .option(\"dbtable\", dbtable) \\\r\n", + " .option(\"url\", url) \\\r\n", + " .option(\"dbtable\", dbtable) \\\r\n", + " .option(\"principal\", user) \\\r\n", + " .option(\"keytab\", keytab).load()\r\n", + "\r\n", + "jdbcDF.show(5)\r\n", + "\r\n", + "print(\"MSSQL Spark Connector read done\")" + ], + "metadata": { + "azdata_cell_guid": "e3e19e1f-1325-47ea-87d1-1170f316d2d8" + }, + "outputs": [ + { + "output_type": "stream", + "name": "stdout", + "text": "read data from SQL server table \n+---+----------------+------+---------+-------------+------------------+-----------------+-------------+-----+------+------------+------------+--------------+--------------+------+\n|age| workclass|fnlwgt|education|education_num| marital_status| occupation| relationship| race| sex|capital_gain|capital_loss|hours_per_week|native_country|income|\n+---+----------------+------+---------+-------------+------------------+-----------------+-------------+-----+------+------------+------------+--------------+--------------+------+\n| 39| State-gov| 77516|Bachelors| 13| Never-married| Adm-clerical|Not-in-family|White| Male| 2174| 0| 40| United-States| <=50K|\n| 50|Self-emp-not-inc| 83311|Bachelors| 13|Married-civ-spouse| Exec-managerial| Husband|White| Male| 0| 0| 13| United-States| <=50K|\n| 38| Private|215646| HS-grad| 9| Divorced|Handlers-cleaners|Not-in-family|White| Male| 0| 0| 40| United-States| <=50K|\n| 53| Private|234721| 11th| 7|Married-civ-spouse|Handlers-cleaners| Husband|Black| Male| 0| 0| 40| United-States| <=50K|\n| 28| Private|338409|Bachelors| 13|Married-civ-spouse| Prof-specialty| Wife|Black|Female| 0| 0| 40| Cuba| <=50K|\n+---+----------------+------+---------+-------------+------------------+-----------------+-------------+-----+------+------------+------------+--------------+--------------+------+\nonly showing top 5 rows" + } + ], + "execution_count": 7 + } + ] +} \ No newline at end of file