This issue affects SAS/ACCESS Interface to Spark on the SAS® Viya® platform (LTS 2023.10) when you use the CData JDBC driver for Databricks.
When unloading data from the Databricks database to SAS Viya, some rows might not be returned to SAS. There is no error message in the SAS Viya log when this issue occurs. The returned row count of the SAS data set is accurate but might be smaller than the row count of the source table. Depending on the number of columns, this issue might occur when reading five million rows or more from Databricks.
This issue is caused by a bug in the CData JDBC driver for Databricks. Specifically, the problem occurs when you use cloud fetch (UseCloudFetch=true). This option is turned on by default when using SAS/ACCESS Interface to Spark with the CData JDBC driver for Databricks.
To circumvent this issue, disable the cloud fetch option. For SAS/ACCESS Interface to Spark with the CData JDBC driver for Databricks, turn off cloud fetch by setting UseCloudFetch=false in the URL option (not in the Properties option).
The following libname statement turns off cloud fetch in the CData JDBC for Databricks. You need to use your own Databricks credentials.
%let MYDBRICKS=adb-566190778XXXXXX.19.azuredatabricks.net;
%let MYPWD=dapi3baa47a423f0c5df30fc8f88f4XXXXXX-X;
%let MYHTTPPATH=sql/protocolv1/o/566190778127239/1201-203936-XXXXXXXX;
/*
SAS Viya Platform
SAS/ACCESS Interface to Spark libname statement
CData JDBC for Databricks with cloud fetch turned off
*/
libname dbricks spark platform=databricks
user=token pwd="&MYPWD"
url="jdbc:cdata:databricks:Server=&MYDBRICKS;
Database=default;HTTPPath=&MYHTTPPATH;
UseCloudFetch=false;DefaultColumnSize=1024;ConnectRetryWaitTime=20"
bulkload=no
;