CSV file from HDFS to Oracle BLOB using Spark

Question

I'm working on Java app that uses Spark 2.3.1 to load data from Oracle to HDFS and vice versa. I want to create CSV file in HDFS and then load it to Oracle (12.2) BLOB. The code.. I'm new to Spark.. so any ideas please how to convert JavaRDD to BufferedInputStream, or get rid of mess above and put Dataset

Accepted Answer

Finally.. after couple days of fighting with Oracle, Hadoop and Spark, I found solution for my task:        try {        String trgtFolderPath = "tmp/ETLFramework/csv/form_name";        Configuration conf = new Configuration();        String hdfsUri = "hdfs://" + /*nameNode*/ + ":" + /*hdfsPort*/;        FileSystem fileSystem = FileSystem.get(URI.create(hdfsUri), conf);        RemoteIterator<LocatedFileStatus> fileStatusListIterator = fileSystem.listFiles(new Path(trgtFolderPath), true);        while(fileStatusListIterator.hasNext()){            LocatedFileStatus fileStatus = fileStatusListIterator.next();            String fileName = fileStatus.getPath().getName();            if (fileName.contains(".csv") && fileStatus.getLen()>0) {                log.info("fileName=" + fileName);                log.info("fileStatus.getLen=" + fileStatus.getLen());                BufferedInputStream bis = new BufferedInputStream(fileSystem.open(new Path(trgtFolderPath + "/" + fileName)), 500);                ETLParams param = ETLParams.getParams();                Connection conn = tbl.getJdbcConnection();                String apiPackageInsertLOB = ETLService.replaceParams(tbl.getConnection().getFullSchema() + "." + tbl.getApiPackage().getDbTableApiPackageInsertLOB(), param.getParamsByName());                log.info(String.format("Call %s(%s, %s, %s);", apiPackageInsertLOB, tbl.getFullTableName(), trgtFolderPath + "/" + fileName, "p_nInsertedRows"));                CallableStatement cstmt = conn.prepareCall(String.format("{call %s(?, ?, ?, ?)}", apiPackageInsertLOB));                cstmt.setString(1, tbl.getFullTableName());                cstmt.setString(2, trgtFolderPath + "/" + fileName);                cstmt.setBlob(3, bis, fileStatus.getLen());                cstmt.registerOutParameter(4, Types.INTEGER);                cstmt.execute();                int rowsInsertedCount = cstmt.getInt(3);                log.info("Inserted " + rowsInsertedCount + " rows into table blob_file");                cstmt.close();            }        }        fileSystem.close();    }    catch (IOException |           SQLException exc){        exc.printStackTrace();    }Writing of 2 Gb CSV from Spark Dataset into HDFS, and following reading of this CSV from HDFS into Oracle BLOB took about 5 minutes..

Advertisement

Answer