Azure Data Factory - Copy multiple files from HTTP website dynamically to Data Lake using ADF
Here we are copying multiple files dynamically from HTTP website . We are using GitHub as source website to copy multiple files to Azure Data Lake Gen2
Previous video to copy single HTTP file: https://youtu.be/PNN5VPoP2zQ
Here source is having incoming values with some date value, we are keeping a watermark table which keeps track of the data copied and this helps to incrementally load the data to the destination preventing copy of the old data.
Refer below to setup self hosted IR to use SQL from local Machine:
Azure Data Factory - Copy multiple files from AWS S3 to Azure Data Lake Storage using ADF pipeline
Here we are getting all the files in a folder in Amazon S3. We are copying all files dynamically using Azure Data Factory
To prevent hard coded values in sensitive information, its recommended to use Azure Key Vault. Refer the below video for reference to how to use Azure Key vault in ADF:
Azure Data Factory - Delete specific folders and specific files using lookup file in ADF
Here we are making a lookup file as reference which tells the specific files and folders to be deleted. We are using lookup , forEach and delete activity to execute this logic to delete the required files
Azure Data Factory - Create Managed VNET Integration runtime and creating Managed Private Endpoint
In this we are first creating Managed VNET Integration runtime and we are establishing a private endpoint connection with ADLS Gen2 and doing a copy activity using Managed VNET Integration Runtime
Add additional column dynamically with Row count of files in Copy activity
In this we are dynamically adding a column which tells us the row count that each file holds in the Azure Data Lake and we are copying them using Copy activity to SQL Database
Azure Data Factory - Adding Keyvault secret for password in SSIS Package in Azure Data Factory
Storing the sensitive information is an importing thing while developing a solution. Using Azure KeyVault we can store our secrets and access them using a linked service in Azure Data Factory.
Azure Data Factory - Run SSIS packages in ADF as File System (Project Model) using SSIS-IR
In this we are using Project Deployment Model, we are taking the .ispac file that is getting generated after build. From that we are uploading the .ispac file and the package(.dtsx) to Azure File Share.
Azure Data Factory - Copy files which are not empty and last modified is today to ADLS container
This approach helps us to get only files which are not empty
Expression is @greaterOrEquals(activity('Lookup1').output.count,2)
This checks file along with the header, Sometimes files can also contain header but it could be an empty file with no further rows. The above expressions evaluates row count from 2 or more and that will be treated as non empty file.
For True conditions,
Take a copy Activity,
Source is same dataset given to lookup activity
FileName as parameter, should have @item().name as value
To access one premise file system, we need to setup a Self Hosted Integration runtime.
Default Azure runtime will have only scope up to Azure environment, to access files/resources outside azure we need to configure Self Hosted Integrated Runtime.
Here to stimulate an On premise environment we had a Windows 10 virtual machine and we installed Self Hosted IR in that
Using Lookup activity and forEach activity we can copy multiple SQL tables to multiple files dynamically taking their tablename as file name into Azure Datalake Storage.
This helps to modify or generate a new column based on the condition we define. This can also helps to generate a new column based on the data of the existing columns
SORT Transformation
As name defines this helps to sort the data of the column based on the column we provide to it
Below video have the complete explanation of both with example
Conditional Split Transformation - This helps to make different data streams based on the matching condition that we give.
This is similar to CASE statement that we use in our Programming languages. We can give multiple conditions and split the data to different streams. The following video have a detailed example with explanation for Conditional Split Transformations
An Azure Data Engineer and having vast experience on Azure Data Engineering Services and building ETL Pipelines.
With passion on learning Cloud Technologies with hands-on learning and Certified with
- Microsoft Azure Data Engineer (DP-203)
- Microsoft Certified Power BI Data Analyst (PL-300)
- Microsoft Certified Azure Administrator (AZ-104)
- Data bricks Certified Lakehouse Fundamentals
- AWS Certified Solutions Architect - Associate
- AWS Certified Cloud Practitioner
- Microsoft Certified Azure Fundamentals (AZ-900)
- Microsoft Certified Azure Data Fundamentals (DP-900)
- Microsoft Certified Azure Security, Compliance, and Identity Fundamentals (SC-900)
"Evolve ourselves along with the trending technology by learning and enhance the skill set to master it"