Coderoad Data Engineer at Coderoad responsible for creating and maintaining data pipeline architecture and building ETL/ELT processes on GCP. Supports data analytics and integrates diverse data ecosystems.
Responsibilities
Create and maintain optimal data pipeline architecture
Assemble large, complex data sets meeting business requirements
Identify and implement internal process improvements
Design and build robust ETL/ELT processes on GCP
Build analytics tools for actionable business insights
Integrate diverse data ecosystems with CDPs, ERPs, and external systems
Assist stakeholders with data-related technical issues
Create data tools for analytics and data scientist teams
Qualification
Minimum requirementsCloud ExpertiseAnalytics frameworks and languagesFamiliar with big data toolsPipeline ManagementExperience with Relational DB MySQLUnstructured DataWhat you’ll loveExperience leading teams
Required
Minimum requirements:
4+ years as a Data Engineer experience building processes supporting data transformation, data structures, metadata, dependency, and workload management.
Cloud Expertise: Proven experience designing and implementing ETLs on Databricks and/or GCP
Analytics frameworks and languages, such as: Databricks, GCP, Python, Pyspark (Scala is a plus)
Familiar with big data tools: Apache Spark, Big Query, GCS, Cloud Composer, Unity Catalog
Pipeline Management: Experience with Big Data pipelines via Streams and/or Batches for high-intensive applications.
Experience with Relational DB MySQL .
Experience performing root cause analysis on internal and external data and processes to answer specific business questions and identify opportunities for improvement.
Unstructured Data: Strong analytic skills related to working with unstructured datasets.