Required education: Bachelor's degree (4-year) in a STEM field, preferably Information Technology, Computer Science, or Management Information Systems/Technology.
Intermediate to advanced SQL scripting skills and experience working with relational databases and/or data lakes such as Databricks (preferred), Snowflake, PostgreSQL, and Oracle.
Advanced experience in schema design and a deep understanding of star schema, entity-attribute modeling, dimensional modeling, metadata design, data dictionaries, and overall data governance.
Intermediate to advanced Python skills (PySpark preferred) for building and optimizing big-data pipelines, architectures, and datasets.
Experience performing root-cause analysis on internal and external data to identify gaps, answer business questions, and find opportunities for improvement.
Job Description:
Support the development and maintenance of workflows, transformation designs, and integration activities for the GS Enterprise Data Lake & Analytics Platform (GSIE).
Support the establishment and enforcement industry standards and best practices for data engineering across new and existing projects and datasets.
Work independently and collaboratively with data professionals, analytics specialists, and technical and business leaders to gather functional and non-functional requirements, assemble large, sophisticated datasets, and ensure system functionality.
Support the creation and maintenance of standardized governance for data modeling and transformation by defining and refining data relationships and structures (conceptual, logical, and physical models).
Understand and maintain data lineage, metadata, and relationships.
Ensure the Entity Relationship Diagram (ERD) is kept up to date.
Support the identification, design, and implementation of internal process improvements — for example, automating manual processes, optimizing data delivery, and redesigning infrastructure for greater scalability.
Support building the infrastructure for optimal extraction, transformation, and loading (ETL) from diverse data sources.
Support building analytics tools that leverage the data pipeline to deliver insights into customer experience, operational efficiency, and process effectiveness.
Support the creation of data-modeling assets for visualization analysts and data scientists to support the development of effective and innovative data products.
Contribute creative and innovative ideas to user-experience and technical design discussions.
Support the development, documentation, training, and other requirements as it emerge.
Perform other work-related duties as assigned such as but not limited to run, maintain, modify, diagnose and publish SQL code, Alteryx/Dataiku workflows, Tableau/PowerBI dashboards, update documentation of projects, etc