Grok Glue Logstash Scala - GregLinthicum/From-Logistic-Regression-to-Long-short-term-memory-RNN GitHub Wiki

CLI AWS Glue

https://docs.anaconda.com/anaconda-scale/howto/spark-basic/

aws glue start-job-run --job-name my-job

Using Python in Lambda (layer)

Create an AWS Lambda Layer for Python Runtime

Accessing Python Packages from AWS Glue

Wheel on Web

Wheel from the S3 bucket

{ "--additional-python-modules" : "psutil==5.7.2,scikit-learn==0.23.1,geopy==2.0.0,Shapely==1.7.1,googleads==25.0.0,nltk==3.5", "--python-modules-installer-option" : "--no-index --find-links=http://MY-BUCKET.s3-website-us-east-1.amazonaws.com/wheelhouse --trusted-host MY-BUCKET.s3-website-us-east-1.amazonaws.com" }

S3fs with wheel

using INDEX.html pip install awesomepy rattlesnake==0.1.1 --extra-index-url= --trusted-host=

The following command uploads a requirements.txt file to an Amazon S3 bucket.

aws s3 cp requirements.txt s3://YOUR_S3_BUCKET_NAME/requirements.txt

You can use S3 console instead


get current directory. os.getcwd()

import os.path

file1 = open("PythonWrittenName.txt", "w")

toFile = raw_input("Write what you want into the field")

file1.write("content to be written by Python")

file1.close()

^Z


Debian 9 and Ubuntu 16.04 or newer:

sudo apt install s3fs

https://medium.com/spritle-software/mounting-an-s3-bucket-using-fuse-1c1d3cc594c4


https://s3fs.readthedocs.io/en/latest/install.html https://github.com/fsspec/s3fs

import s3fs

fs = s3fs.S3FileSystem(anon=True)

fs.ls('my-bucket')

with fs.open('my-bucket/my-file.txt', 'rb') as f: print(f.read()) b'Hello, world'


BDD & TDD

Testing data products: BDD for data engineers

Specs2 for Scala

BDD with Scala BDD with Scala

Apache Spark with Cucumber for Behavioral-Driven Development

Testing strategies for Apache Spark based projects

Test Driven Development in Grok

Test-Driven Development (TDD) tests with ScalaTest

Scala SBT TDD

AWS Glue offers an optimized Apache Parquet writer

Glue:: l'extraction-l'enrichissement- le chargement et l'organisation DataBrew

Python Shell for Glue

PyGreSQL

PySpark Cheat Sheet

Amazon EMR to run a PySpark job using Python 3.x

Cribl: grok-patterns-library Cribl on AWS

whats-new-elasticsearch-7-12

logstash-grok-tutorial

plugins-filters-grok

introduction-to-logstash-grok-patterns

Dates in Scala on Spark

Glue ETL Scala:: gluecontext , AWS

EMR vs AWS Glue

AWS Lake Formation Workshop

ShiftLeft for Scala

Akka (PaaS)

gRPC vs REST

gRPC on AWS

Cucumber Open

Scala jobs go with Cats, Cats-Effect, Akka, Kafka, Spark

AWS Glue examples

Creating a PySpark project with pytest, pyenv, and egg files

FAQs Les cibles pour Glue Elastic Views actuellement prises en charge sont Amazon Redshift, Amazon S3 et Amazon OpenSearch Service (successeur d'Amazon Elasticsearch Service). Amazon Aurora, Amazon RDS et Amazon DynamoDB seront prochainement pris en charge.

Testing

Testing PySpark DataFrame transformations

Streaming

« connectionType » : « dynamodb » comme collecteur Migrating to DynamoDB using Parquet and Glue

DynamoDB Hybrid

Apache Kafka-> Glue -> Redshift

AWS Amplify, Pinpoint and Kinesis Firehose-> AWS Glue -> Spark ETL -> AWS Athena

IBM ETL vs ELT

scipt cspan Boto3

PySpark write/read to JSON file

Glue Jupyter Glue Jupyter Glue Jupyter

ETL

Oracle to PostgreSQL

Oracle to DynamoDB (Financial Ledger and Accounting Systems Hub (FLASH)). The Financial Ledger and Accounting Systems Hub (FLASH) is a suite of microservices that ingests these financial transactions, performs complex and business-critical functions to substantiate the general ledger, and generates financial statements such as balance sheet, cash flow, and income. ( moved from 90 instances of Oracle to a managed DynamoDB).

MySQL to DynamoDB

DynamoDB to DynaamoDB CLI

Relating AWS Schema Conversion Tool to AWS Glue ETL Actuellement, seules les conversions Oracle ETL, Microsoft SSIS et Teradata BTEQ vers AWS Glue sont pris en charge.

Evaluation of Fault Finding Ability of Assertions https://github.com/GregLinthicum/Attachments/blob/main/ETL%20Standard%20Functional%20Analysis.PNG

SQL-Based ETL with Apache Spark on Amazon EKS Implementation Guide

Recent

ReInvent 2020: Glue Studio, Partition Indexes, Custom Connectors, DataBrew, Schema Registry

AWSLabs - Glue on Ubuntu

Desarrollo y pruebas locales de scripts de ETL mediante la biblioteca de ETL de AWS Glue

pip3 install --upgrade boto3 pytz tzlocal

pip show boto3

PS C:\Users\Greg> gh repo clone awslabs/aws-glue-libs fatal: destination path 'aws-glue-libs' already exists and is not an empty directory. exit status 128

================

https://issueexplorer.com/issue/awslabs/aws-glue-libs/100

I installed each prerequisites and still getting No module named 'awsglue' error.

AWS Glue version 3.0,
Apache Maven from the following location: https://aws-glue-etl-artifacts.s3.amazonaws.com/glue-common/apache-maven-3.6.0-bin.tar.gz
AWS Glue version 3.0: https://aws-glue-etl-artifacts.s3.amazonaws.com/glue-3.0/spark-3.1.1-amzn-0-bin-3.2.1-amzn-3.tgz
SPARK_HOME is setup
run glue-setup.sh from \\wsl$\Ubuntu-20.04\home\my_user\aws_ds\glue_libs\aws-glue-libs\bin

Please help on debbuging this as I don't know where to start else. marcin2x4 wrote this answer on 2021-10-23

Working solution:

where the content of glue-setup.sh is here https://github.com/awslabs/aws-glue-libs/blob/master/bin/glue-setup.sh and rub from git bash is >$ bash glue-setup.sh

================ Glue in Docker (once again ) https://www.youtube.com/watch?v=2-zOXnSDdZ0

=================== Policy for AWS Glue to access S3 ========================

{ "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": [ "glue:", "s3:GetBucketLocation", "s3:ListBucket", "s3:ListAllMyBuckets", "s3:GetBucketAcl", "ec2:DescribeVpcEndpoints", "ec2:DescribeRouteTables", "ec2:CreateNetworkInterface", "ec2:DeleteNetworkInterface", "ec2:DescribeNetworkInterfaces", "ec2:DescribeSecurityGroups", "ec2:DescribeSubnets", "ec2:DescribeVpcAttribute", "iam:ListRolePolicies", "iam:GetRole", "iam:GetRolePolicy", "cloudwatch:PutMetricData" ], "Resource": [ "" ] }, { "Effect": "Allow", "Action": [ "s3:CreateBucket", "s3:PutBucketPublicAccessBlock" ], "Resource": [ "arn:aws:s3:::aws-glue-" ] }, { "Effect": "Allow", "Action": [ "s3:GetObject", "s3:PutObject", "s3:DeleteObject" ], "Resource": [ "arn:aws:s3:::aws-glue-/", "arn:aws:s3:::/aws-glue-/" ] }, { "Effect": "Allow", "Action": [ "s3:GetObject" ], "Resource": [ "arn:aws:s3:::crawler-public", "arn:aws:s3:::aws-glue-" ] }, { "Effect": "Allow", "Action": [ "logs:CreateLogGroup", "logs:CreateLogStream", "logs:PutLogEvents", "logs:AssociateKmsKey" ], "Resource": [ "arn:aws:logs:::/aws-glue/" ] }, { "Effect": "Allow", "Action": [ "ec2:CreateTags", "ec2:DeleteTags" ], "Condition": { "ForAllValues:StringEquals": { "aws:TagKeys": [ "aws-glue-service-resource" ] } }, "Resource": [ "arn:aws:ec2:::network-interface/", "arn:aws:ec2:::security-group/", "arn:aws:ec2:::instance/*" ] } ] }

=================== IAM Role for AWS Glue ======================

To create an IAM role for AWS Glue

  1. Sign in to the AWS Management Console and open the IAM console at https:// console.aws.amazon.com/iam/.
  2. In the left navigation pane, choose Roles.
  3. Choose Create role.
  4. For role type, choose AWS Service, find and choose Glue, and choose Next: Permissions.
  5. On the Attach permissions policy page, choose the policies that contain the required permissions; for example, the AWS managed policy AWSGlueServiceRole for general AWS Glue permissions and the AWS managed policy AmazonS3FullAccess for access to Amazon S3 resources. Then choose Next: Review.

===========================================================

{ "Version":"2012-10-17", "Statement":[ { "Effect":"Allow", 16 AWS Glue Developer Guide Step 3: Attach a Policy to IAM Users That Access AWS Glue "Action":[ "kms:Decrypt" ], "Resource":[ "arn:aws:kms:*:account-id-without-hyphens:key/key-id" ] } ] }

=====================================================