Skip to main content

Command Palette

Search for a command to run...

AWS Series #5: Lambda and API Gateway

Updated
•15 min read•View as Markdown
AWS Series #5: Lambda and API Gateway

I utilized slides from AWS to enhance my explanations of the topics, strictly for educational purposes and without any commercial intent. I used Gemini to create images alongside the original file. For any related concerns, feel free to contact me via my LinkedIn profile.

Let's talk about the history of serverless:

  • In the early stages, when we had hardware devices, building applications on physical servers took a lot of time.

  • Gradually, applications were built using containers (such as Docker). To manage multiple Docker containers, Kubernetes could be used. This opened up many opportunities because microservices required fewer physical resources, and applications could be tested and built more easily. They could be turned off when not needed and activated when necessary. (However, with containers, we still needed to manage the infrastructure.)

  • Later, when Lambda was announced, the concept of serverless emerged. Essentially, it's a small piece of code used to run a particular solution. However, infrastructure management is handled by AWS, which greatly supports scaling up according to application needs.

What is Serverless?

  • Firstly, you don't have to provision and manage your infrastructure

  • Automatic scaling

  • Pay for what you use

  • Highly available and secure

Currently, following the development trend, serverless is no longer limited to Lambda. In the past, when mentioning serverless, people often only thought of Lambda.

The image above describes AWS services; the more you move to the left, the more physical infrastructure is needed, while moving to the right involves services that don't require infrastructure management.

High Level View:

  • Lambda is just a function code that can be written in various languages of your choice.

  • Lambda receives events from multiple sources to process and then sends results to different destinations depending on the case.

  • Supported languages: Python, JavaScript, Java, C#, Golang, BYOL, Container images

Customer View:

  • You need to consider who in your application is concerned with Lambda.

  • No need to configure firewall or network

  • When configuring Lambda, keep in mind:

    • Event Source Mapping: a service triggers events sent to Lambda for processing (can be triggered from API calls, message queues)

    • You can upload a zip file for Lambda to extract code

    • Lambda can be assigned IAM Roles and Permissions to access and interact with specific services

  • You can configure multiple environments for Lambda, for example, in canary deployment, you can create multiple Aliases for different environments, splitting traffic, such as 20% running in production and the remaining 80% in the test environment

  • Version: Lambda can have multiple versions, and you can roll back to a previous version if needed

Question: When deploying on Lambda, is my function shared with other customers?

The answer is no, because each Lambda is created in a completely independent environment, built within AWS's infrastructure sandbox

Invocation Modes:

  • To execute Lambda, there are three methods: Event Source Mapping, Synchronous Request, and Asynchronous.

    • Synchronous: For synchronous, send a request and wait for the response.

    • Asynchronous: Send a request without waiting for the result.

    • Event Source Mapping: Lambda is mapped to event sources (when an event changes, it triggers Lambda to run)

Synchronous invocation mode:

Useful when an immediate response from the function is needed; errors are sent back to the caller or throttles are sent when you hit the concurrency limit.

There are two scenarios here: a timeout when exceeding 15 minutes, or throttling when the number of functions configured is around 100, but calls reach 200.

Flow of synchronous request:

  • The request is sent to Lambda's frontend service, which then authenticates to see if the user is authorized to use Lambda and checks if there are any quota limits.

  • If everything is fine, it will allocate a suitable VM and initialize the configuration to execute the code.

Asynchronous invocation mode:

  • The caller only receives acknowledgment from the Lambda service, supporting retries (up to 2 retries, or 3 total invokes, including destinations and DLQ).

Flow of asynchronous request:

Go through the front-end → authenticate → place in queue (processing can take up to 6 hours) → when a request is made → Lambda polls the queue (if an error occurs, it can be sent to a dead-letter queue for handling)

An asynchronous request can integrate with SNS and S3 to handle transactions through a queue.

Event Source Mapping:

  • Event source mapping is quite similar to Async → poll data → process (the difference is it allows batch processing, stacking up to a certain number before Lambda processes it).

    For example, once there are 10 messages, Lambda can run.

  • Additionally, you can set up event handling, sending errors to a dead-letter queue for processing.

  • Note that to optimize costs, you don't need to call continuously; you can wait for enough messages to process instead of handling each one individually.

Lambda Function Isolation:

Each function runs in a dedicated sandbox

  • Built on purpose-built virtualization technology: Firecracker

  • Provides enhanced security and workload isolation

AWS maintains the execution environment and runtime

  • Patching, etc.

  • AWS does not maintain a runtime for custom runtimes or container-packaged functions

In the case of using libraries or runtime packages, you need to take care of them yourself

AWS Lambda Execution Environment:

  • Occurs on the first call ⇒ Subsequent runs are warm starts

    When implementing with Lambda, you need to optimize your cold start

Lambda concurrency:

  • What if it executes in parallel? Initially, it will only support one event
  • For example, at the initial moment, there are 3 requests calling Lambda simultaneously.

  • Since no execution environment exists yet, Lambda must create 3 new environments.

Process for each request:

  • Initialization: create environment, load runtime, and load function code

  • Execution: Run function

After Lambda has been running for a while and the execution environments have been created, Lambda retains these environments in memory.

  • New requests can reuse the old environments. For example, Requests 1-5 create 5 environments, and by the 6th request, Lambda reuses an existing environment (no need for Initialization -> direct Execution [Warm Start]).

  • Provisioned concurrency: for instance, initialize 20 environments → configure to retain 10 environments to use those 10 for subsequent calls → maintain warm start.

  • Reserved Concurrency: the opposite of the concept of provisioned concurrency = the maximum number of Lambdas that can run in parallel for a function. It's like reserving a slot to run Lambda for that function. If there are 3 invocations, the first 2 are okay, but the 3rd gets throttled (if sent synchronously, the request fails; asynchronously, it will retry).

💡
Question: Why is this needed? When the number of Lambdas is just two (N=2), it's small, but if the environment reaches around N=10,000, can the backend handle all of this? This method is used to protect the infrastructure behind it, for example, when requests call on-prem, SAP, and the processing time is very long.

Cold Start - Execution Environment:

  • The speed of my function depends on the cold start.

  • A cold start can occur when hardware fails or after a scaling phase, requiring resources to be rebalanced from different availability zones.

The facts:

  • <1% of production workloads

  • Varies from <100ms to >1s

Cold starts occur when:

  • The environment is reaped

  • Failure in underlying resources

  • Rebalancing across AZs

  • Updating code/config flushes

  • Scaling up: Sometimes upgrading code/config leads to scaling up, resulting in a cold start

Influenced by:

  • Memory allocation

  • Size of function package

  • How often a function is called

  • Internal algorithms

Solutions:

  • Optimize dependencies to only the necessary modules

  • Don't load if you don't need it

  • Establishing connections: should not place connections outside the Lambda but can use them in the initialization phase of the Lambda

  • Use provisioned concurrency

Optimize Dependencies:

Only use libraries that are specific to your workload
// const AWS = require('aws-sdk')
const DynamoDB = require('aws-sdk/clients/dynamodb') // 125ms faster

// const AWSXRay = require('aws-xray-sdk')
const AWSXRay = require('aws-xray-sdk-core') // 5ms faster

// const AWS = AWSXRay. captureAWS(require('aws-sdk'))
const dynamodb = new DynamoDB. DocumentClient()
AWSXRay. captureAWSClient(dynamodb.service) // 140ms faster

Provisioned Concurrency:

On-demand pricing (us-east-1)
Requests $0.20 per 1M requests
Invocation Duration $0.0000166667 per GB-second
Provisioned Concurrency pricing (us-east-1)
Requests $0.20 per 1M requests
Provisioned Concurrency $0.0000041667 per GB-second
Invocation Duration $0.0000097222 per GB-second

Maintaining 10 environments [warm start] for a month costs about $19 → AWS will then initialize around 10 environments.

On demand vs Provisioned Concurrency, x86:

  • If application traffic is low, provisioned will be expensive → but when traffic reaches millions → 16% cheaper

Analyze traffic patterns → Apply fixed levels for stability → Automate for cost optimization

The image illustrates setting a fixed level of available resources (green section) throughout the day.

  • Analysis: The green level represents the amount of Lambda functions kept "warm" to avoid Cold Start delays.

  • The yellow columns above the green area are handled by On-demand Concurrency (running based on actual demand).

  • Advantages: Simple, ensures a minimum amount of resources is always ready to respond immediately.

Upgrade to auto-scaling available resources based on actual hourly traffic patterns:

  • Analysis: The current green area "tracks" the yellow columns. When demand is high (midday), the system automatically increases readiness; when demand is low (night), it automatically decreases

  • Advantages: Thorough cost optimization by avoiding payment for excess resources during low demand, while minimizing Cold Start during peak hours

Lambda pricing:

  • Based on memory and execution time

  • AWS Lambda Power Tuning: Application for testing

  • Finding the right balance between cost and execution time

  • Increasing the Lambda function memory configuration also scales the allocated CPU

  • AWS Lambda Power Tuning can help optimize your Lambda functions for cost and/or performance in a data-driven way

  • Option to find the right CPU architecture

Example:

Decision tree to use Lambda:

How many times have you asked yourself when to use Lambda, containers, ECS, EC2, Fargate, or EKS?

  • First, if you want to run your application both on-premises and in the cloud → EKS

  • If you’re not running hybrid, using Windows or .NET without hardware support (ARM, GPU, etc.), then go with Fargate

  • If you need to use AI for heavy processing and require a lot of CPU, you might chooseEC2

  • If on-prem is not needed → and there are no special hardware requirements, you can choose Fargate (supports up to 128 GB). 99% of the time, Fargate can be used (as a serverless solution, you don’t need to worry much about infrastructure)

  • Organizations favor Fargate because it complies with security regulations (when infrastructure needs upgrading, many rules must be followed)

  • No need to go all in → you can combine them

API Gateway:

  • In serverless architecture, Lambda contains business logic in the form of small functions

  • To expose Lambda functions externally as a Web API, we use API Gateway

  • API Gateway acts as an intermediary and security layer, handling requests from clients before invoking Lambda

  • This service is fully managed and serverless, allowing you to build APIs without managing servers

Types of API GW:

Restful APIs:

  • Request/Response

  • HTTP Methods like GET, HEAD, POST, OPTIONS, DELETE

  • Short-lived communication

  • Stateless

Divided into two more types:

RESTful APIs types:

1. REST API (v1)

  • Feature-rich

  • Supports advanced configurations like request transformation, API keys, usage plans, caching, etc

  • Suitable for systems requiring extensive control and advanced features

2. HTTP API (v2)

  • Built from the ground up to optimize for serverless workloads

  • Faster—up to 60% faster

  • Lower cost—up to 71% cheaper

  • Easier to configure and use

Notes:
Currently, REST API still has more features, and some of those features are not yet fully available in HTTP API

Endpoint types:

  • REST Edge-Optimized:

    • Utilizes CloudFront to reduce TLS connection overhead (reduces roundtrip time)

    • Designed for a globally distributed set of clients

  • Regional:

    • Recommended API type for general use cases

    • Designed for building APIs for clients in the same region

  • Private (REST):

    • Only accessible from within VPC (and networks connected to VPC)

    • Designed for building APIs used internally or by private microservices

Supported Protocols

RESTful APIs

  • Request / Response

  • HTTP Methods like GET, POST, etc

  • Short-lived communication

  • Stateless

WebSocket APIs:

  • Two-way connection, like chatting or phone calls

  • Stateful

  • Long-lived communication

Pattern #1: API with AWS services

  • API Gateway supports a wide range of different services

  • Runs hybrid API Gateway + Endpoint in VPC + AWS Direct Connect

Pattern #3: More complex architecture

  • API Gateway can also support caching

API Gateway protection mechanisms:

  • Typically, I would be concerned with WAF (Web Application Firewall, where I can use third-party rules)

Types of authorization:

  1. Resource Policy (REST)

  2. Mutual TLS

  3. WAF (REST)

  4. Cognito User Pools (REST)

  5. JWT (HTTP)

  6. IAM

  7. Lambda Authorizer

💡
Why is REST API slower than HTTP API? Typically, each request through the REST protocol differs (reading each header, configuring authentication mechanisms according to policy) → slower → however, it's not about performance since it's measured in milliseconds.
  • Can integrate with Cognito: SAML, OAuth2, … third party

Resource Policies (REST API):

Resource policies allow an API Gateway REST API to control who can invoke the API based on AWS accounts or source IP addresses before the request reaches the backend services.

Web Application Firewall (WAF):

JWT/Cognito authorizer:

  • OAuth2 compliant (part of OpenID Connect - OIDC)

  • Allows or denies access based on token validity and optional scopes .

  • Any required scopes for the route are validated in the token

IAM Authorizer:

  • Clients must use Signature Version 4 to sign their requests with AWS credentials

  • The authorization token is decoded . User is verified against the Identity & Access Management (IAM) service

  • User must have execute-api access on the route to proceed

Lambda authorizer:

  • Your custom logic to validate the request

  • 2 payload options:

    • Payload 1: must return an IAM policy that allows or denies access to your API route

    • Payload 2: can return IAM policy or Boolean

  • Authorization can be cached

Throttling and usage plans:

  • Protect backend systems

  • Prevents one customer from consuming all your backend system's capacity

  • Let you decide how to allocate capacity among your API consumers with quotas and request rates.

  • Example:

  • I can configure usage plans for different clients, similar to business subscriptions with multiple tiers.

  • If API calls exceed the limit, throttling will occur.

Four levels of Throttling:

1. User Classification (Usage Plans)

  • The system categorizes users into different groups with varying priority levels:

  • User’s Usage Plan: For general users (Web/Mobile)

  • Partner Usage Plan: For strategic partners (usually with higher limits)

  • Services Usage Plan: For internal or system services.

2. Control Layers (Throttling Levels)

The flow goes from left to right, and if any layer's limit is exceeded, the request will be blocked (error 429 Too Many Requests):

  • Per client & method (Yellow/Orange): The most detailed limit. For example, a specific user can only call the "Login" function 5 times per second.

  • Per Client (Pink): The total limit for a user (based on API Key), regardless of which function they are calling.

  • Per method (Blue): The limit for an entire function (API Method). For example, everyone worldwide can call the "Search" function a maximum of 1000 times per second.

  • Per account (Purple): The highest limit at the AWS account level to ensure your entire infrastructure doesn't crash.

Usage Plans and API Keys:

  • API Keys

    • Alphanumeric string values that you distribute to clients (per user/client)

    • Generated by API Gateway or you can imported from a CSV file

  • Use API keys together with Usage Plans or Lambda authorizers to control access to your APIs

  • Usage Plans

    • Specifies who can access API stages and methods

    • How much and how fast they can access the resource

      • Rate limit

      • Quota limit

      • Uses API keys to identify API clients

Stages:

  • API Gateway enables you to set stage variables, allowing the same API to point to different backends.

  • Your APIs are versioned and can be rolled back.

    • APIs are deployed to staging environments.

    • You choose what to name them.

    • For example, these environments:

      • Dev (e.g., example.com/dev)

      • Beta (e.g., example.com/beta)

      • Prod (e.g., example.com/prod)

Canary Releases (REST):

  • Canary Releases: của các big tech

  • With the same API Gateway, traffic gradually shifts from the prod environment to production + 1 environment

Caching (REST):

  • Effective for quick responses and minimizing load to your backend

  • Leverage cache keys to optimize response

    • Path, headers, query strings
  • Can be set up per Stage or Method

SAM:

to be continued ... :D