Skip to main content
Full course — free access

AI Automation
Pro

The most comprehensive business-automation course, with AI. You'll learn to build advanced automation systems with Make.com and n8n, connect AI tools via API, and deploy 10 real business projects to production.

20 hours of content
10 projects
Full lifetime access
5 modules
Your progress
Module 1 of 5
1 module completed of 5
1
Module 1 — 4 hours

Advanced Make.com

Active
1.1

Webhooks in Make.com

Quick Win

What is a Webhook and why is it useful?

A Webhook is a unique URL that lets external systems "knock on the door" and pass you data in real time — without frequent polling. Instead of Make.com checking every two minutes whether something new happened, the third-party system sends an HTTP request the moment the event occurs. The result: much faster automation, with far fewer wasted API quotas.

Examples of common uses: receiving a new lead from a Typeform form, receiving a payment from Stripe, receiving a new message from Slack, updating an order status from WooCommerce, and hundreds of other services that support sending Webhooks.

Creating a Webhook in Make.com — step by step

1
Add a Trigger — Custom Webhook

Open a new Scenario in Make.com. Choose the first Trigger in Add a module, search for "Webhooks" and choose Custom Webhook. Click Add and give it a descriptive name like new-lead-webhook.

2
Set up Instant Trigger

Make.com generates a unique URL for you. Make sure the Instant Trigger option is checked — that is what makes the Scenario run immediately with every request, instead of waiting for polling.

3
Send a sample POST request

Click Determine data structure so Make.com learns your data structure. Send it a sample request (see code below) — the system will automatically analyze the JSON and generate a Schema.

4
Processing the data

After the structure is learned, all your fields (name, email, phone, etc.) will be available for mapping in the following modules in the Scenario. Drag them into the relevant fields.

Code: sending data to a Webhook in Make.com

If you want to test your Webhook from JavaScript code (Node.js, browser, or any other environment), here is a full example:

send-to-webhook.js
// sending data to a Webhook in Make.com
const webhookUrl = 'https://hook.eu1.make.com/YOUR_WEBHOOK_ID';

const data = {
  name: 'John Doe',
  email: 'israel@example.com',
  phone: '050-1234567',
  source: 'landing_page',
  timestamp: new Date().toISOString()
};

const response = await fetch(webhookUrl, {
  method: 'POST',
  headers: { 'Content-Type': 'application/json' },
  body: JSON.stringify(data)
});

console.log('Status:', response.status); // 200 = success
const text = await response.text();
console.log('Response:', text); // "Accepted" = Make.com received successfully
Tip: Make.com response

Make.com returns 200 Accepted the moment it receives the request — even before the Scenario finished. If you need an asynchronous response with results, use the Webhook Response module at the end of the Scenario.

Common mistakes and how to avoid them

Missing Content-Type

You forgot to add the header Content-Type: application/json. Make.com will not parse the Body and will throw a parsing error.

Scenario not active

The Webhook will return 200 even when the Scenario is off, but the requests will pile up in a Queue. Make sure the Scenario is active before testing.

Using GET instead of POST

A Custom Webhook in Make.com expects a POST request with a Body. Using GET will fail to read the form fields.

Testing with RequestBin

Before connecting Make.com, test your Webhook with requestbin.com to see exactly what is being sent to you.

1.2

Advanced Routers and Filters

Deep Dive

Router Module — routing by conditions

The Router is one of the most powerful tools in Make.com. It lets you split the Scenario into several branches, each of which runs under different conditions. Think of it like an switch-case in code — each Branch gets a Filter with an activation condition.

Example: a lead coming from Facebook is sent to an immediate WhatsApp chat, a lead from the website gets a detailed email, a lead from LinkedIn is fed directly into the CRM. Each Branch runs independently and in parallel.

Setting Filter Conditions

Text operators
  • Equal to
  • Contains
  • Starts with
  • Matches pattern (RegEx)
Number operators
  • Greater than
  • Less than or equal
  • Between
Date and state
  • Date is before/after
  • Exists / Is empty
  • Is true / Is false

Fallback Route — "otherwise"

Every Router needs at least one Branch with a Fallback (Otherwise). This is the Branch that runs all the data that did not pass any other Filter. Without a Fallback, data that does not match any Branch is silently lost — and you won't know about it. The right way: add a final Branch as Otherwise → send yourself an alert + save to an Error Log.

Nested Routers

You can nest Routers within Routers to create complex logic. For example: a first Router splits by country (Israel / abroad). The Israel Branch goes to a second Router that splits by city. Main use: complex Escalation processes in Sales Pipelines.

Note: operations quota

Each Branch in a Router counts as a separate operation, including the operations inside it. A Router with 5 Branches, each containing 3 modules = 15 operations per run. Plan the architecture according to your quota.

1.3

Error Handling in Make.com

Deep Dive

Automation in production will encounter errors — it's not a question of if, but when. Make.com offers a powerful error-handling mechanism that lets you control exactly what happens when a module fails.

The four directives

B Break

Stop the run and mark it as failed. The data is saved in Incomplete Executions for manual handling. The default directive.

R Resume

Skip the failed record and finish the run as a success. Useful when some records are non-critical and you can continue without them.

I Ignore

Ignore the error completely and mark as success. Used when the error is expected and unimportant (for example, a permitted Duplicate in the DB).

C Commit

Save all operations completed up to the point of the error. Used with Roll Back Transactions to prevent duplicates.

Error Handler Routes

Right-click any module and choose Add error handler. This adds a special Branch that runs only when the module fails. There you can put: sending an email to the admin, adding a row to Google Sheets with the error details, sending a Slack message, and then choosing the appropriate directive.

Retry — automatic retry

For temporary network errors, set up Retry inside the Break directive. Set the number of attempts (up to 5) and the interval in seconds between each attempt. This is ideal for API calls that can fail due to Rate Limiting.

error-handler-config.json
{
  "errorHandler": "resume",
  "retryCount": 3,
  "retryInterval": 60,
  "fallbackValue": null
}
Recommended production approach

Set an Error Handler on every critical module (Google Sheets Write, CRM Update, sending email). In the Error Handler: save to a Log row ← send a Slack message with the run ID ← use Commit ← continue to the next record. This way you never lose data silently.

1.4

Data Stores and Aggregators

Deep Dive

Data Stores — a built-in database in Make.com

A Data Store is a minimalist database built directly into Make.com. It lets you save and retrieve data between different runs of the Scenario — something you cannot do with regular variables. Main uses: preventing duplicate processing (Deduplication), storing statuses, data cache, tracking a lead over time.

Each Data Store is defined with a Schema — the structure of the fields it will contain. For example: a field email (Text, Unique), status (Text), last_seen (Date).

Add / Update Record

Add a new record or update an existing one by a unique Key

Get Record / Search

Retrieve a record by Key or search by filters

Delete Record

Delete a specific record or clear the entire Store

Aggregators — collecting items from an Iterator

When an Iterator produces multiple items (for example: 50 rows from Google Sheets), an Aggregator collects them back into a single item. The three common Aggregators:

Array Aggregator

Collects all the items into a single Array. Useful for sending a list of leads as a JSON Body to an external service, or building a consolidated Payload.

Text Aggregator

Merges multiple text values into a single string with a defined separator. Useful for building an email body from a list, or summarizing multiple fields.

Numeric Aggregator

Computes Sum, Average, Min, Max, Count across all items. Useful for daily reports, revenue calculation, counting leads by campaign.

🎯
Module 1 project

Automatic CRM — end to end

In this project you'll build a complete Scenario in Make.com that receives a new lead from the site, filters it by source, saves it, sends messages to the team and creates a new Contact in the CRM — all automatically within seconds.

The Scenario architecture

1
Trigger — Custom Webhook

Receives a new lead from the site. Fields: name, email, phone, source (facebook/website/linkedin), budget

2
Data Store — duplicate check

Search by email in the Data Store. If it exists — update the last date and stop. If not — continue to Step 3.

3
Router — sorting by source

Branch A: source = "facebook" ← send immediate WhatsApp. Branch B: source = "linkedin" ← direct CRM. Branch C: Otherwise ← welcome email.

4
Google Sheets — saving the lead

Add a new row in the "Leads" sheet. Include: name, email, phone, source, date, budget, status (NEW).

5
Gmail — welcome email

Send a personalized email with the lead's name, an initial offer, and a link to schedule a call (Calendly). Include the rep's details in the signature.

6
Slack — alert to the sales team

Send a message to the #new-leads channel: lead name, source, budget, link to the row in Sheets. Include an @mention of the relevant rep.

7
HubSpot / Pipedrive — creating a Contact

Create a new Contact in the CRM with all the lead's details. Add a Deal/Opportunity with a "New Lead" stage and link it to the appropriate Pipeline.

Bonus challenge

Add an OpenAI module before the Google Sheets step. Let it assess the lead's seriousness based on the "message" field they left in the form — and ask it for a 1-10 score + a short reason. Save the score in the sheet and the CRM as a Custom Property.

2
Module 2 — 4 hours

n8n Self-Hosted

2.1

Installing n8n on Docker

Quick Win

n8n is the most powerful open-source automation tool available. Unlike Make.com, when you run n8n Self-Hosted — there are no quotas, no per-operation cost, and your data stays with you. The fastest way to install n8n in a production environment is Docker Compose.

Prerequisites

Docker + Docker Compose VPS / server with Linux A domain name (for SSL) 2GB RAM minimum

Configuring docker-compose.yml

docker-compose.yml
# create a folder
mkdir n8n-data && cd n8n-data

# docker-compose.yml
cat > docker-compose.yml << 'EOF'
version: '3.1'
services:
  n8n:
    image: docker.n8n.io/n8nio/n8n
    restart: always
    ports:
      - "5678:5678"
    environment:
      - N8N_BASIC_AUTH_ACTIVE=true
      - N8N_BASIC_AUTH_USER=admin
      - N8N_BASIC_AUTH_PASSWORD=your-secure-password
      - WEBHOOK_URL=https://your-domain.com/
      - N8N_HOST=your-domain.com
      - N8N_PORT=5678
      - N8N_PROTOCOL=https
    volumes:
      - ~/.n8n:/home/node/.n8n
EOF

docker compose up -d

Explanation of the variables

N8N_BASIC_AUTH_ACTIVE

Enables basic authentication. Mandatory in production so not everyone can reach your interface.

N8N_BASIC_AUTH_PASSWORD

Choose a strong password. After going live, change it to an environment Variable.

WEBHOOK_URL

Your public address. n8n will use it to build Webhook URLs. Must be HTTPS.

volumes: ~/.n8n

Store the data outside the Container. Without this you'll lose all the Workflows on every Restart.

HTTPS with Nginx + Certbot

For production with a real domain, add Nginx as a Reverse Proxy with SSL from Let's Encrypt. The command: certbot --nginx -d your-domain.com. The course includes a full Nginx config with Rate Limiting.

2.2

Complex Workflows and Code Nodes

Deep Dive

n8n offers a Code Node that lets you write JavaScript (or Python) directly inside the Workflow. This makes anything that can be coded — possible. Data cleaning, complex calculations, building dynamic Payloads, and processing business logic that does not exist as a built-in Node.

n8n Code Node — lead processing
// n8n Code Node — lead processing from CRM
const items = $input.all();
const processed = items.map(item => {
  const data = item.json;

  // clean an Israeli phone number
  const phone = data.phone?.replace(/\D/g, '');
  const formattedPhone = phone?.startsWith('972')
    ? `+${phone}`
    : `+972${phone?.replace(/^0/, '')}`;

  // compute Lead Score
  let score = 0;
  if (data.email?.endsWith('.co.il')) score += 20;
  if (data.company) score += 30;
  if (data.budget > 10000) score += 50;

  return {
    json: {
      ...data,
      phone: formattedPhone,
      leadScore: score,
      processedAt: new Date().toISOString()
    }
  };
});

return processed;

The code receives all the items ($input.all()), iterates over them with map, normalizes the phone, computes a Lead Score according to business criteria, and returns the enriched data to the next Nodes. Note that you must always return an array of objects with a field json.

2.3

AI Nodes in n8n

Deep Dive

n8n 1.x includes full support for LangChain directly in the interface. You can build AI Agents, Prompt chains, and conversations with Memory — all without a line of code.

AI Agent Node

A Node that runs a full Agent with Tools. Define the System Prompt, connect Tools (Webhooks, APIs, SQL), and it will automatically decide what to use.

Memory Nodes

Window Buffer Memory keeps the last N messages. Redis Chat Memory for long-term memory. Postgres/SQLite for permanent storage.

OpenAI Chat Model

A direct connection to GPT-4o, GPT-4 Turbo and GPT-3.5. Includes Temperature, Max Tokens, and System Message settings.

LangChain Integration

Every LangChain Node is available: Document Loaders, Text Splitters, Vector Store Retrievers, Embeddings, and more.

2.4

Security in n8n Self-Hosted

When you run n8n in production with access to sensitive business data, security is not optional. Here is the list of essential checks:

Basic Auth / OAuth2

Protect the n8n interface with Basic Auth. For multi-user environments — enable n8n User Management and move to OAuth2 with Google.

Credentials Encryption

Set N8N_ENCRYPTION_KEY — a long, random string. n8n will use it to encrypt all Credentials. Back up this key separately!

Firewall + IP Whitelist

If n8n does not need public access, restrict access to Port 5678 to your specific IP only. Keep Port 443 open for Webhooks only.

Automatic backup

Set up a daily Cron Job that pushes the ~/.n8n to S3 or Google Drive. A backup without a restore plan is not a backup.

🎯
Module 2 project

Lead Management Pipeline with AI Scoring

In this project you'll build a Workflow in n8n that handles incoming leads, passes them through a Code Node for processing and only then scores them with GPT-4o — and then routes them automatically to the right rep.

1
Webhook Trigger

Receiving a lead from Typeform/HubSpot Forms with all the form fields

2
Code Node — data normalization

Cleaning the phone, analyzing the email domain to identify the company, computing an initial Lead Score

3
OpenAI Node — AI Score

Send the inquiry content to GPT-4o: "Rate this lead 1-10 by level of seriousness" + structured JSON in the response

4
IF Node — routing by score

Score ≥ 8: senior rep + immediate WhatsApp. Score 5-7: regular email process. Score < 5: newsletter only

5
Postgres / Airtable

Save all the data including AI Score, a manual score to be set later, inquiry history

Modules 3–5

Click any module to navigate to its full content

4
Module 4 — 5 hours

10 Business Automation Projects

Ten full business projects ready for production. Each project includes documentation, downloadable code, and a Walk-through video. The projects cover the areas: E-commerce, customer service, HR, sales, marketing and content.

P1
Order-to-Delivery Tracker

WooCommerce + Shipping API + WhatsApp updates

P2
Customer Support Triage

Classifying customer inquiries with GPT-4 + handoff to a rep

P3
Social Media Content Engine

Weekly content creation with AI + publishing schedule

P4
CV Screening Pipeline

Reading resumes + scoring + a report to the manager

P5
Invoice Processing Bot

Extracting data from PDF invoices + recording in the ERP

P6
Competitor Intelligence

Scraping + AI Summary + an automatic weekly report

P7
Meeting Notes Automation

Zoom/Meet transcription + summary + tasks in Asana

P8
Inventory Alert System

Inventory check + automatic ordering + a report to the manager

P9
Personalized Email Sequences

A personalized email series with GPT-4 + A/B Testing

P10
Full Sales Dashboard

Collects CRM data + calculations + Google Data Studio

5
Module 5 — 2 hours

Production, Monitoring & Scale

Moving the automations from development to production professionally. Ongoing monitoring with Dashboards, error management, scaling and a course certificate.

Course completion certificate

After completing all 5 modules and 10 projects you'll receive a digital certificate from Automation4MI. Shareable on LinkedIn.

3
Module 3 — 5 hours

AI API Integration

3.1

OpenAI API from 0

Fundamentals

So far you have used AI through prepared nodes in Make and n8n. The node does three things for you: it holds the key, builds the request, and parses the response. The moment you call the API directly, all three move to you — and that is exactly what gives you control over cost, format and error handling.

The key, and where it must never be

A leaked key is a bill you pay until you revoke it. The three places keys actually leak: code pushed to GitHub, client-side code in a browser, and a screenshot in a group chat. One rule covers all of them — the key lives in an environment variable, and the file holding it is in .gitignore.

setup.sh
pip install openai python-dotenv

# .env — never enters git
echo 'OPENAI_API_KEY=sk-proj-...' > .env
echo '.env' >> .gitignore
first_call.py
import os
from dotenv import load_dotenv
from openai import OpenAI

load_dotenv()
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])

resp = client.chat.completions.create(
    model="gpt-5-mini",
    messages=[
        {"role": "system", "content": "You are terse. Two lines maximum."},
        {"role": "user",   "content": "What does a webhook do?"},
    ],
)

print(resp.choices[0].message.content)
print(resp.usage)   # prompt_tokens, completion_tokens, total_tokens

Note the last line. resp.usage is your bill, and it comes back on every call. People who never look at it discover their consumption at the end of the month.

Which model

The common mistake is reaching for the strongest model for every task. In most automations the majority of calls are simple — classify, extract a field, rewrite a sentence — and a small model does them at the same quality for a fraction of the cost and half the latency.

Small model

Classification, extraction, tagging, short rewriting, translation. That is 80% of the calls in a typical automation.

Large model

Reasoning over a long document, code, decisions that need judgement. Reach for it only when the small one fails.

Estimate cost before running it a thousand times

A token is roughly four characters of English — and fewer in most other languages, because the tokenizer is less efficient there. The same text in Hebrew can cost close to twice what it costs in English. Do not guess: measure one real example and multiply.

estimate.py
PRICE_IN  = 0.25 / 1_000_000    # $ per input token — check the price list
PRICE_OUT = 2.00 / 1_000_000    # $ per output token

u = resp.usage
per_call = u.prompt_tokens * PRICE_IN + u.completion_tokens * PRICE_OUT

for runs in (100, 1_000, 10_000):
    print(f"{runs:>6} runs: ${per_call * runs:.2f}")
Constant part first, variable part last

Most providers discount the portion of a prompt that repeats between calls, but only if it sits at the start of the request and does not change. So: system instructions, examples and reference documents at the top — the user's specific question at the bottom. It is purely a reordering, and it cuts the bill of anything running at volume.

Exercise

Take a real prompt from an automation you built in module 1 or 2, run it once through the API, and print usage. Work out what 1,000 runs a month costs. If the number surprises you, run the same input through a smaller model and compare both cost and quality.

3.2

Structured Outputs and Function Calling

Critical

This is the lesson that separates a demo from an automation that runs overnight. A model returning free text is a liability: today it answers {"status":"hot"}, tomorrow it prefixes "Sure! Here's the JSON:", and the day after it writes "Hot" with a capital H. The next node breaks, and nobody finds out until someone checks.

The fix is not asking more politely. It is enforcing a schema, at the API level.

structured.py
from pydantic import BaseModel, Field
from typing import Literal

class Lead(BaseModel):
    company:  str
    contact:  str
    intent:   Literal["hot", "warm", "cold"]
    budget:   int | None = Field(None, description="Budget if stated, else null")
    next_step: str

resp = client.chat.completions.parse(
    model="gpt-5-mini",
    messages=[
        {"role": "system", "content":
         "Extract lead details from the enquiry. Never invent a field "
         "that is not present."},
        {"role": "user", "content": email_body},
    ],
    response_format=Lead,          # schema enforced server-side
)

lead = resp.choices[0].message.parsed   # a Lead object, not a string
print(lead.intent, lead.budget)

Two details decide whether this survives production. First, Literal instead of str on a classification field — the model now cannot return a fourth category. Second, | None on anything optional. Without it the model will invent a budget to satisfy the schema, which is worse than an empty field because it looks like data.

Function calling — when the model decides what to do

Structured outputs hand you data. Function calling lets the model choose which action to take and with what arguments. The distinction matters: in the first you know in advance what comes back, in the second you are delegating a decision.

tools.py
tools = [{
    "type": "function",
    "function": {
        "name": "create_deal",
        "description": ("Open a new CRM deal. Use ONLY when both a company "
                        "name and a contact are present. Do not use for "
                        "general questions."),
        "parameters": {
            "type": "object",
            "properties": {
                "company": {"type": "string"},
                "amount":  {"type": "number"},
                "stage":   {"type": "string",
                            "enum": ["lead", "proposal", "negotiation"]},
            },
            "required": ["company", "stage"],
            "additionalProperties": False,
        },
    },
}]

resp = client.chat.completions.create(
    model="gpt-5-mini", messages=messages, tools=tools,
)

calls = resp.choices[0].message.tool_calls
if calls:
    args = json.loads(calls[0].function.arguments)
    create_deal(**args)          # your code executes. the model only chose.
The description is the prompt

The model picks a tool by its description, not its name. A lazy description — "creates a deal" — produces calls in the wrong places. A description that also says when not to use it improves accuracy more than any change to the system prompt.

Retry — and what not to retry

API errors split into two kinds and the handling is opposite. 429 and 5xx are transient — back off and try again. 400 and 401 are your bug — retrying only burns calls and bills you for them.

retry.py
import time, random
from openai import RateLimitError, APIStatusError

def call_with_retry(fn, attempts=4):
    for i in range(attempts):
        try:
            return fn()
        except RateLimitError:
            pass                                  # transient — worth retrying
        except APIStatusError as e:
            if e.status_code < 500:
                raise                             # ours — do not retry
        wait = 2 ** i + random.random()           # jitter avoids a thundering herd
        print(f"attempt {i+1} failed, waiting {wait:.1f}s")
        time.sleep(wait)
    raise RuntimeError("failed after all attempts")

The random() in the backoff is not decoration. Without it, a hundred calls that failed together retry together at exactly the same instant, and fail together again.

3.3

Vision API — image processing

Quick Win

This is the fastest payback in the module. Invoices, receipts, order screenshots, delivery notes — all of it arrives at a business as images, and somebody retypes it by hand. A vision model reads them and returns JSON.

Two ways to send an image

A URL if the image is already on the web and reachable without auth. Base64 if it is yours — a local file, a file from Drive, or an image that arrived on a webhook. The second is the common case in automation.

invoice_ocr.py
import base64
from pydantic import BaseModel

def to_data_url(path: str) -> str:
    mime = "image/png" if path.endswith(".png") else "image/jpeg"
    with open(path, "rb") as f:
        return f"data:{mime};base64," + base64.b64encode(f.read()).decode()

class Invoice(BaseModel):
    supplier:   str
    invoice_no: str
    date:       str          # YYYY-MM-DD
    total:      float
    tax:        float | None
    line_items: list[str]

resp = client.chat.completions.parse(
    model="gpt-5-mini",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text":
             "Extract the invoice fields. If a field is not clearly present, "
             "return null. Never estimate a number."},
            {"type": "image_url",
             "image_url": {"url": to_data_url("invoice.jpg"), "detail": "high"}},
        ],
    }],
    response_format=Invoice,
)

inv = resp.choices[0].message.parsed
Documents in Hebrew and other RTL scripts

Three things break, and no English-language guide covers them. Digits inside RTL text: an invoice number embedded in Hebrew can come back with its digit order flipped — always cross-check it against another field. Dates: the local format is dd/mm/yyyy, which is ambiguous to a model trained mostly on US documents; ask for YYYY-MM-DD explicitly and state the source order. Tax: Israeli invoices sometimes show VAT as its own line and sometimes fold it into the total — ask for both and reconcile.

detail — the parameter that changes cost several-fold

low

A small fixed token count. Enough for "what is in this image", for classification, for identifying a document type. Substantially cheaper.

high

Tiles the image. Required to read small print — invoices, forms. Cost scales with image size.

A flow that saves money: run low first just to classify what the image is, and only run high if it turns out to be an invoice. On a mailbox that also receives spam, that skips most of the expensive calls.

A warning that has to be said

Do not feed extraction output straight into an accounting system. The error rate is not zero, and on monetary figures a single mistake costs more than the automation saved. The correct flow is: extract → row in a sheet with status "pending review" → a human approves → then continue.

3.4

Embeddings and Vector Search

Keyword search finds words. Semantic search finds meaning. The query "how do I cancel an order" will match a document titled "Returns policy" without sharing a single word — and that is precisely what you need to build a chatbot over company documents.

The mechanism: an embedding converts text into a list of numbers that represents its meaning. Two texts with similar meaning get lists that sit close together. Search is measuring that closeness.

embed.py
def embed(texts: list[str]) -> list[list[float]]:
    # batch it — one call for 100 chunks, not 100 calls
    r = client.embeddings.create(
        model="text-embedding-3-small",
        input=texts,
    )
    return [d.embedding for d in r.data]

vecs = embed(["Returns policy", "Opening hours", "Product warranty"])
print(len(vecs[0]))     # vector dimensions

Pinecone — storing and querying

pinecone_store.py
from pinecone import Pinecone

pc = Pinecone(api_key=os.environ["PINECONE_API_KEY"])
index = pc.Index("company-docs")

# store the text in metadata, otherwise you have numbers with no source
index.upsert([
    {"id": f"doc-{i}",
     "values": v,
     "metadata": {"text": t, "source": "policy.pdf", "page": i}}
    for i, (t, v) in enumerate(zip(chunks, embed(chunks)))
])

q = embed(["how do I cancel an order?"])[0]
hits = index.query(vector=q, top_k=5, include_metadata=True)

for h in hits.matches:
    print(round(h.score, 3), h.metadata["source"], h.metadata["text"][:60])
Two traps that will cost you a day

One: the embedding model used for writing and for querying must be identical. Vectors from different models are not comparable, and the system will return plausible-looking nonsense without raising an error — the hardest failure here to diagnose. Two: always keep the original text in metadata. Without it you have a similarity score and nothing to feed the model in the next step.

Non-English content

Modern embedding models are multilingual and handle Hebrew reasonably, but two things are worth knowing. First, do not mix languages in one index without a reason — a Hebrew and an English document on the same subject will not necessarily land near each other. Second, test it for real: take five actual customer questions and run the search. If the top hit is not the right document, the problem is almost always your chunking rather than the model — which is exactly the next lesson.

3.5

Full RAG Pipeline

Module peak

RAG is four steps: split documents into chunks, turn them into vectors, retrieve the ones relevant to a question, and hand them to the model as context. You already have the first three from the previous lesson. This lesson connects them — and dwells on the step that decides whether any of it works.

Chunking — where most systems fail

Chunking is the highest-impact decision in the pipeline and gets the least attention. A chunk that is too large carries noise that dilutes relevance; too small and it loses context and returns a stranded sentence. The working rule: a chunk should answer one question completely.

chunk.py
def chunk(text: str, size=900, overlap=150) -> list[str]:
    """Split on paragraph boundaries, never mid-sentence."""
    paras = [p.strip() for p in text.split("\n\n") if p.strip()]
    out, cur = [], ""
    for p in paras:
        if len(cur) + len(p) < size:
            cur += ("\n\n" if cur else "") + p
        else:
            if cur:
                out.append(cur)
            # overlap: the tail of the previous chunk opens the next
            cur = (cur[-overlap:] + "\n\n" + p) if cur else p
    if cur:
        out.append(cur)
    return out

The overlap exists so that an answer sitting on the seam between two chunks is not cut in half. Without it, a question whose answer starts at the end of one chunk and finishes in the next is fully present in neither.

The whole pipeline

rag.py
SYSTEM = """Answer only from the sources below.
If the answer is not in them, say "I don't have that in the documents".
Do not fill gaps from your general knowledge.
Cite the source in brackets after each claim."""

def ask(question: str, top_k: int = 5) -> str:
    q_vec = embed([question])[0]
    hits  = index.query(vector=q_vec, top_k=top_k, include_metadata=True)

    # relevance floor — better to not answer than to answer from noise
    good = [h for h in hits.matches if h.score > 0.35]
    if not good:
        return "I don't have that in the documents."

    context = "\n\n---\n\n".join(
        f"[{h.metadata['source']}]\n{h.metadata['text']}" for h in good
    )

    r = client.chat.completions.create(
        model="gpt-5-mini",
        messages=[
            {"role": "system", "content": SYSTEM},
            {"role": "user",
             "content": f"Sources:\n{context}\n\nQuestion: {question}"},
        ],
    )
    return r.choices[0].message.content

Three lines in that code are the entire difference between a system you can trust and one that invents. The grounding instruction, the relevance floor that prefers "I don't know" over answering from irrelevant chunks, and the citation requirement — which makes every answer checkable by the person reading it.

When answers are bad, debug in this order

Do not touch the prompt first. Print what was retrieved before it reaches the model. If the right chunks are not in there, the problem is retrieval — chunking or the embedding model. If they are in there and the answer is still wrong, then it is the prompt. Most people spend an hour on prompt wording and discover retrieval never returned the right document at all.

What we did not cover, and should be on your radar

Basic RAG is enough for most cases. When it is not, the next two upgrades are reranking — retrieve 20 chunks and reorder them with a dedicated model before sending five — and hybrid search, combining semantic with keyword search to catch model numbers and exact terms that semantic search misses.

🎯
Module 3 project

A chatbot with custom knowledge

You will build a chatbot that answers customer questions from company documents only — and admits when it has no answer. It is exposed through a Make.com webhook, so you can wire it to a form on the site, to WhatsApp, or to an internal chat.

What you build

1. Indexing (runs once, and on every document update)

Read the policy files, chunk them, embed, write to Pinecone with source and page in metadata.

2. Query (runs on every question)

A webhook takes a question, retrieves chunks above the floor, produces a cited answer, returns JSON.

3. The gap log

Every question the system could not answer is written to a sheet. That is your list of missing documents.

4. Wiring in Make

Webhook → HTTP to your service → Router on answered/unanswered → reply to the customer or alert the team.

app.py — the endpoint
from fastapi import FastAPI
from pydantic import BaseModel

app = FastAPI()

class Q(BaseModel):
    question: str
    user_id: str | None = None

@app.post("/ask")
def ask_endpoint(q: Q):
    answer = ask(q.question)
    unanswered = answer.startswith("I don't have that")

    if unanswered:
        log_gap(q.question, q.user_id)   # to the sheet — the gap to close

    return {"answer": answer, "answered": not unanswered}

Acceptance criteria

Ten real customer questions — at least eight answered correctly, citing the right source.
A question with no answer in the documents returns "I don't have that" rather than an invention. This is the most important criterion of the four.
Every answer names the document it came from, so the customer can verify it.
Unanswered questions accumulate in a sheet with a timestamp.
Before you expose this to customers

A chatbot reading internal documents has two exposures. First: confirm every document you indexed is genuinely meant for customer eyes — an internal policy file that drifted into the index will surface in an answer. Second: if documents come from somewhere users can write to, text embedded in them can steer the model. Run it over documents you control.

5
Module 5 — 2 hours

Production, Monitoring & Scale

5.1

Production Deployment Checklist

An automation that works on your machine and one that works in production are two different things. The difference is not the code — it is what happens when something changes: a rotated key, a service that goes down, load that grows, or someone editing the flow by accident.

Secrets do not belong in the flow

The worst thing you can do is paste an API key into an HTTP node. It is then stored in the flow backup, present in every export, visible in a screenshot, and travels to anyone who receives a copy. All three platforms have a proper place for it:

n8n

Credentials, encrypted with N8N_ENCRYPTION_KEY. Back that key up separately — without it your backup is worthless.

Make.com

Connections at the account level. Never a key in the text field of an HTTP module.

Your own code

Environment variables only. .env in .gitignore, production secrets in your deploy platform's secret store.

Rotation

A leaked key gets rotated, not quietly deleted. And document where it was used, or you will find out from a crash.

Staging is not a luxury

An automation that touches real customers is not tested on real customers. The minimum separation that works: a copy of the flow with a single environment variable that swaps the destinations — a test sheet instead of the real one, a private Slack channel instead of the team's, your inbox instead of the customer's.

config.py
import os

ENV = os.getenv("APP_ENV", "staging")   # safe default

TARGETS = {
    "staging":    {"sheet": "leads-test", "slack": "#bot-test",
                   "notify": "me@example.com"},
    "production": {"sheet": "leads",      "slack": "#sales",
                   "notify": "sales@example.com"},
}[ENV]

if ENV == "production" and not os.getenv("I_MEAN_IT"):
    raise SystemExit("production requires an explicit acknowledgement")

Note the default. If the variable is missing, the system runs against staging — not production. The other way round is a bug that surfaces late and expensively.

Pre-deploy checklist

Every secret is in credentials or an environment variable, and none is in the flow itself.
The flow is exported and stored — in git if possible. n8n exports JSON, which lets you diff changes and roll back.
The failure path is tested, not just the happy path. Break the API deliberately and watch what happens.
A failure alert exists. Without it the rest of the list is pointless.
Anything that writes or sends passes a duplicate check — a unique key, and a lookup before the write.
CI/CD — the minimum version

You do not need a full pipeline. Two things cover most automations: flows kept in git so you can see what changed and when, and an automated validation step before deploy — a short script that checks the file parses and the required fields exist. Deploy without a check step and your users find the problem for you.

5.2

Monitoring & Alerting

Most important

This is the last lesson in the course and the most important one. Not because monitoring is interesting, but because the dangerous automation is the one that fails quietly. A flow that crashes and shouts gets fixed within the hour. A flow that stops running without saying anything keeps "working" in everyone's mind, and surfaces when a customer asks why nobody got back to them.

A real incident from this site's own systems

The Automation4MI payment pipeline had a step designed to verify a secret was configured, and to fail if it was not. The secret was never set, the step failed, and the deploy stopped before reaching the code that ships. The result: production froze on an old build for hours while the site was taking real money, because nobody was reading the run log. The lesson is not to write fewer checks — it is that a step that fails silently is worse than one that does not exist. Every failure has to land somewhere a person will see it.

Three layers, in order of importance

1. Failure alert

The flow errored — notify immediately. Five minutes of work, and 80% of the value of all monitoring.

2. Heartbeat

The flow did not run at all — which raises no error. The layer most people skip, and the one that catches silent failure.

3. Metrics

How often it ran, how often it failed, how long it took, what it cost. For spotting trends, not for reacting.

What not to do

Alert on everything. A team receiving forty alerts a day stops reading them, and the important one gets buried with the rest.

Error Workflow in n8n

n8n lets you define one flow that runs whenever another flow fails. Build it once, then point every flow at it from settings.

Error Trigger → Code node
const e = $input.first().json;

return [{ json: {
  text:
    `🔴 *${e.workflow.name}* failed\n` +
    `node: ${e.execution.lastNodeExecuted}\n` +
    `error: ${e.execution.error?.message ?? 'unknown'}\n` +
    `<${e.execution.url}|open the execution>`
}}];

Those three fields — which flow, which node, what error — are the difference between an alert you can act on and "something failed". The execution link saves the hunt.

Heartbeat — catching what does not shout

A scheduled flow that got disabled, a server that went down, a trigger that disconnected — none of these produce an error, because nothing ran. The fix inverts it: the flow reports that it is alive, and an external service alerts when the reports stop arriving.

at the end of every scheduled flow
# HTTP Request node, last in the flow, after success
GET https://uptime.example.com/api/push/<token>

# Uptime Kuma configured as a "Push" monitor with a 65-minute
# heartbeat interval for a flow that runs hourly.
# No report arrives → alert.
Set the window wider than the interval

A flow that runs hourly wants a 65–70 minute window, not 60. A run that slipped by a minute produces a false alarm, and two of those teach the team to ignore the channel. Better to alert five minutes late than to alert when nothing is wrong.

What to measure — and what not to

Once alerting is in place, four numbers are worth accumulating in a sheet or a dashboard: runs per day, failure rate, median duration, and API cost. Those four surface trends: a failure rate creeping upward, a duration that has doubled, or a cost that jumped after a prompt change.

What is not worth measuring: everything else. A dashboard with twenty charts does not get looked at, and what does not get looked at is not monitoring.

The last exercise in the course

Take one automation you built in an earlier module and add the first two layers: a failure alert and a heartbeat. Then break it on purpose. Change an API key to a wrong value and confirm the alert arrives. Disable the flow entirely and confirm the heartbeat fires. Monitoring that has never been tested is an assumption, not a defence.

You finished Module 1

Ready for the next project?

Modules 1 through 5 are open in full — from advanced Make.com, through self-hosted n8n and calling model APIs directly, to shipping into production with monitoring. The ten projects are waiting in module 4.

To the ten projects
The private Discord community

Connect to an exclusive community of AI Automation Pro students. Ask questions, share projects, get Feedback from mentors and senior students.