Skip to content

Latest commit

 

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

RealTime Collaborative Editor

A multi-user, real-time collaborative code editor built with React, Monaco Editor, Yjs (CRDT) and Socket.IO. Multiple people open the same document, type at the same time, and every keystroke is merged conflict-free and broadcast to everyone else — along with a live list of who is currently connected.


Table of Contents


Features

  • Conflict-free concurrent editing — powered by Yjs CRDTs, so simultaneous edits from any number of users converge without locks or last-write-wins data loss.
  • Full-featured code editor — Monaco (the editor that powers VS Code) with syntax highlighting, bracket matching and a dark theme.
  • Live presence — an awareness-backed sidebar showing every user currently connected, updated on join and disconnect.
  • Username-based join flow — users pick a display name, which is persisted in the URL query string (?username=...) so a session can be shared or reloaded.
  • Single-origin deployment — the built React bundle is served as static assets by the Express server, so the app runs as one container on one port with no CORS or cross-origin WebSocket setup.
  • Container-ready — a multi-stage Dockerfile produces a single lean runtime image.

Tech Stack

Layer Technology
UI React 19, Tailwind CSS 4, Vite 7
Editor Monaco Editor (@monaco-editor/react)
CRDT / Sync Yjs, y-monaco, y-socket.io
Transport Socket.IO 4 (WebSocket with HTTP long-polling fallback)
Server Node.js 20, Express 5
Packaging Docker (multi-stage build)
Cloud AWS ECR + ECS Fargate + Application Load Balancer

Architecture

                        ┌───────────────────────────────────────┐
   Browser A            │           Node.js Container           │
 ┌───────────────┐      │                                       │
 │ Monaco Editor │      │  ┌─────────────────────────────────┐  │
 │      ▲        │      │  │ Express 5                       │  │
 │      │        │      │  │  • static /public (React build) │  │
 │ MonacoBinding │      │  │  • GET /health                  │  │
 │      ▲        │      │  └─────────────────────────────────┘  │
 │   Y.Doc       │◄────►│  ┌─────────────────────────────────┐  │
 │      ▲        │  WS  │  │ Socket.IO Server                │  │
 │ SocketIOProv. │      │  │   └── YSocketIO                 │  │
 └───────────────┘      │  │        • doc sync (room "monaco")│ │
                        │  │        • awareness / presence   │  │
   Browser B ◄─────────►│  └─────────────────────────────────┘  │
                        └───────────────────────────────────────┘

How a keystroke travels:

  1. The user types in Monaco. MonacoBinding translates the editor delta into a Yjs transaction on the shared Y.Doc (ydoc.getText("monaco")).
  2. SocketIOProvider serialises the Yjs update and emits it over the Socket.IO connection.
  3. YSocketIO on the server applies the update to the server-side document for that room and fans it out to every other connected client.
  4. Each peer applies the remote update to its own Y.Doc; MonacoBinding reflects it in the editor. Because Yjs is a CRDT, arrival order does not matter — all replicas converge to the same state.
  5. Awareness (presence) rides the same connection on a separate ephemeral channel. Each client publishes { user: { username } }, and the change event rebuilds the users sidebar.

The client connects with new SocketIOProvider("/", "monaco", ydoc) — a relative origin. This is deliberate: the frontend is served by the same Express process, so the app works identically on localhost:3000, in Docker, and behind an ALB without any environment-specific URL configuration.


Project Structure

RealTime-Collaborative-Editor/
├── client/                     # React + Vite frontend
│   ├── src/
│   │   ├── app/
│   │   │   ├── App.jsx         # Join screen, editor, presence sidebar, Yjs wiring
│   │   │   └── App.css
│   │   └── main.jsx            # React entry point
│   ├── index.html
│   ├── vite.config.js          # React + Tailwind plugins
│   └── package.json
├── server/                     # Express + Socket.IO backend
│   ├── public/                 # Built frontend is copied here at image build time
│   ├── server.js               # HTTP server, static hosting, YSocketIO bootstrap
│   └── package.json
├── dockerfile                  # Multi-stage build (frontend builder → runtime)
├── .dockerignore
└── README.md

Getting Started (Local)

Prerequisites

  • Node.js 20+ and npm
  • Docker (optional, for the containerised workflow)

1. Clone

git clone https://github.com/Naman501/RealTime-Collaborative-Editor.git
cd RealTime-Collaborative-Editor

2. Start the backend

cd server
npm install
npm run dev          # nodemon, restarts on change
# or: npm start      # plain node

The server listens on http://localhost:3000 and exposes both the Socket.IO endpoint and GET /health.

3. Start the frontend

In a second terminal:

cd client
npm install
npm run dev

Vite serves the app on http://localhost:5173.

Note on the dev proxy. In production the client connects to "/" because Express serves the bundle. During development the Vite dev server runs on a different port, so add a proxy to client/vite.config.js (or point SocketIOProvider at http://localhost:3000):

export default defineConfig({
  plugins: [react(), tailwindcss()],
  server: {
    proxy: {
      "/socket.io": { target: "http://localhost:3000", ws: true },
    },
  },
})

4. Try it

Open two browser windows at http://localhost:5173, join with different usernames, and type in either editor. Edits and the users list stay in sync.

5. Production-style local run

Build the frontend into the server's static directory and run a single process:

cd client && npm run build

# macOS / Linux
cp -r dist/* ../server/public/

# Windows (PowerShell)
Copy-Item -Recurse -Force dist\* ..\server\public\

cd ../server && npm start

Then open http://localhost:3000.


Available Scripts

client/

Script Description
npm run dev Vite dev server with HMR
npm run build Production bundle into client/dist
npm run preview Serve the built bundle locally
npm run lint ESLint across the client

server/

Script Description
npm run dev Start with nodemon
npm start Start with plain node

Endpoints

Method Path Description
GET / Serves the built React application from server/public
GET /health Health probe — returns { "message": "Ok", "success": true } with HTTP 200. Used by the ALB target group and container health checks.
WS /socket.io/ Socket.IO endpoint (WebSocket, with HTTP long-polling fallback) carrying Yjs document sync and awareness traffic

Docker

The repository ships a multi-stage dockerfile:

  1. Stage 1 (frontend-builder) — node:20-alpine, installs the client dependencies and runs vite build.
  2. Stage 2 (runtime) — node:20-alpine, installs only the server dependencies and copies /app/dist from stage 1 into /app/public, which Express serves statically.

The frontend build toolchain never reaches the final image, so the runtime layer stays small.

Build

docker build -f dockerfile -t realtime-collab-editor:latest .

The -f flag is needed on case-sensitive systems because the file is named dockerfile (lowercase) rather than Dockerfile.

Run

docker run --rm -p 3000:3000 --name editor realtime-collab-editor:latest

Open http://localhost:3000.

Recommended Dockerfile hardening

The current Dockerfile is intentionally minimal. For production, consider adding to the runtime stage:

ENV NODE_ENV=production
EXPOSE 3000
RUN npm ci --omit=dev            # deterministic, prod-only installs
USER node                        # drop root privileges
HEALTHCHECK --interval=30s --timeout=3s --start-period=10s \
  CMD wget -qO- http://localhost:3000/health || exit 1

Also copy package*.json before the rest of the source in each stage so Docker's layer cache is not invalidated on every code change.

Multi-architecture builds

Fargate on Graviton (ARM64) is cheaper. If you build on an x86 machine for an ARM64 task, build for the target platform explicitly:

docker buildx build --platform linux/arm64 -t realtime-collab-editor:latest --push .

Mismatched architectures are the most common cause of exec format error in ECS.


Deploying to AWS

The reference deployment is ECR → ECS Fargate → Application Load Balancer, which is the natural fit for a single long-lived container serving both HTTP and WebSocket traffic.

Why not S3 + CloudFront for the frontend? You can split them, but this app is designed to be served from the same origin as the WebSocket endpoint, which removes CORS handling and origin configuration entirely. Serving both from ECS keeps the deployment to one moving part.

Why not API Gateway / Lambda? Yjs document state is held in server memory for the lifetime of a room. That requires a persistent process, which rules out Lambda for the sync layer.

1. Prerequisites

  • AWS account with the AWS CLI v2 installed and configured (aws configure)
  • A VPC with at least two public subnets in different Availability Zones (the ALB requires two AZs) and two private subnets for the tasks
  • Docker with buildx
  • Optionally: a registered domain in Route 53 and an ACM certificate for HTTPS

Set the shell variables used throughout:

export AWS_REGION=ap-south-1
export ACCOUNT_ID=$(aws sts get-caller-identity --query Account --output text)
export APP=realtime-collab-editor
export ECR_URI=$ACCOUNT_ID.dkr.ecr.$AWS_REGION.amazonaws.com/$APP

2. IAM Permissions

Three distinct identities are involved. Keeping them separate is what makes the deployment least-privilege.

a) Deployer identity (you, or your CI runner)

The principal that builds and ships the image and updates the service. Attach a customer-managed policy:

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "EcrPush",
      "Effect": "Allow",
      "Action": [
        "ecr:GetAuthorizationToken",
        "ecr:BatchCheckLayerAvailability",
        "ecr:CompleteLayerUpload",
        "ecr:InitiateLayerUpload",
        "ecr:PutImage",
        "ecr:UploadLayerPart",
        "ecr:DescribeRepositories",
        "ecr:DescribeImages"
      ],
      "Resource": "*"
    },
    {
      "Sid": "EcsDeploy",
      "Effect": "Allow",
      "Action": [
        "ecs:RegisterTaskDefinition",
        "ecs:DescribeTaskDefinition",
        "ecs:UpdateService",
        "ecs:DescribeServices",
        "ecs:DescribeClusters",
        "ecs:ListTasks",
        "ecs:DescribeTasks"
      ],
      "Resource": "*"
    },
    {
      "Sid": "PassRolesToEcsOnly",
      "Effect": "Allow",
      "Action": "iam:PassRole",
      "Resource": [
        "arn:aws:iam::ACCOUNT_ID:role/ecsTaskExecutionRole",
        "arn:aws:iam::ACCOUNT_ID:role/realtimeEditorTaskRole"
      ],
      "Condition": {
        "StringEquals": { "iam:PassedToService": "ecs-tasks.amazonaws.com" }
      }
    }
  ]
}

ecr:GetAuthorizationToken must be scoped to "*" — it is an account-level action and does not accept a repository ARN. The other ECR actions can be narrowed to arn:aws:ecr:REGION:ACCOUNT_ID:repository/realtime-collab-editor once the pipeline works.

The PassRole condition is the important guardrail: without it, anyone who can deploy could attach any role in the account to a task.

For CI (GitHub Actions), prefer OIDC federation over long-lived access keys — create an IAM OIDC identity provider for token.actions.githubusercontent.com and a role whose trust policy is restricted to your repository and branch.

b) ECS Task Execution Role (ecsTaskExecutionRole)

Used by the ECS agent, not by your code — it pulls the image and writes the log stream.

  • Trust policy principal: ecs-tasks.amazonaws.com
  • Attach the AWS-managed policy AmazonECSTaskExecutionRolePolicy, which grants:
    • ecr:GetAuthorizationToken, ecr:BatchCheckLayerAvailability, ecr:GetDownloadUrlForLayer, ecr:BatchGetImage
    • logs:CreateLogStream, logs:PutLogEvents
  • Add secretsmanager:GetSecretValue / ssm:GetParameters only if you inject secrets into the task definition.
aws iam create-role --role-name ecsTaskExecutionRole \
  --assume-role-policy-document '{
    "Version":"2012-10-17",
    "Statement":[{"Effect":"Allow","Principal":{"Service":"ecs-tasks.amazonaws.com"},"Action":"sts:AssumeRole"}]
  }'

aws iam attach-role-policy --role-name ecsTaskExecutionRole \
  --policy-arn arn:aws:iam::aws:policy/service-role/AmazonECSTaskExecutionRolePolicy

c) ECS Task Role (realtimeEditorTaskRole)

The role your application code assumes at runtime. This app currently calls no AWS APIs, so create it with no attached policies — it exists so that adding S3 document persistence or ElastiCache later does not require re-plumbing the task definition. Attaching nothing is the correct least-privilege posture today.

If you enable ECS Exec for shell access into a running task, add to the task role:

{
  "Effect": "Allow",
  "Action": [
    "ssmmessages:CreateControlChannel",
    "ssmmessages:CreateDataChannel",
    "ssmmessages:OpenControlChannel",
    "ssmmessages:OpenDataChannel"
  ],
  "Resource": "*"
}

d) Service-linked role

ECS needs AWSServiceRoleForECS to register targets with the ALB. It is created automatically on first cluster creation; in a locked-down account, create it explicitly:

aws iam create-service-linked-role --aws-service-name ecs.amazonaws.com

3. Push the Image to Amazon ECR

# Create the repository (once), with image scanning enabled
aws ecr create-repository \
  --repository-name $APP \
  --image-scanning-configuration scanOnPush=true \
  --image-tag-mutability IMMUTABLE \
  --region $AWS_REGION

# Authenticate Docker to ECR
aws ecr get-login-password --region $AWS_REGION \
  | docker login --username AWS --password-stdin $ACCOUNT_ID.dkr.ecr.$AWS_REGION.amazonaws.com

# Build, tag and push
SHA=$(git rev-parse --short HEAD)
docker build -f dockerfile -t $APP:$SHA .
docker tag $APP:$SHA $ECR_URI:$SHA
docker push $ECR_URI:$SHA

Tag images with the git SHA rather than latest — immutable tags make rollbacks a one-line task-definition change and make it unambiguous which build is live. Add an ECR lifecycle policy to expire untagged images after ~14 days so storage costs stay flat.


4. Networking (VPC, Subnets, Security Groups)

Layout

Component Placement
Application Load Balancer Public subnets, ≥ 2 AZs
ECS Fargate tasks Private subnets, ≥ 2 AZs
NAT Gateway Public subnet — needed for private tasks to pull from ECR (or use VPC endpoints)

Security groups — two of them, chained so the tasks are reachable only through the load balancer:

SG Inbound Outbound
alb-sg TCP 80 and 443 from 0.0.0.0/0 TCP 3000 to app-sg
app-sg TCP 3000 from alb-sg only (source = security group, not a CIDR) All (for ECR pulls, CloudWatch)
aws ec2 authorize-security-group-ingress \
  --group-id $APP_SG --protocol tcp --port 3000 --source-group $ALB_SG

Referencing alb-sg as the source instead of a CIDR block means the task port is never reachable from the internet, even if a task briefly lands in a public subnet.

Cost note: to avoid the NAT Gateway hourly charge, place tasks in private subnets with VPC interface endpoints for ecr.api, ecr.dkr and logs, plus an S3 gateway endpoint (ECR layers are stored in S3). For a small project, a single NAT Gateway — or public subnets with assignPublicIp=ENABLED — are both acceptable simplifications.


5. Application Load Balancer (ALB)

The ALB needs the most care here, because this app is WebSocket-first.

Create the target group

aws elbv2 create-target-group \
  --name realtime-editor-tg \
  --protocol HTTP --port 3000 \
  --vpc-id $VPC_ID \
  --target-type ip \
  --health-check-protocol HTTP \
  --health-check-path /health \
  --health-check-interval-seconds 30 \
  --health-check-timeout-seconds 5 \
  --healthy-threshold-count 2 \
  --unhealthy-threshold-count 3 \
  --matcher HttpCode=200
  • --target-type ip is required for Fargate (awsvpc networking); instance targets will not register.
  • /health is the endpoint defined in server.js and returns 200 with a JSON body, which matches the default success matcher.

Critical target group attributes

aws elbv2 modify-target-group-attributes \
  --target-group-arn $TG_ARN \
  --attributes \
    Key=stickiness.enabled,Value=true \
    Key=stickiness.type,Value=lb_cookie \
    Key=stickiness.lb_cookie.duration_seconds,Value=86400 \
    Key=deregistration_delay.timeout_seconds,Value=60 \
    Key=load_balancing.algorithm.type,Value=least_outstanding_requests

Why stickiness is mandatory. Socket.IO's HTTP long-polling fallback performs a handshake followed by further requests carrying a session id. If those land on a different task, the server returns 400 {"code":1,"message":"Session ID unknown"} and the client enters a reconnect loop. Sticky sessions pin a client to one task for the whole session. This matters as soon as you run more than one task.

deregistration_delay — the default is 300 s. Lowering it to 60 s speeds up deployments; raising it gives in-flight WebSocket connections longer to drain. Choose based on how disruptive a mid-edit disconnect is for your users.

Create the load balancer and listeners

aws elbv2 create-load-balancer \
  --name realtime-editor-alb \
  --type application --scheme internet-facing \
  --subnets $PUBLIC_SUBNET_A $PUBLIC_SUBNET_B \
  --security-groups $ALB_SG

# Redirect all HTTP to HTTPS
aws elbv2 create-listener --load-balancer-arn $ALB_ARN \
  --protocol HTTP --port 80 \
  --default-actions '[{"Type":"redirect","RedirectConfig":{"Protocol":"HTTPS","Port":"443","StatusCode":"HTTP_301"}}]'

# HTTPS listener terminating TLS with an ACM certificate
aws elbv2 create-listener --load-balancer-arn $ALB_ARN \
  --protocol HTTPS --port 443 \
  --certificates CertificateArn=$ACM_CERT_ARN \
  --ssl-policy ELBSecurityPolicy-TLS13-1-2-2021-06 \
  --default-actions Type=forward,TargetGroupArn=$TG_ARN

WebSocket-specific ALB settings

Setting Value Reason
Idle timeout 300 seconds (default 60) The ALB closes connections idle for longer than this. Socket.IO pings roughly every 25 s so 60 s technically survives, but a longer window absorbs network stalls and browser tab throttling.
HTTP/2 to targets Leave the target group protocol version at HTTP/1.1 The WebSocket Upgrade handshake is an HTTP/1.1 mechanism; an HTTP/2 target group breaks it.
Cross-zone load balancing Enabled (default on ALB) Even task distribution across AZs.
Access logs Enabled, to an S3 bucket The only reliable way to debug 502/504s at the edge.
aws elbv2 modify-load-balancer-attributes \
  --load-balancer-arn $ALB_ARN \
  --attributes \
    Key=idle_timeout.timeout_seconds,Value=300 \
    Key=routing.http.drop_invalid_header_fields.enabled,Value=true \
    Key=access_logs.s3.enabled,Value=true \
    Key=access_logs.s3.bucket,Value=$LOG_BUCKET

ALB supports the WebSocket Upgrade handshake natively — no special listener rule, protocol setting, or annotation is needed. It simply must not be fronted by anything that strips the Upgrade and Connection headers.

Optional: path-based routing

If you later split traffic across target groups, note that /socket.io/* must reach the same target group as /, or sync breaks:

aws elbv2 create-rule --listener-arn $HTTPS_LISTENER_ARN --priority 10 \
  --conditions Field=path-pattern,Values='/socket.io/*' \
  --actions Type=forward,TargetGroupArn=$TG_ARN

6. ECS Fargate Service

Task definition (task-definition.json)

{
  "family": "realtime-collab-editor",
  "networkMode": "awsvpc",
  "requiresCompatibilities": ["FARGATE"],
  "cpu": "512",
  "memory": "1024",
  "runtimePlatform": {
    "cpuArchitecture": "X86_64",
    "operatingSystemFamily": "LINUX"
  },
  "executionRoleArn": "arn:aws:iam::ACCOUNT_ID:role/ecsTaskExecutionRole",
  "taskRoleArn": "arn:aws:iam::ACCOUNT_ID:role/realtimeEditorTaskRole",
  "containerDefinitions": [
    {
      "name": "app",
      "image": "ACCOUNT_ID.dkr.ecr.REGION.amazonaws.com/realtime-collab-editor:GIT_SHA",
      "essential": true,
      "portMappings": [{ "containerPort": 3000, "protocol": "tcp" }],
      "environment": [{ "name": "NODE_ENV", "value": "production" }],
      "healthCheck": {
        "command": ["CMD-SHELL", "wget -qO- http://localhost:3000/health || exit 1"],
        "interval": 30,
        "timeout": 5,
        "retries": 3,
        "startPeriod": 15
      },
      "logConfiguration": {
        "logDriver": "awslogs",
        "options": {
          "awslogs-group": "/ecs/realtime-collab-editor",
          "awslogs-region": "REGION",
          "awslogs-stream-prefix": "ecs",
          "awslogs-create-group": "true"
        }
      }
    }
  ]
}

containerPort is 3000 because server.js hard-codes httpServer.listen(3000). If you make the port configurable via process.env.PORT, keep the task definition, target group and security group in agreement.

Register and create the service

aws ecs register-task-definition --cli-input-json file://task-definition.json

aws ecs create-cluster --cluster-name realtime-editor-cluster

aws ecs create-service \
  --cluster realtime-editor-cluster \
  --service-name realtime-editor-svc \
  --task-definition realtime-collab-editor \
  --desired-count 1 \
  --launch-type FARGATE \
  --network-configuration "awsvpcConfiguration={subnets=[$PRIV_A,$PRIV_B],securityGroups=[$APP_SG],assignPublicIp=DISABLED}" \
  --load-balancers "targetGroupArn=$TG_ARN,containerName=app,containerPort=3000" \
  --health-check-grace-period-seconds 60 \
  --deployment-configuration "maximumPercent=200,minimumHealthyPercent=100"

--health-check-grace-period-seconds prevents the ALB from killing a task that is still booting — without it, a slow cold start can cause a restart loop.

Deploying a new version

SHA=$(git rev-parse --short HEAD)
docker build -f dockerfile -t $ECR_URI:$SHA . && docker push $ECR_URI:$SHA

# Update the image field in task-definition.json, re-register, then:
aws ecs update-service --cluster realtime-editor-cluster \
  --service realtime-editor-svc \
  --task-definition realtime-collab-editor:NEW_REVISION

Rolling deployments disconnect active editors as old tasks drain. Clients reconnect automatically via Socket.IO and Yjs re-syncs from the server document — but see Scaling Notes regarding in-memory state.


7. HTTPS, DNS and CDN

  1. ACM — request a public certificate for editor.example.com in the same region as the ALB and validate it via DNS.
  2. Route 53 — create an A record with Alias pointing at the ALB rather than a CNAME to the ALB DNS name; alias records are free and work at the zone apex.
  3. HTTPS is not optional here — browsers refuse a ws:// connection from an https:// page (mixed content), so the WebSocket must be wss://. Serve over TLS end to end.
  4. CloudFront (optional) — if you front the ALB with CloudFront to cache the static bundle, you must:
    • Use the AllViewer origin request policy so Upgrade, Connection and Sec-WebSocket-* headers reach the origin.
    • Set the cache policy for /socket.io/* to CachingDisabled.
    • Keep /assets/* on a long-TTL cache policy — Vite emits content-hashed filenames, so they are safely immutable.

8. Logging and Monitoring

Signal Source Why it matters
Application logs CloudWatch Logs /ecs/realtime-collab-editor Startup, crashes, Socket.IO errors
HealthyHostCount ALB CloudWatch metric Alarm on < 1 — the earliest signal of a bad deploy
TargetResponseTime, HTTPCode_ELB_5XX_Count ALB metrics Edge-level failures the app never sees
CPUUtilization, MemoryUtilization ECS Service metrics Yjs holds documents in memory; memory growth is the metric to watch
ALB access logs S3 Per-request forensics for WebSocket upgrade failures

Set log group retention explicitly (aws logs put-retention-policy --log-group-name /ecs/realtime-collab-editor --retention-in-days 30) — the default is never expire, which quietly accrues cost.


9. Alternative: EC2 + Docker

For a low-traffic or demo deployment, a single EC2 instance is materially cheaper and simpler than Fargate + ALB:

# On an Amazon Linux 2023 instance
sudo dnf install -y docker && sudo systemctl enable --now docker

aws ecr get-login-password --region $AWS_REGION | sudo docker login \
  --username AWS --password-stdin $ACCOUNT_ID.dkr.ecr.$AWS_REGION.amazonaws.com

sudo docker run -d --restart unless-stopped -p 80:3000 --name editor $ECR_URI:latest

Attach an instance profile with AmazonEC2ContainerRegistryReadOnly so the instance can pull from ECR without stored credentials. Put Nginx or Caddy in front for TLS, and make sure the reverse proxy forwards the WebSocket upgrade:

location / {
    proxy_pass http://127.0.0.1:3000;
    proxy_http_version 1.1;
    proxy_set_header Upgrade $http_upgrade;
    proxy_set_header Connection "upgrade";
    proxy_set_header Host $host;
    proxy_read_timeout 300s;
}

Trade-off: no rolling deploys, no multi-AZ redundancy, and you own the patching. Appropriate for a portfolio deployment, not for production.


Scaling Notes and Known Limitations

These are honest constraints of the current implementation, and worth reading before scaling desired-count above 1.

  1. Document state is in-memory and per-task. YSocketIO keeps each room's Y.Doc in the process heap. Two tasks serving the same room maintain two independent documents — clients on task A will not see edits from clients on task B. Sticky sessions do not solve this, because stickiness pins clients, not rooms.
    • Mitigation for now: run desired-count = 1 and rely on ECS to replace an unhealthy task.
    • Proper fix: add a Socket.IO Redis adapter (ElastiCache for Redis) so document updates fan out across tasks, and persist documents to DynamoDB or S3 so state survives task replacement.
  2. No persistence. When the last client leaves and the task recycles, the document is gone. A persistence layer backed by S3 or EFS is the next meaningful feature.
  3. The room name is hard-coded to "monaco". Every user shares one global document regardless of the username they join with. Deriving the room from the URL is a small change with a large usability payoff.
  4. CORS is origin: "*". Acceptable while the client is same-origin, but should be tightened to the deployed domain before going public.
  5. No authentication or authorisation. Anyone with the URL can read and edit. Put the ALB behind Cognito (authenticate-cognito listener action) or add application-level auth before exposing anything sensitive.
  6. Autoscaling. Once the Redis adapter is in place, target-tracking on ALBRequestCountPerTarget or CPU is appropriate. Until then, scaling out will silently split rooms.

Troubleshooting

Symptom Likely cause Fix
Editor loads but edits do not sync WebSocket upgrade blocked Check ALB access logs; ensure nothing between client and task strips Upgrade/Connection headers
Session ID unknown (400) in the console Long-polling requests hitting different tasks Enable target group stickiness (see §5)
Connection drops every ~60 seconds ALB idle timeout at its default Raise idle_timeout.timeout_seconds to 300
Target group stuck in unhealthy Wrong port, path, or security group Confirm target port 3000, health check path /health, and that app-sg allows 3000 from alb-sg
Task exits with exec format error Image architecture ≠ task architecture Rebuild with --platform linux/amd64, or set cpuArchitecture: ARM64
ECS cannot pull the image Execution role missing ECR permissions, or no route to ECR Attach AmazonECSTaskExecutionRolePolicy; verify NAT Gateway or VPC endpoints
Mixed-content error in the browser Page on HTTPS, socket on ws:// Terminate TLS at the ALB and serve the page over HTTPS
Blank page, 404 on assets Frontend was not built into server/public Verify the COPY --from=frontend-builder /app/dist /app/public stage succeeded

Roadmap

  • Multiple rooms via URL-derived document names
  • Redis adapter (ElastiCache) for horizontal scaling
  • Document persistence to S3 / DynamoDB
  • Remote cursors and selection highlighting with per-user colours
  • Language selector and file tabs
  • Authentication (Cognito or OAuth) and per-room access control
  • Infrastructure as Code (Terraform or AWS CDK) for the whole stack
  • CI/CD via GitHub Actions with OIDC federation to AWS

License

ISC


Author: Naman Mittal

About

A multi-user, real-time collaborative code editor built with React, Monaco Editor, Yjs (CRDT) and Socket.IO.Deployment using Docker and AWS.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages