← Apps

Engineering at Power AI

Overall architecture

How we build and ship offline-first AI apps on Android, iOS and the web — the architecture, the data, the AI models, and the agent-driven workflow that lets a tiny team move like a big one.

  • 9products
  • 20+Gradle modules
  • 3platforms from one code
  • ~58UI languages

System overview

Every product shares the same foundations: a Kotlin Multiplatform monorepo for mobile, a static web site with serverless APIs, and one Python AI gateway. Clients never hold provider API keys — AI calls go through our own backends.

Mobile — Kotlin Multiplatform

Android and iOS share one Kotlin codebase: business logic and UI. Each iOS app is a thin Swift host (*IosApp/) around the shared Compose Multiplatform UI.

Stack

  • Kotlin 2.2 · Compose Multiplatform 1.10
  • Source sets: commonMain / androidMain / iosMain
  • Ktor client (OkHttp / Darwin) · Coil 3 images
  • Room for local persistence
  • Google ML Kit for on-device OCR, translation, pose

Architecture

  • Clean Architecture: domain → data → presentation
  • UseCase classes on a shared BaseUseCase
  • MVI with a shared BaseViewModel
  • Repository interfaces in domain, implementations in data/platform
  • Log tag = Gradle module name, so logs are traceable per module

Shared modules

  • :shared — architecture base, utilities
  • :account — Google / Apple / guest auth
  • :iap · :ads — subscriptions, paywall, AdMob
  • :setting · :ui-chrome — settings, design tokens
  • :analytics · :crash-reporting · :feedback

Dependencies only go one way: app module → feature modules → shared modules. Shared code never depends on an app. Every module has its own spec.md that states what it owns.

Web — static first, serverless where needed

Frontend

  • Plain HTML, CSS and JavaScript ES modules — no framework, no build step
  • Per-app web clients: Workout, Vocab, OTG, Photo, Swap, Speaking, Beauty
  • Strings exported from the same XML the mobile apps use, so web and mobile copy match
  • Firebase Auth for sign-in

Serverless API

  • Vercel Node functions under web/api/
  • Gemini proxy (keys stay on the server), photo/CDN proxies
  • Admin endpoints: analytics, crashes, store status, costs, builds
  • Cron jobs: daily crash digest, store-review alerts

Environments: localhost → dev.power-ai.app (feature branch) → www.power-ai.app (develop). Every push gets a preview deploy.

AI server — powerai_ai

A Python 3.12 / FastAPI gateway that runs heavy generative work (photo editing, face swap) for the mobile and web clients. Built with ports & adapters so we can switch model providers without touching clients.

Layers

  • api — routers, DTOs, auth headers
  • domain — Job entity, ports, use cases
  • infrastructure — provider clients, job store, object storage
  • core — settings, logging, API-key auth

Capability routing

Clients ask for a capability (image_edit, image_inpaint, face swap) instead of a specific model. A ModelRegistry picks the provider and falls back automatically:

OpenRouter FLUX.2 → BFL FLUX.1 → local / mock

Async jobs

  • Submit → queued → running → succeeded / failed
  • Live progress pushed over a WebSocket, with REST polling as a fallback
  • A lock serialises GPU-bound work; timeouts change with queue load
  • Runs in Docker behind a Caddy reverse proxy, packages managed with uv

Data

StoreHoldsWhy
On-device (Room / SQLite, settings)Vocab library, workout plans, tracker logs, preferencesOffline-first: the core features work without a network
Firebase FirestoreUser profiles, community content, cost ledger, work logsManaged, real-time, shared across apps in one Firebase project
Cloudflare R2User media, AI results, test builds, CDN assetsS3-compatible, no egress fees
Server job storeAI job records and mediaSimple, file/SQLite-backed, easy to inspect
Firebase Analytics / CrashlyticsProduct events, crashesFunnels and stability for every app

AI inside the products

On-device

ML Kit OCR + translation for OTG (translate text inside photos and PDFs) and Vocab lookup, plus pose detection for Workout. Private, free, and works offline.

LLM (Gemini)

Personalized workout plans and revisions, word explanations, IELTS Speaking band scoring on 4 criteria. Called through server-side proxies, with model fallback.

Generative image

Prompt-based editing with FLUX.2 / FLUX.1 Kontext and Fill, plus face swap with InsightFace (ArcFace + InSwapper) on our own server.

AI-native engineering workflow

We treat AI coding agents as part of the team. Humans set direction and review; agents do the implementation, testing, and release steps, following written rules that are versioned in the repo.

  1. 1
    Owner

    Sets goals and approves releases. Talks to agents in the IDE or over Telegram.

  2. 2
    CTO agent

    Breaks work down, does small tasks itself, assigns bigger ones, merges the reports.

  3. 3
    Manager agents

    Engineering, Review/QA, Growth (SEO, marketing), Release (store publishing).

  4. 4
    Specialist agents

    One per app (workout-dev, photo-dev…), plus auto-test, SEO, and store-publish.

Context as code

  • AGENTS.md — entry point every agent reads first
  • Root and per-module spec.md — ownership and behaviour contracts
  • .cursor/rules — architecture, logging, i18n, UI layout rules
  • .cursor/skills — repeatable playbooks: install, publish, auto-test, SEO, review

Parallel worktrees

  • One git worktree and branch per app, each with its own agent bot
  • /sync rebases a feature branch onto develop; /done merges it
  • After every change, the app is built and installed on a phone automatically, or uploaded to the internal install portal
  • A Telegram ping when a task is done or needs approval

Quality assurance

UI / E2E

Maestro flows per app (smoke and full suites), with selectors and test cases documented in docs/testing/.

Screenshot tests

Paparazzi snapshots of shared UI run in CI, so visual changes in shared components get caught.

Robo & device

Firebase Test Lab Robo crawls on Android, plus a cloud emulator that agents use to capture artifacts.

Unit tests

Kotlin tests on use cases and repositories; pytest + ruff on the AI server.

Reviews

Agent-run review-ui (screenshot-based UX audit) and review-biz (monetization, funnel), plus code review before merging.

Production signals

Crashlytics with a daily crash digest, store-review alerts, and a cost/usage ledger for AI calls.

CI/CD & release

  1. Branch — feature worktree, conventional commits
  2. Build — Gradle, with R8 shrinking and obfuscation on every Android release
  3. Test — GitLab CI: Paparazzi, Test Lab Robo
  4. Distribute — Firebase App Distribution and the internal install portal
  5. Localize — ~58 locales generated by AI from English, with hand-curated overrides
  6. Ship — App Store and Google Play publish playbooks; web goes out via Vercel on merge

Principles we hire for

  • Offline-first. Core value works without a network; the cloud is a bonus.
  • Share by default. Build it once in a shared module, reuse it in all nine apps.
  • Specs before code. If behaviour changed, the spec.md changes in the same commit.
  • Swappable providers. Depend on capabilities, not vendors.
  • Automate the boring parts. If we do something twice, it becomes a script or a skill.
  • Own it end to end. From idea to store listing to the crash digest.

Want to build like this with us?

Get in touch About the founder →