Working record / The Lab

A working practice, made visible.

The Lab records what we build, test, reject and still need to prove. A deployed product, a controlled alpha and an internal proof of concept are different stages. The notes keep those distinctions visible.

Sky Link Solutions
Trials, decisions and limitations
Evidence checked 21 September 2026

Field notes

Follow the work beyond the first draft.

Three dated notes describe current work without turning unfinished evaluation into a success claim. Peninsula is in a hosted sales-team trial. Jev is an inactive evaluation candidate. The anonymous remodeling POC has only exercised synthetic inputs; real walkthrough footage remains untested.

Revision record

The original seven records.

These are the existing Lab notes and their recorded stages. Open a record to read its engineering notes. A status describes that piece of work; it is not a guarantee that another project will produce the same result.

VALIDATING

Hermes / FieldThread: Constrained Agent Architecture

Turning everyday field messages, photos, and voice notes into structured project records without giving AI unrestricted control.

FieldThread uses the Hermes agent to interpret existing field communication and convert it into traceable, organization-scoped records. Development has focused on source attribution, anti-fabrication rules, constrained write commands, project routing, audit history, append-oriented corrections, and recovery. Real incident testing exposed identity-attribution and stale-deployment failure modes, which were converted into explicit verification controls. The system is in controlled alpha and still requires broader tenant-isolation and adversarial testing before wider use.

REJECTED

Unverified Live Gmail Automation

Evaluating whether an AI-assisted mail application should be trusted with live mailbox changes.

The answer was no—not until its provider boundary and recovery behavior could be proven. Live Gmail reads and writes were disabled by default, while unqualified send, draft, trash, and label operations were made to fail closed. One folder-filing operation was rebuilt as a typed, journaled, restart-resumable command with read-back verification and Undo. Seventeen deterministic tests now pass, but valuable Gmail accounts remain blocked until authentication, synchronization, message creation, attachment, and recovery workflows complete the qualification process.

DEPLOYED

Occasly: Secure Invitation & RSVP Platform

Taking a design-led product from early prototype to a live, mobile-first application with real access and data boundaries.

Occasly combines invitation creation, public RSVP flows, guest management, media uploads, private-access controls, and provider-ready email workflows. Its production foundation includes database row-level security, server-controlled changes, schema validation, unguessable event links, duplicate-RSVP protection, upload restrictions, security headers, and endpoint rate limits. The result is more than a visual prototype: it is a working product with operational controls and a deliberate path from private testing to broader availability.

VALIDATING

PacketLens: Evidence-First Network Diagnostics

Determining where AI can assist network analysis without being allowed to invent packet-level facts.

PacketLens ingests packet captures and diagnostic screenshots, preserves their provenance, and extracts deterministic flow and protocol observations before any model is involved. Findings distinguish supporting evidence, counter-evidence, limitations, and next steps. Optional AI synthesis receives bounded, normalized observations rather than raw packet payloads and cannot overwrite the underlying facts. The analyzer also uses non-root container execution, reduced system capabilities, a read-only root filesystem, and explicit confidence limits where encrypted traffic prevents a definitive conclusion.

VALIDATING

Realtime Voice Architecture Benchmark

Comparing modular and native voice-agent architectures for natural conversation, latency, control, and provider flexibility.

A purpose-built evaluation lab was created to compare OpenAI Realtime, Gemini Live, xAI Voice, and modular speech-to-text, language-model, and text-to-speech stacks using LiveKit, Deepgram, Anthropic, OpenAI, and Rime. Twenty-eight scored sessions have been recorded with provider configuration, transcripts, latency observations, quality ratings, and exportable findings. Long-lived provider credentials remain server-side; browser sessions receive scoped short-lived credentials, while same-origin checks, request limits, and access controls protect paid provider routes.

VALIDATING

AI Search Growth: Evidence Before Automation

Building a local-visibility operations system that converts scattered website and search signals into prioritized, reviewable work.

The system combines bounded website crawling, technical and trust-signal analysis, service-and-location targeting, local-rank context, competitor evidence, historical comparison, and implementation planning. Its core audit completes deterministically before optional AI enrichment is introduced. Business context can inform recommendations but cannot overwrite what the public evidence actually shows. Proposed actions remain reviewable and approval-controlled rather than being published automatically. The current MVP is being developed as an operator system—not a content factory or an opaque AI scoring tool.

REJECTED

Prompt-Only Agent Security

Testing whether instructions alone can safely prevent an AI agent from performing administrative operations.

Rejected. Telling an agent not to administer a system is not equivalent to preventing it. In FieldThread, administrative code, credentials, and capabilities were removed from the agent environment rather than protected only by prompt language. The agent receives a constrained set of auditable domain actions, while account, organization, and configuration management remain separate and human-controlled. The result is a smaller failure surface and a general design rule: behavioral instructions can guide a model, but enforceable capability boundaries must exist outside the model.

Into practice

Connect a technical decision to the operation.

A useful result from the Lab may be a working application, a better test or a decision to keep an automation disabled. The common thread is making the evidence and the boundary clear enough to review.

Start a conversation

Bring a problem worth investigating.

Start a conversation

Founder-led technology consultancy and custom solutions partner.
Pleasanton, CA · Bay Area and remote.