←Back to projects
Web DevelopmentFull StackBackendSaaSInfrastructure

Horizon Labs

Full-Stack & AI Engineering

Horizon Labs is an Australian AI and software development company where our team built the company website and contributed to the development of self-hosted Retrieval-Augmented Generation (RAG) systems. The website was built with Next.js and TypeScript on the frontend and Go on the backend, while the AI platform combined Python-based services, vector search, embedding and reranking models, and modern language models. Our team designed the architecture to support different deployment models, including fully self-hosted and hybrid environments, allowing businesses to work with private documents while maintaining control over sensitive information.

Technologies

Next.jsTypeScriptGoPythongRPCProtocol BuffersPostgreSQLQdrantRedisDockerLlamaMistralQwenOpenAIClaudeGemini

Team

Full-stack, Backend, AI/ML and Infrastructure teams

Timeline

Long-term

Challenges

The challenges

Supporting different AI models and providers through a unified architecture.

Keeping private client documents and sensitive business information protected.

Providing reliable retrieval from large collections of private documents.

Improving retrieval accuracy for contracts, policies, technical manuals, reports, and internal knowledge bases.

Maintaining acceptable response times across document processing, retrieval, reranking, and answer generation.

Supporting both fully self-hosted and hybrid AI deployment architectures.

Maintaining a clean and strongly typed connection between business applications and AI services.

Solutions

How we approached it

Built model-agnostic AI integrations that allowed different language models to be used without changing the core application architecture.

Designed self-hosted RAG deployments to keep sensitive document processing and retrieval within the client's infrastructure when required.

Used hybrid architectures when clients needed private retrieval combined with external model providers.

Implemented document chunking, embeddings, vector search, metadata filtering, retrieval, and reranking to improve knowledge retrieval.

Used Qdrant as the vector database for efficient semantic search.

Introduced caching through Redis to reduce repeated processing and improve response times.

Used gRPC and Protocol Buffers to provide a fast and strongly typed communication layer between Go applications and RAG services.

Used streaming responses to improve the perceived response time of AI-powered interactions.

Features

What we built

Modern Company Platform

Our team designed and developed the Horizon Labs company website using Next.js and TypeScript, with a responsive interface and a Go-based backend.

Self-Hosted RAG

The team developed RAG solutions that allow organizations to connect AI models to private documents and internal knowledge bases while maintaining control over sensitive data.

Hybrid AI Architecture

The platform supports hybrid deployments where document processing and retrieval remain private while external AI models can be used for final response generation.

Model-Agnostic AI Layer

The architecture abstracts model communication so different language models and providers can be integrated without tightly coupling the application to a single provider.

Document Processing Pipeline

Documents are extracted, cleaned, chunked, converted into embeddings, indexed, retrieved, and reranked before being provided to the answer-generation layer.

Semantic Search

Qdrant-based vector search combined with metadata filtering and reranking enables the system to retrieve relevant information from private document collections.

AI Service Integration

Go applications communicate with dedicated RAG services through gRPC and Protocol Buffers, providing a fast and strongly typed service-to-service interface.

Private Knowledge Bases

Organizations can connect contracts, policies, technical documentation, reports, and other internal knowledge sources to AI-powered search and question-answering workflows.

Architecture

RAG Processing Pipeline

The RAG platform processes private documents through extraction, cleaning, chunking, embedding generation, vector storage, retrieval, and reranking before passing relevant context to the selected language model for answer generation with citations.

01

Documents

Private contracts, policies, reports, technical manuals, and internal knowledge.

02

Text Extraction & Cleaning

Documents are processed and converted into clean searchable text.

03

Chunking

Documents are divided into smaller context-aware sections for retrieval.

04

Embedding Generation

Document chunks are converted into vector representations.

05

Qdrant

Embeddings and metadata are stored for semantic retrieval.

06

Retrieval & Reranking

Relevant document chunks are retrieved, filtered, and reranked to improve contextual accuracy.

07

AI Model

The selected language model generates an answer using the retrieved context.

08

Answer with Citations

The final response includes references to the relevant source material.

Team

Our work

Designed and developed the company website as part of the full-stack team.

Built responsive frontend experiences using Next.js and TypeScript.

Developed backend APIs and business logic using Go.

Integrated PostgreSQL and Redis into backend services.

Contributed to the architecture and development of self-hosted RAG systems.

Worked on document processing and retrieval pipelines.

Integrated Qdrant for vector search and semantic retrieval.

Worked with embedding and reranking models.

Integrated multiple language model providers through a model-agnostic architecture.

Integrated Go backend services with RAG services through gRPC and Protocol Buffers.

Worked on Docker-based deployment and service infrastructure.

Contributed to improving retrieval accuracy, privacy, and AI response performance.

Results

Outcomes

Created a modern full-stack company platform using Next.js, TypeScript, and Go.

Established a flexible RAG architecture capable of working with multiple AI models and providers.

Enabled organizations to use AI with private internal documents and knowledge bases.

Supported both fully self-hosted and hybrid deployment models.

Improved retrieval quality through semantic search, metadata filtering, and reranking.

Created a strongly typed and efficient communication layer between Go applications and AI services.

Improved AI response performance through caching and streaming.

Built an architecture that could adapt to different client privacy, deployment, and model requirements.