Hackathon · Machine learning · 2025

ExoLens

Exoplanet exploration and classification, built in a weekend for NASA Space Apps.

Status
Prototype
Context
Team project · NASA Space Apps 2025
My role
Machine learning and API: I trained the models, built the FastAPI service and added the light-curve upload and interactive chart to the front end.
Stack
  • Python
  • scikit-learn
  • XGBoost
  • pandas
  • FastAPI
  • Next.js
  • TypeScript
  • three.js
  • Recharts
ExoLens lab page: a form with six exoplanet parameters next to a 3D preview.
The form that sends the six measurements to the model.

Overview

A five-person team project built over the NASA Space Apps Challenge weekend (October 4–5, 2025), for a challenge about using AI with NASA’s exoplanet data.

Problem

Many exoplanet candidates in NASA’s catalogs turn out to be false positives. Telling them apart is a classic classification problem.

Approach

We split the work with a simple contract: six numbers in, a classification out. That let the front end move in parallel while I trained the model.

How it works

  1. Data

    Cleaned and merged the K2 and TESS tables and picked six measurements of the planet and its star.

  2. Models

    Compared Random Forest and XGBoost with cross-validation and class weights, since there were far more confirmed planets than false positives.

  3. API

    A FastAPI service with single and batch predictions.

  4. Interface

    On the lab page, I built the CSV light-curve upload and the interactive chart.

From catalog to prediction
  • Data
  • Code
  • AI
  • Outcome
  1. NASA catalogs (Data)K2 and TESS
  2. Preparation (Code)Cleaning and six measurements
  3. Models (AI)Random Forest and XGBoost
  4. API (Code)FastAPI
  5. Interface (Outcome)Next.js, 3D and charts
@app.post("/predict", response_model=PredictionResponse)
def predict(features: ExoplanetFeatures):
    if model is None or scaler is None:
        raise HTTPException(status_code=503, detail="Modelos de classificação não carregados")

    # Converter para array na ordem correta
    feature_array = np.array([[
        features.pl_orbper,
        features.pl_rade,
        features.pl_trandep,
        features.st_teff,
        features.st_rad,
        features.st_logg
    ]])

    feature_scaled = scaler.transform(feature_array)
    prediction = model.predict(feature_scaled)[0]
    probabilities = model.predict_proba(feature_scaled)[0]
Excerpt from the public repository (back-ia/api/api.py), unchanged — the original comments are in Portuguese.

Key decisions

  • A simple model, on purpose

    With only hours, tree-based models on tabular data were the right choice instead of a neural network.

  • Look at each class’s errors

    The test set had 127 false positives against 1,263 confirmed planets. Overall accuracy hid the mistakes on the smaller class.

  • An API between model and interface

    A small contract kept both halves independent until the end.

Result

A working prototype by the end of the weekend, with public source code. The screenshots come from running the repository locally.

Screenshots