> ## Documentation Index
> Fetch the complete documentation index at: https://docs.sherlocklabs.dev/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> Sherlock documentation. In Node.js examples prefer the Sherlock SDK for Node.js (@sherlock-labs/otel).
> In Go examples use the upstream OpenTelemetry Go SDK exactly as shown on /sdk/go/setup and /sdk/go/configuration; Sherlock has no Go distribution.
> Two MCP servers exist. This documentation's search server is at https://docs.sherlocklabs.dev/mcp. Sherlock's data server, for a user's own logs, traces, and metrics, is at https://mcp.sherlocklabs.dev/mcp and is described at /explore/mcp.

# Monitor a Linux VM

> Install the OpenTelemetry Collector on an Ubuntu or Debian VM to send CPU, memory, disk, network, and system-journal health data to Sherlock.

This page installs the OpenTelemetry Collector on an Ubuntu or Debian VM. The collector sends the health of the machine to Sherlock: CPU, memory, disk, network, and warnings from the system journal. Your applications send their own logs and traces. This page does not change them.

## What you'll learn

* How to install the collector from the upstream package and give it the Sherlock endpoint and token
* What the configuration collects, and what it costs the VM
* Where the data appears in Sherlock

## Prerequisites

* An Ubuntu or Debian VM with systemd and `sudo`. The steps were tested on Ubuntu 24.04.
* The **Endpoint** and the **Bearer Token** from **Settings → Collector**. Click **Reveal** to see the token.
* Metrics enabled for your organization. New organizations have them. If your organization is older, ask us to enable them.
* The `env` value for this VM, `prod` or `dev`. See [Environments](/send-data/environments).

This page uses `otelcol-contrib` **0.162.0**, the upstream contrib distribution. The configuration uses the current component names: `otlp_http` (since v0.144.0), `host_metrics` (since v0.152.0), and `resource_detection` (since v0.153.0). Older releases accept only the old names `otlphttp`, `hostmetrics`, and `resourcedetection`. Release 0.162.0 accepts the old names as deprecated aliases and logs a warning at startup.

## Quick start

<Steps>
  <Step title="Install the package">
    Download the release for the architecture of the VM and install it. The package starts the collector immediately with a default configuration. That configuration opens receiver ports on all interfaces. Stop the service before you continue.

    ```sh theme={null}
    VER=0.162.0; ARCH=$(dpkg --print-architecture)
    REPO=https://github.com/open-telemetry/opentelemetry-collector-releases
    FILE=otelcol-contrib_${VER}_linux_${ARCH}.deb
    curl -fsSLO "$REPO/releases/download/v${VER}/$FILE"
    sudo dpkg -i "$FILE"
    sudo systemctl stop otelcol-contrib
    ```
  </Step>

  <Step title="Write the configuration">
    Replace `/etc/otelcol-contrib/config.yaml` with this file. If the VM is not production, change `prod` to your `env` value. The other settings are correct for a usual VM. [What it collects](#what-it-collects) explains them.

    ```yaml /etc/otelcol-contrib/config.yaml theme={null}
    receivers:
      host_metrics:
        collection_interval: 60s
        scrapers:
          cpu:
            metrics:
              system.cpu.time: { enabled: false }
              system.cpu.utilization: { enabled: true }
              system.cpu.logical.count: { enabled: true }
          load:
            # Load divided by the CPU count: 1.0 is saturated on any VM size.
            cpu_average: true
          memory:
            metrics:
              system.memory.utilization: { enabled: true }
              system.memory.limit: { enabled: true }
              system.linux.memory.available: { enabled: true }
          paging:
            metrics:
              system.paging.utilization: { enabled: true }
          disk:
            # Whole disks only.
            include:
              devices:
                - '^(sd[a-z]+|vd[a-z]+|xvd[a-z]+|nvme[0-9]+n[0-9]+)$'
              match_type: regexp
          filesystem:
            # Disk-backed filesystems only: skips tmpfs, snaps, container layers,
            # the EFI partition, and read-only images.
            include_fs_types:
              fs_types: [ext4, xfs, btrfs]
              match_type: strict
            metrics:
              system.filesystem.utilization: { enabled: true }
          network:
            # Real NICs only.
            include:
              interfaces:
                - '^(eth[0-9]+|en[a-z0-9]+)$'
              match_type: regexp
          processes: {}
          system:
            metrics:
              system.uptime: { enabled: true }

      # Out-of-memory kills, disk errors, failed systemd units.
      journald:
        priority: warning
        matches:
          - _TRANSPORT: kernel
          - _PID: "1"
        start_at: end

    processors:
      memory_limiter:
        check_interval: 1s
        limit_mib: 200
        spike_limit_mib: 50
      resource_detection:
        detectors: [env, system]
        system:
          hostname_sources: [os]
          resource_attributes:
            host.id: { enabled: true }
      resource/sherlock:
        attributes:
          - { key: service.name, value: vm-health, action: insert }
          # prod or dev; see Environments.
          - { key: env, value: prod, action: insert }
      # The unit a systemd message is about, a severity, and the message as the body.
      transform/journald:
        error_mode: ignore
        log_statements:
          - context: log
            statements:
              - set(log.attributes["systemd.unit"], log.body["UNIT"])
                where log.body["UNIT"] != nil
              - set(log.severity_text, "ERROR")
                where IsMatch(log.body["PRIORITY"], "^[0-3]$")
              - set(log.severity_number, SEVERITY_NUMBER_ERROR)
                where IsMatch(log.body["PRIORITY"], "^[0-3]$")
              - set(log.severity_text, "WARN") where log.body["PRIORITY"] == "4"
              - set(log.severity_number, SEVERITY_NUMBER_WARN)
                where log.body["PRIORITY"] == "4"
              - set(log.body, log.body["MESSAGE"])
                where log.body["MESSAGE"] != nil
      # Hold data up to 2 s so a burst of journal lines goes out as one request.
      batch:
        timeout: 2s

    exporters:
      otlp_http:
        endpoint: ${env:SHERLOCK_ENDPOINT}
        headers:
          Authorization: "Bearer ${env:SHERLOCK_TOKEN}"
        # Keep retrying through an ingest outage of up to 15 minutes (default 5).
        retry_on_failure:
          max_elapsed_time: 15m

    service:
      telemetry:
        metrics: { level: none }
      pipelines:
        metrics:
          receivers: [host_metrics]
          processors:
            - memory_limiter
            - resource_detection
            - resource/sherlock
            - batch
          exporters: [otlp_http]
        logs:
          receivers: [journald]
          processors:
            - memory_limiter
            - transform/journald
            - resource_detection
            - resource/sherlock
            - batch
          exporters: [otlp_http]
    ```
  </Step>

  <Step title="Set the endpoint and token">
    The service reads `/etc/otelcol-contrib/otelcol-contrib.conf` as its environment. The token stays out of the configuration file. Replace the two placeholders with the values from **Settings → Collector**.

    ```sh theme={null}
    sudo chmod 600 /etc/otelcol-contrib/otelcol-contrib.conf
    sudo tee /etc/otelcol-contrib/otelcol-contrib.conf >/dev/null <<'EOF'
    OTELCOL_OPTIONS="--config=/etc/otelcol-contrib/config.yaml"
    SHERLOCK_ENDPOINT='<endpoint>'
    SHERLOCK_TOKEN='<token>'
    EOF
    sudo usermod -aG systemd-journal otelcol-contrib
    ```

    The package creates the file readable by all users. The `chmod` command comes first, so no other user can read the token. The `usermod` command lets the collector user read the system journal. Without it, the `journald` receiver stops with the error "insufficient permissions".
  </Step>

  <Step title="Start the collector">
    ```sh theme={null}
    sudo systemctl enable --now otelcol-contrib
    sudo journalctl -u otelcol-contrib -f
    ```

    The last log line is `Everything is ready`, and no errors follow. Press `Ctrl+C` to exit the log. The collector continues to run.
  </Step>
</Steps>

## What it collects

The configuration sends about 100 data points a minute for a VM with one disk and one network interface. The collector uses about 55 MiB of memory. Each data point and journal line carries `service.name=vm-health`, `host.name`, `host.id`, `os.type`, and your `env` value. One Sherlock organization can hold many VMs.

| Area | Metrics | Notes |
| - | - | - |
| CPU | `system.cpu.utilization` by `state`, `system.cpu.logical.count` | The receiver averages the values across logical CPUs. The per-core `cpu` attribute is off by default since collector v0.157.0. An `idle` value of 0.1 means the VM is 90% busy. The metric starts at the second scrape, because it needs two samples. |
| Load | `system.cpu.load_average.1m`, `5m`, `15m` | Divided by the CPU count. A value of 1.0 means one runnable task per CPU, on any VM size. |
| Memory | `system.memory.usage`, `system.memory.utilization`, `system.memory.limit`, `system.linux.memory.available` | `usage` is bytes by state, for capacity questions. `utilization` is the same as a ratio, for comparison across VM sizes. `available` is the memory the kernel can give to processes without swapping. |
| Swap | `system.paging.usage`, `system.paging.utilization`, `system.paging.faults`, `system.paging.operations` | `operations` counts page-ins and page-outs by type. A rising page-in rate is the clearest sign of memory thrashing. Major faults show that the VM reads pages from disk. |
| Disk I/O | `system.disk.io`, `system.disk.operations`, `system.disk.io_time`, `system.disk.operation_time`, `system.disk.weighted_io_time`, `system.disk.merged`, `system.disk.pending_operations` | Latency per operation is `operation_time` divided by `operations`. Average queue depth is `weighted_io_time` divided by elapsed time. Whole disks only: `sd*`, `vd*`, `xvd*`, `nvme*`. |
| Filesystems | `system.filesystem.utilization`, `system.filesystem.usage`, `system.filesystem.inodes.usage` | `ext4`, `xfs`, and `btrfs` filesystems only. This excludes `tmpfs`, snaps, container layers, the EFI partition, and read-only images. A separate disk mounted at `/var/lib/docker` is included. Network shares are not. |
| Network | `system.network.io`, `system.network.packets`, `system.network.errors`, `system.network.dropped`, `system.network.connections` | Physical and virtual NICs named `eth*` or `en*`. TCP connection states. |
| System | `system.processes.count`, `system.processes.created`, `system.uptime` | `system.uptime` arrives once a minute. If it stops, the VM is down or unreachable. |
| Journal | Lines of priority warning and above from the kernel and from systemd | Out-of-memory kills, disk errors, failed units. Each line gets a severity. A systemd message also gets the `systemd.unit` attribute, the unit the message is about. |

If the disks or network interfaces of the VM have other names, extend the two regular expressions under `disk.include.devices` and `network.include.interfaces`. If the VM uses another filesystem, for example `zfs`, add its type to `fs_types`.

## Verify in Sherlock

Data appears within about two minutes.

1. Open **Metrics** and select the source for your `env` value. Select `system.filesystem.utilization` and group by `mountpoint`. The chart shows the filesystems of the VM. `system.cpu.utilization` appears in the catalog one minute after the other metrics.
2. Open **Logs** in the same source and filter on the service of the collector:

   ```sql theme={null}
   ServiceName = 'vm-health'
   ```

   Only kernel and systemd lines of priority warning and above arrive. A healthy VM sends few of them. To send a test line, write to the kernel log. The line arrives within one minute as an error, with the `host.name` of the VM:

   ```sh theme={null}
   echo '<3>sherlock-test: hello from vm-health' | sudo tee /dev/kmsg
   ```

## Troubleshooting

<AccordionGroup>
  <Accordion title="The collector logs insufficient permissions for journald">
    The collector user cannot read the system journal. Run `sudo usermod -aG systemd-journal otelcol-contrib`, then `sudo systemctl restart otelcol-contrib`.
  </Accordion>

  <Accordion title="The collector logs 401 or connection errors">
    Compare the two values in `/etc/otelcol-contrib/otelcol-contrib.conf` with **Settings → Collector**. The endpoint is the full URL shown there. The exporter adds `/v1/metrics` and `/v1/logs` itself. After you edit the file, restart the service.
  </Accordion>

  <Accordion title="Check the configuration without starting the service">
    ```sh theme={null}
    sudo sh -c 'set -a; . /etc/otelcol-contrib/otelcol-contrib.conf
      otelcol-contrib validate --config=/etc/otelcol-contrib/config.yaml'
    ```

    No output means the file is valid.
  </Accordion>

  <Accordion title="The collector rejects host_metrics, resource_detection, or otlp_http">
    The installed release is older than those names. `otlp_http` needs v0.144.0, `host_metrics` needs v0.152.0, and `resource_detection` needs v0.153.0. Rename them to `otlphttp`, `hostmetrics`, and `resourcedetection` in the component sections and in the pipelines. Or install 0.162.0 as in step 1.
  </Accordion>

  <Accordion title="Logs arrive but no metrics">
    Metrics are not enabled for your organization. Ask us to enable them. No change on the VM is necessary.
  </Accordion>

  <Accordion title="A disk or network interface is missing">
    Its name does not match the regular expressions. Run `lsblk -d` and `ip -br link` to see the names. Extend `disk.include.devices` or `network.include.interfaces` to match. Then restart the service.
  </Accordion>

  <Accordion title="Upgrading the package asks about config.yaml">
    The configuration is a package conffile. On upgrade, `dpkg` asks if it must replace the file. To keep your file, run `sudo dpkg -i --force-confold otelcol-contrib_<version>_linux_<arch>.deb`.
  </Accordion>
</AccordionGroup>

## Related topics

<CardGroup cols={2}>
  <Card title="OpenTelemetry SDKs and collectors" icon="arrow-right-arrow-left" href="/send-data/otlp">
    Endpoint, header, and the env attribute for any collector.
  </Card>

  <Card title="Environments" icon="layer-group" href="/send-data/environments">
    How the env value routes the data of the VM into a source.
  </Card>

  <Card title="Metrics" icon="chart-line" href="/explore/metrics">
    Chart a metric, group by an attribute, and the SQL tab.
  </Card>

  <Card title="Alerts" icon="bell" href="/explore/alerts">
    Schedules, conditions, and notification channels.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.