2014年4月23日 星期三

GCC 4.9 Release Note

以下是 4.9 更新的部份, 主要整理自官方 Release 頁面, 但由於個人能力有限, target 相關僅翻譯 x86, arm, aarch64 以及 nds32, 語言相關方面則只翻譯 C/C++ 部份:

提醒!

  • 移除 mudflap run time checker, mudflap 相關的選項還會留著但不會有任何作用.
  • 許多老舊系統以及年久沒維護或測試的 target 在 GCC 4.9 將會被宣佈為過時的 (obsolete), 除非有善心人士幫忙不然在下個 Release 的時候就會直接砍掉.

    下列的系統將將被列為 obsoleted:
    • Solaris 9 (*-*-solaris2.9). 詳情請洽公告
  • 更多 Porting 到 GCC 4.9 的相關資訊請洽 porting guide.

一般性的最佳化改善

  • AddressSanitizer, 快速記憶體錯誤偵測器, 在 ARM 上面可以動了!!
  • UndefinedBehaviorSanitizer (ubsan), 快速未定義行為偵測器, 可以透過 -fsanitize=undefined 這個 flag 開啟. 許多計算(computations) 都會被塞入(instrumented)偵測未定義行為的東西, 目前這個東東可以在 C/C++ 上運作.

連結時期最佳化(Link-time optimization LTO)的改善:

  • 型態合併的部份砍掉重練. 新的實作更快更好更省記憶體!
  • 更好的切割演算法, 連結時期使用更少的 streaming.
  • 透過即早砍掉 Vritual Method 來減少 Object File Size, 以及改善連結時間跟編譯時間.
  • 函數內容 (Function bodies) 需要時才載入 (loaded on-demand), 以及即早釋放來改善連結時期所消耗的記憶體.
  • C++ 隱藏(hidden) 函數們現在可以被最佳化掉.
  • 當使用 linker plugin 時, 開 -flto 會產生只放 GCC IR 的苗條 (slim) Object File (.o), 使用 -ffat-lto-objects 的時候才會產生肥大 (Fat) Object File, 裡面就會裝 IR 以及 machine code. 要產生可以餵給 LTO 的 static library 時, 請使用 gcc-ar 以及 gcc-ranlib; 要看 symbol 時則請用 gcc-nm. (需要 ar, ranlib 以及 nm 有開啟 plugin 支援).
  • 以 LTO build FireFox 為例, 記憶體使用量從 15G 變 3.5G, 連結時間由 1700 秒變 350 秒.

跨函數最佳化(Inter-procedural optimization) 的改善:

  • 新的型別繼承分析模組使得 devirtualization 有所改善!
  • Devirtualization 現在會把匿名命名空間以及 C++11 的 final 關鍵字一併考慮進去分析.
  • 新的 speculative devirtualization pass (可透過 -fdevirtualize-speculatively 開啟).
  • Function Call 被 speculation 後若太貴的話也可能會變回 indierct call.
  • Local aliases are introduced for symbols that are known to be semantically equivalent across shared libraries improving dynamic linking times. 引入 Local Alias, 讓 gcc 可以判斷 symbol 穿越過 shared librrary 後是否還是否語意等價 (譯註:沒理解錯的話應該是指 Escape analysis XD).

回饋導向最佳化(Feedback directed optimization) 的改善:

  • C++ inline function 的 Profiling 資訊現在變得更可靠.
  • 新的 time profiling 會吐一個 Function 執行時間的排序.
  • 新的 Function Reorder pass (可透過 -freorder-functions 開啟) 可以大幅的減少大型應用程式的起始時間, 這項功能目前只對於 LTO 有效, 非 LTO 須等到 binutils 全面支援,
  • 即日起開啟 LTO 即可享有回饋導向的 Inderct Call 移除以及跨模組的 devirtualization.

新語言以及語言相關的改善:

  • gcc/g++ 現在支援 OpenMP 4.0, 加入新的編譯選項 -fopenmp-simd 來開啟 OpenMP 的 SIMD directive. 另外也加入了新的選項 -fsimd-cost-model= 來選擇迴圈中的 annotatation 以及 Cilk Plus simd directive 來作為向量化的 Cost Model; 若只有下 -Wopenmp-simd 則會提醒使用者 Cost Model 被 simd directive 給覆蓋.
  • 新的警告選項 -Wdate-time, 主要針對 C, C++ 以及 Fortan Compiler, 這個選項開啟時, 若使用到 __DATE__, __TIME__ 或 __TIMESTAMP__ Marco 則會發出警告, 提醒使用者會無法生出完全一模一樣 (bit-wise-identical) 的編譯結果.

C 家族

  • GCC 現在也支援多采多姿的診斷訊息, -fdiagnostics-color=auto 時會會自動偵測是否書出到終端機, 是的話就會多采多姿. -fdiagnostics-color=always 則會強迫多采多姿輸出, 另外 GCC_COLORS 這個環境變數也可以客製化顏色或關閉多采多姿輸出, 如果 GCC_COLORS 有設定的話, 會採用 -fdiagnostics-color=auto 為預設值, 否則預設值是 -fdiagnostics-color=never. 範例診斷訊息輸出
        $ g++ -fdiagnostics-color=always -S -Wall test.C
        test.C: In function ‘int foo()’:
        test.C:1:14: warning: no return statement in function returning non-void [-Wreturn-type]
         int foo () { }
                      ^
        test.C:2:46: error: template instantiation depth exceeds maximum of 900 (use -ftemplate-depth= to increase the maximum) instantiating ‘struct X<100>’
         template <int N> struct X { static const int value = X<N-1>::value; }; template struct X<1000>;
                                                      ^
        test.C:2:46:   recursively required from ‘const int X<999>::value’
        test.C:2:46:   required from ‘const int X<1000>::value’
        test.C:2:88:   required from here
    
        test.C:2:46: error: incomplete type ‘X<100>’ used in nested name specifier
        
  • 增加新的 pragma: #pragma GCC ivdep, 讓使用者告訴 GCC 迭代間無相依性, 來避免產生過度保守的 SIMD 指令.
  • 支援 Cilk Plus 可透過 -fcilkplus 來使用, Cilk Plus 是一組 C/C++ 語言的語言延伸(extension) 讓其支援資料層級或任務層級的平行化, 目前的實作是遵照 ABI 1.2 版, 除了 _Cilk_for 外已經全數實作完成.

C

  • 支援 ISO C11 的 Atomic (_Atomic type specifier, qualifier 以及 stdatomic.h).
  • 支援 C11 _Generice 關鍵字.
  • 支援 C11 thread-local storage (_Thread_local, 類似 GNU C 中的 __thread)
  • 幾乎完全支援 C11, 完成度大概就跟 C99 的支援度差不多:
    修正一堆 bug 以及新增 C11 相關的關鍵字 (除了一些 Corner Case 外幾乎全部支援, 可透過 fextended-identifiers 開啟), 浮點數相關運算 (大部分但不是完整的 C99 功能, 可參考 C99 的附錄 F 及 G) 以及一些選擇性功能 Annexes K (Bounds-checking interface) 以及 L (Analyzability).
  • 增加新的 C extension: __auto_type, 可提供部份類似 C++11 中 auto 的功能.

C++

  • 目前 G++ 實作 C++1y 中的 return type deduction 已經更新到 N3638, 這個提案目前已經被接受到 working paper 中. Most notably, 可透過 decltype(auto) 來取得 decltype semantics 的回傳型別, 相較之下一般的 auto 則是取得 template argument deduction semantics.
    int& f();
             auto  i1 = f(); // int
    decltype(auto) i2 = f(); // int&
    
  • G++ 支援 C++1y 的 lambda capture initializers:
    [x = 42]{ ... };
    
    事實上他已經在 GCC 4.5 中被支援了, 使用 -std=c++1y 的話 compiler 就不會抱怨, 目前支援 parenthesized 初始化以及 brace-enclosed 初始化的形式.
  • G++ 支援 C++1y 的可變長度陣列 (VLA: variable length arrays), 事實上 G++ 支援 GNU/C99-style VLA 很久了~但現在支援了 initializers 以及 lambda capture by reference. 在 C++1y 模式下 G++ 會抱怨某些 VLA 還沒被正式收錄(permitted) 到 Draft Standard 中的功能, 例如指向 VLA 的 type 或著是對 VLA 使用 sizeof operator. 這部份的功能不會出現在 C++14 中, 但有可能出現在更未來的 C++17 中.
    void f(int n) {
      int a[n] = { 1, 2, 3 }; // throws std::bad_array_length if n < 3
      [&a]{ for (int i : a) { cout << i << endl; } }();
      &a; // error, taking address of VLA
    }
    
  • G++ 支援 C++1y 中的 [[deprecated]] 屬性, 就像 [[gnu::deprecated]] 屬性一樣, 類別或函數都可以套用:
    class A;
    int bar(int n);
    #if __cplusplus > 201103
    class [[deprecated("A is deprecated in C++14; Use B instead")]] A;
    [[deprecated("bar is unsafe; use foo() instead")]]
    int bar(int n);
    
    int foo(int n);
    class B;
    #endif
    A aa; // warning: 'A' is deprecated : A is deprecated in C++14; Use B instead
    int j = bar(2); // warning: 'int bar(int)' is deprecated : bar is unsafe; use foo() instead
    
  • G++ 支援 C++1y 中的 digit separators. 長的數字常數可被 ' 來切開表示藉此提升可讀性:
    int i = 1048576;
    int j = 1'048'576;
    int k = 0x10'0000;
    int m = 0'004'000'000;
    int n = 0b0001'0000'0000'0000'0000'0000;
    
    double x = 1.602'176'565e-19;
    double y = 1.602'176'565e-1'9;
    
  • G++ 支援 C++1y polymorphic(多型) lambda
    // a functional object that will increment any type
    auto incr = [](auto x) { return x++; };
    

Runtime Library (libstdc++)

  • 改善 C++11 的支援, 其中包含:
    • 支援 <regex>
    • <map>, <set>, <unordered_map> 以及 <unordered_set> 現在都符合 allocator-aware 的需求了!
  • 改善許多實驗性的 C++14 支援, 包含:
    • 沒有 const 的 constexpr 成員函數.
    • 實作 std::exchange()
    • 支援透過型別取值的 tuple
    • 實作 std::make_unique
    • 實作 std::shared_lock
    • 讓 std::result_of SFINAE-friendly;
    • integral_constant 新增 operator()
    • 新增標準函式庫中的使用者字定義字面常數 (user-defined literals), 例如 std::basic_string, std::chrono::duration, 以及 std::complex
    • 為 std::equal 及 std::mismatch 新增 non-modifying sequence oprations 的版本 oprations
    • 新增針對 quoted strings 的 IO manipulators.
    • 新增 constexpr member 到 <utility>, <complex>, <chrono> 以及部份容器.
    • 新增編譯時期的 std::integer_sequence
    • 新增 cleaner transformation traits
    • 讓 <functional> 中的 operator functors 更好用而且更一般化 (more generic).
    • 實作 std::experimental::optional.
    • 實作 std::experimental::string_view.
    • std::copy_exception 這個非標準的函數被 deprecate , 在未來會移除, 請換使用 std::make_exception_ptr.

新的 Target 以及 Target 相關改善:

AArch64

  • 新增 ARMv8-A 加密以及 CRC 的 intrinsic, 目前要透過 -march=armv8-a+crc 以及 -march=armv8-a+crypto 來開啟.
  • 初步支援 ILP32, 目前可透過 -mabi=ilp32 來開啟, ILP32 目前是個實驗性質的 ABI, 尚在 beta 階段.
  • 涵蓋更多的 ISA: 現在包含 SIMD extensions 以及相關的 intrinsic.
  • AArch64 目前預設使用 Local Register Allocator (LRA)!
  • 在 AArch64 中 REE (Redundant extension elimination) Pass 預設開啟.
  • 改善 Cortex-A53 以及 Cortex-A57 的微調(Tuning).
  • 初步 big.LITTLE 微調可整合 Cortex-A57 以及 Cortex-A53, 可透過 -mcpu=cortex-a57.cortex-a53 來開啟.
  • 大量的基礎建設改善使得 ARM 以及 AArch64 的 Code Gen 獲得不思議的改善.

ARM

  • 目前不再預設 Advanced SIMD (Neon) 來操作 64-bit scalar, 這部份已經被證實只能改善小幅度的 case, 可透過 -mneon-for-64bits 來開啟.
  • 更多的 ARMv8-A architecture 支援, 目前 Thumb32 指令的嚴格限制版的 IT block 可透過 -mrestrict-it 來開啟, 並且搭配 -march=armv7-a 或 -march=armv7ve 來產生對 ARMv8-A deprecated instructions 更佳的相容性.
  • 支援一些 ARMv7ve 相關的架構, 可透過 -march=armv7ve 開啟.
  • 新增 ARMv8-A 加密以及 CRC 的 intrinsic, 目前要透過 -march=armv8-a+crc 以及 -mfpu=crypto-neon-fp-armv8 來開啟.
  • ARM 現在預設開啟 LRA, 可以透過 -mno-lra 來關掉, 這個選項指示拿來方便發現 bug 或 performance regressions 用的, 未來會拔除.
  • 加入新選項 -mslow-flash-data: 可改善在 ARMv7-M profile cores 上的程式效能.
  • 加入新選項 -mpic-data-is-text-relative 來讓 text segment 可透過相對位址存取 data segment, 目前除了 VxWorks RTP 外預設開啟.
  • 大量的基礎建設改善使得 ARM 以及 AArch64 的 Code Gen 獲得不思議的改善.
  • GCC 現在支援 Cortex-A12 以及 Cortex-R7: -mcpu=cortex-a12 及 -mcpu=cortex-r7 選項.
  • GCC 現在支援 Cortex-A57 以及 Cortex-A53: -mcpu=cortex-a57 及 -mcpu=cortex-a53 選項.
  • 初步 big.LITTLE 微調可整合 Cortex-A57 以及 Cortex-A53, 可透過 -mcpu=cortex-a57.cortex-a53 來開啟.
  • 新增更多針對 Cortex-A15 以及 the Cortex-M4 的效能改善
  • 針對 M-profile processors 改善 Thumb2 的 Code Gen

IA-32/x86-64

  • 支援 AVX-512 指令集, 支援 inline assembly, 新的 registers 以及延伸舊有的東西, 新的 intrinsics (包含相對應的 testsuite), 以及基本的自動向量化. AVX-512 指令目前可透過以下 GCC 選項開啟: AVX-512 foundation instructions: -mavx512f, AVX-512 prefetch instructions: -mavx512pf, AVX-512 exponential 以及 reciprocal instructions: -mavx512er, AVX-512 conflict detection instructions: -mavx512cd.
  • 現在可以針對特定函數打上相對應的 target attribute 而不用整個檔案用 -mxxx 選項來使得整的檔案來指定特定 target, 這使得單一檔案中容易實作 Function Multiversioning.
  • GCC 現在支援 Silvermont: -march=silvermont.
  • GCC 現在支援 Broadwell: -march=broadwell.
  • -march 部份重新命名: -march=nehalem, westmere, sandybridge, ivybridge, haswell, bonnell
  • -march=generic 現在針對 Intel core and AMD Bulldozer architectures 提供更好的效能, AMD K7, K8, Intel Pentium-M, and Pentium4 based 的 CPU 效能不再重要.
  • -mtune=intel 可以針對大部分的 Intel cpu 產生更好的 Code.
  • 支援將 32-bit assembly instructions encode 成 16-bit, 可透過 -m32 開啟.
  • 更好的 memcpy 及 memset, 會針對長度以及內容物產生更短的 alignment prologues.
  • -mno-accumulate-outgoing-args 現在會產生更精確的 unwind 資訊. 針對 -Os, Argument accumulation 將預設關閉.
  • 支援 AMD 家族的 15h processors (Excavator core): -march=bdver4 以及 -mtune=bdver4.

NDS32

  • 新的 nds32 port, 一個來自台灣晶心科技的 32-bit 架構.
  • 目前的 Porting 初步支援 V2, V3, V3m instruction set architectures.

2013年3月23日 星期六

GCC 4.8 Release Note

隔了整整一年的 GCC 4.8註1 終於 Release 了!

以下大略整理一下此次更新的部份, 主要整理自官方 Release 頁面[1], 但由於個人偏好, 非ARM, C/C++ 將不會出現在下面:

一般性

  1. GCC 切換使用 C++ 實作, 回不去了, 但要注意 GCC 中目前政策是只使用 C++2003 的 subset, 禁止列表可參考 [3]
  2. GCC 現在採用更激烈(aggressive)註2的最佳化, 同時也增加 -fno-aggressive-loop-optimizations 這個選項來避免過激行為
  3. DWARF4 為預設產生的 Debug Info
  4. 新的選項 -Og 可以產生有最佳化但又方便 Debug 的 Code, 以後可以不用下 -O0 -g3 可以換下 -Og -g3 了
  5. 新的最佳化選項 -ftree-partial-pre , 預設開啟於 -O3
  6. Struct 及 Matrix 的 reorg 功能被拔掉了(-fipa-struct-reorg 及 -fipa-matrix-reorg), 因為有機率產生出爛掉的東西, 而且 lto 時會爛掉
  7. LTO 大幅改善
  8. AddressSanitizer 及 ThreadSanitize 正式加入, 分別使用 -fsanitize=address 及 -fsanitize=thread 來開啟, AddressSanitizer 功能簡介可參照[5]

語言相關

  1. 錯誤訊息會用 ^ 來標示出錯誤地點, 例如以下是少分號的範例
    #include<stdio.h>
    #include<stdlib.h>
    
    int main(){
      return 0
    }
    
    GCC 4.8 的錯誤訊息
    test.c: In function ‘main’:
    test.c:6:1: error: expected ‘;’ before ‘}’ token
     }
     ^
    
    但跟 clang 比起來似乎還是沒做的很好, clang 的輸出範例:
    test.c:5:11: error: expected ';' after return statement
      return 0
              ^
              ;
    
    更多比較可參考[6]
  2. 新的選項 -ftrack-macro-expansion= 可以在讓你在顯示錯誤訊息時展開 Marco N 層, 範例:
    #define MARCO_(a, b) a b
    #define MARCO(a, b) MARCO_(a, b)
    
    void f(int a, int b) {
      MARCO(a, b);
    }
    
    GCC 4.7.2 的錯誤訊息
    marco.c: In function ‘f’:
    marco.c:5:3: error: expected ‘;’ before ‘b’
    
    GCC 4.8 的錯誤訊息
    marco.c: In function ‘f’:
    marco.c:6:12: error: expected ‘;’ before ‘b’
       MARCO(a, b);
                ^
    marco.c:2:24: note: in definition of macro 'MARCO_'
     #define MARCO_(a, b) a b
                            ^
    marco.c:6:3: note: in expansion of macro 'MARCO'
       MARCO(a, b);
       ^
    
    clang 的錯誤訊息
    marco.c:6:12: error: expected ';' after expression
      MARCO(a, b);
               ^
    marco.c:3:31: note: expanded from macro 'MARCO'
    #define MARCO(a, b) MARCO_(a, b)
                                  ^
    marco.c:2:24: note: expanded from macro 'MARCO_'
    #define MARCO_(a, b) a b
                           ^
    
  3. 新的警告 -Wsizeof-pointer-memaccess, 當該填 size 時你直接 sizeof(pointer) 就會跟你抱怨, 範例如下:
    #include <string.h>
    
    void f(int *p){
      memset(p, 1, sizeof(p));
    }
    
    然後開 -Wall 或 -Wsizeof-pointer-memaccess 會跟你抱怨
    memset.c: In function ‘f’:
    memset.c:4:22: warning: argument to ‘sizeof’ in ‘memset’ call is the same expression as the destination; did you mean to dereference it? [-Wsizeof-pointer-memaccess]
       memset(p, 1, sizeof(p));
                          ^
    
    不過還是要說一下 clang 吐出來的資訊比較清楚
    memset.c:4:23: warning: 'memset' call operates on objects of type 'int' while the size is based
          on a different type 'int *' [-Wsizeof-pointer-memaccess]
      memset(p, 1, sizeof(p));
             ~            ^
    memset.c:4:23: note: did you mean to dereference the argument to 'sizeof' (and multiply it by the
          number of elements)?
      memset(p, 1, sizeof(p));
                          ^
    

C++11專區

  1. thread_local 關鍵字實作了!
  2. 的 attribute 及 alignment 實作, 範例:
    [[noreturn]] void f();
    
    alignas(double) int i;
    

ARM

  1. AArch 64 支援
  2. ARM 的 AAPCS ABI 有稍微更動, 有關 vector types 部份的 ABI 將無法與舊的 GCC 產生的 Code 相容
  3. 現在 GCC 會想辦法產生 VFMA, VFMS, REVSH 及 REV16 指令
  4. 新的 Inst Scheduler 會考慮 Register Pressure, 不喜歡的話可透過 -fno-sched-pressure 關掉
  5. 支援 Marvell 的 iWMMX2 SIMD, 可透過指定 -mcpu=iwmmxt2 來開啟
註1: GCC 4.7 於 2012-03-22 Release 而 4.8 是 2013-03-22, 剛好一年
註2: Aggressive 這個字眼在 Compiler 領域中通常代表需要花較多時間分析及最佳化, 並且也有可能不會比較好(指 Code Size 或 Performance)

Reference

[1] GCC 4.8 Release 官方網頁
[2] C++ Conversion(GCC Wiki)
[3] http://gcc.gnu.org/wiki/CppConventions
[4] GCC and C vs C++ Speed, Measured
[5] Address-sanitizer : 新的 Android 快速記憶體錯誤偵測工具
[6] C++ Diagnostic Survey

2013年1月11日 星期五

Address-sanitizer : 新的 Android 快速記憶體錯誤偵測工具


簡介

Android 在 Jerry Bean 的時候加入了 Address-sanitizer 的支援

這東西是什麼勒, 說白話一點就是個可以偵測一些記憶體使用錯誤的一個小工具,

有點類似 valgrind 中 memcheck 的功能但是更快更簡單(但相對的某些錯誤無法偵測)

偵測錯誤的類型

廢話不多說, 直接看範例程式跟範例輸出

越界存取(Out of bound)

Heap 越界存取
#include 
#include 
int main(int argc, char **argv) {
  char *x = (char*)malloc(10 * sizeof(char));
  memset(x, 0, 10);
  int res = x[argc * 10];  // BOOOM
  free(x);
  return res;
}
範例執行輸出:
=================================================================
==799== ERROR: AddressSanitizer heap-buffer-overflow on address 0x41255cca at pc 0x2a00055b bp 0xbeff6b0c sp 0xbeff6b08
READ of size 1 at 0x41255cca thread T0
    #0 0x40022a4b (/system/lib/libasan_preload.so+0x8a4b)
    #1 0x40023e77 (/system/lib/libasan_preload.so+0x9e77)
    #2 0x4001c947 (/system/lib/libasan_preload.so+0x2947)
    #3 0x2a000559 (/system/bin/heap-out-of-bounds+0x559)
    #4 0x4114371d (/system/lib/libc.so+0x1271d)
0x41255cca is located 0 bytes to the right of 10-byte region [0x41255cc0,0x41255cca)
allocated by thread T0 here:
    #0 0x40022a4b (/system/lib/libasan_preload.so+0x8a4b)
    #1 0x40022aff (/system/lib/libasan_preload.so+0x8aff)
    #2 0x2a0004e9 (/system/bin/heap-out-of-bounds+0x4e9)
    #3 0x4114371d (/system/lib/libc.so+0x1271d)
Shadow byte and word:
  0x0824ab99: 2
  0x0824ab98: 00 02 fb fb
More shadow bytes:
  0x0824ab88: 00 00 00 00
  0x0824ab8c: 04 fb fb fb
  0x0824ab90: fa fa fa fa
  0x0824ab94: fa fa fa fa
=>0x0824ab98: 00 02 fb fb
  0x0824ab9c: fb fb fb fb
  0x0824aba0: fa fa fa fa
  0x0824aba4: fa fa fa fa
  0x0824aba8: fa fa fa fa
Stats: 0M malloced (0M for red zones) by 36 calls
Stats: 0M realloced by 0 calls
Stats: 0M freed by 0 calls
Stats: 0M really freed by 0 calls
Stats: 2M (642 full pages) mmaped in 5 calls
  mmaps   by size class: 7:4095; 8:2047; 11:255; 12:128; 13:64; 
  mallocs by size class: 7:27; 8:4; 11:2; 12:2; 13:1; 
  frees   by size class: 
  rfrees  by size class: 
Stats: malloc large: 0 small slow: 5
==799== ABORTING
全域變數越界存取
int global_array[100] = {-1};
int main(int argc, char **argv) {
  return global_array[argc + 100];  // BOOM
}
範例執行輸出:
=================================================================
==7161== ERROR: AddressSanitizer global-buffer-overflow on address 0x2a002194 at pc 0x2a00051b bp 0xbeeafb0c sp 0xbeeafb08
READ of size 4 at 0x2a002194 thread T0
    #0 0x40022a4b (/system/lib/libasan_preload.so+0x8a4b)
    #1 0x40023e77 (/system/lib/libasan_preload.so+0x9e77)
    #2 0x4001c98f (/system/lib/libasan_preload.so+0x298f)
    #3 0x2a000519 (/system/bin/global-out-of-bounds+0x519)
    #4 0x4114371d (/system/lib/libc.so+0x1271d)
0x2a002194 is located 4 bytes to the right of global variable 'global_array (external/test/global-out-of-bounds.cpp)' (0x2a002000) of size 400
Shadow byte and word:
  0x05400432: f9
  0x05400430: 00 00 f9 f9
More shadow bytes:
  0x05400420: 00 00 00 00
  0x05400424: 00 00 00 00
  0x05400428: 00 00 00 00
  0x0540042c: 00 00 00 00
=>0x05400430: 00 00 f9 f9
  0x05400434: f9 f9 f9 f9
  0x05400438: 00 00 00 00
  0x0540043c: 00 00 00 00
  0x05400440: 00 00 00 00
Stats: 0M malloced (0M for red zones) by 35 calls
Stats: 0M realloced by 0 calls
Stats: 0M freed by 0 calls
Stats: 0M really freed by 0 calls
Stats: 2M (642 full pages) mmaped in 5 calls
  mmaps   by size class: 7:4095; 8:2047; 11:255; 12:128; 13:64; 
  mallocs by size class: 7:26; 8:4; 11:2; 12:2; 13:1; 
  frees   by size class: 
  rfrees  by size class: 
Stats: malloc large: 0 small slow: 5
==7161== ABORTING
區域變數越界存取
int main(int argc, char **argv) {
  int stack_array[100];
  stack_array[1] = 0;
  return stack_array[argc + 100];  // BOOM
}
範例執行輸出:
=================================================================
==7165== ERROR: AddressSanitizer stack-buffer-overflow on address 0xbed0bad4 at pc 0x2a000557 bp 0xbed0b914 sp 0xbed0b910
READ of size 4 at 0xbed0bad4 thread T0
    #0 0x40022a4b (/system/lib/libasan_preload.so+0x8a4b)
    #1 0x40023e77 (/system/lib/libasan_preload.so+0x9e77)
    #2 0x4001c98f (/system/lib/libasan_preload.so+0x298f)
    #3 0x2a000555 (/system/bin/stack-out-of-bounds+0x555)
    #4 0x4114371d (/system/lib/libc.so+0x1271d)
Address 0xbed0bad4 is located at offset 436 in frame 
of T0's stack: This frame has 1 object(s): [32, 432) 'stack_array' HINT: this may be a false positive if your program uses some custom stack unwind mechanism (longjmp and C++ exceptions *are* supported) Shadow byte and word: 0x17da175a: f4 0x17da1758: 00 00 f4 f4 More shadow bytes: 0x17da1748: 00 00 00 00 0x17da174c: 00 00 00 00 0x17da1750: 00 00 00 00 0x17da1754: 00 00 00 00 =>0x17da1758: 00 00 f4 f4 0x17da175c: f3 f3 f3 f3 0x17da1760: 00 00 00 00 0x17da1764: 00 00 00 00 0x17da1768: 00 00 00 00 Stats: 0M malloced (0M for red zones) by 35 calls Stats: 0M realloced by 0 calls Stats: 0M freed by 0 calls Stats: 0M really freed by 0 calls Stats: 2M (642 full pages) mmaped in 5 calls mmaps by size class: 7:4095; 8:2047; 11:255; 12:128; 13:64; mallocs by size class: 7:26; 8:4; 11:2; 12:2; 13:1; frees by size class: rfrees by size class: Stats: malloc large: 0 small slow: 5 ==7165== ABORTING

釋放後使用 (Use after free)

#include 
int main() {
  char *x = (char*)malloc(10 * sizeof(char*));
  free(x);
  return x[5];
}
範例執行輸出:
=================================================================
==7169== ERROR: AddressSanitizer heap-use-after-free on address 0x41255cc5 at pc 0x2a0004cd bp 0xbece1b1c sp 0xbece1b18
READ of size 1 at 0x41255cc5 thread T0
    #0 0x40022a4b (/system/lib/libasan_preload.so+0x8a4b)
    #1 0x40023e77 (/system/lib/libasan_preload.so+0x9e77)
    #2 0x4001c947 (/system/lib/libasan_preload.so+0x2947)
    #3 0x2a0004cb (/system/bin/use-after-free+0x4cb)
    #4 0x4114371d (/system/lib/libc.so+0x1271d)
0x41255cc5 is located 5 bytes inside of 40-byte region [0x41255cc0,0x41255ce8)
freed by thread T0 here:
    #0 0x40022a4b (/system/lib/libasan_preload.so+0x8a4b)
    #1 0x40022a97 (/system/lib/libasan_preload.so+0x8a97)
    #2 0x2a0004b1 (/system/bin/use-after-free+0x4b1)
    #3 0x4114371d (/system/lib/libc.so+0x1271d)
previously allocated by thread T0 here:
    #0 0x40022a4b (/system/lib/libasan_preload.so+0x8a4b)
    #1 0x40022aff (/system/lib/libasan_preload.so+0x8aff)
    #2 0x2a0004ab (/system/bin/use-after-free+0x4ab)
    #3 0x4114371d (/system/lib/libc.so+0x1271d)
Shadow byte and word:
  0x0824ab98: fd
  0x0824ab98: fd fd fd fd
More shadow bytes:
  0x0824ab88: 00 00 00 00
  0x0824ab8c: 04 fb fb fb
  0x0824ab90: fa fa fa fa
  0x0824ab94: fa fa fa fa
=>0x0824ab98: fd fd fd fd
  0x0824ab9c: fd fd fd fd
  0x0824aba0: fa fa fa fa
  0x0824aba4: fa fa fa fa
  0x0824aba8: fa fa fa fa
Stats: 0M malloced (0M for red zones) by 36 calls
Stats: 0M realloced by 0 calls
Stats: 0M freed by 1 calls
Stats: 0M really freed by 0 calls
Stats: 2M (642 full pages) mmaped in 5 calls
  mmaps   by size class: 7:4095; 8:2047; 11:255; 12:128; 13:64; 
  mallocs by size class: 7:27; 8:4; 11:2; 12:2; 13:1; 
  frees   by size class: 7:1; 
  rfrees  by size class: 
Stats: malloc large: 0 small slow: 5
==7169== ABORTING


使用方式

使用方式出乎意料的簡單
只要在你的 Android.mk 裡面加一行
LOCAL_ADDRESS_SANITIZER := true
主要原因在於 Address-sanitizer 也是 google 自家弄的東西, 所以整合的還算不錯

另外有開 Address-sanitizer 有個副作用就是會強迫開啟 clang 模式

什麼意思勒? 就是你的 code 變得不是使用 gcc 編譯而是 clang/llvm

若 code 有用到 gcc 特有的東西的話可能就會炸掉:P

另外 libasan_preload.so 也記得要推到 /system/lib 的地方, 這是 Address-sanitizer run-time library, 可以去 out/target/product/*/system/lib/ 裡面撈撈

如果找不到開了 LOCAL_ADDRESS_SANITIZER := true 但建置系統跟你抱怨 libasan_preload.so 或 libasan.a 的話,
就切到 external/compiler-rt/lib/asan/ 自行建置一下, 或著是用 mmm 讓建置系統自己拉相依關係


錯誤報告的解讀

接著這部份最重要的就是他吐出那團錯誤報告要怎麼讀
就直接從最複雜的釋放後使用 (Use after free)的報告來講解

=================================================================
==7169== ERROR: AddressSanitizer heap-use-after-free on address 0x41255cc5 at pc 0x2a0004cd bp 0xbece1b1c sp 0xbece1b18
首先第一個部份是會跟你報告什麼類型的錯誤, 如這邊是 heap-use-after-free, 另外還有 stack-buffer-overflow, global-buffer-overflow, heap-buffer-overflow 這幾種錯誤類型, 分別對到 stack, global 及 heap 的越界存取

on address 0x41255cc5 部份是你要存取的記憶體位置
at pc 0x2a0004cd bp 0xbece1b1c sp 0xbece1b18 部份則是說發生錯誤的 program counter 值是多少, 以及 frame pointer 跟 stack pointer 的內容

READ of size 1 at 0x41255cc5 thread T0
    #0 0x40022a4b (/system/lib/libasan_preload.so+0x8a4b)
    #1 0x40023e77 (/system/lib/libasan_preload.so+0x9e77)
    #2 0x4001c947 (/system/lib/libasan_preload.so+0x2947)
    #3 0x2a0004cb (/system/bin/use-after-free+0x4cb)
    #4 0x4114371d (/system/lib/libc.so+0x1271d)
0x41255cc5 is located 5 bytes inside of 40-byte region [0x41255cc0,0x41255ce8)
接著很貼心的也會將你炸掉的地方的 call stack 吐出來
只有 pc 值不知道在哪的話就呼叫 addr2line 這個工具來幫忙就可以了
例如你想查 /system/bin/use-after-free+0x4cb 對應的行號是多少則輸入以下指令即可
addr2line -e out/target/product/*/symbols/system/bin/use-after-free -a 0x4cb
* 自行帶入目前建置的 device name, -e 後面塞檔案, -a 後面塞位址, call stack 加號後面的東西

freed by thread T0 here:
    #0 0x40022a4b (/system/lib/libasan_preload.so+0x8a4b)
    #1 0x40022a97 (/system/lib/libasan_preload.so+0x8a97)
    #2 0x2a0004b1 (/system/bin/use-after-free+0x4b1)
    #3 0x4114371d (/system/lib/libc.so+0x1271d)
然後也會吐出 free 的 call stack, 解讀方式同上
previously allocated by thread T0 here:
    #0 0x40022a4b (/system/lib/libasan_preload.so+0x8a4b)
    #1 0x40022aff (/system/lib/libasan_preload.so+0x8aff)
    #2 0x2a0004ab (/system/bin/use-after-free+0x4ab)
    #3 0x4114371d (/system/lib/libc.so+0x1271d)
malloc 的地方也會有 call stack 吐出來
Shadow byte and word:
  0x0824ab98: fd
  0x0824ab98: fd fd fd fd
More shadow bytes:
  0x0824ab88: 00 00 00 00
  0x0824ab8c: 04 fb fb fb
  0x0824ab90: fa fa fa fa
  0x0824ab94: fa fa fa fa
=>0x0824ab98: fd fd fd fd
  0x0824ab9c: fd fd fd fd
  0x0824aba0: fa fa fa fa
  0x0824aba4: fa fa fa fa
  0x0824aba8: fa fa fa fa
Stats: 0M malloced (0M for red zones) by 36 calls
Stats: 0M realloced by 0 calls
Stats: 0M freed by 1 calls
Stats: 0M really freed by 0 calls
Stats: 2M (642 full pages) mmaped in 5 calls
  mmaps   by size class: 7:4095; 8:2047; 11:255; 12:128; 13:64; 
  mallocs by size class: 7:27; 8:4; 11:2; 12:2; 13:1; 
  frees   by size class: 7:1; 
  rfrees  by size class: 
Stats: malloc large: 0 small slow: 5
==7169== ABORTING
最後這團東西看不懂就算了, Address-sanitizer 的內部資訊, 有興趣去挖他們簡報出來就有詳細一點的範例加解釋了

副作用及使用限制

Address-sanitizer 使用時主要有一些額外的負擔:
  1. 每次 new/malloc 會多出額外的 128-255 byte 作為警戒區(Red Zone), 也就是越界時爆炸的引發器:P
  2. Stack 則是每個區域變數會吃掉 32~63 byte 作為警戒區(Red Zone)
  3. 全域變數也是每個變數會吃掉 32~63 byte 作為警戒區(Red Zone)
  4. 一般而言會吃掉 2~4 倍的記憶體, 但最差情況也有可能高達 20 倍的記憶體消耗
  5. Stack 會幾乎長三倍大以上
  6. Address Space 會直接被吃掉八分之一, 32 位元情況下會吃掉 0.5 G, 64 位元則會吃掉 16T, 但僅消耗定址空間, 不會實際耗費掉記憶體 (該區域被設定為不可讀不可寫, 存取到會直接炸掉)
其最大的使用限制是當 Address-sanitizer 偵測到任何記憶體錯誤時就會馬上中斷程式的執行,

這是正常現象, 也因為這樣的設計使得此工具簡單快速又方便

與 Valgrind 的比較表格

Valgrind Address-sanitizer
Heap out-of-bounds
Heap 越界存取
Yes Yes
Stack out-of-bounds
Stack 越界存取
No Yes
Global out-of-bounds
全域變數越界存取
No Yes
Use-after-free
釋放(free/delete)後使用
Yes Yes
Use-after-return
回傳後使用
(例如回傳區域變數指標)
No Sometimes/YES
Uninitialized reads
讀取未初始化的值
Yes No(註1)
Overhead
對程式影響(變慢幾倍)
10x-30x 1.5x-3x
Platforms
可使用平台
Linux, Mac Same as GCC/LLVM (註2)
註1: 這部份有 Address-sanitizer 的好兄弟 Memory-sanitizer 可以處理, 但尚未納入 Android 中
註2: GCC 部份於 2012/11/01 正式納入 Address-sanitizer[4], 可用度沒玩過不太確定, 在 Android 中要開 Address-sanitizer 就是要用 clang 就是了:P...

小結

Address-sanitizer 算是一個輕量級的好用小工具, 根據開發者所釋放出來的簡報 Google 內部自己也有使用此工具並且抓出了上千個 bug, 包含 llvm, gcc, vim, firefox 等大型 Open Source 軟體的記憶體使用 bug, 目前 Address-sanitizer 還有一些相關的兄弟如 Memory-sanitizer 及 Thread-sanitizer, 分別可以偵測 Uninitialized reads 及 Race condition, 不過在 Android 中目前還沒引進, 但都 Google 自家人弄的, 可以期待之後應該是會引進到 Android 中, 讓一些常見的記憶體錯誤可以早期發現:)

參考鍊結


[1] http://code.google.com/p/address-sanitizer
[2] http://llvm.org/devmtg/2011-11/Serebryany_FindingRacesMemoryErrors.pdf
[3] http://address-sanitizer.googlecode.com/files/address_sanity_checker.pdf
[4] http://gcc.gnu.org/ml/gcc/2012-11/msg00016.html

2013年1月8日 星期二

快快樂樂 Makefile 入門教學 Part.II 泛用規則/萬用字元應用

這邊文章主要是 快快樂樂 Makefile 入門教學 的續集, 主要針對的對象依然初學者

要談的主題則是在 Makefile 中萬用字元跟泛用規則的應用, 可以讓整個 Makefile 盡量變得簡潔有力

並且一窺 Makefile 的內部推導邏輯

一堆 .o 一堆 rule

當使用 Makefile 一段時間後, 大致上會碰到的問題就是如果你有一堆 .c 要編譯成 .o 檔的話

手刻一堆 xxx.c -> xxx.o 的 rule 會讓人挺困擾的, 尤其是通常都是 gcc xxx.c -c -o xxx.o 之類的

但如果你是個懶惰牌的好 Programmer 那麼你可能會嘗試寫一些萬用字元的 rule

於是你寫下了以下的 rule
*.o: *.c
 gcc -o *.o  -c *.c
嗚呼! Let's Make

然後你會發現如果有兩個 .c 以上的話就會發現:「幹!不能動」
gcc -o *.o  -c *.c
gcc: fatal error: cannot specify -o with -c, -S or -E with multiple files
compilation terminated.
make: *** [*.o] Error 4
錯誤訊息跟我們抱怨如果指定 -o 加上 -c, -S 或 -E 的話就不能輸入多個檔案

現在假設你還沒有任何的 .o 擋在該資料夾並且有 a.c b.c 要編譯

shell 真實看到的指令變成這樣
gcc -o  -c a.c b.c
然後就引發剛才的錯誤了XD

救命 難道只能寫一堆 rule 了嗎?

萬用字元不能動的主要原因在於 Makefile 為了避免與 shell 使用的萬用字元 * 衝突而選擇了另一個

並且 *.c 跟 *.o 的 * 的部份哪知道會不會一樣, 所以 Makefile 有自己的萬用字元系統
*.o: *.c
 gcc -o *.o  -c *.c
所以在 Makefile 中是使用 % 作為萬用字元, 因此第一直覺會下出類似下面的建置規則
%.o: %.c
 gcc -o %.o  -c %.c
敲入 Make 後發現建置規則是有抓到了, 但是建置指令並沒有將 % 替換程式適當的字串
gcc -o %.o  -c %.c
gcc: error: %.c: No such file or directory
gcc: fatal error: no input files
compilation terminated.
make: *** [a.o] Error 4
主要原因在於建置規則中不是用 % 來表達該萬用字元是用 $@ 及 $^
(當然還有一堆其它 $ 開頭的神奇變數, 不過常用的大概就這些)

Makefile 萬用字元使用方式 %, $@, $^


$@, $^這幾個符號各代表不同意義, 但直接先看範例再來解說吧

%.o: %.c
 gcc -o $@  -c $^
其中 $@ 代表建置目標, 所以以要建置 a.o 來看的話 $@ 等於 a.o
接著 $^ 則代表相依目標, 建置 a.o 時會替換成 a.c

從此之後就只要寫一個 .c -> .o 的建置規則在加上一個 link 成執行檔的規則即可!



這些萬用規則也可以拿來作一些方便的應用, 例如有時候你會將 debug 版本跟一般版跟分開放置
則你寫兩個 rule 則可以建置出不同編譯參數且在不同資料夾生成
debug/%.o: %.c
 gcc -o $@  -c $^ -g

release/%.o: %.c
 gcc -o $@  -c $^ -O2


小結

Makefile 會了萬用字元後, 建置規則就可以寫的簡潔又好看

在加上 .c/.cpp 與 .h/.hpp 相依性自動建立就很完整了

預計會在弄一篇文章寫產生 .c/.cpp 與 .h/.hpp 相依性的:P

2012年11月3日 星期六

使用 buildroot 建置可用於 SimpleScalar 的 ARM toolchain

0. 前言

剛好最近有人來信詢問,就順便整理先前的一些文件一下,雖然 SimpleScalar 相對於現階段 ARM 的發展而言還停留在 ARMv4 的指令集,但若是在學術用途方面的話,還算是個相當有用的工具。

但在研究初期最容易遇到的問題就是 toolchain 的問題,官方附贈的過舊,要自行建置則網路上資訊過於零散,建立出來也不見得能用...

而在這篇文章中會透過 buildroot 來建立一串基礎的 ARM toolchain ,來簡化整個建置的過程,若一般用途而言這樣的方式就夠了

另外此篇文章假設你已經有一個已經建立好的 SimpleScalar ARM 可以用了


1. 使用 buildroot 建置 ARM toolchain

buildroot 的官方網站是 http://buildroot.uclibc.org/download.html
以下操作以 2012.08 的版本為例子,其它版本操作上大同小異,但部份路徑可能會有不同,第一次建議使用相同版本操作較為保險

目前作過測試是到 gcc 4.6 都可以在 SimpleScalar ARM 上運作沒有問題

gcc 4.7 因為不支援 ARM oabi 的部份所以會有點問題 (需要對 SimpleScalar 作一些修改)

i. 抓取 buildroot

wget http://buildroot.uclibc.org/downloads/buildroot-2012.08.tar.bz2

ii. 解壓縮 buildroot

tar -jxf buildroot-2012.08.tar.bz2

iii. 組態 buildroot

打入 make menuconfig 進入選單,會有 console 版本的選單出現
make menuconfig
# 接著會進入選單
Target Architecture 按 enter 選 ARM (little endian)
Target Architecture Variant 保留 generic_arm ,simplescalar 只支援 arm v4
Target ABI 選 oabi <--這是大重點,eabi 會不能動
Toolchain 點進去可以選 gcc 跟 binutils 版本
  uClibc C library Version 選 uClibc 0.9.31.x ,太新會沒辦法用(需要用到 bx 指令, armv4 不支援)
  需要 g++ 的話在  Enable C++ support 那邊按空白打 *
  需要使用 hard floating point 則 Use software floating point by default 的 * 按空白拿掉
其它看無就不用動

選擇最下面 Save an Alternate Configuration File
你會看到有一串
/home/kito/xxxxxxx/buildroot-2012.08/.config 之類的路徑
不要改
直接按 enter
然後 exit (左右可以移動下面的按鈕)
接著 make
然後轉身去泡杯咖啡,需要一點時間下載跟建置

build 完後在 buildroot 資料夾底下的
./output/host/usr/
這個資料夾
裡面就有 arm toolchain 了

iii. 測試 buildroot

先弄一個 Hello World
echo "int main() {printf(\"hello world\");}" > hello.c
編譯 Hello World ,-static 一定要加, simplescalar 只支援 static link
./output/host/usr/bin/arm-linux-gcc hello.c -o hello -static
使用 SimpleScalar ARM 測試,前面的 path 記得自行替換成 simplescalar 的路徑
/<path-to-simplescalar>/sim-uop hello
看到下面那一沱就代表你成功了!

sim-uop: SimpleScalar/ARM Tool Set version 3.0 of November, 2000.
Copyright (c) 1994-2000 by Todd M. Austin.  All Rights Reserved.
This version of SimpleScalar is licensed for academic non-commercial use only.

sim: command line: /home/kito/simulators/simplesim-arm-v2/sim-uop hello

sim: simulation started @ Mon Oct 22 20:25:18 2012, options follow:

sim-safe: This simulator implements a functional simulator.  This
functional simulator is the simplest, most user-friendly simulator in the
simplescalar tool set.  Unlike sim-fast, this functional simulator checks
for all instruction errors, and the implementation is crafted for clarity
rather than speed.

# -config                     # load configuration from a file
# -dumpconfig                 # dump configuration to a file
# -h                    false # print help message
# -v                    false # verbose operation
# -i                    false # start in Dlite debugger
-seed                       1 # random number generator seed (0 for timer seed)
# -q                    false # initialize and terminate immediately
# -chkpt               <null> # restore EIO trace execution from <fname>
# -redir:sim           <null> # redirect simulator output to file
(non-interactive only)
# -redir:prog          <null> # redirect simulated program output to file
-nice                       0 # simulator scheduling priority
-max:inst                   0 # maximum number of inst's to execute
-trigger:inst               0 # trigger instruction

sim: ** starting functional simulation **
warning: unsupported ioctl call: ioctl(21505, ...)
warning: unsupported ioctl call: ioctl(21505, ...)
hello world
sim: ** simulation statistics **
sim_num_insn                   5276 # total number of instructions executed
sim_num_uops                   7095 # total number of UOPs executed
sim_avg_flowlen              1.3448 # uops per instruction
sim_num_refs                   1259 # total number of loads and stores executed
sim_elapsed_time                  1 # total simulation time in seconds
sim_inst_rate             5276.0000 # simulation speed (in insts/sec)
ld_text_base           0x00008094 # program text (code) segment base
ld_text_bound          0x0000c54c # program text (code) segment bound
ld_text_size                  17592 # program text (code) size in bytes
ld_data_base           0x0000c54c # program initialized data segment base
ld_data_bound          0x000177d0 # program initialized data segment bound
ld_data_size                  45700 # program init'ed `.data' and
uninit'ed `.bss' size in bytes
ld_stack_base          0xc0000000 # program stack segment base
(highest address in stack)
ld_stack_size                 16384 # program initial stack size
ld_prog_entry          0x000080b0 # program entry point (initial PC)
ld_environ_base        0xbfffc000 # program environment base address address
ld_target_big_endian              0 # target executable endian-ness,
non-zero if big endian
mem.page_count                   11 # total number of pages allocated
mem.page_mem                    44k # total size of memory pages allocated
mem.ptab_misses                  11 # total first level page table misses
mem.ptab_accesses             63794 # total page table accesses
mem.ptab_miss_rate           0.0002 # first level page table miss rate

2. 將 GCC, Bintuils 及 uClibc 拉出來重新建置

這個段落主要是給需要修改到編譯器的屠龍人士看的,如果你只是要 ARM Toolchain 或著是要修改 Simulator 的話這段可以直接略過。

i. 建立一個要放等一下所有東西的資料夾

mkdir ~/arm-linux-gcc

ii. 把整串 toolchain 拉出來

cp -a buildroot-2012.08/output/host/usr ~/arm-linux-gcc/

iii. 資料夾重新命名為 arm-linux-gcc

mv ~/arm-linux-gcc/usr ~/arm-linux-gcc/arm-linux-gcc

iv. gcc, binutils, uclibc 等 source code 拉出來

cp -a buildroot-2012.08/output/toolchain/gcc-4.5.4 ~/arm-linux-gcc/
cp -a buildroot-2012.08/output/toolchain/uClibc-0.9.31.1/ ~/arm-linux-gcc/
tar -jxf buildroot-2012.08/dl/binutils-2.21.tar.bz2 -C ~/arm-linux-gcc

v. 建立 bulid-* 資料夾準備

cd ~/arm-linux-gcc
mkdir build-gcc
mkdir build-binutils

vi. 重新 build bintuils

cd build-binutils

vii. 開啟bintuils 的 config log 看一下怎麼下

vim ~/buildroot-2012.08/output/build/host-binutils-2.21/config.log
大約在第七行會有類似的東西
  $ ./configure --prefix=/home/kito/buildroot-2012.08/output/host/usr
--sysconfdir=/home/kito/buildroot-2012.08/output/host/etc
--enable-shared --disable-static --disable-multilib --disable-werror
--target=arm-unknown-linux-uclibc --disable-shared --enable-static
--with-sysroot=/home/kito/buildroot-2012.08/output/host/usr/arm-unknown-linux-uclibc/sysroot

viii. 產生新的 configure 參數

接著把 configure 中所有路徑更新一下, 然後 --sysconfdir 可以拿掉
note : 這邊路徑用我家示範,--prefix 跟 --with-sysroot 都要改
../binutils-2.21/configure
--prefix=/home/kito/arm-linux-gcc/arm-linux-gcc --enable-shared
--disable-static --disable-multilib --disable-werror
--target=arm-unknown-linux-uclibc --disable-shared --enable-static
--with-sysroot=/home/kito/arm-linux-gcc/arm-linux-gcc
/arm-unknown-linux-uclibc/sysroot

ix. 建置並安裝 buildroot

make -j8
make install

x. 重新 build gcc

cd build-gcc

xi. 偷看 buildroot config 怎麼下

~/arm-linux-gcc/arm-linux-gcc/bin/arm-linux-gcc -v
應該會吐出下面訊息,看到 Configured with: 那一段,抄過來修改
Using built-in specs.
COLLECT_GCC=/home/kito/buildroots/buildroot-2012.08/output/host/usr/bin/arm-linux-gcc
COLLECT_LTO_WRAPPER=/home/kito/buildroots/buildroot-2012.08/output/host/usr/libexec/gcc/arm-unknown-linux-uclibc/4.5.4/lto-wrapper
Target: arm-unknown-linux-uclibc
Configured with:
/home/kito/buildroots/buildroot-2012.08/output/toolchain/gcc-4.5.4/configure
--prefix=/home/kito/buildroots/buildroot-2012.08/output/host/usr
--build=x86_64-unknown-linux-gnu --host=x86_64-unknown-linux-gnu
--target=arm-unknown-linux-uclibc --enable-languages=c
--with-sysroot=/home/kito/buildroots/buildroot-2012.08/output/host/usr/arm-unknown-linux-uclibc/sysroot
--with-build-time-tools=/home/kito/buildroots/buildroot-2012.08/output/host/usr/arm-unknown-linux-uclibc/bin
--disable-__cxa_atexit --enable-target-optspace --disable-libquadmath
--disable-libgomp --with-gnu-ld --disable-libssp --disable-multilib
--disable-tls --enable-shared
--with-gmp=/home/kito/buildroots/buildroot-2012.08/output/host/usr
--with-mpfr=/home/kito/buildroots/buildroot-2012.08/output/host/usr
--with-mpc=/home/kito/buildroots/buildroot-2012.08/output/host/usr
--disable-nls --enable-threads --disable-decimal-float
--with-float=soft --with-abi=apcs-gnu --disable-largefile
--with-pkgversion='Buildroot 2012.08'
--with-bugurl=http://bugs.buildroot.net/
Thread model: posix
gcc version 4.5.4 (Buildroot 2012.08)

xii. 產生新的 gcc 的 configure 參數

接著把 configure 中所有路徑更新一下, 然後 --with-bugurl, --with-pkgversion 可以拿掉 build 跟 host 也可以拔掉,它會自己抓,這樣整串指令就不會太長
note1 : mpfr, gmp, mpc 如果你系統有裝的話可以使用系統的就好, 如果使用系統的話--with-gmp --with-mpfr --with-mpc 這三個可以拔掉
note2 : 這邊路徑用我家示範,--prefix --with-sysroot 跟 --with-build-time-tools= 都要改
note3 : 需要 g++ 的話,把--enable-languages=c 改成 --enable-languages=c,c++
../gcc-4.5.4/configure --prefix=/home/kito/arm-linux-gcc/arm-linux-gcc
--target=arm-unknown-linux-uclibc --enable-languages=c
--with-sysroot=/home/kito/arm-linux-gcc/arm-linux-gcc/arm-unknown-linux-uclibc/sysroot
--with-build-time-tools=/home/kito/arm-linux-gcc/arm-linux-gcc/arm-unknown-linux-uclibc/bin
--disable-__cxa_atexit --enable-target-optspace --disable-libquadmath
--disable-libgomp --with-gnu-ld --disable-libssp --disable-multilib
--disable-tls --enable-shared --disable-nls --enable-threads
--disable-decimal-float --with-float=soft --with-abi=apcs-gnu
--disable-largefile

xiii. 建置 gcc !

make! 另外注意一下 gcc 4.4 不支援 make -jx 的功能
make -j8
make install

xiv. 檢查 gcc

~/arm-linux-gcc/arm-linux-gcc/bin/arm-linux-gcc -v
輸出會跟第一次很像但 Configured with: 後面已經換成新的參數,這樣就代表有覆蓋掉舊的了
Using built-in specs.
COLLECT_GCC=/home/kito/arm-linux-gcc/arm-linux-gcc/bin/arm-linux-gcc
COLLECT_LTO_WRAPPER=/home/kito/arm-linux-gcc/arm-linux-gcc/libexec/gcc/arm-unknown-linux-uclibc/4.5.4/lto-wrapper
Target: arm-unknown-linux-uclibc
Configured with: ../gcc-4.5.4/configure
--prefix=/home/kito/arm-linux-gcc/arm-linux-gcc
--target=arm-unknown-linux-uclibc --enable-languages=c
--with-sysroot=/home/kito/arm-linux-gcc/arm-linux-gcc/arm-unknown-linux-uclibc/sysroot
--with-build-time-tools=/home/kito/arm-linux-gcc/arm-linux-gcc/arm-unknown-linux-uclibc/bin
--disable-__cxa_atexit --enable-target-optspace --disable-libquadmath
--disable-libgomp --with-gnu-ld --disable-libssp --disable-multilib
--disable-tls --enable-shared --disable-nls --enable-threads
--disable-decimal-float --with-float=soft --with-abi=apcs-gnu
--disable-largefile : (reconfigured) ../gcc-4.5.4/configure
--prefix=/home/kito/arm-linux-gcc/arm-linux-gcc
--target=arm-unknown-linux-uclibc --enable-languages=c,c++
--with-sysroot=/home/kito/arm-linux-gcc/arm-linux-gcc/arm-unknown-linux-uclibc/sysroot
--with-build-time-tools=/home/kito/arm-linux-gcc/arm-linux-gcc/arm-unknown-linux-uclibc/bin
--disable-__cxa_atexit --enable-target-optspace --disable-libquadmath
--disable-libgomp --with-gnu-ld --disable-libssp --disable-multilib
--disable-tls --enable-shared --disable-nls --enable-threads
--disable-decimal-float --with-float=soft --with-abi=apcs-gnu
--disable-largefile
Thread model: posix
gcc version 4.5.4 (GCC)

xv. 重新 build uClibc

cd ~/arm-linux-gcc/uClibc-0.9.31.1

xvi. 將 buildroot 中 linux header 複製過來

cp -a ~/buildroot-2012.08/output/toolchain/linux/ .

xvii. 更新 uClibc 組態

vim .config
# 修改以下地方
KERNEL_HEADERS="/home/kito/buildroot-2012.08/output/toolchain/linux/include"
改成
KERNEL_HEADERS="/home/kito/arm-linux-gcc/uClibc-0.9.31.1/linux/include"

RUNTIME_PREFIX="/"
改成
RUNTIME_PREFIX="/home/kito/arm-linux-gcc/arm-linux-gcc/"

DEVEL_PREFIX="/usr/"
改成
DEVEL_PREFIX="/home/kito/arm-linux-gcc/arm-linux-gcc/"

CROSS_COMPILER_PREFIX="/home/kito/buildroots/buildroot-2012.08/output/host/usr/bin/arm-unknown-linux-uclibc-"
改成
CROSS_COMPILER_PREFIX="/home/kito/arm-linux-gcc/arm-linux-gcc/bin/arm-unknown-linux-uclibc-"
# 然後存檔離開

xviii. 重新 build uClibc !

make -j8
make install

xix. 往後更動 gcc 步驟

大功告成,往後如果有更動到 gcc
uClibc 必須清掉重新 build 一次
make clean
make -j8
make install

2012年10月30日 星期二

LLVM 程式員手冊 - 重要及有用的 LLVM API

目錄

  1. 簡介
  2. 背景知識
  3. 重要跟有用的 LLVM API
  4. 為你的程式挑個正確的資料結構
  5. 常用操作的小提示集
  6. 執行緒與 LLVM
  7. 進階議題
  8. LLVM 核心類別的族譜
本文翻譯自 LLVM Programmer's Manual
Written by Chris Lattner, Dinakar Dhurjati, Gabor Greif, Joel Stanley, Reid Spencer and Owen Anderson

Translated by Kito Cheng (kito at 0xlab.org)
WIKI 版本

重要跟有用的 LLVM API

這邊會列出一些有用且玩弄 LLVM 前最好知道的一些 LLVM API。


有關 isa<>, cast<> and dyn_cast<> templates

在 LLVM 大量的使用自製的 RTTI,這些 templates 的功能主要類似於 dynamic_cast<> operator,但是 LLVM 自製版本的沒有 C++ 內建版本的一些缺點 (主要是 dynamic_cast<> 只對於有 v-table 譯註1 的 class 有用,沒有就不能動)。在 LLVM 這類東西經常會用到,所以你最好知道它是怎麼運作的。所有相關的 template 都定義在 llvm/Support/Casting.h 這個檔案 (通常不用自己去 include 這檔案 譯註2)

譯註1: 一個 Class 只有在有 Virtual Function 的時候才會有 v-table ,所以換句話說,沒 Virtual Function 的 Class 家族就完全不能用 dynamic_cast<>

譯註2:幾乎每個 LLVM Header 都會 include 到它,所以基本上你也不用自己去 include

isa<>

isa<> operator 的功能就跟 Java 中的 instanceof operator一樣,它會根據你丟進去的 pointer 或 reference 並且檢查是不是你所預期的類別來回傳 true 或 false ,在許多情況下這傢伙很好用(下面有例子)

cast<>

cast<> operator 主要是拿來轉型用的,並且會作檢查,當你從父類別 (base class) 轉型到子類別 (derived class) 失敗的時候會直接 assertion failure 炸掉,所以只能在你非常確定它真的可以正確的向下轉型的時候使用,下面則是一個使用 isa<> 跟 cast<> template 的例子:

/* 檢查一個 Value 是不是 Loop Invariant */
static bool isLoopInvariant(const Value *V, const Loop *L) {
  if (isa<Constant>(V) || isa<Argument>(V) || isa<GlobalValue>(V))
    return true;

  /* 不是 Constant 、 Argument 或 GlobalValue 則一定是一個 Instruction
      如果不是存在該迴圈中則代表是 Loop Invariant */
  return !L->contains(cast<Instruction>(V)->getParent());
}

註:不要使用 isa<> 然後接著 cast<>,這種情況請直接使用 dyn_cast<> operator

dyn_cast<>

dyn_cast<> operator 主要是拿來轉型用的,並且會作檢查,當你從父類別 (base class) 轉型到子類別 (derived class) 失敗的時候會回傳 NULL pointer,所以上你不能餵 Reference 進去,而它整個功能就跟 C++ 的 dynamic_cast<> operator 非常類似,而且使用情境一樣,通常 dyn_cast<> operator 可以直接拿來塞在 if 判斷式或著是其它塞條件判斷式的地方,下面舉個例子:
/* 如果 Val 可以轉型成 AllocationInst */
if (AllocationInst *AI = dyn_cast<AllocationInst>(Val)) {
  /* ㄎㄎ,可以玩弄 AllocationInst 了 */
}

這樣子就可以有效的結合 isa<> 及 cast<> 變成一個 statement,方便吧~

註:dyn_cast<> operator 就像 C++ 的 dynamic_cast<> 或著是 Java 的 instanceof operator,常被濫用。千萬不要用一串的 dyn_cast<> + if/then/else 去檢查一堆類別。這種情況通常你可以直接用 InstVisitor 這個傢伙會比較方便又好看。

cast_or_null<>

cast_or_null<> operator 功能就跟 cast<> operator 一樣,唯一差別在於它可以塞 NULL pointer 進去,在某些情況下它還滿有用的。

dyn_cast_or_null<>

dyn_cast_or_null<> operator 功能就跟 dyn_cast<> operator 一樣,唯一差別在於它可以塞 NULL pointer 進去,在某些情況下它還滿有用的。

有關 isa<>, cast<> and dyn_cast<> templates 的結語

以上五個 template 能夠能來運作在任何 Class 上,不論它有沒有 v-table。如果你寫的 Class 也想要支援這些 template 的話參考這份文件: How to set up LLVM-style RTTI for your class hierarchy


字串傳遞 (StringRef 及 Twine Class)

雖然在 LLVM 中一般而言不用太多字串的操作,但在 LLVM 中一些重要的 API 參數中是用字串來傳遞,其中兩個重要的例子:
  1. Value Class :拿來命名指令或著函數之類的
  2. StringMap Class :在 LLVM 跟 Clang 經常被用到
這兩個 Class 基本上可以接受任何可能塞有 Null 字元的字串,不過它們不能直接轉換成 const char * 或著是 const string &,而許多 LLVM API 的參數通常是吃 StringRef 或著是 const Twine&。

The StringRef class


StringRef 是拿來表示常數字串(字元陣列加上一個長度)用的,支援許多 std::string 的操作,並且大部分不需要額外的 heap 空間。

它可以透過 Implicitly Constructor 來直接吃 C style null-terminated 字串或著是 std::string,或著是一個字元陣列加上一個長度。
例如 StringRef 的 find 函數宣告如下:
  iterator find(StringRef Key);
然後呼叫方可以用下面任意一個方式呼叫
  Map.find("foo");                 // Lookup "foo"
  Map.find(std::string("bar"));    // Lookup "bar"
  Map.find(StringRef("\0baz", 4)); // Lookup "\0baz"
而通常 API 也是回傳 StringRef ,如果你需要轉換成 std::string 的話要使用 str 函數詳細的資訊自己去爬一下llvm/ADT/StringRef.h。

大部分情況下請直接使用 StringRef ,主要是它字串跟物件本身是分離的,也因此在 LLVM 程式碼或 API 中可以發現它幾乎都是直接 pass by value 傳遞。

Twine Class


Twine Class 是一個高效能的字串串接 API。例如在 LLVM 慣例中指令名稱的結尾通常是令一個指令的名稱,範例如下:
    New = CmpInst::Create(..., SO->getName() + ".cmp");

而 Twine Class 則是一個高效率且輕量級建立於 stack 上的 rope 譯註1,Twine 可以由兩個字串的 operator+ 來隱式建構 (例如 C-style strings, std::string 或著是 StringRef)。Twine 主要把實際的字串串接動作延遲到實際要需要的時候才進行,這樣可以有效避免不必要中間暫存結果的 Heap 分配譯註2。詳細可以去挖 llvm/ADT/Twine.h 檔案來啃。

如果是跟 StringRef 互動的話 Twine只會紀錄指標並且幾乎不需要額外的記憶體,他們兩譯註3主要就是設計來快速有效的傳遞串接字串。

譯註1:一種拿來實作大量字串儲存的資料結構,詳見 wiki 說明 Rope
譯註2:直接看下面範例可以了解 Twine 省了啥:
/* 僅紀錄 abc 及 def 的 pointer, 不實際進行串接動作 */
Twine t1 = "abc" + "def";

/* Twine t1 跟 const char * "xyz" 串接, 但也不進行實際串接動作 */
Twine t2 = t1 + "xyz";

/* 實際到需要的時候內部才會進行串接動作!, 可避免到中間 abcdef 這個暫存字串出現 */
std::cout << t2.str();
譯註3:指 StringRef 跟 Twine 這對好兄弟


DEBUG() macro 跟 -debug 選項


通常在撰寫你的 pass 的時候你會放一堆拿來 debug 用的輸出程式碼,當它正式運作的時候又會想砍掉那串,但等到某天你發現它有 bug 或著是又要開始寫新功能的時候又要加進去那串 debug 用的輸出程式碼。。。

所以很自然的你會不希望砍掉那堆程式碼,但你又不想要它隨時的輸出一堆訊息,一些常見的作法就是把它註解掉,然後要的時候又把那個註解拿掉譯註1。

譯註1:直接看下面 code
/* 例如把輸出的部份用個 ifdef 包裝 */
#ifdef DEBUG
 fprintf(stderr, "Debug Debug Debug");
#endif

/* 或著好看一點用 marco 包起來 */
#ifdef DEBUG
#define D(arg...) fprintf(stderr, __VA_ARGS__)
#else
#define D(arg...)
#endif

D("Debug Debug Debug");

在 "llvm/Support/Debug.h" 這個檔案中提供了一個 DEBUG() 來漂亮的解決這一類的問題,基本上你可以塞任何程式碼到 DEBUG 當參數,而包在裡面的程式碼只會在執行 opt 加上 -debug 參數的時候會吐出東西:

  DEBUG(errs() << "媽!我在這裡!\n");

所以你可以像這樣去跑你的 pass :

$ opt < a.bc > /dev/null -mypass
<沒有輸出>
$ opt < a.bc > /dev/null -mypass -debug
媽!我在這裡!

使用 DEBUG() Marco 取代自幹解法讓你不用弄一堆命令列參數譯註2,在你使用最佳化類型建置 LLVM 時,DEBUG() Marco 則會整個關閉,進而不會影響任何的效能(所以你也不要在 DEBUG 裡面有 side-effects譯註3!)。

譯註2:GCC 就是這樣幹。。。內部使用的 Debug 命令列參數無敵多。。。可以去 <gcc-source>/gcc/common.opt 參觀所有的命令列列表XD…

譯註3:大致上就是都不要更動到任何變數的值,不論區域或全域。

另一個 DEBUG() Marco 的方便東東就是當你在 gdb 中 debug LLVM 的時候只要輸入 "set DebugFlag=0" 或著是 "set DebugFlag=1" 就可以控制 DEBUG 的開關。

使用 DEBUG_TYPE 及 -debug-only 選項來細部控制 debug 資訊

有些時候你只想要 debug 自己的程式,而 -debug 又吐出全世界的錯誤訊息(例如在 Code Gen 的階段的時候),如果你想要細部控制 debug 資訊的話你就需要定義 DEBUG_TYPE 這個 marco 以及 -debug-only 選項,下面是使用範例:

#undef  DEBUG_TYPE
DEBUG(errs() << "No debug type\n");
#define DEBUG_TYPE "foo"
DEBUG(errs() << "'foo' debug type\n");
#undef  DEBUG_TYPE
#define DEBUG_TYPE "bar"
DEBUG(errs() << "'bar' debug type\n"));
#undef  DEBUG_TYPE
#define DEBUG_TYPE ""
DEBUG(errs() << "No debug type (2)\n");
然後接著你可以這樣跑你的 pass :
$ opt < a.bc > /dev/null -mypass
<沒有輸出>
$ opt < a.bc > /dev/null -mypass -debug
No debug type
'foo' debug type
'bar' debug type
No debug type (2)
$ opt < a.bc > /dev/null -mypass -debug-only=foo
'foo' debug type
$ opt < a.bc > /dev/null -mypass -debug-only=bar
'bar' debug type

當然在實務上你只需要再程式碼的最上方定義 DEBUG_TYPE 即可,這樣就可以為你的整個模組定義 debug type,(記得要放在#include "llvm/Support/Debug.h" 之前,通常你應該不會想用到醜不拉機的 #undef),然後最好把名稱取的有意義一點,不要用 foo 或 bar 這類沒營養的名字,主要是因為目前沒有任何機制去避免 DEBUG_TYPE 撞名的問題,如果兩個不同的模組使用同樣的 DEBUG_TYPE 名稱,則它們會被一起啟動,例如所有在 instruction scheduling 的 debug 資訊都會在 -debug-type=InstrSched 的時候一起噴出來,而那堆程式碼是散落在許多檔案當中。

DEBUG_WITH_TYPE 這個 Marco 則可以用在你想為某些 DEBUG 資訊設定特定 DEBUG_TYPE 時可以用,這個 Marco 比 DEBUG 多一個參數,第一個參數可以指定 DEBUG_TYPE,下面則是它的使用範例:
DEBUG_WITH_TYPE("", errs() << "No debug type\n");
DEBUG_WITH_TYPE("foo", errs() << "'foo' debug type\n");
DEBUG_WITH_TYPE("bar", errs() << "'bar' debug type\n"));
DEBUG_WITH_TYPE("", errs() << "No debug type (2)\n");


Statistic Class 及 -stats 選項


在 llvm/ADT/Statistic.h 這個檔案中提供一個叫做 Statistic 的 Class,他是專門拿來提供 LLVM 來紀錄各種最佳化對於程式有無實質上的改進。

你會在你的 pass 中處理一些東西,然後通常你會對於某些最佳畫到底執行幾次感興趣,雖然你可以直接在某些重要的函數中插入一些 code 去統計,但這樣的方式實在是有點鳥,而使用 Statistic Class 則可以讓你可以很簡單的去追蹤一些資訊,然後統一的在 pass 執行完後輸出。

下面是一些使用 Statistic class 的範例,他們基本上可以這樣用:

定義一個你的 statistic :
#define DEBUG_TYPE "mypassname"   // 這行 code 記得塞在所有 #include 前面
STATISTIC(NumXForms, "The # of times I did stuff");
STATISTIC Macro 定義了一個全域的靜態變數,其變數名稱如第一個參數,然後這個 Pass 的名稱它會直接從 DEBUG_TYPE 拿,它的描述則是放在第二個參數,這個變數實際上就像是個 unsigned integer 一 當你要執行一些最佳化或轉換的時候,遞增一下這個變數:
++NumXForms;   // 我做了某些事!
然後接著你只要在執行 opt 時加入 -stats 參數:
$ opt -stats -mypassname < program.bc > /dev/null
... statistics output ...

當你用 opt 跑某些測試時他會出現類似下面的統計報告:

   7646 bitcodewriter   - Number of normal instructions
    725 bitcodewriter   - Number of oversized instructions
 129996 bitcodewriter   - Number of bitcode bytes written
   2817 raise           - Number of insts DCEd or constprop'd
   3213 raise           - Number of cast-of-self removed
   5046 raise           - Number of expression trees converted
     75 raise           - Number of other getelementptr's formed
    138 raise           - Number of load/store peepholes
     42 deadtypeelim    - Number of unused typenames removed from symtab
    392 funcresolve     - Number of varargs functions resolved
     27 globaldce       - Number of global variables removed
      2 adce            - Number of basic blocks removed
    134 cee             - Number of branches revectored
     49 cee             - Number of setcc instruction eliminated
    532 gcse            - Number of loads removed
   2919 gcse            - Number of instructions removed
     86 indvars         - Number of canonical indvars added
     87 indvars         - Number of aux indvars removed
     25 instcombine     - Number of dead inst eliminate
    434 instcombine     - Number of insts combined
    248 licm            - Number of load insts hoisted
   1298 licm            - Number of insts hoisted to a loop pre-header
      3 licm            - Number of insts hoisted to multiple loop preds (bad, no loop pre-header)
     75 mem2reg         - Number of alloca's promoted
   1444 cfgsimplify     - Number of blocks simplified
由上面的統計輸出可以看出,程式執行了許多最佳化,而統一的界面讓這件事變得很容易,在你的 pass 中使用這個統一的界面將會使得你的程式碼更好維護!


在 Debug 程式的時候觀看某些 Graph

在 LLVM 當中許多重要的資料結構都是 Graph:例如 CFG 由一堆 Basic Block 組成,在 Instruction Selection 時使用的 DAG,在對 Compiler 除錯的情況下,如果能視覺化的看到內部的 Graph 則會使得除錯變得容易許多。

LLVM 提供許多的 Callback 提供 Debug 的時候用,例如你呼叫 Function::viewCFG() 這個函數,目前的 LLVM 會跳出一個視窗上面畫有該函數精美的 CFG,圖中的節點還會放置著 Basic Block 中的所有指令,而 Function::viewCFGOnly() 則可以讓你只看 Basic Block,不要顯示裡面的指令,類似的東西還有 MachineFunction::viewCFG() , MachineFunction::viewCFGOnly() 以及SelectionDAG::viewGraph() 這幾個函數,在 GDB 中你只要使用 DAG.viewGraph() 就會跳出視窗並且顯示出來,所以你也可以試著將那些函數呼叫塞到你正在 Debug 的部份。

要讓這個功能動起來事實上你可能需要一些額外的設定,例如在 Unix-linke 系統上需要安裝 graphviz 套件,並且確定 dot 跟 gv 這兩隻程是在你的 PATH 中,如果你在 Mac OS/X 的話,可以下載並安裝 Mac OS/X 的 graphviz 套件,然後加到 /Applications/Graphviz.app/Contents/MacOS/ (或任何你安裝的地方)到你的 PATH,一旦你系統的 PATH 設定好,在重新執行一次 LLVM configure script,並且重新建置 LLVM 就可以啟動這個好用的功能了!

SelectionDAG 部份則有一些方便你定位 Graph 中某些 Node 的功能,在 GDB 中如果你先呼叫 DAG.setGraphColor(node, "color"),再呼叫 DAG.viewGraph() 就會將你想看的 Node 標上指定的顏色(你可以在這個網頁找到color 的列表),事實上你還可以呼叫 DAG.setGraphAttrs(node, "attributes") 更詳細的去設定 Node 的屬性(可參考graphviz 的網頁),如果你想要回復預設 Graph 屬性的話可以呼叫DAG.clearGraphAttrs() 。


LLVM 程式員手冊 - 簡介及背景知識

目錄

  1. 簡介
  2. 背景知識
  3. 重要跟有用的 LLVM API
  4. 為你的程式挑個正確的資料結構
  5. 常用操作的小提示集
  6. 執行緒與 LLVM
  7. 進階議題
  8. LLVM 核心類別的族譜
本文翻譯自 LLVM Programmer's Manual
Written by Chris Lattner, Dinakar Dhurjati, Gabor Greif, Joel Stanley, Reid Spencer and Owen Anderson

Translated by Kito Cheng (kito at 0xlab.org)
WIKI 版本

簡介

這份文件主要列出一些 LLVM 中重要的類別跟界面,而這邊並不會解釋 LLVM 是甚麼東西,它內部怎麼運作以及 LLVM 的程式碼看起來如何。對於本文件的閱讀者我們假設你對於 LLVM 已經有一些基礎的了解,並且對於寫最佳化、分析或玩弄程式碼有興趣。

這份文件主要引導你如何擴充 LLVM 來達到你想要作的事。另外閱讀這份文件並不能取代啃 Source Code。如果你想看某個 Class 有哪些 Method 並且在幹嘛,那建議可以直接去看線上 doxygen 文件比較符合你的需求。
接下來的第一個章節主要介紹一些背景知識,第二章節則列出一些 LLVM 中核心的一些 Class ,未來這份文件會撰寫有關如何擴充整個 LLVM ,例如使用 Dominator 的資訊, Control Flow Graph 的走訪以及一些有用的小工具例如 InstVisitor template。

背景知識

這個章節放了一些有幫助你玩弄 LLVM 相關資訊的連結,但裡面沒提到 LLVM 相關的 API。

譯註:會寫 C++ 的話直接跳過吧,另外沒用過 STL 的話不算在會 C++ 的範圍內

The C++ Standard Template Library

LLVM 大量使用 C++ 的 Standard Template Library (STL),所以基本上你需要一些對於 C++ STL 的基礎知識以及一些相關使用慣例,下面提供一些相關資訊的連結可以給你惡補一下。
下面是惡補專區:
開始玩弄 LLVM 前最好也先閱讀一下 LLVM Coding Standards guide ,這份文件主要是讓你寫出好維護又好讀的 Code,而不是去規定你 { 跟 } 要怎麼放。

其它有用的連結

Using static and shared libraries across platforms